A system and method for processing real-time video from a medical imaging device and detecting objects within the video

The integration of an object detector network with an adversarial generation network in a computer-implemented system addresses the limitations of existing medical imaging object detection systems by enhancing accuracy, reducing false positives, and enabling real-time processing.

JP7692510B2Active Publication Date: 2025-06-13COSMO ARTIFICIAL INTELLIGENCE AI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024037840
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-28
Filing Date
2024-03-12
Publication Date
2025-06-13
Estimated Expiration
2039-06-11

AI Technical Summary

Technical Problem

Existing object detection systems in medical imaging struggle with accurate location identification of objects, reliance on manual annotation for training, high false positive rates, and delayed processing, making them inefficient for real-time applications.

Method used

A computer-implemented system that combines an object detector network with an adversarial generation network to improve object detection and location accuracy, reduce false positives, and enable real-time processing by using a two-loop training technique and integrating the networks for simultaneous feature detection and true/false positive discrimination.

Benefits of technology

The system achieves improved object detection accuracy, reduces false positives, and enables real-time processing of medical images, enhancing the efficiency and reliability of medical imaging analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692510000001
    Figure 0007692510000001
  • Figure 0007692510000002
    Figure 0007692510000002
  • Figure 0007692510000003
    Figure 0007692510000003
Patent Text Reader

Abstract

To provide computer-implemented systems and methods for training generative adversarial networks for use in medical image analysis.SOLUTION: A system includes an input port for receiving real-time video obtained from a medical image device 103, a first bus for transferring the received real-time video, and at least one processor configured to receive the real-time video from the first bus, perform object detection by applying a trained neural network to frames of the received real-time video, and overlay a border indicating a location of at least one detected object in the frames. The system also includes a second bus for receiving the video with the overlaid border, an output port for outputting the video with the overlaid border from the second bus to an external display, and a third bus for directly transmitting the received real-time video to the output port.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of neural networks and the use of such networks for image analysis and object detection. More particularly and without limitation, the present disclosure relates to computer-implemented systems and methods for training adversarial generative networks and processing real-time video. The systems and methods disclosed herein and the trained neural networks can be used in a variety of applications and vision systems, such as systems that benefit from medical image analysis and accurate object detection capabilities.

Background Art

[0002] In many object detection systems, objects are detected within an image. The object of interest may be a person, place, or thing. In some applications, such as medical image analysis and diagnosis, the location of the object is equally important. However, computer-implemented systems that utilize image classifiers typically cannot identify or provide the location of the detected object. Thus, existing systems that use only image classifiers are not very useful.

[0003] Furthermore, training techniques for object detection may rely on manually annotated training sets. Such annotation takes time when the detection network to be trained is based on bounding boxes, such as the You Only Look Once (YOLO) architecture, the Single Shot Detector (SSD) architecture, or the like. Thus, large datasets are difficult to annotate for training, which often results in neural networks being trained on relatively small datasets, resulting in reduced accuracy.

[0004] In the case of a computer-implemented system, existing medical imaging is typically built on a single detector network. Thus, once detection is performed, the network simply outputs the detection to, for example, a physician or other medical personnel. However, such detections may be false positives, such as non-polyp or similar ones in endoscopy. Such a system does not provide a separate network for discriminating true positives from false positives.

[0005] Furthermore, object detectors based on neural networks typically supply features identified by the neural network to the detector, and the detector may have a second neural network. However, such networks are often inaccurate because feature detection is performed by a generalized network and only the detector part is specialized.

[0006] Finally, many existing object detectors function with delays. For example, medical images may be captured and saved before analysis. However, some medical procedures, such as endoscopy, are diagnosed in real time. As a result, these systems are usually difficult to apply in the required real-time manner. SUMMARY OF THE INVENTION

[0007] In view of the above, embodiments of the present disclosure provide computer-implemented systems and methods for training adversarial generation networks and using them for applications such as medical image analysis. The systems and methods of the present disclosure provide benefits including improved object detection and location information in comparison with existing systems and techniques.

[0008] According to some embodiments, a computer-implemented system is provided that includes, along with the location, an object detector network that identifies a feature of an object (i.e., an abnormality or an object of the object), and an adversarial network that discriminates true positives from false positives. Further, embodiments of the present disclosure also provide a two-loop technique for training the object detector network. This training process uses annotation based on detection consideration so that manual annotation can occur much faster and thus with a relatively large dataset. Further, this process can be used to train an adversarial generation network to discriminate true positives from false positives.

[0009] In addition, a disclosed system combines an object detector network with an adversarial generation network. By combining such networks, false positives are discriminated from true positives, and thus a relatively accurate output can be provided. By reducing false positives, a physician or other medical practitioner can apply increased attention to the output from the network due to the increased accuracy.

[0010] Further, embodiments of the present disclosure do not use general feature identification by one neural network combined with a specialized detector. Rather, a single, seamless neural network is trained for the object detector portion, resulting in not only relatively high specialization but also increased accuracy and efficiency.

[0011] Finally, embodiments of the present disclosure are configured to display real-time video (such as an endoscopic examination video or other medical images) along with object detection on a single display. Thus, embodiments of the present disclosure provide a video bypass to minimize potential problems resulting from errors associated with object detectors and other potential drawbacks. Further, object detection can be displayed in a special manner designed to relatively well attract the attention of a physician or other medical personnel.

[0012] In one embodiment, a system for processing real-time video may include an input port for receiving real-time video, a first bus for transferring the received real-time video, at least one processor configured to receive the real-time video from the first bus, perform object detection on frames of the received real-time video, and overlay a boundary notifying the location of at least one detected object within the frame, a second bus for receiving the video with the overlaid boundary, an output port for outputting the video with the overlaid boundary from the second bus to an external device, and a third bus for directly transmitting the received real-time video to the output port.

[0013] In some embodiments, the third bus may be activated upon receipt of an error signal from at least one processor.

[0014] In any of the embodiments, at least one detected object may be an abnormality. In such embodiments, the abnormality can have a formation on human tissue or a formation of human tissue. In addition to, or instead of, this, the abnormality can have a change in human tissue from one type of cell to another type of cell. In addition to, or instead of, this, the abnormality can also have an absence of human tissue from where the human tissue is expected.

[0015] In any of the embodiments, the abnormality can have a lesion. For example, the lesion can have a polypoid lesion or a non-polypoid lesion.

[0016] In any of the embodiments, the overlaid boundary may have a graphical pattern around the region of the image that includes at least one detected object, in which case the pattern is displayed in a first color. In such an embodiment, after a predetermined time has elapsed, at least one processor changes the pattern to be displayed in a second color when at least one detected object is a true positive, and further changes the pattern to be displayed in a third color when at least one detected object is a false positive. In such an embodiment, at least one processor can be further configured to send a command to one or more speakers to generate a sound when the pattern of the boundary is changed. In such an embodiment, at least one of the duration, tone, frequency, and amplitude of the sound can depend on whether at least one detected object is a true positive or a false positive. In addition to or instead of the sound, at least one processor can be further configured to send a command to at least one wearable device to vibrate when the pattern of the boundary is changed. In such an embodiment, at least one of the duration, frequency, and amplitude of the vibration depends on whether at least one detected object is a true positive or a false positive.

[0017] In one embodiment, a system for processing real-time video includes an input port that receives real-time video, and at least one processor configured to receive the real-time video from the input port and apply a trained neural network to frames of the received real-time video to perform object detection and overlay a boundary that notifies the location of at least one detected object in the frame. The system can also include an output port that outputs the video with the overlaid boundary from the at least one processor to an external display, and an input device that receives sensitivity settings from a user. The processor can be further configured to adjust at least one parameter of the trained neural network in response to the sensitivity settings.

[0018] In some embodiments, at least one detected object can be an abnormality. In such embodiments, the abnormality can have a formation on human tissue or a formation of human tissue. Additionally or alternatively, the abnormality can have a change in human tissue from one type of cell to another type of cell. Additionally or alternatively, the abnormality can have an absence of human tissue from where the human tissue is expected.

[0019] In any of the embodiments, the abnormality can have a lesion. For example, the lesion can have a polypoid lesion or a non-polypoid lesion.

[0020] Further objects and advantages of the present disclosure will in part be apparent from the following detailed description, and in part will be learned by practice of the present disclosure. The objects and advantages of the present disclosure will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.

[0021] It should be understood that the foregoing general description and the following detailed description are merely illustrative and explanatory and are not restrictive of the disclosed embodiments.

[0022] The accompanying drawings, which form a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles and features of the disclosed embodiments. The accompanying drawings are as follows.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 8C

DETAILED DESCRIPTION OF THE INVENTION

[0024] The disclosed embodiments relate to computer-implemented systems and methods for training and using adversarial generation networks. Advantageously, the exemplary implementations can provide an improved trained network and fast and efficient object detection. Also, the embodiments of the present disclosure can provide improved object detection for medical image analysis with reduced false positives.

[0025] The embodiments of the present disclosure can be implemented and used in various applications and vision systems. For example, the embodiments of the present disclosure can be implemented for medical image analysis systems and other types of systems that can benefit from object detection where an object can be a true positive or a false positive. The embodiments of the present disclosure are described herein with a general reference to medical image analysis and endoscopy, but it should be understood that the embodiments can be applied to other medical imaging procedures such as upper endoscopy such as gastroscopy, colonoscopy, enteroscopy, and esophagoscopy. Furthermore, the embodiments of the present disclosure are not limited to other environments and vision systems such as those for LIDAR, surveillance, autopilot, and other imaging systems or those including them.

[0026] According to one aspect of the present disclosure, a computer-implemented system for training an adversarial generation network using an image including an expression of a feature of an object is provided. The system can include at least one memory configured to store instructions and at least one processor configured to execute the instructions (see, e.g., FIGS. 1 and 6). The at least one processor can provide a first plurality of images. For example, the at least one processor can extract the first plurality of images from one or more databases. In addition to, or instead of, this, the first plurality of images can also have a plurality of frames extracted from one or more videos.

[0027] As used herein, the term "image" means any digital representation of a scene or view. The digital representation can be encoded in any suitable format, such as the JPEG (Joint Photographic Experts Group) format, GIF (Graphic Interchange Format), bitmap format, SVG (Scalable Vector Graphics) format, EPS (Encapsulated PostScript) format, or the like. Similarly, the term "video" also means any digital representation of an object scene or area composed of a plurality of consecutive images. The digital representation can be encoded in any suitable format, such as the MPEG (Moving Picture Experts Group) format, flash video format, AVI (Audio Video Interleave) format, or the like. In some embodiments, a series of images can be paired with audio.

[0028] The first plurality of images can include a representation of a feature of an object (i.e., an abnormality or an object of the object) and an indicator of the location of the feature of the object within the image of the first plurality of images. For example, the feature of the object can have an abnormality on human tissue or an abnormality of human tissue. In some embodiments, the feature of the object can have an object, such as a vehicle, a person, or other entity.

[0029] According to the present disclosure, "abnormality" can include the formation on human tissue or the formation of human tissue, the change in human tissue from one type of cell to another type of cell, and / or the absence of human tissue from the location where the human tissue is expected. For example, the growth of a tumor or other tissue may have an abnormality because there are more cells than expected. Similarly, a wound or other change in cell type may have an abnormality because blood cells are present at a location outside the expected location (i.e., outside the capillary). Similarly, a depression within human tissue may also have an abnormality because cells are not present within the expected location, resulting in the depression.

[0030] In some embodiments, the abnormality can have a lesion. The lesion can have a lesion of the gastrointestinal mucosa. The lesion may be classified histologically (e.g., based on the Vienna classification), morphologically (e.g., based on the Paris classification), and / or structurally (e.g., as serrated or non-serrated). The Paris classification includes polypoid and non-polypoid lesions. The polypoid lesion can have a protruding, pedunculated and protruding, or sessile lesion. The non-polypoid lesion can have a lesion with a raised surface, a flat surface, a shallowly depressed surface, or a dug-in surface.

[0031] In relation to the detected abnormalities, serrated lesions can have sessile serrated adenomas (SSA), traditional serrated adenomas (TSA), hyperplastic polyps (HP), fibroblastic polyps (FP), or mixed polyps (MP). According to the Vienna classification, the abnormalities are divided into five categories: (category 1) negative for neoplasia / dysplasia, (category 2) indeterminate for neoplasia / dysplasia, (category 3) non-invasive low-level neoplasia (low-level adenoma / dysplasia), (category 4) high-level adenoma / dysplasia, non-invasive carcinoma (carcinoma in situ), or suspicion of invasive carcinoma, i.e., mucosal high-level neoplasia, and (category 5) invasive neoplasia, intramucosal carcinoma, submucosal carcinoma, or the like.

[0032] Indicators of the location of an abnormality or a feature of the subject can have points (e.g., coordinates) or regions (e.g., rectangles, squares, ellipses, or any other regular or irregular shape). The indicators can have manual annotations on the image or manual annotation of the image. In some embodiments, the first plurality of images can have medical images, such as images of gastrointestinal organs or other organs or areas of human tissue. The images can be generated from a medical imaging device, such as those used during upper endoscopy procedures such as endoscopy, gastroscopy, colonoscopy, enteroscopy, or esophagoscopy. In such embodiments, when the feature of the subject is a lesion or other abnormality, a physician or other healthcare provider can annotate the image to place an indicator of the abnormality within the image.

[0033] One or more processors of the system can use a first plurality of images and indicators of the target feature to train an object detection network to detect the target feature. For example, the object detection network can have a neural network having one or more layers configured to receive an image as input and output an indicator of the location of the target feature. In some embodiments, the object detection network can have a convolutional network.

[0034] Training the object detection network can include adjusting the weights of one or more nodes of the network and / or adjusting the activation (or, transfer) function of one or more nodes of the network. For example, the weights of the object detection network can be adjusted to minimize a loss function associated with the network. In some embodiments, the loss function can have a squared loss function, a hinge loss function, a logistic loss function, a cross-entropy loss function, or any other suitable loss function, or a combination of loss functions. In some embodiments, the activation (or, transfer) function of the object detection network can be changed to improve the fit between one or more models of one or more nodes and the input to one or more nodes. For example, one or more processors may increase or decrease the exponent of a polynomial function associated with one or more nodes, change the associated function from one type to another (e.g., from a polynomial to an exponential function, from a logarithmic function to a polynomial function, or the like), or perform any other adjustment to one or more models of one or more nodes.

[0035] One or more system processors can further provide a second plurality of images that include representations of the features of interest. For example, one or more processors can extract the first plurality of images from one or more databases, whether the one or more databases are the same as those in which the first plurality of images were stored or one or more different databases. In addition to, or instead of, this, the second plurality of images can have a plurality of frames extracted from one or more videos, whether the one or more videos are the same as those used to extract the first plurality of images or one or more different videos.

[0036] In some embodiments, the second plurality of images can have medical images, such as images from an endoscopic device. In such embodiments, the features of interest can have lesions or other abnormalities.

[0037] In some embodiments, the second plurality of images can have a greater number of images than those included within the first plurality of images. For example, the second plurality of images can include at least a hundred times more images than the first plurality of images. In some embodiments, the second plurality of images can at least partially include the first plurality, or can be images different from the first plurality. In embodiments where the second plurality of images are at least partially extracted from one or more videos from which at least a portion of the first plurality of images were extracted, the second plurality of images can have frames different from the first plurality from one or more of the same videos.

[0038] One or more processors can apply a trained object detection network to a second plurality of images to generate a first plurality of detections of target features. For example, in embodiments where the trained object detection network has a neural network, at least one processor can input the second plurality of images into the network and receive detections. The detections can have indicators of the locations of the target features within the second plurality of images. If the second plurality of images do not contain the target features, the indicator can have a null indicator or other indicator indicating that the target features are absent.

[0039] One or more processors can further provide manually set verification of true positives and false positives in relation to the first plurality of detections. For example, the verification can be extracted from one or more databases or received as input. In embodiments where the target feature has a lesion or other abnormality, the verification can be input by a physician or other healthcare provider. For example, one or more processors can output the detections to a physician or other healthcare provider for display and receive the verification in response to the displayed detections.

[0040] One or more system processors can use the verification of true positives and false positives in relation to the first plurality of detections to train an adversarial generation network. For example, the generation branch of the network can be trained to generate artificial representations of the target features. Thus, the generation branch can have a convolutional neural network.

[0041] Similar to the object detection network, the training of the generation branch can include the step of adjusting the weights of one or more nodes of the network and / or the step of adjusting the activation (or transfer) function of one or more nodes of the network. For example, as described above, the weights of the generation branch can be adjusted to minimize the loss function associated with the network. In addition to this, or instead of this, the activation (or transfer) function of the generation branch can be changed to improve the fit between one or more models of one or more nodes and the input to one or more nodes.

[0042] Furthermore, the adversarial branch of the network can be trained to discriminate true positives from false positives based on manual verification. For example, the adversarial branch can have a neural network that receives an image and one or more corresponding detections as input and generates a verification as output. In some embodiments, one or more processors can further retrain the generation network by providing a verification of false negatives for missed detections of features of an object in two or more images. By providing an artificial representation from the generation branch as input to the adversarial branch and recursively using the output from the adversarial branch, the adversarial branch and the generation branch can perform unsupervised learning.

[0043] Similar to the generation branch, the training of the adversarial branch can include the step of adjusting the weights of one or more nodes of the network and / or the step of adjusting the activation (or transfer) function of one or more nodes of the network. For example, as described above, the weights of the adversarial branch can be adjusted to minimize the loss function associated with the network. In addition to this, or instead of this, the activation (or transfer) function of the adversarial branch can be changed to improve the fit between one or more models of one or more nodes and the input to one or more nodes.

[0044] Thus, in embodiments where the feature of interest has a lesion or other abnormality, the generative branch can be trained to generate a representation of normality that looks like the abnormality, and the adversarial branch can be trained to discriminate artificial normality from the abnormality in the second plurality of images.

[0045] One or more system processors can retrain the adversarial generation network by using at least one additional set of images and detection of the feature of interest, along with additional manually set verification of true positives and false positives in relation to further detection of the feature of interest. For example, one or more processors can extract an additional set of images from one or more databases, whether the one or more databases are the same as those storing the first plurality of images and / or the second plurality of images or one or more different databases. In addition to, or instead of, this, the additional set of images can have a plurality of frames extracted from one or more videos, whether the one or more videos are the same as those used to extract the first plurality of images and / or the second plurality of images or one or more different videos. Similar to training, retraining of the adversarial branch can include further adjustment to the weights of one or more nodes of the network and / or further adjustment to the activation (or transmission) function of one or more nodes of the network.

[0046] According to another aspect of the present disclosure, a computer-implemented method is provided for training a neural network system to detect abnormalities in images of human organs. The method can be implemented by at least one processor (e.g., see processor 607 of FIG. 6).

[0047] According to an exemplary method, one or more processors can store a plurality of videos including manifestations of anomalies in a database. For example, the videos can include endoscopic examination videos. The videos can be encoded in one or more formats such as the MPEG (Moving Picture Expers Goup) format, the Flash video format, the AVI (Audio Video Interleave) format, or the like.

[0048] The method can further include the step of selecting, by one or more processors, a first subset of the plurality of videos. For example, one or more processors can randomly select the first subset. Alternatively, instead, one or more processors can use one or more indexes of the database to select the first subset. For example, one or more processors can select the first subset as videos indexed as including manifestations of anomalies.

[0049] The method can further include the step of applying, by one or more processors, a perception branch of an object detection network to frames of a first subset of the plurality of videos to generate detections of a first plurality of anomalies. For example, the object detection network can have a neural network trained to receive an image as an input and output a first plurality of detections. The first plurality of detections can have indicators of locations of anomalies within the frame, such as points or regions of detected anomalies. The absence of anomalies can result in a null indicator or other indicator of non - anomalies. The perception branch can have a neural network (e.g., a convolutional neural network) configured to detect polyps and output indicators of locations of any detected anomalies.

[0050] The method can further include the step of selecting, by one or more processors, a second subset of the plurality of videos. In some embodiments, the second subset may at least partially include the first subset, or may be videos different from the first subset.

[0051] The method can further include using, for training a generator network to generate a plurality of artificial representations of anomalies, the first plurality of detections and frames from a second subset of the plurality of videos. For example, the generator network can have a neural network configured to generate artificial representations. In some embodiments, the generator network can have a convolutional neural network. The plurality of artificial representations can be generated through residual learning.

[0052] As described above, training the generator network can include adjusting the weights of one or more nodes of the network and / or adjusting the activation (or transfer) function of one or more nodes of the network. For example, as described above, the weights of the generator network can be adjusted to minimize a loss function associated with the network. In addition to, or instead of, this, the activation (or transfer) function of the generator network can also be changed to improve the fit between one or more models of one or more nodes and the input to the one or more nodes.

[0053] The method can further include training, by one or more processors, an adversarial branch of a discriminator to discriminate between an artificial representation of an anomaly and a true representation of the anomaly. For example, the adversarial branch can have a neural network that receives a representation as input and outputs a notification as to whether the input representation is artificial or true. In some embodiments, the neural network can have a convolutional neural network.

[0054] Similar to the generation branch, the training of the adversarial branch of the discriminator network can include the step of adjusting the weights of one or more nodes of the network and / or the step of adjusting the activation (or transfer) function of one or more nodes of the network. For example, as described above, the weights of the adversarial branch of the discriminator network can be adjusted to minimize the loss function associated with the network. In addition to this, or instead of this, the activation (or transfer) function of the adversarial branch of the discriminator network can be changed to improve the fit between one or more models of one or more nodes and the input to one or more nodes.

[0055] The method can further include the step of applying the adversarial branch of the discriminator network to a plurality of artificial representations to generate a difference indicator between the artificial representation of the abnormality and the true representation of the abnormality contained within the frames of the second subset of the plurality of videos, by one or more processors. For example, the artificial representation can have a non-abnormal representation that appears similar to the abnormality. Thus, each artificial representation can provide a false representation of the abnormality that is very similar to the true representation of the abnormality. The adversarial branch can learn to identify the difference between non-abnormality (false representation) and abnormality (true representation), particularly non-abnormality similar to the abnormality.

[0056] The method can further include the step of applying the perceptual branch of the discriminator network to the artificial representation to generate a second plurality of abnormality detections, by one or more processors. Similar to the first plurality of detections, the second plurality of detections can have an indicator of the location of the abnormality in the artificial representation, such as the detected abnormality points or regions. The absence of an abnormality can result in a null indicator or other indicator of non-abnormality.

[0057] The method can further include the step of retraining the perception branch based on the difference indicator and the second plurality of detections. For example, retraining the perception branch can include the step of adjusting the weights of one or more nodes of the network and / or the step of adjusting the activation (or, transfer) function of one or more nodes of the network. For example, as described above, the weights of the perception branch can be adjusted to minimize the loss function associated with the network. In addition to this, or instead of this, the activation (or, transfer) function of the perception branch can be changed to improve the fit between one or more models of one or more nodes and the difference indicator and the second plurality of detections.

[0058] The exemplary method of training described above can generate a trained neural network system. The trained neural network system can form a part of a system used to detect features of a subject in an image of a human organ (for example, the neural network system can be implemented as a part of the overlay device 105 in FIG. 1). For example, such a system can include at least one memory configured to store instructions and at least one processor configured to execute the instructions. The at least one processor can select a frame from a video of a human organ. For example, the video can have an endoscopic examination video.

[0059] One or more system processors can apply the trained neural network system to the frame to generate at least one detection of a feature of the subject. In some embodiments, the feature of the subject can have an abnormality. The at least one detection can include an indicator of the location of the feature of the subject. For example, the location can have a point of the detected feature of the subject or a region including this. The neural network system can be trained to detect an abnormality as described above.

[0060] In some embodiments, one or more processors can further apply one or more additional classifiers and / or neural networks to the detected features of the object. For example, if the features of the object have a lesion, at least one processor can classify the lesion into one or more types (e.g., cancerous or non-cancerous, or the like). In addition to, or instead of, this, the neural network system can further output whether the detected features of the object are false positives or true positives.

[0061] One or more system processors can generate an indicator of the location of at least one detection on one of the frames. For example, the location of the features of the object can be extracted from the location indicators and graphical indicators arranged on the frame. In embodiments where the location has a point, the graphical indicator can have a circle, star, or any other shape arranged on the point. In embodiments where the location has a region, the graphical indicator can have a boundary around the region. In some embodiments, the shape or boundary may be animated, and thus the shape or boundary can be generated for a plurality of frames such that it not only tracks the location of the features of the object across the frame, but also appears in an animated state when the frames are shown in sequence. As will be described later, the graphical indicator can be paired with other indicators such as sound and / or vibration indicators.

[0062] Any aspect of the indicator may depend on the classification of the characteristics of the object, such as, for example, as one or more types, or as false or true positives, etc. Thus, the color, shape, pattern, or other aspect of the graphical indicator may depend on the classification. Also, in embodiments using sound and / or vibration indicators, the duration, frequency, and / or amplitude of the sound and / or vibration may depend on the classification.

[0063] One or more system processors can re-encode the frame as a video. Thus, after generating a (graphic) indicator and overlaying it on one or more frames, the frame can be reassembled as a video. Thus, one or more processors of the system can output the re-encoded video along with the indicator.

[0064] According to another aspect of the present disclosure, a computer-implemented system for processing real-time video (see, for example, FIGS. 1 and 6) will be described. The system can have an input port for real-time video. For example, the input port can have a video graphics array (VGA) port, a high-definition multimedia interface (HDMI (registered trademark)) port, a digital visual interface (DVI) port, a serial digital interface (SDI), or the like. The real-time video can include medical video. For example, the system can receive real-time video from an endoscope device.

[0065] The system can further have a first bus for transmitting the received real-time video. For example, the first bus can have a parallel connection or a serial connection and can be wired in a multi-drop topology or a daisy-chain topology. The first bus can have a PCI Express (Peripheral Component Interconnect Express) bus, a Universal Serial Bus (USB), an IEEE 1394 interface (FireWire (registered trademark)), or the like.

[0066] The system can have at least one processor configured to receive real-time video from the first bus, perform object detection on the frames of the received real-time video, and overlay a boundary notifying the location of at least one detected object within the frame. One or more processors can perform object detection by using a neural network system trained to generate at least one detection of an object. In some embodiments, at least one object can have a lesion or other abnormality. Thus, the neural network system can be trained to detect abnormalities as described above.

[0067] One or more processors can overlay a boundary as described above. For example, the boundary can surround a region containing an object, in which case the region is received by one or more processors along with at least one detection.

[0068] The system can further have a second bus to receive video, along with an overlaid boundary. For example, similar to the first bus, the second bus can have a parallel connection or a serial connection and can be wired in a multi-drop topology or a daisy-chain topology. Thus, similar to the first bus, the second bus can have a PCI Express (Peripheral Component Interconnect Express) bus, a Universal Serial Bus (USB), an IEEE 1394 interface (FireWire (registered trademark)), or something similar. The second bus may have the same type of bus as the first bus, or may have a different type of bus.

[0069] The system can further have an output port to output video to an external display, along with an overlaid boundary. The output port can have a VGA port, an HDMI (registered trademark) port, a DVI port, an SDI port, or something similar. Thus, the output port may be the same type of port as the input port, or may be a different type of port.

[0070] The system can have a third bus to directly send the received real-time video to the output port. The third bus can passively convey the real-time video from the input port to the output port so as to be enabled even when the entire system is turned off. In some embodiments, the third bus can be a predefined bus that is enabled when the entire system is in the off state. In such embodiments, the first and second buses may be activated when the entire system is started, and thus the third bus may be stopped. The third bus may be activated when the entire system is turned off or when an error signal is received from one or more processors. For example, when object detection implemented by a processor malfunctions, one or more processors can activate the third bus, thereby allowing continuous output of the real-time video stream without interruption due to the malfunction.

[0071] In some embodiments, the overlaid boundary can change across frames. For example, the overlaid boundary may have a two-dimensional shape displayed around the region of an image containing at least one detected object, in which case the boundary is the first color. After a predetermined time has elapsed, one or more processors can change the boundary to the second color if at least one detected object is a true positive, and to the third color if at least one detected object is a false positive. In addition to, or instead of, this, one or more processors can also change the boundary based on the classification of the detected object. For example, when the object has a lesion or other abnormality, the change may be based on whether the lesion or formation is cancerous or an abnormality in some other way.

[0072] In any of the above embodiments, the overlaid indicator can be paired with one or more additional indicators. For example, one or more processors can send commands to one or more speakers to generate sound when at least one object is detected. In embodiments where the boundary is changed, one or more processors can send commands when the boundary is changed. In such embodiments, at least one of the sound duration, tone, frequency, and amplitude can depend on whether at least one detected object is a true positive or a false positive. In addition to, or instead of, this, at least one of the sound duration, tone, frequency, and amplitude can also depend on the classification of the detected object.

[0073] In addition to, or instead of, this, one or more processors can send commands to at least one wearable device to vibrate when at least one object is detected. In embodiments where the boundary is changed, one or more processors can send commands when the boundary is changed. In such embodiments, at least one of the vibration duration, frequency, and amplitude can depend on whether at least one detected object is a true positive or a false positive. In addition to, or instead of, this, at least one of the vibration duration, frequency, and amplitude can also depend on the classification of the detected object.

[0074] According to another aspect of the present disclosure, a system for processing real-time video will be described. Similar to the processing system described above, the system has an input port for receiving real-time video, and at least one processor configured to receive the real-time video from the input port, apply a trained neural network to frames of the received real-time video to perform object detection, and overlay a boundary for notifying the location of at least one detected object within the frame, and an output port for outputting the video from the processor to an external display together with the overlaid boundary.

[0075] The system can further have an input device for receiving sensitivity settings from a user. For example, the input device can have a knob, one or more buttons, or any other device suitable for receiving one command for increasing the setting and another command for decreasing the setting.

[0076] One or more system processors can adjust at least one parameter of the trained neural network in response to the sensitivity setting. For example, one or more processors can adjust one or more weights of one or more nodes of the network to increase or decrease the number of detections generated by the network based on the sensitivity setting. In addition to or instead of this, one or more thresholds applied to the output layer of the network and / or to the detections received from the output layer of the network can be increased or decreased in response to the sensitivity setting. Thus, when the sensitivity setting is increased, one or more processors can decrease one or more thresholds so as to increase the number of detections generated by the network. Similarly, when the sensitivity setting is decreased, one or more processors can increase one or more thresholds so as to decrease the number of detections generated by the network.

[0077] FIG. 1 is a schematic representation of an exemplary system 100 that includes a pipeline for overlaying object detection on a video feed, which is consistent with an embodiment of the present disclosure. As shown in the example of FIG. 1, system 100 includes an operator 101 who controls an imaging device 103. In embodiments where the video feed has medical video, operator 101 can be a physician or other medical personnel. Imaging device 103 can be a medical imaging device such as an X-ray device, a computed tomography (CT) device, a magnetic resonance imaging (MRI) device, an endoscopy device, or any other medical imaging device that generates video or one or more images of a human body or a part thereof. Operator 101 can control imaging device 103, for example, by controlling the capture rate of device 103 and / or the movement of device 103 through or in relation to the human body. In some embodiments, imaging device 103 can have a Pill-Cam (trademark) device or other form of capsule endoscopy device instead of an external imaging device such as an X-ray device or an imaging device inserted through a cavity of the human body such as an endoscopy device.

[0078] As further depicted in FIG. 1, imaging device 103 can transmit the captured video or image to an overlay device 105. Overlay device 105 can have one or more processors for processing the video, as described above. Also, in some embodiments, operator 101 can control overlay device 105 in addition to imaging device 103, for example, by controlling the sensitivity of an object detector (not shown) of overlay device 105.

[0079] As depicted in FIG. 1, the overlay device 105 can expand the video received from the image device 103 and then transmit the expanded video to the display 107. In some embodiments, the expansion can have the above-described overlaying. Also, as further depicted in FIG. 1, the overlay device 105 can also be configured to directly relay the video from the image device 103 to the display 107. For example, the overlay device 105 can perform direct relaying under a predetermined condition, such as when an object detector (not shown) included in the overlay device 105 malfunctions. In addition to or instead of this, the overlay device 105 can perform direct relaying when the operator 101 inputs a command to the overlay device 105 to perform direct relaying. The command can be received via one or more buttons included on the overlay device 105 and / or through an input device such as a keyboard or the like.

[0080] FIG. 2 is a schematic representation of a two-phase training loop 200 for an object detection network, which is consistent with an embodiment of the present disclosure. The loop 200 can be implemented by one or more processors. As shown in FIG. 2, Phase I of the loop 200 can use a database 201 of images containing features of the object. In embodiments where the images are medical images, the features of the object can include abnormalities such as lesions.

[0081] As described above, the database 201 may store individual images and / or one or more videos, in which case each video includes a plurality of frames. In phase I of loop 200, one or more processors can extract a subset 203 of images and / or frames from the database 201. The one or more processors can select the subset 203 randomly or at least in part by using one or more patterns. For example, when the database 201 stores videos, the one or more processors can select one, two, or fewer similar numbers of frames from each video included in the subset 203.

[0082] As further depicted in FIG. 2, the feature indicator 205 can have an annotation for the subset 203. For example, the annotation can include a point of the feature of interest or a region containing it. In some embodiments, the operator can observe the video or image and manually input an annotation to one or more processors via an input device (e.g., any combination of a keyboard, mouse, touch screen, and display). The annotation can be stored as a data structure separate from the image in a format such as JSON, XML, text, or the like. For example, in embodiments where the image is a medical image, the operator may be a doctor or other medical professional. Although depicted as being added to the subset 203 after extraction, the subset 203 may be annotated before storage in the database 201 or at another previous time point. In such embodiments, the one or more processors can select the subset 203 by selecting an image in the database 201 having the feature indicator 205.

[0083] Subset 203 has a training set 207, along with a feature indicator 205. One or more processors can train a discriminator network 209 by using the training set 207. For example, the discriminator network 209 can have an object detection network as described above. As further described above, the training of the discriminator network can include steps of adjusting the weights of one or more nodes of the network and / or adjusting the activation (or, transmission) function of one or more nodes of the network. For example, the weights of the object detection network can be adjusted to minimize a loss function associated with the network. In another example, the activation (or, transmission) function of the object detection network can be changed to improve the fit between one or more models of one or more nodes and the inputs to the one or more nodes.

[0084] As shown in FIG. 2, in Phase II of loop 200, one or more processors can extract a subset 211 of images (and / or, frames) from the database 201. The subset 211 can have, at least in part, some or all of the images from the subset 203, or can have different subsets. In an embodiment where the subset 203 has multiple frames from one or more videos, the subset 211 can include adjacent or other frames from one or more of the same videos. The subset 211 can have, for example, a greater number of images than the subset 203, such as at least 100 times more images, a large number of images, etc.

[0085] One or more processors can apply discriminator network 209’ (representing discriminator network 209 after completion of the Phase I training) to subset 211 to generate a plurality of feature indicators 213. For example, the feature indicator 213 can have a point of the feature of the object detected by the discriminator network 209’ or a region including the same.

[0086] As further depicted in FIG. 2, verification 215 can have an annotation for the feature indicator 213. For example, the annotation can include an indicator as to whether each feature indicator is a true positive or a false positive. An image that did not have the detected feature of the object but included the feature of the object can be annotated as a false negative.

[0087] Subset 211 has a training set 217 together with feature indicator 213 and verification 215. One or more processors can train the adversarial generation network 219 by using the training set 217. For example, the adversarial generation network 219 can have a generation network and an adversarial network as described above. The training of the adversarial generation network can include training the generation network to generate an artificial representation of the feature of the object or a false feature of the object that appears similar to the true feature of the object, and training the adversarial network to discriminate the artificial representation from an actual representation such as those included within subset 211.

[0088] Although not depicted in FIG. 2, verification 215 can be further used to retrain discriminator network 209'. For example, the weights and / or activation (or transfer) functions of discriminator network 209' can be adjusted to remove detections in images annotated as false positives and / or can also be adjusted to generate detections in images annotated as false negatives.

[0089] FIG. 3 is a flowchart of an exemplary method 300 for training an object detection network. Method 300 can be executed by one or more processors. In step 301 of FIG. 3, at least one processor can provide a first plurality of images including representations of features of interest and indicators of the locations of the features of interest within the images of the first plurality of images. The indicators can have manually set indicators. The manually set indicators may be extracted from a database or may be received as input from an operator.

[0090] In step 303, at least one processor can train an object detection network to detect features of interest by using the first plurality of images and the indicators of the features of interest. For example, the object detection network can be trained as described above.

[0091] In step 305, at least one processor may provide a second plurality of images including representations of features of interest, where the second plurality of images has a greater number of images than those included within the first plurality of images. In some embodiments, the second plurality of images can at least partially overlap the first plurality of images. Alternatively, instead, the second plurality of images can be composed of images different from those within the first plurality.

[0092] In step 307, at least one processor can apply a trained object detection network to a second plurality of images to generate a first plurality of detections of the target feature. In some embodiments, as described above, the detections can include indicators of the locations of the detected target features. For example, the object detection network can optionally output one or more matrices, each matrix defining the coordinates and / or regions of any detected target feature, along with one or more associated confidence scores for each detection.

[0093] In step 309, at least one processor can provide manually set verifications of true positives and false positives in relation to the first plurality of detections. For example, at least one processor may extract the manually set verifications from a database or receive them as input from an operator.

[0094] In step 311, at least one processor can train an adversarial generation network by using the verifications of true positives and false positives in relation to the first plurality of detections. For example, the adversarial generation network can be trained as described above.

[0095] In step 313, at least one processor can retrain the adversarial generation network by using a further set of at least one image and detection of the target feature, along with further manually set verification of true positives and false positives in relation to further detection of the target feature. In some embodiments, the further set of images can overlap at least partially with the first plurality of images and / or the second plurality of images. Alternatively, instead, the further set of images may be composed of images different from those within the first plurality and those within the second plurality. Accordingly, step 313 can include applying the trained object detection network to a further set of images to generate further detection of the target feature, providing manually set detection of true positives and false positives in relation to the further detection, and retraining the adversarial generation network using the verification in relation to the further detection.

[0096] In a state consistent with the present disclosure, the exemplary method 300 can include further steps. For example, in some embodiments, the method 300 can include retraining the adversarial generation network by providing verification of false negatives for missed detection of the target feature in two or more images. Accordingly, the manually set verification extracted from the database or received as input can include verification of false negatives as well as true positives and false positives. False negatives can be used to retrain the adversarial generation network. In addition to or instead of this, false negatives can also be used to retrain the object detection network.

[0097] Figure 4 is a schematic representation of object detector 400. Object detector 400 can be implemented by one or more processors. As shown in Figure 4, object detector 400 can use a database 401 of videos containing features of interest. In embodiments where the images are medical images, the features of interest can include abnormalities such as lesions. In the example of Figure 4, database 401 has an endoscopy video database.

[0098] As further depicted in Figure 4, detector 400 can extract a subset 403 of videos from database 401. As described above in relation to Figure 2, subset 403 can be selected randomly and / or by using one or more patterns. Detector 400 can apply the perception branch 407 of discriminator network 405 to the frames of subset 403. Perception branch 407 can have an object detection network, as described above. Perception branch 407 can be trained to detect features of interest and to identify the locations (e.g., points or regions) associated with the detected features of interest. For example, perception branch 407 can detect abnormalities and output a bounding box containing the detected abnormalities.

[0099] As shown in FIG. 4, the perception branch 407 can output a detection 413. As described above, the detection 413 can include a point or region that identifies the location of the detected feature of the object within the subset 403. As further depicted in FIG. 4, the detector 400 can extract a subset 411 of the video from the database 401. For example, the subset 411 can at least partially overlap with the subset 403, or can be composed of different videos. The subset 411 can have, for example, a greater number of videos than the subset 403, such as at least 100 times as many videos. The detector 400 can use the subset 411 and the detector 413 to train the generator network 415. The generator network 415 can be trained, for example, to generate an artificial representation 417 of the feature of the object, such as an abnormality. The artificial representation 417 can have a false representation of the feature of the object that appears similar to the true representation of the feature of the object. Thus, the generator network 415 can be trained to deceive the perception branch 407 into making a determination that it is a false positive.

[0100] As further depicted in FIG. 4, once trained, the generator network 415 can generate the artificial representation 417. The detector 400 can use the artificial representation 417 to train the adversarial branch 409 of the discriminator network 405. As described above, the adversarial branch 409 can be trained to discriminate the artificial representation 417 from the subset 411. Thus, the adversarial branch 409 can determine a difference indicator 419. The difference indicator 419 can represent any feature vector or other aspect of the image that is present in the artificial representation 417 but not in the subset 411, that is present in the subset 411 but not in the artificial representation 417, or a subtraction vector or other aspect that represents the difference between feature vectors, or other aspects of the artificial representation 417, as well as those of the subset 411.

[0101] As depicted in FIG. 4, detector 400 can re-train perception branch 407 by using difference indicator 419. For example, in embodiments where artificial representation 417 has a false representation of the target feature, detector 400 can re-train perception branch 407 such that the false representation does not result in the detection of true representations within subset 411.

[0102] Although not depicted in FIG. 4, detector 400 can further use recursive training to improve generator network 415, perception branch 407, and / or adversarial branch 409. For example, detector 400 can re-train generator network 415 using difference indicator 419. Thus, the output of adversarial branch 409 can be used to re-train generator network 415 such that the artificial representation appears relatively similar to the true representation. In addition to this, the re-trained generator network 415 can generate a new set of artificial representations that are used to re-train adversarial branch 409. Thus, adversarial branch 409 and generator network 415 may perform unsupervised learning, in which case the output of each is used to re-train the other in a recursive manner. This recursive training can be repeated until a threshold number of cycles is reached and / or until the loss function associated with generator network 415 and / or the loss function associated with adversarial branch 409 reaches a threshold. Furthermore, in this recursive training, perception branch 407 can also be re-trained by using each new output of the difference indicator such that a new subset with new detections can be used to further re-train generator network 415.

[0103] FIG. 5 is a flowchart of an exemplary method 500 for detecting features of interest using a discriminator network and a generator network. Method 500 can be executed by one or more processors.

[0104] In step 501 of FIG. 5, at least one processor can store in a database a plurality of videos containing representations of features of interest, such as anomalies. For example, the videos may have been captured during an endoscopic procedure. As part of step 501, at least one processor can further select a first subset of the plurality of videos. As described above, at least one processor can select randomly and / or by using one or more patterns.

[0105] In step 503, at least one processor can apply the perception branch of an object detection network to frames of the first subset of the plurality of videos to generate detections of a first plurality of anomalies. In some embodiments, as described above, the detections can include indicators of the locations of the detected anomalies. Also, in some embodiments, the perception branch can have a convolutional neural network, as described above.

[0106] In step 505, at least one processor can select a second subset of the plurality of videos. As described above, at least one processor can select randomly and / or by using one or more patterns. Using the first plurality of detections and frames from the second subset of the plurality of videos, at least one processor may further train a generator network to generate artificial representations of the plurality of anomalies, where the plurality of artificial representations are generated through residual learning. As described above, each artificial representation provides a false representation of the anomaly that is very similar to the true representation of the anomaly.

[0107] In step 507, at least one processor can train the adversarial branch of the discriminator network to discriminate between artificial representations of anomalies and true representations of anomalies. For example, as described above, the adversarial branch can be trained to identify the difference between the artificial representation and the true representation within the frame. In some embodiments, the adversarial branch can have a convolutional neural network, as described above.

[0108] In step 509, at least one processor can apply the adversarial branch of the discriminator network to a plurality of artificial representations to generate a difference indicator between the artificial representation of the anomaly and the true representation of the anomaly contained within the frames of the second subset of the plurality of videos. For example, as described above, the difference indicator can represent any feature vector of the image or other aspect that exists in the artificial representation but not in the frame, exists in the frame but not in the artificial representation, or represents the difference between feature vectors or other aspects of the artificial representation, or is of the frame.

[0109] In step 511, at least one processor can apply the perceptual branch of the discriminator network to the artificial representation to generate a second plurality of anomaly detections. Similar to the first plurality of detections, the detections can include an indicator of the location of the detected anomaly within the artificial representation.

[0110] In step 513, at least one processor can re-train a perception branch based on a difference indicator and a second plurality of detections. For example, in embodiments where each artificial representation provides a false representation of an anomaly that is very similar to the true representation of the anomaly, at least one processor can re-train the perception branch to reduce the number of detections returned from the artificial representation and, thus, increase the number of null indicators or other indicators of non-anomaly returned from the artificial representation.

[0111] In a state consistent with the present disclosure, the exemplary method 500 can include additional steps. For example, in some embodiments, method 500 can include a step of re-training a generation network based on a difference indicator. In such embodiments, method 500 can further include applying the generation network to generate a further plurality of artificial representations of anomalies and re-training an adversarial branch based on the further plurality of artificial representations of anomalies. Such re-training steps can be recursive. Further, method 500 can include applying the re-trained adversarial branch to a further plurality of artificial representations to generate a further difference indicator between the further artificial representations of anomalies and the true representations of anomalies included in the frames of the second subset of the plurality of videos, and re-training the generation network based on the further difference indicator. As described above, this recursive re-training can be repeated until a threshold number of cycles is reached and / or until a loss function associated with the generation network and / or a loss function associated with the adversarial branch reaches a threshold.

[0112] FIG. 6 is a schematic representation of a system 600 having a hardware configuration for a video feed, which is consistent with an embodiment of the present disclosure. As shown in FIG. 6, the system 600 may be communicatively coupled to an imaging device 601, such as a camera or other device that outputs a video feed. For example, the imaging device 601 may include a medical imaging device, such as a CT scanner, an MRI device, an endoscopic examination device, or the like. The system 600 may further be communicatively coupled to a display 615 or other device for displaying or storing the video. For example, the display 615 may include a monitor, a screen, or other device for displaying images to a user. In some embodiments, the display 615 may be replaced or supplemented by a storage device (also not shown) communicatively connected to a cloud-based storage system (not shown) or a network interface controller (NIC).

[0113] As further depicted in FIG. 6, the system 600 may include not only an input port 603 that receives a video feed from the camera 601, but also an output port 611 that outputs the video to the display 615. As described above, the input port 603 and the output port 611 may include a VGA port, an HDMI (registered trademark) port, a DVI port, or the like.

[0114] System 600 further includes a first bus 605 and a second bus 613. As shown in FIG. 6, the first bus 605 can transmit video received through input port 603 through at least one processor 607. For example, one or more processors 607 can implement either the object detector network and / or the discriminator network described above. Accordingly, one or more processors 607 can overlay one or more indicators, such as the exemplary graphical indicator of FIG. 8, on the video received via the first bus 602, for example, by using the exemplary method 700 of FIG. 7. The processor 607 can then transmit the overlaid video to the output port 611 via the third bus 609.

[0115] In certain situations, the object detector implemented by one or more processors 607 may malfunction. For example, the software implementing the object detector may crash, or otherwise cease to operate properly. In addition to, or instead of, this, one or more processors 607 may also receive a command (e.g., from an operator of the system 600) to stop the overlay operation of the video. In response to the malfunction and / or the command, one or more processors 607 can activate the second bus 613. For example, one or more processors 607 can transmit a command or other signal to activate the second bus 613, as depicted in FIG. 6.

[0116] As depicted in FIG. 6, the second bus 613 can directly transmit the received video from the input port 603 to the output port 611, thereby allowing the system 600 to function as a pass-through for the imaging device 601. The second bus 613 can also allow for seamless presentation of the video from the imaging device 601, even if the software implemented by the processor 607 malfunctions or if the operator of the hardware overlay 600 decides to stop the overlay operation in the middle of a video feed.

[0117] FIG. 7 is a flowchart of an exemplary method 700 for overlaying object indicators on a video feed using an object detector network that is consistent with an embodiment of the present disclosure. Method 700 can be executed by one or more processors. In step 701 of FIG. 7, at least one processor can provide at least one image. For example, the at least one image may be extracted from a database or received from an imaging device. In some embodiments, the at least one image can have a frame within a video feed.

[0118] In step 703, at least one processor may overlay a boundary having a two-dimensional shape around a region of the image detected as containing a feature of interest, where the boundary is rendered in a first color. In step 705, after a predetermined time has elapsed, at least one processor can change the boundary to appear in a second color if the feature of interest is a true positive and to appear in a third color if the feature of interest is a false positive. The elapse of the predetermined time may represent a pre-set period (e.g., a threshold number of frames and / or seconds) and / or the elapsed time between the detection of the feature of interest and its classification as a true or false positive.

[0119] In addition to, or instead of, this, at least one processor can change the boundary to a second color if the feature of interest is classified in the first category, and can change the boundary to a third color if the feature of interest is classified in the second category. For example, if the feature of interest is a lesion, the first category can have a cancerous lesion and the second category can have a non-cancerous lesion.

[0120] In a state consistent with the present disclosure, the exemplary method 700 can include additional steps. For example, in some embodiments, the method 700 includes sending a command to one or more speakers to generate sound and / or sending a command to at least one wearable device to vibrate when the boundary is changed. In such embodiments, at least one of the duration, tone, frequency, and amplitude of the sound and / or vibration can depend on whether at least one detected object is a true positive or a false positive.

[0121] FIG. 8A shows an exemplary overlay 801 for object detection in a video, consistent with an embodiment of the present disclosure. Not only in FIG. 8A, but also in the examples of FIGS. 8B and 8C, the illustrated video samples 800a and 800b are from a colonoscopy procedure. It will be understood from the present disclosure that other procedures and videos from imaging devices can be used when implementing embodiments of the present disclosure. Thus, the video samples 800a and 800b are non-limiting examples of the present disclosure. In addition, by way of example, the video displays of FIGS. 8A-8C may be presented on a display device, such as the display 107 of FIG. 1 or the display 615 of FIG. 6.

[0122] Overlay 801 represents an example of a graphical boundary used as an indicator for detected anomalies or features of an object within a video. As shown in FIG. 8A, images 800a and 800b have frames of a video that include features of the detected object. Image 800b includes the graphical overlay 801 and corresponds to a frame that is later in sequence or later in time than image 800a.

[0123] As shown in FIG. 8A, images 800a and 800b have video frames from a colonoscopy, and the features of the object have lesions or polyps. In other embodiments, as described above, images from other medical procedures such as gastroscopy, enteroscopy, esophagoscopy, or the like, or those similar thereto, can be utilized, and graphical indicators such as overlay 801 can be overlaid. In some embodiments, indicator 801 can be overlaid after detection of an anomaly and after the passage of time (e.g., a specific number of frames and / or seconds between image 800a and image 800b). In the example of FIG. 8A, overlay 801 has an indicator in the form of a rectangular boundary with a predefined pattern (i.e., solid corner angles). In other embodiments, overlay 801 can have a different shape (whether regular or irregular). In addition to this, overlay 801 can be displayed in a predefined color or can transition from one color to another.

[0124] In the example of FIG. 8A, overlay 801 has an indicator with solid corner angles that surround the detected location of the feature of the object within the video frame. Overlay 801 appears within video frame 800b, which can be subsequent to video frame 800a in sequence.

[0125] Figure 8B shows another example of a display with an overlay for object detection in a video according to an embodiment of the present disclosure. Figure 8B depicts an image 810a (similar to image 800a) and an image 810b (similar to image 800b) after being overlaid with an indicator 811. In the example of Figure 8B, the overlay 811 has a rectangular boundary with solid lines on all sides. In other embodiments, the overlay 811 may have a first color and / or a different shape (regardless of whether it is regular or irregular). In addition to this, the overlay 811 may be displayed in a predefined color or may transition from a first color to another color. As shown in Figure 8B, the overlay 811 is placed on the detected abnormality or the feature of the object in the video. The overlay 811 appears within the video frame 810b, and the video frame 810b may be subsequent to the video frame 810a in sequence.

[0126] Figure 8C shows another example of a display with an overlay for object detection in a video according to an embodiment of the present disclosure. Figure 8C depicts an image 820a (similar to image 800a) and a subsequent image 820b (similar to image 800b) after being overlaid with an indicator 821. In the example of Figure 8C, the overlay 821 has a rectangular boundary with dashed lines on all sides. In other embodiments, the indicator 821 may have a different shape (regardless of whether it is regular or irregular). In addition to this, the overlay 821 may be displayed in a predefined color or may transition from a first color to another color. As shown in Figure 8C, the overlay 821 is placed on the detected abnormality or the feature of the object in the video. The overlay 821 appears within the video frame 820b, and the video frame 820b may be subsequent to the video frame 820a in sequence.

[0127] In some embodiments, the graphical indicator (i.e., overlay 801, 811, or 821) can change pattern and / or color. For example, the color of the pattern and / or the boundary of the pattern can be changed in response to the passage of time (e.g., a predetermined number of frames and / or seconds between image 800a and image 800b, image 810a and image 810b, or image 820a and image 820b). In addition to or instead of this, the pattern and / or color of the indicator can also be changed in response to a specific classification of the feature of interest (e.g., when the feature of interest is a polyp, the classification of the polyp as cancerous or non-cancerous). Furthermore, the pattern and / or color of the indicator can also depend on the classification of the feature of interest. Thus, the indicator may have a first pattern or color when the feature of interest is classified in a first category, a second pattern or color when the feature of interest is classified in a second category, and so on. Or, instead of this, the pattern and / or color of the indicator can also depend on whether the feature of interest is identified as a true positive or a false positive. For example, the feature of interest can be detected by an object detection network (or the perception branch of the discriminator network) as described above, which can consequently result in an indicator, but then be determined to be a false positive by an adversarial branch or network as described above, which can result in the indicator being in a first pattern or color. Instead, when the feature of interest is determined to be a true positive by an adversarial branch or network, the indicator can be displayed in a second pattern or color.

[0128] The above description is presented for illustrative purposes. It is not exhaustive and is not limited to the disclosed forms or embodiments as such. Modifications and adaptations of the embodiments will become apparent from a consideration of this specification and the practice of the disclosed embodiments. For example, while the described implementations include hardware, the systems and methods consistent with the present disclosure can be implemented by hardware and software. In addition, although specific components are described as being coupled to each other, such components may be integrated with each other or distributed in any suitable manner.

[0129] Furthermore, although exemplary embodiments are described herein, the scope includes any and all embodiments having equivalent elements, modifications, omissions, combinations, adaptations, and / or variations based on the present disclosure (e.g., aspects spanning various embodiments). The elements in the claims are to be construed broadly based on the language employed in the claims and are not limited to the examples described herein or in the practice of the application, which examples are to be construed as non-exclusive. Furthermore, the steps of the disclosed methods can be modified in any manner, including reordering of the steps and / or insertion or deletion of steps.

[0130] The features and advantages of the present disclosure will be apparent from the above detailed description, and accordingly, the appended claims are to be construed to cover all systems and methods that fall within the true spirit and scope of the present disclosure. As used herein, the indefinite articles "a" and "an" mean "one or more." Similarly, the use of plural terms does not necessarily denote a plurality unless it is clear in a given context. Terms such as "and" or "or" mean "and / or" unless specifically stated otherwise. Furthermore, many modifications and variations will occur to those of ordinary skill in the art from a consideration of the present disclosure, and it is not desirable to limit the present disclosure to the exact structures and operations illustrated and described, and accordingly, all suitable modifications and equivalents are intended to be included within the scope of the present disclosure.

[0131] For other embodiments, it will be apparent from a consideration of the specification and practice of the embodiments disclosed herein. The description and examples are to be considered as illustrative only, and the true scope and spirit of the disclosed embodiments are to be determined by the appended claims. The above embodiments may be described as follows, but are not limited thereto. [Configuration 1] A computer-implemented system for processing real-time video, an input port for receiving real-time video obtained from a medical imaging device, a first bus for transferring the received real-time video, at least one processor configured to receive the real-time video from the first bus, perform object detection by applying a trained neural network to frames of the received real-time video, and overlay a boundary notifying the location of at least one detected object within the frame; a second bus for receiving the video having the overlaid boundary, An output port that outputs the video with the overlaid boundary from the second bus to an external display, A third bus that directly transmits the received real-time video to the output port, A system having the same. [Configuration 2] The system according to Configuration 1, wherein the third bus is activated when receiving an error signal from the at least one processor. [Configuration 3] The system according to Configuration 1 or 2, wherein at least one of the first plurality of images and the second plurality of images has an image from an imaging device used in at least one of gastroscopy, colonoscopy, enteroscopy, or, optionally, upper endoscopy including an endoscope. [Configuration 4] The system according to any one of Configurations 1 to 3, wherein the at least one detected object is an abnormality, and the abnormality optionally has a formation on human tissue or a formation of human tissue, a change in human tissue from one type of cell to another type of cell, and / or a lack of the human tissue from a location where the human tissue is expected. [Configuration 5] The system according to Configuration 4, wherein the abnormality optionally has a polypoid lesion or a non-polypoid lesion, and has a lesion. [Configuration 6] The system according to any one of Configurations 1 to 5, wherein the overlaid boundary has a graphical pattern around the region of the image including the at least one detected object, and the pattern is displayed in a first color. [Configuration 7] The system according to Configuration 6, further configured such that after a predetermined time has elapsed, the at least one processor changes the pattern to be displayed in a second color when the at least one detected object is a true positive, and further changes the pattern to be displayed in a third color when the at least one detected object is a false positive. [Configuration 8] The at least one processor is further configured to send a command to one or more speakers to generate sound and / or send a command to at least one wearable device to vibrate when the pattern of the boundary is changed, as described in Configuration 7. [Configuration 9] At least one of the duration, tone, frequency, and amplitude of the sound depends on whether the at least one detected object is a true positive or a false positive, and / or at least one of the duration, frequency, and amplitude of the vibration depends on whether the at least one detected object is a true positive or a false positive, as described in Configuration 8. [Configuration 10] A computer-implemented system for processing real-time video, An input port for receiving real-time video obtained from a medical imaging device, At least one processor configured to receive the real-time video from the input port, perform object detection by applying a trained neural network to frames of the received real-time video, and overlay a boundary notifying the location of at least one detected object within the frame, An output port for outputting the video with the overlaid boundary from the at least one processor to an external display, An input device for receiving a sensitivity setting from a user, And having, The processor is further configured to adjust at least one parameter of the trained neural network in response to the sensitivity setting. [Configuration 11] At least one of the first plurality of images and the second plurality of images has an image from an imaging device used during at least one of a gastroscopy, a colonoscopy, a small intestine endoscopy, or, optionally, an upper endoscopy including an endoscopy device, the system according to configuration 10. [Configuration 12] The at least one detected object is an abnormality, and the abnormality optionally has a formation on human tissue or a formation of human tissue, a change in human tissue from one type of cell to another type of cell, and / or an absence of the human tissue from a location where the human tissue is predicted, the system according to any one of configurations 1 to 11. [Configuration 13] The abnormality optionally has a lesion, including a polypoid lesion or a non-polypoid lesion, the system according to configuration 12.

Claims

1. 1. A computer-implemented system for processing real-time video, comprising: an input port for receiving real-time video having a plurality of frames acquired from a medical imaging device; a first bus for transmitting the received real-time video; At least one processor configured to execute instructions for object detection and boundary overlay, the instructions comprising: receiving the real-time video from the first bus; feeding the plurality of frames of the received real-time video directly to a trained neural network; performing object detection by applying the trained neural network to the plurality of frames of the received real-time video; overlaying a boundary indicating a location of at least one detected object in the plurality of frames determined by applying the trained neural network to the plurality of frames of the received real-time video, the overlaid boundary having a graphical indicator of a first pattern and / or color displayed around an area of ​​the plurality of frames that includes the at least one detected object; At least one processor having a second bus for receiving the plurality of frames having the overlaid boundaries; an output port for outputting the plurality of frames having the overlaid borders from the second bus to an external display; wherein the at least one processor is further configured to modify the graphical indicator to be displayed in a second pattern and / or color if the at least one detected object is a true positive and to further modify the graphical indicator to be displayed in a third pattern and / or color if the at least one detected object is a false positive.

2. The system described in claim 1, further comprising a third bus that transmits the received real-time video directly to the output port, the third bus being activated when the entire system is turned off.

3. 3. The system of claim 1 or 2, wherein the real-time video comprises images from an imaging device used during at least one of a gastroscopy, a colonoscopy, a small intestine endoscopy, or, optionally, an upper endoscopy including an endoscopy device.

4. 4. The system of claim 1, wherein the at least one detected object is an abnormality, and the abnormality optionally comprises a formation on or in human tissue, a change in human tissue from one type of cell to another type of cell, and / or an absence of human tissue from a location where human tissue is expected.

5. The system of claim 4 , wherein the abnormality comprises a lesion, optionally comprising a polypoid lesion or a non-polypoid lesion.

6. 6. The system of claim 1, wherein the at least one processor is further configured to: send a command to one or more speakers to generate a sound when the boundary pattern is changed; and / or send a command to at least one wearable device to vibrate when the boundary pattern is changed.

7. 7. The system of claim 6, wherein at least one of the duration, tone, frequency, and amplitude of the sound is dependent on whether the at least one detected object is a true positive or a false positive, and / or at least one of the duration, frequency, and amplitude of the vibration is dependent on whether the at least one detected object is a true positive or a false positive.

8. The system of claim 1 , wherein the plurality of frames comprises images from the medical imaging device used during at least one of a gastroscopy, a colonoscopy, and an enteroscopy.

9. The system of claim 1 , wherein the trained neural network is trained using multiple frames of a video having indicators of locations of features of interest.

10. 2. The system of claim 1, wherein the trained neural network comprises a generative network and a discriminator network, the generative network being trained to generate a plurality of artificial representations of features of the subject, and the discriminator network being trained to discriminate between the plurality of artificial representations of the feature of the subject and a true representation of the feature of the subject.

11. 11. The system of claim 10, wherein the discriminator network has an adversarial branch and a perceptual branch, the adversarial branch trained to generate a difference indicator between the plurality of artificial representations of the target feature and a true representation of the target feature, and the perceptual branch trained to generate a second plurality of detections of the target feature.

12. The system of claim 1 , wherein the at least one processor is further configured to perform classification of the detected objects by applying the trained neural network.

13. The system of claim 12 , wherein the classification is based on at least one of a histological classification, a morphological classification, and a structural classification.

Citation Information

Patent Citations

  • System and method for image segmentation using a multi-stage classifier

    JP2008535566A

  • Endoscope system, processor device, and operation method for endoscope system

    JP2016144507A

  • System and methods for automatic polyp detection using convulutional neural networks

    US20180075599A1

  • A system and method for detection of suspicious tissue regions in an endoscopic procedure

    WO2017042812A2

  • Method and system for classification of endoscopic images using deep decision networks

    WO2017055412A1