Apparatus, method, and computer-readable storage medium for detecting a subject in a video signal based on visual evidence using the output of a machine learning model
By generating detection chains and applying heuristic filtering to the output of a machine learning model, the mechanism addresses the challenge of temporally consistent object detection in video signals, significantly reducing false positives and maintaining detection accuracy.
Patent Information
- Application Number
- JP2024019756
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-13
- Filing Date
- 2024-02-13
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2040-11-26
AI Technical Summary
Existing object detection methods, such as SSD, struggle to achieve temporally consistent detection in video signals, often resulting in spurious, lost, or unstable detections due to the reliance on single-frame analysis.
A mechanism that utilizes the output of a machine learning model to generate a detection chain by associating current detections with prior detections based on location overlap, and employs heuristic filtering stages to suppress flicker, apply hysteresis thresholding, and smooth detection locations, thereby improving detection stability and consistency across video frames.
The proposed solution effectively reduces false positives by 62% and maintains detection accuracy, while minimizing the impact on true positives, thus achieving temporally consistent and visually improved object detection in video signals without the need for extensive video data training.
Smart Images

Figure 0007696031000001 
Figure 0007696031000002 
Figure 0007696031000003
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus, a method, and a computer-readable storage medium for detecting a subject in a video signal based on visual evidence using the output of a machine learning model.
Background Art
[0002] Conventional machine learning can be useful for finding a decision function that maps features regarding a detection target in an image to class labels. The machine learning algorithm must go through a training phase in which the decision function is modified to minimize errors on the training data. After the training phase is completed, the decision function is fixed and used to predict data not seen before.
[0003] Deep learning, which is a technique capable of automatically discovering appropriate features, is adopted to give appropriate features (for example, color distribution, gradient histogram, etc.) related to the detection target to the machine learning algorithm.
[0004] Deep learning typically utilizes a deep convolutional neural network. Compared with a conventional neural network, the first layer is replaced with a convolution operation. Thereby, the convolutional neural network can learn an image filter capable of extracting features. Since the filter coefficients are part of the decision function here, the training process can also optimize feature extraction. Therefore, the convolutional neural network can automatically discover useful features.
[0005] It is necessary to distinguish between classification and subject class detection. For classification, the input is an image and the output is a class label. Thus, classification can answer the question "Does this image contain a detection target such as a polyp? (Yes / No)". In contrast, subject class detection provides not only the class label but also the location of the subject in the form of a bounding box. It is possible to consider a subject detector as a classifier applied to many different image patches.
[0006] A well-known method for subject detection is the Single Shot MultiBox Detector (SSD) disclosed by W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C-Y. Fu, A.C. Berg: ‘‘SSD: Single Shot MultiBox Detector’’, European Conference on Computer Vision 2016.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] The basic principle of SSD is to place the so-called virtual grid of anchor boxes across the image. At every location, there are multiple anchor boxes with different scales and aspect ratios. When detecting a detection target, for example, a polyp, the question is "Does this anchor box contain a detection target such as a polyp (yes / no)?". Therefore, a neural network with two output neurons is required for each anchor box. Depending on which of the two output neurons is more strongly activated, the anchor box is classified as positive or negative.
[0009] Detectors such as SSD provide a framework for object detection on still images.
[0010] An object of the present invention is to provide an object detection and display mechanism capable of achieving temporally consistent detection in a video signal based on the output from a machine learning model.
[0011] This object is solved by an apparatus, a method, and a computer-readable storage medium as defined by the appended claims.
[0012] According to one aspect of the present invention, an apparatus is provided, the apparatus comprising means for obtaining one or more current detections output from a machine learning model for at least one current video frame input into the machine learning model from a series of consecutive video frames of a video signal, wherein the current detections among the one or more current detections include a confidence value indicating a probability that the current detection includes a detection target to be detected by the machine learning model and a location of the detection target within at least one current video frame; means for generating a detection chain by associating detections output from the machine learning model, wherein the current detections among the one or more current detections are associated with prior detections among one or more prior detections obtained from the machine learning model for at least one prior video frame in the series that precedes the at least one current video frame and is input into the machine learning model, the prior detections among the one or more prior detections include a confidence value indicating a probability that the prior detection includes a detection target and a location of the detection target within at least one prior video frame, and the current detection is associated with the prior detection based on the locations of the current detection and the prior detection; means for causing a display of at least one current detection in the video signal based on the location, confidence value, and location of the current detection in the detection chain; and means for repeating obtaining, generating, and causing a display for at least one next video frame in the series as at least one current video frame.
[0013] According to one embodiment of the present invention, when an overlap between the locations of the current detection and the prior detection satisfies a predetermined condition, the current detection is associated with the prior detection such that the current detection and the prior detection belong to the same detection chain.
[0014] According to one embodiment of the present invention, the display of the current detection is caused when the current detection belongs to N+M detections of the detection chain, where N and M are positive integers greater than or equal to 1, N represents the N earliest detections in time of the detection chain, and the display of the current detection is not caused when the current detection belongs to the N earliest detections in time of the detection chain.
[0015] According to one embodiment of the present invention, the display of the current detection is caused when the confidence value of the current detection is greater than or equal to a first threshold.
[0016] According to one embodiment of the present invention, the display of the current detection is caused when the confidence value of the current detection is greater than or equal to a second threshold smaller than the first threshold, and when the confidence value of a preceding detection belonging to the same detection chain as the current detection is greater than or equal to the first threshold.
[0017] According to one embodiment of the present invention, the apparatus further comprises means for performing smoothing over the locations of the detections in the detection chain.
[0018] According to one embodiment of the present invention, the video signal is captured by an endoscope during an inspection process.
[0019] According to one embodiment of the present invention, the detection target is a polyp.
[0020] According to one embodiment of the present invention, there is provided a subject detection and display mechanism that uses the output of a machine learning model to achieve temporally consistent detections in a video signal based on visual evidence within video frames of the video signal.
[0021] According to an exemplary embodiment, the subject detection and display mechanism processes a video signal of a moving image by using the output of a machine learning model, and the subject detection and display mechanism can suppress artifacts such as spurious detections, lost detections, and unstable localizations, as described later, while suppressing the load of training the machine learning model.
[0022] According to one embodiment of the present invention, a heuristic approach is employed to perform subject detection in a video's video signal using the output of a machine learning model, thereby visually improving the quality of the detection.
Advantages of the Invention
[0023] According to the present invention, for example, when performing an endoscopic examination such as a colonoscopy screening, it is possible to assist a doctor in concentrating on relevant image regions including tissues that match the appearance of polyps.
[0024] Hereinafter, the present invention will be described by way of its embodiments with reference to the accompanying drawings.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Best Mode for Carrying Out the Invention
[0026] According to the present invention, the output of a machine learning model is used. The machine learning model outputs one or more detections for each video frame of the video signal input to the machine learning model. For example, the video signal is captured by an endoscope during an inspection process.
[0027] In particular, the machine learning model outputs a confidence value and a location for each detection. The confidence value indicates the probability that the detection includes the detection target to be detected by the machine learning model, and the location indicates the area of the detection target within the video frame. For example, the detection target is a polyp.
[0028] For example, as the machine learning model, a neural network having two output neurons for each anchor box as described above is adopted. The anchor box is classified as positive or negative according to which of the two output neurons is more strongly activated. The location of the detection target is based on the location of the anchor box. The outputs from the two neurons form the confidence value.
[0029] The machine learning model is trained using teacher data for object detection, that is, teacher data for the detection of detection targets such as images in the form of bounding boxes and polyps with annotated objects.
[0030] To objectively evaluate the performance improvement of the dataset and filtering technology of the machine learning model, standard metrics are used for the object detection task. The related metrics used are precision, recall, and average precision (AP). Precision is defined as the ratio of correctly detected instances compared to the total number of detections returned by the machine learning model. Recall is defined as the ratio of correctly detected instances compared to the total number of instances to be detected.
[0031] Therefore, precision and recall can be formally defined as follows.
[0032] Precision = TP / (TP + FP) Recall = TP / (TP + FN)
[0033] Where TP represents the number of true positives (correct detections), FP represents the number of false positives (incorrect detections), and FN represents the number of false negatives (missed detections).
[0034] To classify a detection as "true" or "false", it is necessary to measure the quality of the localization. To measure the localization quality, the Intersection over Union (IoU) criterion is adopted.
[0035] IoU(A, B) = |A ∩ B| / |A ∪ B|
[0036] The Intersection over Union is 1 only in the case of a complete localization. Penalties are imposed for both underdetections and overdetections. If the IoU between a detection and an annotation is 0.5 or more, the detection is classified as correct. Figure 1 shows examples of insufficient localization, minimum acceptable localization, and complete localization.
[0037] Precision and recall are useful tools for evaluating the performance of an algorithm, but they have significant drawbacks. A classifier outputs a confidence value that measures the probability that an image region contains a detection target such as a polyp. It is necessary to apply a threshold for the final decision of whether to display a detection. However, the values of precision and recall depend on this threshold. For example, it is always possible to increase precision at the expense of recall by increasing the threshold.
[0038] Therefore, in FIGS. 5 and 7 described below, precision (P) and recall (R) are evaluated over all possible thresholds in order to plot the precision-recall curve. The area under the curve is called the average precision (AP) and serves as an indicator of how well different classifiers generally perform. This value can be used to compare different classifiers to each other.
[0039] In the following, it is assumed that the machine learning model in which the output of the present invention is used is trained to achieve good performance when generating detections based on video frames of a video signal. However, information from past video frames may be able to further improve performance.
[0040] When generating detections based only on the current video frame, the following artifacts may occur:
[0041] - Spurious detections: False positives that appear for a single image frame of the video signal and tend to disappear in the next frame of the video signal. - Lost detections: When the machine learning model detects a detection target, e.g., a polyp, the detection is usually very stable over a plurality of consecutive frames of the video signal. However, sometimes the confidence of the detection temporarily falls below the detection threshold, and the detection may blink. - Unstable localization: The machine learning model estimates a bounding box to localize each detection. When the input image changes slightly, the localization also changes accordingly. However, this change may not look smooth to the user.
[0042] A detector that can consider past video frames can have a good chance of reducing these artifacts. However, in order to train such a detector, it is necessary to collect video sequences as a dataset. This places a heavy burden on the doctor because the doctor has to label all single frames in the video signal.
[0043] To avoid training a machine learning model using an image sequence, according to the present invention, a heuristic solution for visually improving the quality of detection is adopted. For this purpose, a filtering heuristic for dealing with the above-mentioned artifacts is introduced. FIG. 2 is a diagram schematically showing an “ideal” solution and the solution according to the present invention.
[0044] The “ideal” solution is shown on the left side of FIG. 2. For example, a long short-term memory (LSTM) architecture of a deep convolutional neural network (DCNN) takes a plurality of video frames as inputs and outputs detections based on visual evidence over the plurality of frames.
[0045] The solution according to the present invention is shown on the right side of FIG. 2. The prediction is based on individual frames filtered by a heuristic.
[0046] The difference between both solutions is that a true multi-frame detector can rely on visual evidence from a plurality of video frames. The heuristic solution according to the present invention relies on the visual evidence of the current frame to generate a detection. As described above, the detection includes a location and a confidence value. Therefore, the heuristic can operate with these values.
[0047] According to an embodiment of the present invention, before the filter heuristic is applied, detections are associated with each other over a plurality of video frames. According to an embodiment of the present invention, detections generally do not tend to move fast over consecutive video frames, and the location of the detections is assumed to be used to associate the detections with each other. According to one example, the aforementioned intersection over union criterion is used to associate detections with each other. For example, detections in consecutive video frames with IoU≧0.3 are considered to be part of the same detection chain. Consecutive detections with IoU<0.3 are considered to be part of different detection chains. The filtering steps described below operate on each of these detection chains.
[0048] Before describing the filtering stage, refer to FIG. 3 showing the process of subject detection and display according to an embodiment of the present invention.
[0049] In step S305 of FIG. 3, one or more current detections for at least one current video frame input to the machine learning model are obtained as an output from the machine learning model. The at least one current video frame belongs to a series of consecutive video frames of a video signal. According to an exemplary embodiment, the video signal is captured from an endoscope device that captures the video signal. For example, the video signal includes a moving image.
[0050] The current detection among the one or more current detections includes a confidence value indicating the probability that the current detection includes a detection target to be detected by the machine learning model, and the location of the detection target within at least one current video frame. In step S305, one or more current detections for at least one current video frame are obtained.
[0051] In step S307, a detection chain is generated by associating the detections output from the machine learning model. The current detection among the one or more current detections is associated with a previous detection among one or more previous detections obtained from the machine learning model for at least one previous video frame in the series that preceded the at least one current video frame and was input to the machine learning model. The previous detection among the one or more previous detections includes a confidence value indicating the probability that the previous detection includes a detection target, and the location of the detection target within at least one previous video frame. According to an embodiment of the present invention, the current detection is associated with the previous detection based on the locations of the current detection and the previous detection. According to an alternative embodiment, or in addition, the current detection is associated with the previous detection based on at least one of the speed and direction of detections within consecutive video frames.
[0052] In step S309, based on the current detection position, the current detection confidence value, and the current detection location in the detection chain, the display of at least one current detection in the video signal is caused.
[0053] In step S311, it is checked whether the end condition is satisfied. If the end condition is satisfied, the process ends. If the end condition is not satisfied, the process returns to step S305 and processes at least one next video frame in the series as at least one current video frame.
[0054] For example, the end condition is satisfied when there is no next video frame in the series.
[0055] According to an exemplary embodiment, in step S307, when the overlap of the current detection and the previous detection locations satisfies a predetermined condition, for example, IoU≧0.3, the current detection is associated with the previous detection so that the current detection and the previous detection belong to the same detection chain.
[0056] Furthermore, according to an exemplary embodiment, in step S309, when the confidence value of the current detection is greater than or equal to the first threshold, the display of the current detection is caused.
[0057] Next, refer to FIG. 4 showing a control unit 40 capable of implementing an example of an embodiment of the present invention. For example, the control unit 40 implements the subject detection and display process of FIG. 3.
[0058] The control unit 40 includes a processing resource (for example, a processing circuit) 41, a memory resource (for example, a memory circuit) 42, and an interface (for example, an interface circuit) 43, which are connected via a link (for example, a bus, a wired connection, a wireless connection, etc.) 44.
[0059] According to an exemplary embodiment, when executed by the processing resource 41, the memory resource 42 stores a program that operates the control unit 40 in accordance with at least some embodiments of the present invention.
[0060] Generally, exemplary embodiments of the present invention can be implemented by computer software stored in the memory resource 42 and executable by the processing resource 41, or by hardware, or by a combination of software and / or firmware and hardware.
[0061] Hereinafter, a filtering stage that operates on the detection chain obtained as described above will be described.
[0062] Filtering stage 1: Flicker suppression Flicker suppression is designed to address the problem of spurious detections. Since spurious detections only appear for a few frames and then disappear again, a solution to mitigate this problem is to suppress the first occurrence of a detection within the image. For example, a detection corresponding to that location is displayed at S309 only if a detection target, such as a polyp, is independently detected in multiple subsequent video frames at the same location.
[0063] There are two different ways to achieve such flicker suppression. One way is suppression without prior knowledge that always suppresses the first N occurrences of a detection. Another way is suppression with prior knowledge that suppresses the first N occurrences of a detection only if the detection disappears in the (N + 1)-th frame.
[0064] Both versions have the effect of enhancing the accuracy of the subject detection and display mechanism. However, since the detection is intentionally suppressed, the recall will be impaired. This reduction in recall is greater when flicker suppression without prior knowledge is used than when flicker suppression with prior knowledge is used. However, there is a delay of N + 1 frames until knowledge of whether to display the detection is obtained. Since such latency is usually unacceptable, it is preferable to use flicker suppression without prior knowledge.
[0065] According to an exemplary embodiment of the subject detection and display process of FIG. 3, in step S309, when the current detection belongs to N + M detections of the detection chain, the display of the current detection is caused, where N and M are positive integers greater than or equal to 1, and N indicates the N temporally first detections of the detection chain. Further, when the current detection belongs to the N temporally first detections of the detection chain, the display of the current detection is not caused.
[0066] In FIG. 5, for all possible thresholds, the precision and recall are evaluated for (1) the original dataset (i.e., without applying flicker suppression in the subject detection and display process of FIG. 3), (2) the dataset with flicker suppression without prior knowledge (wof) applied in the subject detection and display process of FIG. 3, and (3) the dataset with flicker suppression with prior knowledge (wf) applied in the subject detection and display process of FIG. 3, and the precision-recall (PR) curve is plotted.
[0067] As described above, the area under the curve is called the average precision (AP) and serves as an indicator of how well the subject detection and display process of FIG. 3 functions with (1) no flicker suppression, (2) flicker suppression without prior knowledge, and (3) flicker suppression with prior knowledge.
[0068] Figure 5 shows the effect of flicker suppression with and without prior knowledge. While the accuracy of the high-precision part of the detector's (e.g., machine learning model's) characteristics is improved, the achievable maximum recall is reduced. When flicker suppression with prior knowledge is adopted, both effects are not very significant. The improvement in accuracy is very visible to the user, but the decrease in recall is not due to the detector (e.g., machine learning model) having its operating point in the high-precision region of the PR curve in most application scenarios.
[0069] Applying flicker suppression without prior knowledge means that the recall is more strongly reduced, but this lost recall is hardly noticed by the user. Some missed detections after the polyp enters the field of view are much less noticeable than false positives that appear across the image and disappear immediately.
[0070] Filtering Stage 2: Hysteresis Sometimes, reverse flicker detections occur: the detection is lost for a short time in a single frame and then quickly detected again in the next frame. For example, this can happen when motion blur occurs.
[0071] To counter these lost detections, hysteresis thresholding is introduced as shown in Figure 6.
[0072] Hysteresis thresholding uses two thresholds, namely, a first threshold called the high threshold (labeled "High" in Figure 6) and a second threshold called the low threshold (labeled "Low" in Figure 6). First, the confidence value of the detection must exceed the high threshold to be displayed. In other words, first, the detection is displayed when it is detected in a plurality of frames with high reliability at a similar location (e.g., over time as shown in Figure 6). When the detection is displayed at a similar location over several frames, the detection can still be displayed even if it drops below the high threshold. When the detection drops below the low threshold, it is no longer displayed. In Figure 6, the confidence value is shown as "Score".
[0073] According to an exemplary embodiment, in step S309 of FIG. 3, when the confidence value of the current detection is equal to or greater than a second threshold that is smaller than the first threshold, and when the confidence value of a previous detection belonging to the same detection chain as the current detection is equal to or greater than the first threshold, the display of the current detection is caused.
[0074] FIG. 7 shows a typical effect of the application of the hysteresis threshold processing in the subject detection and display process of FIG. 3. At a given accuracy, the recall can be improved. A potential decrease in accuracy is not actually observable.
[0075] Note that the PR curves shown in FIGS. 5 and 7 are obtained based on different data sets.
[0076] With hysteresis threshold processing, more polyps are detected, so the recall can be increased. Potentially, some of these detections may turn out to be incorrect, which may also lead to a decrease in accuracy. However, neural networks are generally very good at assigning high confidence values when polyps are actually present and very low confidence values when polyps are not present, so such problems have not been encountered. In such cases, the confidence values of the network generally do not even exceed a low threshold.
[0077] Filtering Stage 3: Location Smoothing In filtering state 3, smoothing is performed over the location of the detection.
[0078] According to an exemplary embodiment, in step S309 of FIG. 3, when the display of at least one current detection is caused, its location is smoothed based on the location of the detections in the detection chain to which the current detection belongs, and the detections precede the current detection.
[0079] For example, smoothing is performed by executing a weighted average of the detected coordinates. This makes the localization appear more stable than the original state. Alternatively, smoothing can be performed using a more complex filtering structure, for example, by applying a Kalman filter to the location of the detection within the video signal.
[0080] Effect The combined effect of the above-described technique of applying heuristic filtering stages 1 to 3 has been evaluated on a large test dataset of 6000 images. On average, a 62% reduction in false positive detections was observed compared to without heuristic filtering. Similarly, a 16% increase in false negatives was observed. Here too, it should be noted that the reduction in false positives is very noticeable, while the reduction in false negatives is hardly noticeable. The number of frames in which a polyp enters the field of view and is not detected is technically measured as a false negative. However, to a human user, this is hardly visible. However, false positives that appear throughout the video are very prominent to the user.
[0081] At this point, it should also be noted that the 16% increase in false negatives does not mean that 16% more polyps are missed during the colonoscopy. This means that there is a 16% increase in video frames in which a polyp is present but not detected. However, typically, there are many video frames depicting the same polyp. If the network is good at detecting polyps, it is virtually certain to encounter at least one video frame in which a particular polyp is detected. In practice, heuristic filtering does not affect the number of polyps that are detected at least once.
[0082] The above-described subject detection and display mechanism can reliably detect polyps in real time during a colonoscopy.
[0083] The three-stage heuristic filtering method enables filtering of detections over the frames of the video signal, i.e., over time. Thus, the object detection and display mechanism operates on video frames, e.g., a single video frame, but the individual detections appear more stable. This heuristic filtering visually improves the results and enables temporally consistent detections without the need for video data (and corresponding annotations) during training.
[0084] It should be understood that the foregoing description is illustrative of the invention and should not be construed as limiting the invention. Those skilled in the art can envision various modifications and applications without departing from the true spirit and scope of the invention as defined by the appended claims.
Claims
1. The method is inputting to the machine learning model one or more current detections for at least one current video frame of a series of consecutive video frames of a video signal captured by an endoscope during an inspection process, the current detections of the one or more current detections including a confidence value indicating a probability that the current detection includes a polyp to be detected by the machine learning model and a location of the polyp within the at least one current video frame; generating a detection chain by associating detections output from the machine learning model, wherein a current detection of the one or more current detections is associated with a previous detection of one or more previous detections obtained from the machine learning model for at least one previous video frame of the series that precedes the at least one current video frame and that is input to the machine learning model, the previous detection of the one or more previous detections being associated with a confidence value indicative of a probability that the previous detection contains the polyp; a location of the polyp within the at least one preceding image frame; the current detection is associated with the previous detection based on the locations of the current detection and the previous detection; causing an indication of the at least one current detection in the video signal based on a position of the current detection in the detection chain, the confidence value of the current detection, and the location of the current detection; Including, the machine learning model takes as input the series of successive video frames and outputs the current detection based on visual evidence across the successive video frames via a deep convolutional neural network long short-term memory architecture; the steps of obtaining, generating, and causing display are repeated in real time during the inspection process for all video frames constituting the series of video frames, for at least one next video frame in the series as the at least one current video frame; method.
2. The method of claim 1 , wherein if the location overlap of the current detection and the previous detection satisfies a predetermined condition, the current detection is associated with the previous detection such that the current detection and the previous detection belong to the same detection chain.
3. 3. The method of claim 1 or 2, wherein an indication of the current detection is triggered if the current detection belongs to N+M detections of the detection chain, N and M being positive integers equal to or greater than 1, N indicating the N temporally first detections of the detection chain, and wherein an indication of the current detection is not triggered if the current detection belongs to the N temporally first detections of the detection chain.
4. The method of claim 1 , wherein an indication of the current detection is triggered if the confidence value of the current detection is greater than or equal to a first threshold.
5. 5. The method of claim 4, wherein an indication of the current detection is triggered if the confidence value of the current detection is greater than or equal to a second threshold less than a first threshold, and if the confidence value of the previous detection belonging to the same detection chain as the current detection is greater than or equal to the first threshold.
6. The method according to claim 1 , further comprising the step of performing a smoothing over the locations of the detections of the detection chain.
7. A computer-readable non-transitory storage medium storing a program which, when executed by a computer, causes the computer to carry out the method according to any one of claims 1 to 6.
8. An apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code acting in conjunction with the at least one processor to cause the apparatus to: inputting to the machine learning model one or more current detections for at least one current video frame of a series of consecutive video frames of a video signal captured by an endoscope during an inspection process, the current detections of the one or more current detections including a confidence value indicating a probability that the current detection includes a polyp to be detected by the machine learning model and a location of the polyp within the at least one current video frame; generating a detection chain by associating detections output from the machine learning model, wherein a current detection of the one or more current detections is associated with a previous detection of one or more previous detections obtained from the machine learning model for at least one previous video frame of the series that precedes the at least one current video frame and that is input to the machine learning model, the previous detection of the one or more previous detections being associated with a confidence value indicative of a probability that the previous detection contains the polyp; a location of the polyp within the at least one preceding image frame; the current detection is associated with the previous detection based on the locations of the current detection and the previous detection; causing an indication of the at least one current detection in the video signal based on a position of the current detection in the detection chain, the confidence value of the current detection, and the location of the current detection; repeating the steps of obtaining, generating, and causing the display for at least one next video frame in the series as the at least one current video frame during the inspection process in real time for all video frames constituting the series of video frames; and configured to execute at least The machine learning model takes as input the series of consecutive video frames and outputs the current detection based on visual evidence across the consecutive video frames using a deep convolutional neural network long short-term memory architecture. Device.
Citation Information
Patent Citations
Imaging device and program
JP2008288868A
Subject detection apparatus, imaging apparatus, subject detection apparatus control method, subject detection apparatus control program, and storage medium
JP2015104016A
Endoscope image processing device
WO2017073338A1
Image diagnosis assistance apparatus, data collection method, image diagnosis assistance method, and image diagnosis assistance program
WO2019088121A1