Detection Frame Position Accuracy Improvement System and Detection Frame Position Correction Method

The system improves detection frame position accuracy in vehicle cameras by utilizing information from multiple frames to estimate and correct uncertainties, enhancing object detection precision.

JP7698481B2Active Publication Date: 2025-06-25ASTEMO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021099602
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-15
Publication Date
2025-06-25
Estimated Expiration
2041-06-15

AI Technical Summary

Technical Problem

Existing methods for improving detection frame position accuracy in vehicle cameras are limited to using information from a single frame and do not effectively utilize information before and after the target frame.

Method used

A system that includes a time-series image input unit, object detection unit, detection frame position distribution estimation unit, detection frame prediction unit, detection frame uncertainty estimation unit, and detection frame correction unit to estimate and correct the detection frame position using information before and after the target frame.

Benefits of technology

Enhances the accuracy of detection frame position estimation by correcting for uncertainties using information from multiple frames, thereby improving the precision of object detection in vehicle cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698481000001
    Figure 0007698481000001
  • Figure 0007698481000002
    Figure 0007698481000002
  • Figure 0007698481000003
    Figure 0007698481000003
Patent Text Reader

Abstract

To provide a detection frame position accuracy improvement system and a detection frame position correction method which can highly accurately estimate a detection frame position by utilizing information before and after an object frame.SOLUTION: A detection frame position accuracy improvement system comprises: a time-series image input unit 10 which inputs time-series images; an object detection unit 20 which detects an object in the time-series images; a detection frame position distribution estimation unit 30 which estimates a distribution of detection frame position coordinates of a time t from a detection result of the object until a time t-1 (t is a positive integer); a detection frame prediction unit 40 which predicts a position of a detection frame of times t+1 to t+n (n is a positive integer) on the basis of the detection result and the distribution; a detection frame uncertainty estimation unit 50 which updates the distribution of the detection frame position coordinates at the time t with the detection result of the object at the times t+1 to t+n and the overlapping degree with the predicted detection frame, and estimates uncertainty of the detection frame at the time t; and a detection frame correction unit 60 which corrects the detection frame at the time t on the basis of the detection frame and the uncertainty.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a detection frame position accuracy improvement system and a detection frame position correction method.

Background Art

[0002] With the spread of in-vehicle cameras and the like, the diversity of vehicle data that can be obtained has been increasing. As a result, there is an increasing need for objective situation awareness and cause analysis using the information of a recording terminal device that records the acquired vehicle data, especially at the time of an accident or the like. In situation awareness and cause analysis using the image of an in-vehicle camera (camera image), the position accuracy of a detection frame (the size of an object such as a vehicle ahead in the camera image) is important.

[0003] In improving the detection frame position accuracy, there is a technique described in Japanese Patent No. 6614247 (Patent Document 1). This document describes "prediction means for predicting the position of an object in the current frame from the position of the object in the previous frame with respect to the current frame and specifying a prediction area, and determining whether the object exists in a first distance range or in a second distance range farther than the first distance range based on the distance of the object in the previous frame. When it is determined by the determination means that the object exists in the first distance range, template matching is performed using a first template for the object in the previous frame in the prediction area of the current frame to detect the object, and a first matching processing means; when it is determined by the determination means that the object exists in the second distance range, template matching is performed using a second template different from the first template for the object in the previous frame in the prediction area of the current frame to detect the object, and a second matching processing means."

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In Patent Document 1 above, an attempt is made to estimate the detection frame position with high accuracy using only the frame before the target frame. Therefore, the improvement in the detection frame position accuracy of the target frame is limited to the case where the detection frame position accuracy in the previous frame is good, and the improvement in the detection frame position accuracy by correcting the detection frame position using the information before and after the target frame is not assumed.

[0006] Therefore, in view of the above circumstances, an object of the present invention is to provide a detection frame position accuracy improvement system and a detection frame position correction method that can estimate the detection frame position with high accuracy by using information before and after the target frame.

Means for Solving the Problems

[0007] In order to solve the above problems, one of the typical detection frame position accuracy improvement systems of the present invention includes a time-series image input unit that inputs time-series images, an object detection unit that detects an object in the time-series images, a detection frame position distribution estimation unit that estimates the distribution of the detection frame position coordinates at the correction target time from the detection results of the object up to the time before the correction target time, a detection frame prediction unit that predicts the position of the detection frame at a time after the correction target time according to the detection results and the distribution, a detection frame uncertainty estimation unit that updates the distribution of the detection frame position coordinates at the correction target time based on the degree of overlap between the detection result of the object and the predicted detection frame at a time after the correction target time, and estimates the uncertainty of the detection frame at the correction target time, and a detection frame correction unit that corrects the detection frame at the correction target time based on the detection frame and the uncertainty.

Effects of the Invention

[0008] According to the present invention, it is possible to improve the accuracy of the detection frame position.

[0009] Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Modes for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0012] [Example 1] FIG. 1 is a block diagram of Example 1 of the present invention. In this example, the case where it is applied to sensor information obtained from a vehicle will be described. The detection frame position accuracy improvement system 1 shown in FIG. 1 is a system that corrects the detection frame position of an object on an image offline using a time-series image or a distance measurement sensor.

[0013] In the following description, when performing correction (specifically, determining whether correction is necessary and performing correction if it is determined to be necessary), the correction target time is set as time t (t is a positive integer), the time before (past) the correction target time is time t - n (n is a positive integer), and the time after (future) the correction target time is time t + n (n is a positive integer).

[0014] Also, in the following description, although a vehicle such as a preceding vehicle is used as the detection / correction target, it is of course not limited to vehicles only.

[0015] The detection frame position accuracy improvement system 1 shown in FIG. 1 includes a time-series image input unit 10 that inputs time-series images taken and stored by a driving recorder or the like mounted on a vehicle separately from this system, an object detection unit 20 that detects an object (target object) such as a vehicle, a two-wheeler, or a pedestrian in the image input by the time-series image input unit 10, a detection frame position distribution estimation unit 30 that estimates the distribution of the detection frame position coordinates of the image at a certain time t when performing correction, a detection frame prediction unit 40 that predicts the detection frame positions at times t + 1 to t + n based on the outputs of the object detection unit 20 and the detection frame position distribution estimation unit 30, a detection frame uncertainty estimation unit 50 that estimates the uncertainty of the image position (= detection frame) at time t based on the degree of overlap between the predicted detection frame and the detection frame detected in each image by a detector, and a detection frame correction unit 60 that corrects the detection frame using the uncertainty. Details of each of the functions 10, 20, 30, 40, 50, and 60 will be described below.

[0016] The time-series image input unit 10 inputs the images obtained by an imaging device such as a monocular camera or a stereo camera in chronological order.

[0017] Using FIG. 2, the object detection unit 20 will be described. In the object detection unit 20, in each of the time-series images, a region (also referred to as a detection frame) including an object is estimated by a human or a detector. 70 is one image of the time-series images, 80 is the object to be detected, and in FIG. 2, the object is a car. 90 is the detection frame when the object is detected, and the position of the detection frame is determined by specifying (x1, y1) at the upper left of the detection frame and (x2, y2) at the lower right of the detection frame. Here, the detection frame is shown in two dimensions of vertical and horizontal, but a three-dimensional detection frame of vertical, horizontal, and height may also be the target.

[0018] Using FIG. 3, the detection frame position distribution estimation unit 30 will be described. In the detection frame position distribution estimation unit 30, using the detection frame positions up to time t - 1, the probability distribution of the detection frame position coordinates of the image at time t when correction is performed is estimated. 100 is one image of the time-series images, 110 shows the probability distribution of the coordinates on the image where x1 constituting the detection frame exists, 120 shows the probability distribution of the coordinates on the image where x2 exists, 130 shows the probability distribution of the coordinates on the image where y1 constituting the detection frame exists, and 140 shows the probability distribution of the coordinates on the image where y2 exists. Here, a normal distribution is illustrated as the probability distributions of 110, 120, 130, and 140, but the distribution of the coordinates is not limited to the normal distribution. 150 represents the contour line of the bivariate normal distribution of x1 and y1 at the upper left coordinates (x1, y1) of the detection frame of the object. 160 represents the contour line of the bivariate normal distribution of x2 and y2 at the lower right coordinates (x2, y2) of the detection frame of the object. The higher parts of the contour lines of 150 and 160 are the places where the probability is high as the detection frame position coordinates. For the prediction of this probability distribution, statistical methods such as a Kalman filter can be applied.

[0019] The detection frame prediction unit 40 will be described with reference to FIG. 4. The detection frame prediction unit 40 includes a detection frame movement amount acquisition unit 41 that estimates the movement amount of the detection frame from the relative speed of the object at times t to t + n, etc., and a detection frame position sampling unit 42 that samples the upper left coordinate and the lower right coordinate (detection frame position coordinates) of the detection frame at time t based on the probability distribution estimated by the detection frame position distribution estimation unit 30, and a detection frame position prediction output unit 43 that determines the detection frame positions at times t + 1 to t + n from the detection frame movement amount acquisition unit 41 and the detection frame position sampling unit 42. Details of 41, 42, and 43 will be described.

[0020] The detection frame movement amount acquisition unit 41 predicts the change in the size of the detection frame and the direction of the object determining the position (destination of movement) at times t + 1 to t + n, the relative speed between the host vehicle and the object, etc. from the detection information up to times 1 to t - 1 by a Kalman filter or the like, and determines the movement amount of the detection frame. Also, if it is possible to use ranging sensors such as LIDAR and millimeter waves at times t + 1 to t + n, the distance to the object and the object area range may be obtained by these sensors, and the relative speed and direction may be obtained. Furthermore, regarding the movement amount, a method of limiting the upper limit of the movement amount in light of physical laws is also conceivable.

[0021] The detection frame position sampling unit 42 outputs the upper left coordinate and the lower right coordinate (detection frame position coordinates) of the detection frame at time t with high probability based on the probability distribution estimated by the detection frame position distribution estimation unit 30. Furthermore, coordinates with low probability are also randomly output with a certain probability ε so that the detection frame position coordinates can be output globally.

[0022] The detection frame position prediction output unit 43 obtains the position coordinates (also referred to as predicted detection frames) of the detection frame at times t + 1 to t + n with the detection frame at time t (detection frame based on the probability distribution) determined by the detection frame position sampling unit 42 as the initial value and the movement amount by the detection frame movement amount acquisition unit 41 as the constraint condition.

[0023] The detection frame uncertainty estimation unit 50 will be described with reference to FIG. 5. The detection frame uncertainty estimation unit 50 includes a detection frame overlap calculation unit 51 that calculates the degree of overlap between the detection frame (predicted detection frame) predicted by the detection frame prediction unit 40 and the detection frame estimated by the object detection unit 20, a detection frame position distribution update unit 52 that updates the probability distribution estimated by the detection frame position distribution estimation unit 30 based on the degree of overlap, and a detection frame uncertainty output unit 53 that calculates a region (detection frame considering uncertainty) where there may be a detection frame at time t from the estimated probability distribution. Details of 51, 52, and 53 will be described in detail.

[0024] The detection frame overlap calculation unit 51 evaluates how well the detection frame (predicted detection frame) predicted by the detection frame prediction unit 40 and the detection frame estimated by the object detection unit 20 match, using the degree of overlap between the detection frames. As an evaluation index for the degree of overlap, IoU (Intersection over Union) etc. can be considered.

[0025] The detection frame position distribution update unit 52 can use the value (degree of overlap) of the detection frame overlap calculation unit 51 to update the mean and variance of the multivariate normal distribution of the detection frame position coordinates using Bayesian update, or use reinforcement learning to obtain the mean and variance that maximize the reward with the degree of overlap as the reward.

[0026] The detection frame uncertainty output unit 53 outputs a region (detection frame considering uncertainty) where there may be a detection frame at time t, using the standard deviation etc. of the probability distribution of the detection frame position coordinates estimated by the detection frame position distribution update unit 52 of the detection frame uncertainty estimation unit 50. Details will be described later with reference to FIG. 7.

[0027] The part from the detection frame prediction unit 40 to the detection frame overlap calculation unit 51 of the detection frame uncertainty estimation unit 50 will be described with reference to FIG. 6. The time-series image 200 is an image at certain times t+1, t+2, and t+3, and is a case where the detection frame at the upper part of the object is sampled at time t. 170 is composed of the predicted detection frame at time t+1, the detection frame estimated by a detector or the like (that is, the detection frame estimated by the object detection unit 20), and the region where they overlap. 180 is composed of the predicted detection frame at time t+2, the detection frame estimated by a detector or the like, and the region where they overlap. 190 is composed of the predicted detection frame at time t+3, the detection frame estimated by a detector or the like, and the region where they overlap. In the time-series image 200, since the detection frame sampled at time t is at the upper part of the object, even considering the predicted movement amount (detection frame movement amount acquisition unit 41), the predicted detection frames for a relatively short time such as from time t to t+3 are present at the upper part of the object. The time-series image 210 is an image at certain times t+1, t+2, and t+3, and is a case where the detection frame at the lower part of the object is sampled at time t. In the time-series image 210, since the detection frame sampled at time t is at the lower part of the object, even considering the predicted movement amount (detection frame movement amount acquisition unit 41), the predicted detection frames for a relatively short time such as from time t to t+3 are present at the lower part of the object. The time-series image 220 is an image at certain times t+1, t+2, and t+3, and is a case where a large detection frame is sampled for the object at time t. In the time-series image 220, since the detection frame is predicted to be large for the object at time t, even considering the predicted movement amount (detection frame movement amount acquisition unit 41), the predicted detection frames for a relatively short time such as from time t to t+3 become large for the object. Also, since the coordinate values at time t are different in 200, 210, and 220 respectively, the detection frame position coordinates from time t+1 to t+3 are different, but the size of the detection frame (for example, the magnification rate between times t+1, t+2, and t+3) is determined by the detection frame movement amount acquisition unit 41 (the movement amount), so they are all equal in 200, 210, and 220.

[0028] The detection frame uncertainty estimation unit 50 will be described with reference to FIG. 7. In this embodiment, the detection frame 230 visualizing uncertainty is based on the probability distribution obtained by the detection frame position distribution update unit 52, and is the minimum size (standard deviation in the case where the probability distribution is a multivariate normal distribution) of the detection frame of the object that can be allowed as the detection frame of the object set in advance. It is composed of three detection frames: the detection frame 240, the detection frame 250 with the coordinates having the highest probability (average in the case where the probability distribution is a multivariate normal distribution), and the detection frame 260 with the maximum size (standard deviation in the case where the probability distribution is a multivariate normal distribution) that can be allowed as the detection frame of one object set in advance. The sizes of the detection frames 240, 250, and 260 can be determined by the probability distribution of the position coordinates (in other words, the existence range of the detection frame at time t can be limited from the probability distribution of the updated detection frame position coordinates). When assuming a large variation, a large standard deviation is taken. For example, if three times the standard deviation is taken, it is predicted that the detection frame is included with a 99% probability within the range set from 240 to 260.

[0029] The detection frame prediction unit 40 and the detection frame uncertainty estimation unit 50 will be described with reference to the flowchart of FIG. 8. First, in step 270, the movement amount of the object (detection frame) at times t + 1 to t + n is estimated using a Kalman filter or the like from the outputs of the object detection unit 20 and the detection frame position distribution estimation unit 30 (detection frame movement amount acquisition unit 41). Alternatively, the movement amount such as the relative speed is estimated by using a distance measurement sensor. Here, if a large value is set for n, the prediction range becomes too long and the prediction accuracy decreases. On the other hand, if n is too small, when the detection frame is automatically output by the detector, there will be many undetected images (detection frame non-detection images), or a large deviation in the position of one detection frame may become an outlier, which is likely to reduce the correction accuracy. Therefore, it is necessary to determine the value of n in consideration of the frame rate of the obtained image.

[0030] In step 280, the detection frame position coordinates at time t are output according to the probability distribution estimated by the detection frame position distribution estimation unit 30 (detection frame position sampling unit 42). At this time, if only coordinates with high probability are output, the sampling position accuracy will decrease when the estimation accuracy of the detection frame position distribution estimation unit 30 is low. Therefore, coordinates with low probability are also randomly output with probability ε so that the detection frame position coordinates can be output globally.

[0031] In step 290, the detection frame positions (detection frame position coordinates) at times t + 1 to t + n are predicted using the results of steps 270 and 280 (detection frame position prediction output unit 43).

[0032] In step 300, the degree of overlap between the predicted detection frames at times t + 1 to t + n and the detection frames output from the detector at each time is calculated (detection frame overlap calculation unit 51). The degree of overlap is calculated using IoU (Intersection over Union) or the like.

[0033] In step 310, the detection frame position coordinate distribution (probability distribution) at time t is updated according to the degree of overlap (detection frame position distribution update unit 52). That is, the probability is updated to be high for the detection frame position coordinates at time t where the degree of overlap is high, and the probability is updated to be low for the detection frame position coordinates at time t where the degree of overlap is low.

[0034] In step 320, it is determined whether the sampling count has reached a set value pre-set by the user. If the sampling count has been reached, the process ends. If the sampling count has not been reached, the process returns to step 280, and the detection frame position coordinates at time t are sampled again. Since the detection frame position coordinate distribution at time t is updated in step 310, by repeatedly sampling, many coordinates with a high degree of overlap with the detection frames output from the detector at times t + 1 to t + n will be sampled.

[0035] The detection frame correction unit 60 will be described with reference to FIG. 9. 330, 340, 350, and 360 indicate the types of detection frames used in this figure. The solid line 330 is the detection frame output by a human or a detector in each image (in other words, estimated by the object detection unit 20). On the other hand, based on the probability distribution obtained by the detection frame position distribution update unit 52, the two-dot chain line 340 is the detection frame with the minimum size that can be assumed (acceptable as the detection frame of the preset target object), the dashed line 350 is the detection frame with the highest probability, and the one-dot chain line 360 is the detection frame with the maximum size that can be assumed (acceptable as the detection frame of the preset target object) (detection frame uncertainty output unit 53).

[0036] 370 is an image visualizing the uncertainty detection frames of the detection frames 330 and 340, 350, 360. The detection frame 330 output by the detector includes noise 380. The noise 380 corresponds to, for example, the shadow of the object due to backlight. Here, the detection frame 330 includes the noise 380 and is output larger than the detection frame that only detects the object. At this time, the detection frame 330 becomes larger than the maximum detection frame 360 (of uncertainty) and becomes the detection frame to be corrected. When correcting, a method of replacing the detection frame 330 with the detection frame 350 having the maximum probability of the detection frame can be considered. The image 390 is the result of correcting the detection frame 330 in the image 370 (detection frame correction unit 60). After correction, it becomes a detection frame that only detects the object and does not include the noise 400. In the image 370, the case where the detection frame 330 is larger than the maximum detection frame 360 that can be assumed has been described. Conversely, when the detection frame 330 is smaller than the minimum detection frame 340 that can be assumed, it can be corrected (corrected) in the same manner.

[0037] 410 is an image visualizing the uncertainty detection frames of detection frames 330 and 340, 350, 360. The detection frame 330 output by the detector is segmented by noise 420. The noise 420 corresponds to the case where a part of the vehicle ahead is hidden by a wiper, a motorcycle, or the like. In the image 410, since the two detection frames 330 are inside the maximum allowable detection frame 360 and outside the minimum allowable detection frame 340, they are determined to be detection frames for the same object and become the detection frames to be corrected. When correcting, methods such as integrating the two detection frames 330 or replacing them with the detection frame 350 where the probability of the detection frame is maximized can be considered. The image 430 is the result of correcting the detection frame 330 in the image 410 (detection frame correction unit 60). After correction, it is not affected by the noise 440 and becomes the detection frame that detects the object.

[0038] However, the detection frame correction method using the detection frame uncertainty by the detection frame correction unit 60 is not limited to the method described here.

[0039] In the first embodiment of the present invention, with the functional configuration described above, by using the information before and after the image to be corrected to estimate the uncertainty of the detection frame position, the variation of the detection frame due to noise can be corrected with high precision.

[0040] As described above, the detection frame position accuracy improvement system 1 of Example 1 of the present invention includes a time-series image input unit 10 that inputs time-series images, an object detection unit 20 that detects an object in the time-series images, a detection frame position distribution estimation unit 30 that estimates the distribution of the detection frame position coordinates at a correction target time (time t) from the detection results of the object up to a time before the correction target time (time t-1), a detection frame prediction unit 40 that predicts the positions of the detection frames at times after the correction target time (times t+1 to t+n) according to the detection results and the distribution, a detection frame uncertainty estimation unit 50 that updates the distribution of the detection frame position coordinates at the correction target time (time t) based on the degree of overlap between the detection results of the object and the predicted detection frames at times after the correction target time (times t+1 to t+n), and estimates the uncertainty of the detection frame at the correction target time (time t), and a detection frame correction unit 60 that corrects the detection frame at the correction target time (time t) based on the detection frame and the uncertainty.

[0041] Further, the detection frame prediction unit 40 includes a detection frame position sampling unit 42 that samples the position coordinates of the detection frame at the correction target time (time t) from the distribution estimated by the detection results, and a detection frame movement amount acquisition unit 41 that acquires a movement amount including at least one of the relative speed or direction of the object at times after the correction target time (times t+1 to t+n) that determines the movement destination of the detection frame. The detection frame position sampling unit 42 determines the detection frame position at the correction target time (time t), and the detection frame movement amount acquisition unit 41 predicts the positions of the detection frames at times after the correction target time (times t+1 to t+n) based on the movement amount.

[0042] Further, the detection frame uncertainty estimation unit 50 limits the existence range of the detection frame at the correction target time (time t) from the updated distribution of the detection frame position coordinates.

[0043] In addition, the detection frame position correction method according to Embodiment 1 of the present invention inputs time-series images, detects an object in the time-series images, estimates the distribution of the detection frame position coordinates at the correction target time (time t) from the detection results of the object up to the time before the correction target time (time t-1), predicts the positions of the detection frames at times after the correction target time (times t+1 to t+n) according to the detection results and the distribution, updates the distribution of the detection frame position coordinates at the correction target time (time t) according to the degree of overlap between the detection results of the object and the predicted detection frames at times after the correction target time (times t+1 to t+n), estimates the uncertainty of the detection frame at the correction target time (time t), and corrects the detection frame at the correction target time (time t) based on the detection frame and the uncertainty.

[0044] That is, in this Embodiment 1, data such as time-series images before and after the frame to be corrected for the detection frame position and a distance sensor are used to estimate the region (uncertainty) where the current detection frame exists, and the detection result output by a detector or the like is corrected.

[0045] According to this Embodiment 1, it is possible to improve the accuracy of the detection frame position.

[0046] [Embodiment 2] FIG. 10 is a block diagram of Embodiment 2 of the present invention. In this embodiment, it is targeted at the case where a plurality of objects are included in the same image and there are a plurality of detection frames.

[0047] The detection frame position accuracy improvement system 2 shown in FIG. 10 includes a time-series image input unit 10 that inputs time-series images captured and stored by a drive recorder or the like mounted on a vehicle separately from this system, an object detection unit 20 that detects an object (target object) such as a vehicle, a two-wheeler, or a pedestrian in the images input by the time-series image input unit 10, a detection correction target object determination unit 450 that determines a detection frame to be corrected in the time-series images, a detection frame position distribution estimation unit 30 that estimates the distribution of the detection frame position coordinates of the image at a certain time t for which correction is to be performed, a detection frame prediction unit 40 that predicts the detection frame positions from time t+1 to t+n based on the outputs of the object detection unit 20 and the detection frame position distribution estimation unit 30, a detection frame uncertainty estimation unit 50 that estimates the uncertainty of the image position (=detection frame) at time t based on the degree of overlap between the predicted detection frame and the detection frame detected from the image by a detector, and a detection frame correction unit 60 that corrects the detection frame using the uncertainty. 10, 20, 30, 40, 50, and 60 have functions equivalent to those described in the first embodiment.

[0048] The detection correction target object determination unit 450 will be described with reference to FIG. 11. The detection correction target object determination unit 450 includes a detection information extraction unit 451 that extracts feature amounts and the like of the target object (detection frame) used to determine whether the objects are the same, a detection target classification unit 452 that classifies the target objects in the entire time-series images based on the information of the detection information extraction unit 451, and a detection correction target object output unit 453 that outputs the detection frame of the object (detection correction target object) to be corrected.

[0049] Examples of the feature amounts extracted by the detection information extraction unit 451 include labels of the detected objects such as automobiles, humans, and two-wheelers for each detection frame, feature amount descriptors universal to scales and rotations such as SIFT (Scale Invariant Feature Transform), and feature amount descriptors output by applying a learned convolutional neural network a plurality of times.

[0050] In the detection target classification unit 452, for each image and each detection frame, the detection frames are determined and classified for the same object in the time-series images by using the Euclidean distance or cosine similarity for the feature amounts obtained by the detection information extraction unit 451.

[0051] The detection correction target object output unit 453 outputs the detection frames to be corrected. Also, when the detection frames are automatically output by the detector, if a large number of detection omissions occur and the number of detections is small, and correction is difficult or there is a high possibility of a decrease in correction accuracy, a notification is given to the user.

[0052] In the second embodiment of the present invention, with the functional configuration described above, even when a plurality of objects are included in the image, it is possible to narrow down the correction target to one in advance, and by using the information before and after the image of the correction target, the uncertainty of the detection frame position is estimated, so that the variation of the detection frame due to noise can be corrected with high accuracy.

[0053] As described above, the detection frame position accuracy improvement system 2 according to the second embodiment of the present invention includes a detection correction target object determination unit 450 that determines the same object in the time-series images in addition to the above-described first embodiment.

[0054] Further, the detection correction target object determination unit 450 extracts the feature amounts of each detection frame (detection information extraction unit 451), determines the same object in the time-series images from the feature amounts (detection target classification unit 452), and has a detection correction target object output unit 453 that serves as the detection frame correction target object.

[0055] According to the second embodiment, even when a plurality of objects are included in the same image, it is possible to improve the accuracy of the detection frame position.

[0056] Note that the present invention is not limited to the above-described embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Further, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations is possible. Also, each of the above configurations, functions, processing units, processing means, etc. may be realized in hardware, for example, by designing a part or all of them with an integrated circuit. Also, each of the above configurations, functions, etc. may be realized in software by a processor interpreting and executing a program for realizing each function. Information such as a program, table, file, etc. for realizing each function can be placed in a memory, a recording device such as a hard disk, SSD (Solid State Drive), or a recording medium such as an IC card, SD card, DVD. Also, control lines and information lines are shown as those considered necessary for explanation, and not all control lines and information lines are necessarily shown on the product. In practice, it may be considered that almost all configurations are interconnected.

Explanation of Signs

[0057] 1… Detection frame position accuracy improvement system (Example 1), 2… Detection frame position accuracy improvement system (Example 2), 10… Time-series image input unit, 20… Object detection unit, 30… Detection frame position distribution estimation unit, 40… Detection frame prediction unit, 50… Detection frame uncertainty estimation unit, 60… Detection frame correction unit, 450… Detection correction target determination unit (Example 2)

Claims

1. A time-series image input unit that inputs time-series images, An object detection unit that detects an object from the time-series images, A detection frame position distribution estimation unit that estimates the distribution of the detection frame position coordinates at the correction target time from the detection results of the object up to the time before the correction target time, A detection frame prediction unit that predicts the position of the detection frame at a time after the correction target time according to the detection results and the distribution, A detection frame uncertainty estimation unit that updates the distribution of the detection frame position coordinates at the correction target time according to the degree of overlap between the detection result of the object at a time after the correction target time and the predicted detection frame, and estimates the uncertainty of the detection frame at the correction target time, A detection frame correction unit that corrects the detection frame at the correction target time based on the detection frame and the uncertainty, characterized in that it comprises a detection frame position accuracy improvement system.

2. In the detection frame position accuracy improvement system according to Claim 1, The detection frame prediction unit is characterized by having a detection frame position sampling unit that samples the position coordinates of the detection frame at the correction target time from the distribution estimated by the detection results.

3. In the detection frame position accuracy improvement system according to Claim 1, The detection frame prediction unit is characterized by having a detection frame movement amount acquisition unit that acquires a movement amount including at least one of the relative speed or direction of the object at a time after the correction target time that determines the movement destination of the detection frame.

4. In the detection frame position accuracy improvement system according to Claim 1, The detection frame prediction unit comprises a detection frame position sampling unit that samples the position coordinates of the detection frame at the correction target time from the distribution estimated by the detection results, and a detection frame movement amount acquisition unit that acquires a movement amount including at least one of the relative speed or direction of the object at a time after the correction target time that determines the movement destination of the detection frame. The detection frame position sampling unit determines the detection frame position at the correction target time, and the movement amount by the detection frame movement amount acquisition unit predicts the position of the detection frame at a time after the correction target time, characterized in that it comprises a detection frame position accuracy improvement system.

5. In the detection frame position accuracy improvement system according to Claim 1, The detection frame uncertainty estimation unit is characterized by limiting the existence range of the detection frame at the correction target time from the updated distribution of the detection frame position coordinates.

6. In the detection frame position accuracy improvement system according to claim 5, the detection frame uncertainty estimation unit includes, as the detection frame limiting the existence range, a detection frame with the minimum size and a detection frame with the maximum size based on the standard deviation of the distribution of the updated detection frame position coordinates, and a detection frame based on the coordinates with the highest probability in the distribution of the updated detection frame position coordinates. A detection frame position accuracy improvement system characterized by this.

7. In the detection frame position accuracy improvement system according to claim 1, a detection correction target object determination unit for determining the same target object in the time-series images is provided. A detection frame position accuracy improvement system characterized by this.

8. In the detection frame position accuracy improvement system according to claim 7, the detection correction target object determination unit extracts the feature amount of each detection frame, determines the same target object in the time-series images from the feature amount, and has a detection correction target object output unit that uses it as the detection correction target object. A detection frame position accuracy improvement system characterized by this.

9. A computer inputs time-series images, detects an object in the time-series images, estimates the distribution of the detection frame position coordinates at the correction target time from the detection results of the object up to the time before the correction target time, predicts the position of the detection frame at the time after the correction target time according to the detection results and the distribution, updates the distribution of the detection frame position coordinates at the correction target time according to the degree of overlap between the detection result of the object at the time after the correction target time and the predicted detection frame, estimates the uncertainty of the detection frame at the correction target time, corrects the detection frame at the correction target time based on the detection frame and the uncertainty. A detection frame position correction method characterized by this.

10. In the detection frame position correction method according to claim 9, the computer samples the position coordinates of the detection frame at the correction target time from the distribution estimated from the detection results, obtains a movement amount including at least one of the relative speed or direction of the object at the time after the correction target time for determining the movement destination of the detection frame, determines the detection frame position at the correction target time by the sampling, and predicts the position of the detection frame at the time after the correction target time by the obtained movement amount. A detection frame position correction method characterized by this.

11. In the detection frame position correction method according to claim 9, The detection frame position correction method is characterized in that the computer estimates the uncertainty of the detection frame at the correction target time by limiting the existence range of the detection frame at the correction target time from the distribution of the updated detection frame position coordinates.

Citation Information

Patent Citations

  • Image processing device, object recognition device, equipment control system, image processing method, and program

    JP2017151535A

  • Control program, control method, and information processing device

    JP2019036009A

  • Image processing device, object recognition device, equipment control system, image processing method and program

    JP6614247B2