Image processing apparatus, information processing method, and program

The image processing device enhances tracking accuracy by combining detection and tracking results through a weighted average of position and size estimates from multiple units, addressing inaccuracies in subject size estimation.

JP2025182104APending Publication Date: 2025-12-11CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025169644
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-07
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing subject tracking technologies fail to effectively integrate detection results with tracking results, leading to inaccuracies in subject size estimation.

Method used

An image processing device that utilizes a first and second estimation unit to acquire subject position and size from a region of interest in a frame, followed by a determination unit that determines the region of interest in subsequent frames using a weighted average of the acquired positions and sizes from both units.

Benefits of technology

Improves tracking accuracy by integrating detection and tracking results, enhancing subject size estimation precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025182104000001_ABST
    Figure 2025182104000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of tracking by improving the accuracy of a tracker in estimating the size of a subject.SOLUTION: An image processing apparatus uses a first estimation unit that has been learned to estimate the position and size of a subject from a region of interest in each of frames of moving images, to acquire the position and size of a subject from the region of interest in a first frame. The image processing apparatus acquires the position and size of the subject in the first frame that are estimated based on a position acquired from the region of interest in the first frame by a second estimation unit that has been learned to estimate the position and size of a subject in a still image. The image processing apparatus determines a region of interest in a second frame subsequent to the first frame by using the position and size acquired by the first estimation unit and the position and size acquired by the second estimation unit.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an information processing method, and a program. [Background technology]

[0002] The technology of tracking a specific subject by extracting the image of the subject from a time-series of images is used to identify the human face or body area in a moving image. The technology of tracking a subject can be used in many fields, such as teleconferencing, man-machine interfaces, security, monitor systems for tracking a specific subject, and image compression.

[0003] Subject tracking technology is sometimes used to optimize the focus and exposure conditions for a subject. Patent Document 1 discloses technology for automatically tracking a specific subject using template matching. Patent Document 2 discloses technology for detecting a subject using a detection means different from a tracking means, and, if a predetermined condition is met, using the detection result from the detection means as the tracking result of the subject, thereby attempting more accurate tracking. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-60269 [Patent Document 2] JP 2014-7775 A [Non-patent literature]

[0005] [Non-Patent Document 1] Objects as Points, Xingyi Zhou et al., 2019 [Non-patent document 2] High Performance Visual Tracking with Siames Region Proposal Network, Bo Li et al., 2018 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in Patent Document 2, the detection results of the tracking means and the detection means are exclusively selected according to a predetermined criterion, and therefore it is not possible to reflect the detection result of the detection means in the detection result of the tracking means as necessary.

[0007] An object of the present invention is to improve the accuracy of tracking by improving the accuracy of subject size estimation by a tracker. [Means for solving the problem]

[0008] To achieve the object of the present invention, for example, an image processing device according to one embodiment includes the following configuration: a first acquisition means that acquires the position and size of a subject from a region of interest in a first frame using a first estimation unit that has been trained to estimate the position and size of the subject from a region of interest in each frame of a moving image, a second acquisition means that acquires the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in the first frame based on the position acquired by the first acquisition means from the region of interest in the first frame, and a first determination means that determines the position and size of the region of interest in a second frame that follows the first frame using a weighted average of the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means. [Effects of the Invention]

[0009] The accuracy of tracking is improved by improving the accuracy of subject size estimation by the tracker. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram showing an example of the functional arrangement of an image processing apparatus according to a first embodiment. [Figure 2] 4 is a flowchart showing an example of information processing according to the first embodiment. [Figure 3] FIG. 4 is a diagram for explaining a process of determining a region of interest according to the first embodiment. [Figure 4] FIG. 4 is a diagram for explaining a process of determining a region of interest according to the first embodiment. [Figure 5] FIG. 2 is a diagram for explaining a tracking process according to the first embodiment. [Figure 6] FIG. 10 is a diagram showing an example of the functional arrangement of an image processing apparatus according to a second embodiment. [Figure 7] 10 is a flowchart showing an example of information processing according to the second embodiment. [Figure 8] 10A and 10B are diagrams for explaining a process of determining a tracking target according to the second embodiment. [Figure 9] FIG. 1 is a diagram showing an example of the hardware configuration of an image processing apparatus. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0012] [Embodiment 1] The image processing device according to this embodiment acquires the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of a moving image. The image processing device then acquires the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in a still image based on the position acquired by the first estimation unit from the region of interest in the first frame. Furthermore, the image processing device determines the region of interest in the second frame that follows the first frame using the position and size acquired by the first estimation unit and the position and size acquired by the second estimation unit.

[0013] 1 is a block diagram showing an example of the functional configuration of an image processing device 100 according to this embodiment. The image processing device 100 according to this embodiment includes an image acquisition unit 101, an area determination unit 102, a feature acquisition unit 103, a candidate calculation unit 104, a subject detection unit 105, and an object determination unit 106. The image processing device 100 according to this embodiment is assumed to be connected to an imaging device 110 and a result output unit 120 so as to be able to communicate with them.

[0014] The imaging device 110 according to this embodiment is a device with an imaging function, such as a digital camera, a surveillance camera, or a smartphone, and acquires a captured image. The image processing device 100 according to this embodiment may be a device built into the imaging device 110, or may be a device separate from the imaging device 110, such as a personal computer or a server.

[0015] The image acquisition unit 101 acquires an image to be processed. The image acquisition unit 101 according to this embodiment can acquire time-sequential captured images (moving images) captured by the imaging device 110. The image acquisition unit 101 may acquire moving images stored in a storage device, or may acquire moving images via a network. Hereinafter, when simply referring to an "image," this term will be used without distinction between a moving image and a still image included in the moving image.

[0016] The region determination unit 102 determines a region of interest in which to detect a subject to be tracked from an image. Here, the region of interest is a partial region set in an image for detecting a subject, and in this embodiment, it is referred to as a region common to a tracking unit and a detection unit, which will be described later. The process of determining a region of interest according to this embodiment will be described later.

[0017] Furthermore, the region determination unit 102 sets a template region, which is a partial region for extracting template features, on the image. The feature acquisition unit 103 can extract template features to be used when detecting tracking target candidates from the template region. Fig. 4(a) shows an example of the template region set in this embodiment.

[0018] In FIG. 4( a), two detection results 403 are detected for an input image 401 by the subject detection unit 105, which will be described later. In this example, one of the detection results 403 is specified by a user specification 404, and the specified area is clipped and extracted as an image of interest 405. Here, the detection result 403 whose center position is closest to the position of the user specification 404 is set as the template area. The template area set here is an area indicating a subject detected by the area determination unit 102, but the setting method is not particularly limited to this. For example, a score may be assigned to subject candidates using a known detection technique, and the area determination unit 102 may set the area of ​​the candidate with the highest score as the template area.

[0019] FIG. 5(a) is a diagram illustrating a process for extracting template features from an image of interest. The feature acquisition unit 103 according to this embodiment can acquire template features from an image of interest using a process similar to that described in Non-Patent Document 2. Here, the feature acquisition unit 103 acquires intermediate features 502 by inputting the image of interest to a CNN 501 that has been trained in advance to extract features for tracking a subject in an input image. The intermediate features 502 are output maps of the final layer or intermediate layers of the CNN 501. In the example of FIG. 5(a), a portion of the intermediate features 502 (here, a width of 3 and a height of 3 from the center) is acquired as the template features 503. Note that the template features are not limited to these, as long as they reflect features used by the subject detection unit 105 to detect a subject from an image. For example, a color histogram or edge density may also be used.

[0020] The candidate calculation unit 104 and the target determination unit 106 extract features from the search image and track the subject based on the extracted features. The candidate calculation unit 104 according to this embodiment can detect the subject by estimating the position and size of the subject from the region of interest using an estimation unit (tracking unit) that has been trained to estimate the position and size of the subject from the region of interest in each frame of a video. Here, the search image is an image of a certain frame included in a video that is the target of processing for tracking the subject. In addition, hereinafter, the detection result obtained by tracking such a subject will be referred to as the "tracking result."

[0021] The candidate calculation unit 104 can output subject candidates from the search image. The candidate calculation unit 104 according to this embodiment outputs a similarity map indicating the positions of subject candidates in the search image (by likelihood for each position) and a size map indicating the size of the subject. Next, the target determination unit 106 can determine a subject (tracking target) in the search image based on the candidates output by the candidate calculation unit 104. Hereinafter, the subject determination process performed by the candidate calculation unit 104 and the target determination unit 106 according to this embodiment will be described with reference to FIG. 5(b).

[0022] In FIG. 5(b), the candidate calculation unit 104 can first generate intermediate features 504 by inputting features of a search image to a CNN. Next, the candidate calculation unit 104 convolves the generated intermediate features 504 with template features 503, and inputs the result to a CNN 505, thereby outputting a similarity map 506 and a size map 507. As the CNN 505, a fully-convolutional network (FCN) described in Non-Patent Document 2 may be used, or a configuration combined with a fully connected layer may be used. Here, the CNN 505 is trained in advance to minimize the error between the output similarity map 506 and size map 507 and values ​​provided as a teacher signal.

[0023] The target determination unit 106 determines a tracking target based on the tracking target candidates detected by the candidate calculation unit 104. Here, the target determination unit 106 determines the position and size of the subject in the image based on the similarity map 506 and the size map 507. For example, the target determination unit 106 can determine the subject at the position with the highest likelihood value in the similarity map 506 as the tracking target. The tracking target determination process by the target determination unit 106 can use any process based on a known tracking method. For example, a tracking target may be selected and determined from multiple candidates using NMS (Non-Maximum-Suppression) or Cosine-window. The target determination unit 106 may also determine a tracking target by integrating subject candidates.

[0024] Note that the processing in FIG. 5 is an example, and if it is possible to track the subject from the search image, it is possible to output the tracking result of the subject using a different known tracking processing.

[0025] The subject detection unit 105 detects a subject from a region of interest. The subject detection unit 105 can also detect a subject by estimating the position and size of the subject from the region of interest using an estimation unit (detection unit) that has been trained to estimate the position and size of an object from the region of interest in each frame of a still image. Hereinafter, the subject detection result obtained by such a detection unit will be referred to as the "still image detection result."

[0026] The output of the multilayer neural network of the detection unit according to this embodiment is a tensor with 12 columns, 8 rows, and 5 channels. The first channel of this tensor is a similarity map indicating the likelihood that each element is a subject at each position on the image corresponding to that element, the second channel indicates the offset amount in the x direction from each element to the subject center, and the third channel indicates the inferred value of the offset amount in the y direction. The fourth channel of the tensor indicates the width of the subject at each element, and the fifth channel indicates the inferred value of the height. From the information on these five channels, the likelihood of the subject's presence, as well as its center position and size, can be calculated. This tensor is called the object region candidate tensor.

[0027] The multilayer neural network that calculates the object region candidate tensor is trained in advance using a large amount of training data (sets of image, object width and height, and offset amount to the center) as in Non-Patent Document 1. In this embodiment, a multilayer neural network that simultaneously outputs five channels of information is used, but a different format may be used as long as similar information is obtained. For example, five multilayer neural networks that each output one channel may be prepared and the results may be combined. The object detection unit can detect the object based on the value of the obtained object region candidate tensor and the coordinates specified by the user 404 or the center position 408 of the tracking result.

[0028] Comparing a tracking unit (candidate calculation unit 104) that performs detection from moving images with a detection unit (subject detection unit 105) that performs detection from still images, the tracking unit has higher accuracy in estimating the position of the subject, while the detection unit has higher accuracy in estimating the size of the subject. From this perspective, subject detection unit 105 detects a subject from a region of interest using the tracking unit and sets the position as a tracking result, and then detects the subject using the detection unit based on the set position and estimates the size as a still image detection result.

[0029] The candidate calculation unit 104 and the subject detection unit 105 according to this embodiment output a tracking result and a still image detection result from a region of interest in a certain frame (here, frame n-1). Next, the region determination unit 102 according to this embodiment uses the tracking result and the still image detection result in frame n-1 to determine a region of interest in the frame subsequent to that frame (here, frame n).

[0030] The process of determining the region of interest in the next frame using the tracking results and still image detection results of a certain frame will be described with reference to Fig. 3. Images 301 and 304 are both images captured by the imaging device 110 in the (n-1)th frame, and are the same images except for the information displayed on the images. A region of interest 302 determined from information of the previous frame (the (n-2)th frame) is shown by a solid line on image 301. The candidate calculation unit 104 according to this embodiment detects a subject from the region of interest 302, and sets a center position 303 of the detected subject within the region of interest 302.

[0031] The detection unit according to this embodiment detects a subject using the central position 303 set by the tracking unit and estimates the size of the subject. Here, the detection unit can detect, for example, from among subject candidates detected from the image 301, the candidate whose central position is closest to the central position 303 as the subject. Note that, for example, if the detection unit outputs a similarity map indicating the likelihood of the subject for each position, the detection unit may attenuate the likelihood according to the distance from the central position 303 and then detect the candidate including the position with the highest likelihood as the subject. Here, the still image detection result 305 detected by the detection unit is indicated by a dashed dotted line.

[0032] Both images 306 and 309 are images captured by the imaging device 110 in the nth frame, and are the same images except for the information displayed on the images. Image 306 displays an attention area 302, a still image detection result 305, and an attention area 307 (indicated by a dotted line) for the nth frame that is determined using attention area 302 and still image detection result 305. The tracking unit according to this embodiment detects a subject from attention area 307, and sets a center position 308 of the detected subject within attention area 307.

[0033] Here, a process of determining the region of interest 307 by the region determining unit 102 using the region of interest 302 and the still image detection result 305 will be described. The region determining unit 102 according to this embodiment can determine, for example, a region obtained by correcting the region of interest 302 using the still image detection result 305 as the region of interest 307. For example, the region determining unit 102 may correct the coordinates of the four corners of the region of interest 302 using the coordinates of the corresponding four corners of the still image detection result 305, and set these coordinates as the coordinates of the four corners of the region of interest 307.

[0034] Here, the region determination unit 102 determines the coordinates of the four corners of the region of interest 302 and the still image detection result 305 as the coordinates of the four corners of the region of interest 307. Note that the method for determining the region of interest 307 is not particularly limited to this, and the coordinates of the four corners of the region of interest 302 and the still image detection result 305 may be determined as the coordinates of the four corners of the region of interest 307 by weighted averaging. For example, it may be considered that the tracking accuracy of the tracking unit for an object of a predetermined size or less (e.g., 30 pixels or less on a side) significantly decreases due to bias in the learning data. From this perspective, when the size of the object estimated by the tracking unit is a predetermined threshold or less (which can be set arbitrarily), the region determination unit 102 can set the weights in the weighted averaging so that the weight is heavier on the still image detection result 305 side. Furthermore, for example, when a similarity map indicating the positions of candidate objects in the search image is generated as the intermediate feature 504, the region determination unit 102 may perform the above correction according to the value of the similarity map. Furthermore, for example, when the object detection unit 105 generates a likelihood map indicating the positions of object candidates in an image, the region determination unit 102 may perform the above correction according to the value of the similarity map. For example, if the likelihood of the object position in these similarity maps is equal to or greater than a threshold, no correction may be performed, and if the likelihood corresponding to the detection result 403 with the highest IoU (Intersection over Union) is equal to or greater than a threshold, correction may be performed.

[0035] Note that these processes are merely examples, and the method for determining the region of interest 307 is not particularly limited to these, as long as information about the size of the subject included in the still image detection result 305 can be reflected in the region of interest 307. For example, the region determination unit 102 may determine the detection region of the still image detection result as the region of interest directly, rather than correcting the region of interest using the weighted average as described above.

[0036] When center position 308 is set, the detection unit detects the subject using center position 308 and estimates the size of the subject. In image 309, the detection unit detects still image detection result 310 by processing similar to that used to detect still image detection result 305 in image 304. Still image detection result 310 is indicated by a dashed dotted line. When subject imaging and detection continue, a similar processing is used to determine the region of interest in the next frame using region of interest 307 and still image detection result 310. Note that the subject detection processing by the tracking unit and detection unit can be performed by general moving object detection processing or object detection processing, and therefore detailed description thereof will be omitted.

[0037] It should be noted that the process for determining the region of interest described with reference to FIG. 3 may be sufficient if it is performed when the tracking accuracy is insufficient. Therefore, the candidate calculation unit 104 according to this embodiment may determine the region of interest in the next frame by the process shown in FIG. 3 only when the subject cannot be detected from the region of interest. That is, the candidate calculation unit 104 may determine whether or not the subject can be tracked in the region of interest. Next, if it is determined that the subject can be tracked, the candidate calculation unit 104 may determine the detection region of the subject to be tracked as the region of interest in the next frame, and if it is determined that the subject cannot be tracked, the candidate calculation unit 104 may determine the region of interest in the next frame by the process shown in FIG. 3. Furthermore, since there is no frame immediately before to refer to in the process of the first frame, the region of interest is set as the detection region of the subject.

[0038] 2 is a flowchart showing an example of information processing performed by the image processing device 100 according to this embodiment. The processing according to FIG. 2 is realized by the CPU of the image processing device 100 executing a control program. In S200, the image acquisition unit 101 acquires an image to be subjected to tracking processing. In this embodiment, the image acquisition unit 101 acquires the captured image of the nth frame from the imaging device 110. In S201, the image acquisition unit 101 converts the image acquired in S200 for subsequent processing.

[0039] An example of the image conversion process in S201 will be described below. Here, the captured image acquired by the imaging device 110 is assumed to be an RGB image with a width of 6000 pixels and a height of 4000 pixels. The image acquisition unit 101 converts the captured image to a predetermined size. In this example, the image acquisition unit 101 reduces the size of the captured image to one-tenth of its original size in both width and height, and generates an RGB image with a width of 600 pixels and a height of 400 pixels through conversion. However, a different conversion process may be performed. For example, the image acquisition unit 101 may further perform padding processing to convert the image size to one-twelfth of the captured image size, or may crop a predetermined area of ​​the captured image to create the converted image. The converted image has its upper left corner at coordinates (0,0) and its position at row j, column i at coordinates (i,j) (in this case, the lower right corner is (599,399)). Hereinafter, the image to be processed refers to the converted image.

[0040] In S202, the region determination unit 102 sets a template region, which is a region for acquiring template features. Here, the region determination unit 102 sets one of the regions indicating subject candidates detected from the image by the subject detection unit 105 as the template region. Next, in S203, the feature acquisition unit 103 acquires template features (503) from the image within the template region.

[0041] In S204, the region determination unit 102 determines a region of interest. Here, an image of n frames has been acquired in S200, and a region of interest in the image of n frames is determined based on the region of interest determined in the image of n-1 frames and the still image detection result (detected in the processing of S205 described later).

[0042] Fig. 4(b) illustrates an example of a region of interest determined by the region determining unit 102. In Fig. 4(b), a tracking target 407 determined by the target determining unit 106 and its center position 408 are specified for a past image 406 of the n-1th frame. Furthermore, based on the center position 408, the subject detection unit 105 detects a detection result 403 from the past image 406, and here a new region of interest 410 is determined based on the detection result 403 and the tracking result 407.

[0043] In S205, the candidate calculation unit 104 causes the tracking unit to track (detect) the subject from the region of interest determined in S204. In this embodiment, the candidate calculation unit 104 uses the image of the region of interest determined in S204 as a search image, and inputs feature amounts 504 of the search image and template feature amounts 503 acquired in S203 to a CNN 501 to generate a similarity map 506 and a size map 507.

[0044] In S206, the target determination unit 106 determines a tracking target in the region of interest based on the subject tracking result in S205. In S207, the result output unit 120 outputs the tracking result to an external device or the image capture device 110. The tracking result output in this manner can be used for AF functions, such as a phase difference AF (autofocus) function, by sampling several ranging points from the tracking result.

[0045] In S208, the image processing device 100 determines whether or not to end tracking of the subject. If tracking is to be ended, the processing relating to Fig. 2 ends, and if not, the processing returns to S200. Here, it is assumed that tracking of the subject is ended when a predetermined condition is met, such as when the user stops the imaging operation of the imaging device 110 or when a predetermined time has passed since tracking began.

[0046] Note that in S204, a region of interest in the image of frame n is determined based on the region of interest determined in the image of frame n-1 and the still image detection result (detected in the processing of S205 described below). However, as described above, if tracking of the subject is possible, it may not be necessary to determine the region of interest in this manner. From this perspective, the candidate calculation unit 104 may determine whether or not tracking of the subject is possible in the region of interest in the image of frame n-1. In this case, in S204, if tracking is possible, the detection region of the subject is determined as the region of interest; if not, the region of interest is determined by the processing described with reference to FIG. 3.

[0047] This type of processing makes it possible to determine a region of interest in a frame that follows a certain frame based on the results of estimation of the position and size of a subject by the tracking unit and the results of estimation of the position and size of the subject by the detection unit in the region of interest in that frame. Therefore, tracking can be performed that reflects the detection results by the detection unit, which is expected to have higher subject size estimation accuracy than the tracking unit, thereby improving tracking accuracy. As a result, a region of interest can be determined based on tracking and detection results that have already been obtained, thereby improving tracking accuracy even in processes that require real-time performance and have limited calculation time, such as AF functions.

[0048] [Embodiment 2] In the first embodiment, the tracking result and still image detection result of a certain frame (n-1 frame) are used to determine the region of interest in the next frame (n frame). The image processing device 600 according to this embodiment determines the region of interest in the n frame based on the tracking result of the n frame, and determines the tracking result in the region of interest by correcting the tracking result of the n frame by the tracking unit from the determined region of interest using the detection result of the n frame by the detection unit.

[0049] 6 is a block diagram showing an example of the functional configuration of an image processing device 600 according to this embodiment. The image processing device 600 has the same configuration as the image processing device 100 of the first embodiment, except that it includes a region determination unit 601, a subject detection unit 602, and a target determination unit 603 instead of the region determination unit 102, the subject detection unit 105, and the target determination unit 106, respectively. Functional units with the same reference numbers as those of the first embodiment are capable of performing the same processes, and redundant explanations will be omitted.

[0050] The subject detection unit 602 can perform the same processing as the subject detection unit 105, and can calculate subject candidates based on the image acquired from the image acquisition unit 101 and the subject tracking result acquired by the object determination unit 603 based on the image of the previous frame. The region determination unit 601 can also perform the same processing as the region determination unit 102 of the first embodiment, and sets one of the regions indicating subject candidates detected from the image by the subject detection unit 602 as a template region.

[0051] The object determination unit 603 corrects the tracking result by the tracking unit of the candidate calculation unit 104 for an image of a certain frame using the still image detection result by the detection unit of the object detection unit 602 for an image of the same frame. For example, considering that when the object is small, the detection unit is likely to estimate the size more accurately than the tracking unit, the object determination unit 603 may determine the size of the object as the size value estimated by the detection unit rather than the size value estimated by the tracking unit. Also, for example, the object determination unit 603 may determine the size of the object as a weighted average of the size value estimated by the tracking unit and the size value estimated by the detection unit. Note that while only size has been mentioned here, the value estimated by the detection unit may also be used as the position of the object.

[0052] The object determination unit 603 may determine the position and size of the object as values ​​estimated by the detection unit, rather than values ​​estimated by the tracking unit, depending on time-series fluctuations in the size of the object estimated by the tracking unit. Here, for example, if the difference between the average value of the object size estimated by the tracking unit over a predetermined period and the size of the object newly estimated by the detection unit exceeds a certain value, the object determination unit 603 may set the value of the size estimated by the detection unit as the value of the object size. According to this processing, for example, if the tracking unit detects a different object of a different size as the tracking target, the erroneous detection can be detected due to a sudden change in size, and the detection result by the detection unit can be applied.

[0053] The format of the subject position and size calculated here is not particularly limited. For example, the subject position estimated by the tracking unit and the subject position estimated by the detection unit may each be indicated by a similarity map. In other words, a value between 0 and 1 in the similarity map may be used as the "estimated size."

[0054] 8 is a diagram illustrating the process of determining a tracking target performed by the target determining unit 603 according to this embodiment. The target determining unit 603 determines a tracking result 805 using a captured image 801 and a detection result 804 output by the subject detection unit 602 using the captured image 801 as input. Here, for example, if the tracking unit detects a subject tracking candidate 806 as the tracking target instead of the subject corresponding to the tracking result 802, the position and size of the subject estimated by a detection unit with higher size estimation accuracy can be determined as the tracking target value. Therefore, it is possible to suppress erroneous detection of subjects that is obvious at a glance.

[0055] Hereinafter, the information processing performed by the image processing device 600 will be described with reference to Fig. 7. Fig. 7 is a flowchart showing an example of information processing performed by the image processing device 600 according to this embodiment. The processing according to Fig. 7 is realized by the CPU of the image processing device 600 executing a control program. In Fig. 7, the same processing as in Fig. 2 is performed except that S701, S702, and S70 are performed instead of S202, S204, and S206, respectively, and S703 to S704 are performed between S205 and S704, and therefore a duplicated description will be omitted.

[0056] In S701 following S201, the region determination unit 601 sets a template region, which is a region for acquiring template features. Here, the region determination unit 601 sets, as the template region, one of the regions indicating subject candidates output by the subject detection unit 602 based on the image and the tracking result of the previous frame. Post-processing of S701 proceeds to S203.

[0057] In S702 following S203, the region determination unit 601 determines a region of interest. In S734 following S703, the region determination unit 601 sets the region of interest determined in S702 to be used as a region for detecting a subject by the subject detection unit 602 (detection unit) in the image of the same frame (nth frame). In S704, the subject detection unit 602 detects a subject from the region of interest in the image of the same frame by using the detection unit.

[0058] In S705, the target determining unit 106 determines a tracking target in the region of interest based on the result of tracking the subject in S205 and the position and size of the subject acquired in S704.

[0059] [Embodiment 3] In the above-described embodiments, each processing unit shown in, for example, FIG. 1 is realized by dedicated hardware. However, some or all of the processing units of the image processing device 100 may be realized by a computer. In this embodiment, at least some of the processing according to each of the above-described embodiments is executed by a computer.

[0060] FIG. 9 is a diagram showing the basic configuration of a computer. In FIG. 9, a processor 901 is, for example, a CPU, and controls the operation of the entire computer. A memory 902 is, for example, a RAM, and temporarily stores programs, data, etc. A computer-readable storage medium 903 is, for example, a hard disk or a CD-ROM, and stores programs, data, etc. long-term. In this embodiment, a program that realizes the function of each unit, which is stored in the storage medium 903, is read into the memory 902. Then, the processor 901 operates in accordance with the program on the memory 902, thereby realizing the function of each unit.

[0061] 9, an input interface 904 is an interface for acquiring information from an external device. An output interface 905 is an interface for outputting information to an external device. A bus 906 connects the above-mentioned components and enables data exchange.

[0062] The disclosure of this specification includes the following image processing device, information processing method, and program.

[0063] (Item 1) a first acquisition means for acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a second acquisition means for acquiring the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in the still image based on the position acquired by the first acquisition means from a region of interest in the first frame; a first determination means for determining a region of interest in a second frame subsequent to the first frame, using the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means; An image processing device comprising:

[0064] (Item 2) Item 1. The image processing device according to item 1, characterized in that the first determination means determines the position and size of the region of interest in the second frame by correcting the position and size acquired by the first acquisition means with the position and size acquired by the second acquisition means.

[0065] (Item 3) Item 2. The image processing device according to item 2, characterized in that the first determination means determines the position and size of the region of interest in the second frame as a weighted average of the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means.

[0066] (Item 4) The image processing device described in any one of items 1 to 3, characterized in that, when the size acquired by the first acquisition means is equal to or smaller than a predetermined threshold, the first determination means determines the area of ​​interest in a second frame subsequent to the first frame using the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means.

[0067] (Item 5) The image processing device according to any one of items 1 to 4, characterized in that the positions acquired by the first acquisition means and the positions acquired by the second acquisition means are indicated by a similarity map indicating the likelihood of the subject for each position.

[0068] (Item 6) The first determination means when the first estimation unit cannot detect the subject from the region of interest in the first frame, determining a region of interest in a second frame subsequent to the first frame using the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means; 6. The image processing device according to any one of items 1 to 5, characterized in that, when the first estimation unit can detect the subject from the area of ​​interest in the first frame, the position and size acquired by the first acquisition means are determined as the position and size of the area of ​​interest in the second frame.

[0069] (Item 7) The image processing device according to any one of items 1 to 6, further comprising a second determination means for determining the position and size of the subject in the second frame based on the feature amount acquired from the still image of the first frame and the region of interest in the second frame.

[0070] (Item 8) a first acquisition means for acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a second acquisition means for acquiring the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in the still image based on the position acquired by the first acquisition means from a region of interest in the first frame; a third determination means for determining a position and a size of the subject in the first frame by correcting the position and the size acquired by the first acquisition means with the position and the size acquired by the second acquisition means; An image processing device comprising:

[0071] (Item 9) Item 9. The image processing device according to item 8, wherein the third determination means determines the position and size of the subject in the first frame as the position and size acquired by the second acquisition means.

[0072] (Item 10) 10. The image processing device according to item 8 or 9, wherein the position of the subject is indicated by a similarity map indicating the likelihood of the subject for each position.

[0073] (Item 11) acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a step in which a second estimation unit that has been trained to estimate the position and size of a subject in a still image acquires the position and size of the subject in the first frame estimated based on the position acquired by the first estimation unit from a region of interest in the first frame; determining a region of interest in a second frame subsequent to the first frame using the position and size acquired by the first estimation unit and the position and size acquired by the second estimation unit; An information processing method comprising:

[0074] (Item 12) acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a step in which a second estimation unit that has been trained to estimate the position and size of a subject in a still image acquires the position and size of the subject in the first frame estimated based on the position acquired by the first estimation unit from a region of interest in the first frame; determining the position and size of the subject in the first frame by correcting the position and size acquired by the first estimation unit with the position and size acquired by the second estimation unit; An information processing method comprising:

[0075] (Item 13) A program for causing a computer to function as each means of the image processing device according to any one of items 1 to 10.

[0076] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0077] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0078] 100: Image processing device, 110: Imaging device

Claims

1. a first acquisition means for acquiring the position and size of the subject from the region of interest in a first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of a video image; a second acquisition means for acquiring the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in the still image based on the position acquired by the first acquisition means from a region of interest in the first frame; a first determination means for determining a position and a size of a region of interest in a second frame subsequent to the first frame by using a weighted average of the position and the size acquired by the first acquisition means and the position and the size acquired by the second acquisition means; An image processing device comprising:

2. 2. The image processing device according to claim 1, wherein the first determination means determines the position and size of the region of interest in the second frame by correcting the position and size acquired by the first acquisition means with the position and size acquired by the second acquisition means.

3. 2. The image processing device according to claim 1, characterized in that, when the size acquired by the first acquisition means is equal to or smaller than a predetermined threshold, the first determination means determines the area of ​​interest in a second frame subsequent to the first frame using the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means.

4. The image processing device according to claim 1 , wherein the positions acquired by the first acquisition means and the positions acquired by the second acquisition means are indicated by a similarity map indicating the likelihood of the subject at each position.

5. The first determining means when the first estimation unit cannot detect the subject from the region of interest in the first frame, determining a region of interest in a second frame subsequent to the first frame using the position and size acquired by the first acquisition means and the position and size acquired by the second acquisition means; 2. The image processing device according to claim 1, characterized in that, when the first estimation unit can detect the subject from the area of ​​interest in the first frame, the position and size acquired by the first acquisition means are determined to be the position and size of the area of ​​interest in the second frame.

6. 2. The image processing device according to claim 1, further comprising: a second determination means for determining a position and a size of a subject in the second frame based on a feature amount obtained from a still image of the first frame and a region of interest in the second frame.

7. a first acquisition means for acquiring the position and size of the subject from the region of interest in a first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of a video image; a second acquisition means for acquiring the position and size of the subject in the first frame estimated by a second estimation unit that has been trained to estimate the position and size of the subject in the still image based on the position acquired by the first acquisition means from a region of interest in the first frame; a third determination means for determining a position and a size of the subject in the first frame by correcting the position and the size acquired by the first acquisition means by a weighted average using the position and the size acquired by the second acquisition means; An image processing device comprising:

8. 8. The image processing device according to claim 7, wherein the third determining means determines the position and size of the subject in the first frame as the position and size acquired by the second acquiring means.

9. The image processing device according to claim 7, wherein the position of the subject is indicated by a similarity map indicating the likelihood of the subject for each position.

10. acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a step in which a second estimation unit that has been trained to estimate the position and size of a subject in a still image acquires the position and size of the subject in the first frame estimated based on the position acquired by the first estimation unit from a region of interest in the first frame; determining a position and a size of a region of interest in a second frame subsequent to the first frame using a weighted average of the position and the size acquired by the first estimation unit and the position and the size acquired by the second estimation unit; An information processing method comprising:

11. acquiring the position and size of the subject from the region of interest in the first frame using a first estimation unit that has been trained to estimate the position and size of the subject from the region of interest in each frame of the video; a step in which a second estimation unit that has been trained to estimate the position and size of a subject in a still image acquires the position and size of the subject in the first frame estimated based on the position acquired by the first estimation unit from a region of interest in the first frame; determining the position and size of the subject in the first frame by correcting the position and size acquired by the first estimation unit by a weighted average using the position and size acquired by the second estimation unit; An information processing method comprising:

12. A program for causing a computer to function as each of the means of the image processing device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Object tracking method and device

    JP2001060269A

  • Image processing device, image processing method and program

    JP2014007775A