Information processing apparatus and imaging apparatus
The information processing device enhances subject tracking accuracy by using depth information and occlusion estimation to differentiate between similar subjects and handle occlusions, improving tracking precision.
Patent Information
- Application Number
- JP2024102814
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-15
AI Technical Summary
Existing subject tracking technologies erroneously recognize similar subjects as the tracking target or incorrectly estimate occlusion relationships due to similar features, leading to inaccurate tracking.
An information processing device that acquires images in series and depth distance information, detects candidate regions using image features, estimates occlusion states based on time-series distance data, and determines the subject to be tracked by analyzing the occlusion state.
Accurately tracks subjects by minimizing errors caused by similar features and occlusions, ensuring precise subject identification.
Smart Images

Figure 2026004822000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for tracking a subject. [Background technology]
[0002] Technologies for extracting specific subject images from images captured in a time series and tracking the subjects are used to identify human facial and body regions in moving images. Technologies for tracking specific subjects within images include those that use brightness and color information, template matching, and machine learning techniques such as deep neural networks. Non-Patent Document 1 describes a method for identifying the location of a tracked subject in an image by inputting an image containing the tracked subject and an image searching for the tracked subject into a convolutional neural network with the same weights and calculating the correlation between the obtained feature quantities. Patent Document 1 also describes a method for estimating the occlusion relationship between each object detected in an image and other objects, and identifying a correspondence relationship between the object detected in an image captured at a different time from the image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2022-19339 [Non-patent literature]
[0004] [Non-Patent Document 1] L. Bertinetto et al “Fully-Convolutional Siamese Networks for Object Tracking”, ECCV2016. Summary of the Invention [Problem to be solved by the invention]
[0005] However, the method described in Non-Patent Document 1 may erroneously recognize a similar subject as the tracking target when the tracking target subject and a subject with similar features are close to each other. Also, the method described in Patent Document 1 may erroneously estimate the occlusion relationship for subjects with similar features.
[0006] An object of the present invention is to accurately track a subject to be tracked. [Means for solving the problem]
[0007] The information processing device of the present invention is characterized by having an acquisition means for acquiring images captured in time series and depth distance information in multiple regions of the images; a detection means for detecting candidate regions of a subject to be tracked from the images based on image features of the images; an estimation means for estimating an occlusion state indicating whether the subject to be tracked is occluded by another subject different from the tracking target for the candidate regions detected from the images based on time series data of the distance information; and a determination means for determining the candidate regions of the subject to be tracked from the candidate regions detected from the images based on the estimation result of the occlusion state. [Effects of the Invention]
[0008] According to the present invention, it is possible to track a subject to be tracked with high accuracy. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an imaging apparatus. [Figure 2] FIG. 10 is an explanatory diagram of a defocus amount. [Figure 3] 1 is a diagram showing the functional configuration of an imaging device according to a first embodiment. [Figure 4] FIG. 2 is a diagram illustrating a configuration of a first estimation unit. [Figure 5] 4 is a flowchart showing a tracking process according to the first embodiment. [Figure 6]10 is a flowchart showing a first shielding state estimation process. [Figure 7] FIG. 1 is a diagram showing a scene in which a subject to be tracked moves near another subject. [Figure 8] 10 is a flowchart showing a first occlusion determination process. [Figure 9] FIG. 1 is a diagram showing a scene in which a subject to be tracked moves near another subject. [Figure 10] FIG. 10 is a diagram showing the functional configuration of an imaging device according to a second embodiment. [Figure 11] 10 is a flowchart showing a tracking process according to the second embodiment. [Figure 12] 10 is a flowchart showing a second shielded state estimation process. [Figure 13] FIG. 1 is a diagram showing a scene in which a subject to be tracked moves near another subject. [Figure 14] 10 is a flowchart showing a second occlusion determination process. [Figure 15] FIG. 2 is a diagram showing time-series changes in the position of a focus lens of an imaging device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple configurations are described in the embodiments, not all of these multiple configurations are necessarily essential to the invention, and multiple configurations may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar configurations, and redundant explanations will be omitted.
[0011] <Embodiment 1> FIG. 1 shows an example of the hardware configuration of an imaging device including an information processing device according to this embodiment. In FIG. 1, the imaging device 10 is a digital camera with interchangeable lenses. The imaging device may be any electronic device with an imaging function, such as a PTZ (pan-tilt-zoom) camera, a mobile phone (smartphone), or a PC (personal computer). In this embodiment, the information processing device is configured integrally with the imaging device, but may also be configured as a PC or the like externally connected to the imaging device. In this case, the information processing device acquires captured images and information at the time of capture from the imaging device and performs various processes.
[0012] As shown in FIG. 1, an imaging device 10 according to this embodiment is made up of a camera body 100 and a lens unit 200 that guides incident light to an imaging element 101. First, a description will be given of the camera body 100. The camera body 100 has an image sensor 101, a system control unit 102, a shutter 103, a shutter control unit 112 that controls the shutter 103, memory 104, a power switch 105, a mode switching unit 106, and a communication I / F 114.
[0013] The image sensor 101 is configured with a CMOS imaging sensor and converts an optical signal, which is an optical image, into an electrical signal. This electrical signal undergoes predetermined signal processing and is output as a video signal to the system control unit 102. Light rays incident on the photographing lens 201 pass through the aperture 202 and shutter 103 and form an optical image on the imaging surface of the image sensor 101.
[0014] The system control unit 102 is configured with a CPU and other components, and is connected to the lens control unit 205 of the lens unit 200 to control the entire imaging device 10. The system control unit 102 includes an image processing unit for video signals output from the image sensor 101. The system control unit 102 also includes a phase-difference AF unit that performs focus detection processing using a phase-difference detection method based on focus detection image data (signals for phase-difference AF) obtained from the image sensor 101 and the image processing unit. More specifically, the image processing unit generates a pair of image data formed by light beams passing through a pair of pupil regions of the imaging optical system as focus detection image data (first focus detection signal and second focus detection signal). The phase-difference AF unit detects the amount of focus deviation based on the amount of deviation between the pair of image data. In this way, the phase-difference AF unit of this embodiment performs phase-difference AF (image-plane phase-difference AF) based on the output of the image sensor 101 without using a dedicated AF sensor.
[0015] The system control unit 102 is connected to a memory 104, a power switch 105, a mode switching unit 106, and a communication I / F 114. The memory 104 is composed of a volatile memory, a nonvolatile memory, etc. The nonvolatile memory of the memory 104 stores programs for operating the system control unit 102, variables of various parameters, constants, etc. The volatile memory of the memory 104 also temporarily stores setting values of various parameters such as ISO sensitivity. Furthermore, the volatile memory of the memory 104 stores, in chronological order, images captured by the image sensor 101 and a predetermined number of frames of depth information of the images. Details of the depth information will be described later.
[0016] The power switch 105 is a switch for turning the power of the camera body 100 on and off. The mode switching unit 106 is a switch for switching and setting various shooting modes such as live view shooting and video shooting. The communication I / F 114 is an interface for connecting to an external device via a wired or wireless communication path. The system control unit 102 transmits captured images and information at the time of shooting to the external device via the communication I / F 114, and receives control signals and various setting information from the external device.
[0017] The camera body 100 also has a rear monitor 107 and a touch panel 108. The rear monitor 107 and the touch panel 108 are connected to the system control unit 102. The rear monitor 107 is an example of a display unit and is composed of a liquid crystal device, LEDs, etc. Under the control of the system control unit 102, the rear monitor 107 displays an image (live view image) currently being captured by the image sensor 101 and imaging information such as text, figures, and icons representing various information. The system control unit 102 may superimpose rectangular frames corresponding to the areas of the subject to be tracked and other subjects on the live view image. The touch panel 108 is an example of an operation unit and accepts operations from the user. The touch panel 108 is located in an area substantially equivalent to the rear monitor 107, detects contact with the user's finger or pen, and notifies the system control unit 102 of the contact position on the rear monitor 107. The system control unit 102 executes processing associated with the contact position on the touch panel 108.
[0018] The camera body 100 is also equipped with an electronic viewfinder (EVF). The electronic viewfinder (EVF) is composed of a viewfinder display unit 109 and an eyepiece 110. Similar to the rear monitor 107, the viewfinder display unit 109 displays live view images and various types of imaging information under the control of the system control unit 102. An eyecontact detection unit 111 detects whether the user has placed their eye on the camera. The system control unit 102 switches the display of the imaging information between the rear monitor 107 and the viewfinder display unit 109 depending on the detection result of the eyecontact detection unit 111.
[0019] Next, the configuration of the lens unit 200 will be described. The camera body 100 and the lens unit 200 are mechanically and electrically joined via a lens mount mechanism 113, and are detachable. The lens unit 200 has a photographing lens 201, an aperture 202, a lens drive circuit 203, an aperture control circuit 204, and a lens control unit 205. For simplicity, only one photographing lens 201 is shown in FIG. 1, but in reality, the lens unit is made up of a group of multiple photographing lenses including a focus lens.
[0020] The lens control unit 205 controls the entire lens unit 200 under the control of the system control unit 102. The lens control unit 205 also includes a memory (not shown) that stores programs for lens operation, setting values of various parameters, and information specific to the lens unit, such as maximum and minimum aperture values and focal length.
[0021] The system control unit 102 of the camera body 100 calculates the defocus amount using output information from the image sensor 101. Then, the system control unit 102 communicates via the lens control unit 205 of the lens unit 200 based on the calculated defocus amount, and controls the lens drive circuit 203 to achieve focus. The lens control unit 205 also acquires lens drive information related to the drive amount of the focus lens of the photographing lens 201 in the optical axis direction, and outputs the lens drive information to the system control unit 102.
[0022] (Relationship between defocus amount and image shift amount) Here, the relationship between the defocus amount and the image shift amount (phase difference) based on the first focus detection signal and the second focus detection signal output from the image sensor 101 will be described with reference to FIG.
[0023] The image sensor 101 is disposed on the imaging plane 300 in FIG. 2, and the exit pupil of the imaging optical system is divided into a first pupil region 311 and a second pupil region 312. The defocus amount d is defined as |d|, where the distance from the imaging position C of the light beam from the subject to the imaging plane 300 is defined as a positive sign (d>0) in a front-focus state where the imaging position C of the subject is located closer to the subject than the imaging plane 300. The back-focus state where the imaging position C of the subject is located on the opposite side of the subject than the imaging plane 300 is defined as a negative sign (d<0). The in-focus state where the imaging position C of the subject is located on the imaging plane 300 is d=0. In FIG. 2, the subject 321 shows an example of an in-focus state (d=0), and the subject 322 shows an example of a front-focus state (d>0). The front-focus state (d>0) and the back-focus state (d<0) are collectively referred to as a defocus state (|d|>0).
[0024] In a front-focus state (d>0), the light beam from the subject 322 that passes through the first pupil region 311 (second pupil region 312) is first focused and then spreads to a width Γ1 (Γ2) centered at the center of gravity G1 (G2) of the light beam, forming a blurred image on the imaging surface 300. This blurred image is received by each first focus detection pixel (each second focus detection pixel) on the imaging surface 300 of the image sensor 101, and a first focus detection signal (second focus detection signal) is generated. In other words, the first focus detection signal (second focus detection signal) is a signal that represents an image of the subject 322 at the center of gravity G1 (G2) of the light beam on the imaging surface 300, blurred by the blur width Γ1 (Γ2).
[0025] The blur width Γ1 (Γ2) of the subject image increases roughly in proportion to the increase in the magnitude |d| of the defocus amount d. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 in the center of gravity positions of the light beams) between the first focus detection signal and the second focus detection signal also increases roughly in proportion to the increase in the magnitude |d| of the defocus amount d. In the back-focus state (d<0), the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state, but the same is true.
[0026] In this way, the amount of image shift between the first and second focus detection signals increases as the amount of defocus increases. In this embodiment, image-pickup-plane phase-difference detection focus detection is performed, in which the amount of defocus is calculated from the amount of image shift between the first and second focus detection signals obtained using the image sensor 101.
[0027] Therefore, because the magnitude of the image shift amount between the first and second focus detection signals increases as the defocus amount increases, the phase difference AF unit of the system control unit 102 converts the image shift amount into a detected defocus amount using a conversion coefficient calculated based on the base length. Note that the unit of defocus amount in this embodiment is the product [Fδ] of the aperture F-number in the imaging optical system at the time of imaging and the permissible circle of confusion diameter δ.
[0028] Next, the tracking process according to this embodiment will be described. In this embodiment, during tracking processing, in a scene in which a tracking target subject selected by a user during shooting moves near a similar subject, the occlusion state of the tracking target subject is estimated using time-series data of depth information based on the defocus amount. The occlusion state indicates whether the tracking target subject is occluded by another subject. For ease of explanation, the subject will be assumed to be a person in the following description, but the scope of application of this embodiment is not limited to people and can be applied to movable subjects such as animals and vehicles.
[0029] In the following description, a scene will be described in which a similar subject is present near the subject to be tracked and the similar subjects overlap each other.
[0030] 3 shows an example of the functional configuration of the imaging device 10 according to this embodiment. The imaging device 10 functions as an acquisition unit 401, a setting unit 402, a feature extraction unit 403, a detection unit 404, a depth information acquisition unit 405, a first estimation unit 406, and a determination unit 407 when the system control unit 102 executes a program stored in the memory 104 or the like.
[0031] The acquisition unit 401 acquires images captured by the image sensor 101 in time series. Here, frames of a moving image captured by the image capture device 10 are acquired sequentially. The setting unit 402 sets the subject to be tracked based on a user input. For example, when the user touches the touch panel 108 on the captured image displayed on the rear monitor 107, the setting unit 402 detects the subject area closest to the touched position and sets it as the tracking target. When detecting the subject area, for example, a machine learning model trained using a known technology is used. When using a machine learning model, the setting unit 402 applies the machine learning model to the acquired image and detects the human area in the image as the subject area.
[0032] Feature extraction unit 403 extracts image features of the subject region of the tracking target set by setting unit 402. The image features may be, for example, a template image of the subject region of the tracking target, or may be image features extracted by calculation of a trained machine learning model for the subject region of the tracking target, as in Non-Patent Document 1. The extracted image features are stored in memory 104.
[0033] The detection unit 404 detects, from the image acquired by the acquisition unit 401, an area having image features similar to those extracted by the feature extraction unit 403 as a candidate area for a subject to be tracked. For example, as in Non-Patent Document 1, the detection unit 404 performs a correlation calculation between the image features of the tracking target extracted by the feature extraction unit 403 and image features extracted from the image, and detects an area where the matching cost is below a threshold as a subject candidate. When the image features are expressed as n-dimensional feature vectors, the matching cost is given, for example, by the L1 distance between the feature vectors of the image features; the smaller the L1 distance, the more similar the image features are. In this way, the tracking target can be identified by searching for an area where the matching cost is minimum.
[0034] The depth information acquisition unit 405 acquires depth information representing the amount of defocus detected in each focus detection area on the imaging surface, corresponding to the images acquired in time series by the acquisition unit 401. When frames of a moving image are acquired by the acquisition unit 401, depth information for each frame is acquired. Depth information will be described with reference to FIG. 7, so its description will be omitted here. The amount of defocus is an example of distance information in the depth direction. Note that the distance information in the depth direction is not limited to the amount of defocus, and a depth map obtained by stereo matching processing between multiple images may also be used. The first estimation unit 406 uses the depth information acquired by the depth information acquisition unit 405 to estimate an occlusion state that indicates whether or not the subject to be tracked is occluded by another subject.
[0035] As shown in FIG. 4, the first estimation unit 406 includes a depth holding unit 408 , a context estimation unit 409 , and a first occlusion determination unit 410 . The depth storage unit 408 stores the depth information acquired by the depth information acquisition unit 405 in chronological order. For example, the depth storage unit 408 associates tracking information of a subject to be tracked and subjects near the subject to be tracked in images of the past N frames with depth information about the area of each subject, and stores the linked information in the memory 104. The front-to-back relationship estimation unit 409 estimates the front-to-back relationship in the depth direction between the subject to be tracked and subjects near the subject to be tracked (whether the subject to be tracked is located closer to or further from other subjects) using the depth information of the past N frames stored in the depth storage unit 408. The first occlusion determination unit 410 determines the occlusion state of the subject to be tracked using the front-to-back relationship estimated by the front-to-back relationship estimation unit 409 for the past N frames.
[0036] Based on the estimation result of the first estimation unit 406, the determination unit 407 determines the subject to be tracked from the subject candidates detected by the detection unit 404.
[0037] 5 is a flowchart showing the tracking process of the imaging device 10 according to this embodiment. This flowchart is implemented by the system control unit 102 executing a program stored in the memory 104 or the like. Each process (step) is denoted with an S at the beginning to omit the process (step). This flowchart starts when the mode is switched to tracking mode in response to a user instruction.
[0038] In S501, the acquisition unit 401 acquires an image captured by the image sensor 101. The acquired image is displayed as a live view image on the rear monitor 107. In this embodiment, the acquisition unit 401 acquires frames of a moving image. In S502, the setting unit 402 determines whether or not a tracking target has already been set. If a tracking target has already been set, the process proceeds to S504, and if not, the process proceeds to S503. In S503, setting unit 402 detects a subject area that is close to the position specified on touch panel 108 in the image displayed on rear monitor 107, and sets it as the tracking target. Furthermore, feature extraction unit 403 extracts image features of the detected subject area. Then, setting unit 402 records image features and the like of a template of the subject area to be tracked on the image in memory 104. Note that when an image captured by imaging device 10 is distributed to an external device, the subject area to be set as the tracking target may be set based on operation information received from the external device.
[0039] In S504, the feature extraction unit 403 extracts image features from the image acquired in S501. Then, the detection unit 404 reads out the image features of the tracking target from the memory 104 and detects subject candidates by matching the extracted image features with the read image features. As a result, subject regions similar to the tracking target are detected as subject candidates. In this embodiment, the detection unit 404 detects subject candidates for each frame. In S505, the depth information acquisition unit 405 acquires depth information corresponding to the image acquired in S501. In S506, the first estimation unit 406 estimates the occlusion state of the subject to be tracked. Details of the first occlusion state estimation process executed in S505 will be described later with reference to FIGS.
[0040] In S507, the determination unit 407 determines the tracking target from the subject candidates detected in S504 based on the result of estimating the occlusion state of the tracking target in S506. In S508, the system control unit 102 determines whether the tracking mode has ended. As long as the tracking mode continues, the system control unit 102 returns to S501 and acquires images. In this embodiment, frames of a moving image are acquired sequentially. If it is determined that the tracking mode has ended, the series of tracking processes shown in FIG. 5 ends.
[0041] FIG. 6 is a flowchart showing the first occlusion state estimation process executed in S506 of FIG. 5. FIG. 7 shows an example of a scene in which a tracking target subject moves from left to right near a similar subject. The upper part of FIG. 7 shows successive frames. The lower part of FIG. 7 shows depth information corresponding to the frame in the upper part. In image 701 of FIG. 7(a), it is assumed that a subject area 712 is set as a tracking target. First subjects 710, 714, 718, and 722 are subjects of the same person detected as subject candidates in S504. Here, the first subject is a subject to be tracked. Second subjects 711, 715, 719, and 723 are subjects of the same person detected as subject candidates in S504.
[0042] In S601, the first estimation unit 406 acquires subject candidates. For example, the first estimation unit 406 may acquire the subject candidates detected in S504 as the subject candidates as they are, or may narrow down the subject candidates to those that are a certain distance away from the coordinates of the center of gravity of the tracking target in the immediately preceding frame.
[0043] In S602, the first estimation unit 406 extracts subject candidates that are in the vicinity of the tracking target from among the subject candidates acquired in S601. For example, the first estimation unit 406 extracts subject candidates that overlap with the tracking target in the immediately preceding frame. In the image 702 of FIG. 7(b), subject candidate 716 may be extracted as being in an overlapping state with the first subject 710 in the immediately preceding frame, and subject candidate 717 may be excluded from the targets as being in a non-overlapping state.
[0044] Here, IoU (Intersection over Union) is used as an example of an index for evaluating the degree of overlap. For example, the IoU value between rectangular regions surrounding the subject is calculated, and if the IoU value is 0.1 or greater, it is determined that the regions overlap. If there is at least a partial overlap, it may be determined that there is an overlap, and the IoU value threshold for determining that there is an overlap may be adjusted as appropriate.
[0045] In S603, the first estimation unit 406 determines whether the tracking target subject is not occluded and is not close to (not overlapping with) other subject candidate images in the most recent image. If the tracking target subject is not occluded and is not overlapped with other subject candidate images in the most recent image, it is determined that there is no risk of the tracking target subject being occluded, and the process proceeds to S604 and subsequent steps. Alternatively, for example, if it is determined that the tracking target subject is not occluded in the immediately preceding frame and the subject candidate images in the current frame are not overlapping, the process may proceed to S604 and subsequent steps. In other words, the first estimation unit 406 may proceed to S604 and subsequent steps if the tracking target subject is not occluded in the immediately preceding frame and the subject candidate images in the immediately preceding or current frame are not overlapping.
[0046] Furthermore, the first estimation unit 406 proceeds to processing from S608 onwards when the subject to be tracked is not occluded and is close to (overlapped with) other subject candidates in the most recent image. Also, for example, if it is determined that the subject to be tracked is not occluded in the immediately preceding frame and the subject candidates in the current frame overlap, the first estimation unit 406 may proceed to processing from S608 onwards. In other words, the first estimation unit 406 may proceed to processing from S608 onwards when the subject to be tracked is not occluded in the immediately preceding frame and the subject candidates in the immediately preceding or current frame overlap.
[0047] Furthermore, if the subject to be tracked is occluded in the most recent image, the first estimation unit 406 may detect only subject candidates on the foreground and may not be able to calculate the degree of overlap, so the first estimation unit 406 proceeds to processing from S608 onwards. For example, if the subject to be tracked is occluded in the immediately preceding frame, the first estimation unit 406 may proceed to processing from S608 onwards.
[0048] In S604, the first estimation unit 406 determines the tracking target. The tracking target is determined, for example, by associating the subject candidate, whose distance from the coordinates of the center of gravity of the tracking target determined in the immediately preceding frame is within a threshold and whose matching cost with the image features of the template of the tracking target set in S503 is the smallest, with the tracking target. Furthermore, the first estimation unit 406 performs similar association for other subject candidates.
[0049] In the case of image 702 in FIG. 7(b), for example, the first estimation unit 406 performs template matching on each of the candidate subject 716 and candidate subject 717 in the current frame against a template of the first subject 710, which is the tracking target in the immediately preceding frame. Then, the combination with the lowest matching cost is determined to be the tracking target in image 702. Here, the combination of first subject 710 and candidate subject 716 is deemed to have the lowest matching cost. In addition to the matching cost, a penalty may be added to the matching cost value as the distance from the tracking target in the immediately preceding frame increases.
[0050] In S605, first estimation unit 406 acquires tracking information for the most recent N frames. Here, tracking information refers to area information (coordinates, width, and height) and subject ID of the tracking target, and area information (coordinates, width, and height) and subject ID of other subjects different from the tracking target detected as subject candidates. In the following description, subject IDs are assigned such that the subject ID of the tracking target is 0, and the subject IDs of other subjects are 1, . . . , n (n≧1). In the case of FIG. 7, the description will be given assuming that first subject 710 has subject ID=0, and second subject 711 has subject ID=1.
[0051] In S606, the first estimation unit 406 acquires depth information for the most recent N frames. The lower part of FIG. 7 shows defocus maps 705-708, which are examples of depth information obtained by dividing each of images 701-704 into regions in a grid pattern and mapping the defocus amount corresponding to each grid region. A grayscale color bar 709 corresponds to the value of the defocus amount. The defocus map uses a density of 0ΔF as a reference, with a darker color representing a farther location (deeper) and a lighter color representing a closer location (closer). Note that in an actual defocus map, defocus amounts exist in the background region as well, but in the defocus maps 705-708 of FIG. 7, only the defocus amounts of the subject region (person region) are extracted for ease of explanation.
[0052] 7(a), when the focus is on the first object 710 to be tracked, an area 712 of the first object 710 corresponds to an area 726 of the defocus map 705, and therefore it can be read that the defocus amount d indicates a value of d=0. On the other hand, an area 713 of the second object 711 corresponds to an area 727 of the defocus map 705, and therefore it can be read that the defocus amount d indicates a value of d>0. In other words, it can be read that the second object 711 is located farther away (deeper) than the first object 710.
[0053] In S607, the first estimation unit 406 acquires a time series transition of the defocus amount for each of the first object and the second object, and estimates the anteroposterior relationship between the first object and the second object. Hereinafter, for the image 702 in FIG. 7(b), a time series list (queue) of the defocus amount for the coordinates of the object area for the most recent three frames is acquired for the first object (object ID=0) and the second object (object ID=1). Then, the anteroposterior relationship between the first object and the second object is estimated.
[0054] In this embodiment, since there is sensor noise when measuring the defocus amount, it is easy to misjudge the context if only the defocus amount value of the current frame is used, so the context is estimated using multiple defocus amounts arranged in time series, including the immediately preceding frame, which makes it possible to improve robustness.
[0055] First, assume that the chronological list of defocus amounts for a first subject (subject ID=0) is [0Fδ, 0Fδ, 0Fδ]. Also assume that the chronological list of defocus amounts for a second subject (subject ID=1) is [0Fδ, 1Fδ, 2Fδ]. In this case, [0Fδ, -1Fδ, -2Fδ] is acquired as a chronological list of the difference (depth difference) in the defocus amounts of the second subject (subject ID=1) relative to the first subject (subject ID=0).
[0056] In estimating the anteroposterior relationship, it is often sufficient to know the sign of the depth difference, and the list may be divided by the product [Fδ] of the aperture F-number in the imaging optical system and the allowable circle of confusion diameter δ. Therefore, in this embodiment, the anteroposterior relationship is estimated using a list obtained by dividing the list of depth differences by [Fδ].
[0057] An example of a method for estimating the context will be described. For example, if the signs of the depth differences in the most recent M frames are all the same and positive, the first estimation unit 406 estimates that the first subject (subject ID=0) is located further back than the second subject (subject ID=1). Also, for example, if the signs of the depth differences in the most recent M frames are all the same and negative, the first estimation unit 406 estimates that the first subject (subject ID=0) is located closer to the second subject (subject ID=1).
[0058] Here, it is assumed that [0, -1, -2] is acquired as a time-series list of depth differences, and the most recent two frames are used. In this case, the depth differences between the first subject (subject ID = 0) and the second subject (subject ID = 1) all have the same sign and are negative, so the first estimation unit 406 estimates that the first subject (subject ID = 0) is located closer to the second subject (subject ID = 1). On the other hand, when the most recent three frames are used, the signs do not match for three consecutive frames, so the first estimation unit 406 estimates that the anteroposterior relationship between the first subject (subject ID = 0) and the second subject (subject ID = 1) is unknown. The first estimation unit 406 records the estimation result of the anteroposterior relationship with the second subject (subject ID = 1) in the memory 104. If there are two or more subjects other than the tracking target detected as subject candidates, a time-series list of depth differences is obtained for each of the other subjects (subject ID = 1, ..., n), and the anteroposterior relationship of each of the other subjects with respect to the tracking target is estimated.
[0059] Another example of a chronological relationship estimation method will be described. For example, the first estimation unit 406 calculates a weighted average of depth differences for the most recent M frames and determines whether the calculated value is negative (positive) and whether the absolute value is equal to or greater than a predetermined value. If the weighted average is negative (positive) and the absolute value is equal to or greater than a predetermined value, the first estimation unit 406 estimates that the first subject (subject ID=0) is located closer to (farther from) the second subject (subject ID=1). On the other hand, if the absolute value of the weighted average is less than the predetermined value, the first estimation unit 406 estimates that the chronological relationship between the first subject (subject ID=0) and the second subject (subject ID=1) is unknown.
[0060] For example, if the most recent three frames are used, the weight w is set to w=[0.1, 0.3, 0.6] so that the weight decreases as the frame advances. For example, the value used to determine that the front-to-back relationship is unclear is estimated to be 0.5. In this case, the first estimation unit 406 estimates that the tracking target is located further back than the second subject (subject ID=1) if the absolute value of the weighted average is 0.5 or greater and positive, and estimates that the tracking target is located closer to the second subject (ID=1) if the absolute value of the weighted average is 0.5 or greater and negative.
[0061] Assuming that the depth difference time series list is [0,-1,-2] as described above, the weighted average of the depth differences for the most recent three frames can be calculated by taking the inner product using the weight w as shown below. 0×0.1+(-1)×0.3+(-2)×0.6=-1.5 Since the absolute value of the calculated value is 0.5 or more, the first estimation unit 406 estimates that the first subject (subject ID=0) is located in front of the second subject (subject ID=1).
[0062] Next, another example of a context estimation method will be described. For example, the first estimation unit 406 may take a moving average of the depth difference value, and estimate the context when the moving average is negative (positive) and the absolute value of the depth difference is equal to or greater than a predetermined value. The depth difference obtained by taking the moving average is expressed by the following equation (1).
[0063]
number
[0064] As described above, when the first estimation unit 406 determines that the subject to be tracked is not occluded and is not close to (overlaps with) other subject candidates in the most recent image, it estimates the front-to-back relationship between the first subject and the second subject. Then, the process proceeds to S507.
[0065] Next, the processing flow when the process proceeds to S608 will be described. In S608, the first estimation unit 406 obtains the most recent estimation result of the front-to-back relationship between the first subject and the second subject from the memory 104. The estimation result of the front-to-back relationship is calculated by the processes of S604 to S607 described above and stored in the memory 104. It is assumed that in the image 703 of FIG. 7(c), subject candidate 720 and subject candidate 721 are determined to be overlapping. If the subject to be tracked is occluded in this frame, it becomes difficult to estimate the front-to-back relationship, and therefore the estimation result of the front-to-back relationship in the image 702 of the immediately preceding frame is obtained.
[0066] In S609, the first estimation unit 406 performs a first occlusion determination process to determine whether or not the tracking target subject is occluded by another subject for two or more subject candidates that are overlapping in the current frame, using the estimation result of the front-to-back relationship acquired in S608. Thereafter, the process proceeds to S507.
[0067] FIG. 8 shows an example of a flowchart of the first occlusion determination process executed in S609. In S801, if the first estimation unit 406 estimates in the most recent estimation result of the front-to-back relationship between the first subject and the second subject that the first subject is in front of all other subjects other than the tracking target, it proceeds to S802; otherwise, it proceeds to S803. In S802, the first estimation unit 406 turns on the front flag.
[0068] In S803, if the first estimation unit 406 estimates in the most recent estimation result of the front-to-back relationship between the first subject and the second subject that the first subject is located further back than one or more other subjects other than the tracking target, the process proceeds to S804. Otherwise, the process proceeds to S805. In S804, the first estimation unit 406 turns on the occlusion flag. In S805, the first estimation unit 406 turns on an unknown flag indicating whether or not the object is occluded.
[0069] 7(c), if the estimation of the front-to-back relationship in the immediately preceding frame image 702 indicates that the first subject 714 is in front of the second subject 715, as described above, the foreground flag is turned ON. As described above, the first occlusion determination process of S609 is performed.
[0070] Returning to the explanation of S507 in FIG. In S507, the determination unit 407 determines the tracking target. If it is determined in S603 that the tracking target subject is not occluded and does not overlap with other subject candidates in the most recent image, the tracking target has already been determined in S604. On the other hand, when the first occlusion determination process of S609 is executed and the foreground flag is ON, the determination unit 407 determines the matched subject candidate as the tracking target by calculating the matching cost using the image features, as in S604.
[0071] On the other hand, when the first occlusion determination process of S609 is executed and the occlusion flag is ON, it can be determined that the subject to be tracked is occluded, and the determination unit 407 performs control so as not to set the occluded area as the tracking target. In other words, the determination unit 407 does not determine the tracking target from the subject candidates detected in S504. As a result, the system control unit 102 performs control so as not to focus on a similar subject that is occluding the tracking target. Note that the processing performed when the first occlusion determination process of S609 is executed and the unknown flag is ON will be described in embodiment 2.
[0072] FIG. 9 shows a different example from FIG. 7 , depicting a scene in which a tracking target subject moves from left to right near a similar subject. In FIG. 9 , the tracking target subject passes behind another subject. The upper row of FIG. 9 shows consecutive frames. The lower row of FIG. 9 shows depth information corresponding to the frame in the upper row. Assume that a subject region 912 is set as a tracking target in image 901 of FIG. 9( a). First subjects 910, 914, 918, and 922 are subjects of the same person detected as subject candidates in S504. Here, the first subject is the subject to be tracked. Second subjects 911, 915, 919, and 923 are subjects of the same person detected as subject candidates in S504, and are located in front of the first subject in each frame image. Assume that in S603, it is determined that subject candidate 920 and subject candidate 921 overlap in image 903 of FIG. 9( c). In this case, in S608, an estimation result of the front-to-back relationship between the first subject 914 and the second subject 915 in the immediately preceding frame image 902 is acquired. Here, it is assumed that the first subject is estimated to be located further back than the second subject.
[0073] If it is estimated that the first subject is located on the far side, the occlusion flag is turned ON by the processes of S803 and S804. If the occlusion flag is ON, the system control unit 102 controls the lens drive so as not to focus on the subject candidate 921. This makes it possible to continue focusing on the first subject 922 even if the first subject 922 appears in the image 904 of the next frame, crossing the second subject 923 on the far side.
[0074] As described above, when the subject to be tracked is occluded, tracking is temporarily suspended, but when the overlap with the second subject 923 is released, as in the case of the first subject 922 in the image 904, the system control unit 102 performs processing to resume tracking. Specifically, tracking is resumed by setting a subject candidate that is close to the occluded subject candidate and has a defocus amount close to 0 (for example, the set value of the defocus amount is within a predetermined value) as the tracking target.
[0075] According to this embodiment, in a scene where a tracking target subject moves near a similar subject, erroneous tracking can be suppressed by estimating the positional relationship in the depth direction between the tracking target subject and the other subjects. In a tracking method using image features, when a similar subject passes in front of the tracking target subject, the focus may be focused on an obstructing object in front. However, by using the depth distance of the subjects arranged in time series, such erroneous tracking can be suppressed.
[0076] <Embodiment 2> In the second embodiment, a method for estimating the occlusion state of a tracking target will be described even when the depth difference between the subjects is small and it is unclear whether the tracking target subject is occluded by another subject, such as when the unknown flag is turned on in the first occlusion determination process of S609. Note that a description of the content that overlaps with the first embodiment will be omitted.
[0077] 10 shows an example of the functional configuration of the image capture device 10 according to this embodiment. The image capture device 10 has a second estimation unit 1001 and a determination unit 1002 in addition to the functional units shown in FIG.
[0078] The second estimation unit 1001 uses the image features of the subject candidate detected by the detection unit 404 to estimate an occlusion state that indicates whether the subject to be tracked is occluded by another subject. The determination unit 1002 uses the estimation result by the first estimation unit 406 and the estimation result by the second estimation unit 1001 to determine whether the subject to be tracked is occluded by another subject. Furthermore, the determining unit 407 according to this embodiment determines a tracking target from the subject candidates detected by the detecting unit 404 based on the estimation results of the first estimating unit 406 and the estimation results of the second estimating unit 1001 .
[0079] Fig. 11 is a flowchart showing the tracking process of the imaging device 10 according to this embodiment. It differs from the flowchart in Fig. 5 in that the processes of S1101 to S1103 are executed instead of the process of S507. The following description will focus on the processes of S1101 to S1103.
[0080] In S1101, the second estimation unit 1001 performs an obstructed state estimation process using image features. Note that the obstructed state estimation process of S1101 may be executed when the unknown flag is set ON in the obstructed state estimation process of S506, and may not be executed when either the foreground flag or the obstructed flag is set ON in the obstructed state estimation process of S506. When either the foreground flag or the obstructed flag is set ON in the obstructed state estimation process of S506, the system control unit 102 performs the same process as in the first embodiment instead of the processes of S1101 to S1103. Specifically, when the foreground flag is ON, the system control unit 102 determines the matched subject candidate as a tracking target and tracks it, and when the obstructed flag is ON, the system control unit 102 controls so that the detected subject candidate is not set as a tracking target.
[0081] FIG. 12 is a flowchart showing the occlusion state estimation process of S1101. FIG. 13 shows an example of a scene in which a tracking target subject moves from left to right near a similar subject. The upper part of FIG. 13 shows successive frames. The lower part of FIG. 13 shows depth information corresponding to the frame in the upper part. In image 1301 of FIG. 13(a), it is assumed that subject area 1312 is set as the tracking target. First subjects 1310, 1314, and 1322 are subjects of the same person who were detected as subject candidates in S504. Second subjects 1311, 1315, 1319, and 1323 are subjects of the same person who were detected as subject candidates in S504.
[0082] In S1201, the second estimation unit 1001 acquires subject candidates by the same process as in S601. Here, it is assumed that there are n acquired subject candidates. In the example of Fig. 13, n = 2 in images 1301, 1302, and 1304, and n = 1 in image 1303, where the first subject 1318 is not detected as a subject candidate due to the influence of the first subject 1318 being occluded by the second subject 1319.
[0083] In S1202, the second estimation unit 1001 extracts subject candidates near the tracking target by the same process as in S602. Note that if the tracking target subject cannot be detected because it is occluded, for example, the second estimation unit 1001 extracts subject candidates using coordinates in the most recent frame in which the subject was detected.
[0084] In S1203, the second estimation unit 1001 performs matching by calculating a matching cost using image features. Specifically, the second estimation unit 1001 acquires image features (templates) of each of the subject candidates detected in each image of the most recent N frames, performs a correlation calculation with the subject candidates in the image of the current frame, and calculates a matching cost for each template. Note that the calculation of the matching cost may involve correction processing that takes into account the coordinates and size of each of the subject candidates in the image of the most recent frame.
[0085] The second estimation unit 1001 links the image features with the subject candidates and tracks each subject by searching for a pair of the subject candidates and matching costs that minimizes the sum of the matching costs of the n subject candidates, for example. If there is no subject corresponding to the image feature, the second estimation unit 1001 sets the value of the matching cost of the image feature to a value greater than the threshold set in S1205.
[0086] The processing from S1204 to S1210 is processing for determining, for each image feature acquired in each image of the most recent N frames, whether or not the subject corresponding to that image feature is occluded in the current frame. In S1205, the second estimation unit 1001 determines whether the matching cost of the subject candidate associated with the image feature is equal to or less than a threshold value, and if it is equal to or less than the threshold value, proceeds to S1206, and if it is greater than the threshold value, proceeds to S1207.
[0087] In S1206, the second estimation unit 1001 estimates that the tracking of the object corresponding to the image feature has been successful, and sets a non-occluded flag to the object corresponding to the image feature. For example, in an image 1302 of the current frame, matching is performed with object candidates 1316 and 1317 of the image 1302 using the image feature of a first object 1310 in an image 1301 of the immediately preceding frame. As a result, if the matching cost with the object candidate 1316 is equal to or less than a threshold, the second estimation unit 1001 estimates that the object corresponding to the image feature of the first object 1310 is not occluded, and sets a non-occluded flag for the first object 1310 to ON.
[0088] In S1207, the second estimation unit 1001 acquires information on whether the subject associated with the image feature was close to (overlapped with) another subject in the recent past. If the subject was close to (overlapped with) another subject in the recent past, the process proceeds to S1208, and if not, the process proceeds to S1209.
[0089] In S1208, the second estimation unit 1001 estimates that the subject corresponding to the image feature has been lost and is occluded by another subject that was nearby in the past, and sets an occlusion flag to the subject corresponding to the image feature. For example, assume that no subject candidate matching the image feature of the first subject 1314 in the image 1302 of the immediately preceding frame is found in the image 1303 of the current frame. In this case, because the first subject 1314 is nearby the second subject 1315 in the image 1302, the second estimation unit 1001 estimates that the subject corresponding to the image feature of the first subject 1314 is occluded, and sets the occlusion flag for the first subject 1314 to ON.
[0090] In S1209, the second estimation unit 1001 estimates that the subject corresponding to the image feature has been lost, and assigns a lost flag to the subject candidate associated with the image feature. In S1211, the second estimation unit 1001 extracts and stores image features for subject candidates that were not linked during the matching in S1203 as new subject candidates that did not exist in the past frames.
[0091] Through the occlusion state estimation process shown in the flowchart of FIG. 12 as described above, the second estimation unit 1001 uses image features to estimate an occlusion state that indicates whether the subject to be tracked is occluded by another subject.
[0092] Returning to the flowchart of FIG. In S1102, the determination unit 1002 performs a second occlusion determination process to determine whether the subject to be tracked is in an occluded, unoccluded, or lost state, using the occlusion state estimation result from S506 and the occlusion state estimation result from S1101.
[0093] FIG. 14 shows an example of a flowchart of the second occlusion determination process executed in S1102. In S1401, the determination unit 1002 determines whether or not the near flag is ON based on the processing result of the first occlusion determination processing. If the near flag is ON, the process proceeds to S1402, and if not, the process proceeds to S1403. In S1402, the determination unit 1002 determines that the subject to be tracked is not occluded.
[0094] In S1403, the determination unit 1002 determines whether or not the unknown flag is ON based on the result of the first occlusion determination process. If the unknown flag is ON, the process proceeds to S1405, and if not, the process proceeds to S1404. In S1404, the determining unit 1002 determines that the subject to be tracked is located behind the other subject and is therefore occluded.
[0095] If the process proceeds to S1405, that is, if it is difficult to determine whether the subject to be tracked is occluded from the depth information, the determination unit 1002 determines whether the non-occluded flag is ON in S1206 from the result of estimating the occluded state in S1101. If the non-occluded flag is ON for the subject to be tracked (first subject), the process proceeds to S1406, and if not, the process proceeds to S1407. In S1406, the determination unit 1002 determines that the subject to be tracked is not occluded.
[0096] In S1407, the determination unit 1002 determines whether or not an occlusion flag has been set for the subject to be tracked, based on the result of the occlusion state estimation in S1101. If the occlusion flag for the subject to be tracked (first subject) is ON, the process proceeds to S1408, and if not, the process proceeds to S1409. In S1408, the determining unit 1002 determines that the subject to be tracked is located behind the other subjects and is therefore occluded. In S1409, the determination unit 1002 determines that the tracking target has been lost.
[0097] When the second occlusion determination process is executed in S1102, the process proceeds to S1103. In S1103, the determination unit 407 determines a tracking target. If it is determined in the second occlusion determination process of S1102 that the tracking target subject is not occluded, the determination unit 407 determines a subject candidate corresponding to the image features of the tracking target subject as the tracking target. Furthermore, if it is determined in the second occlusion determination process of S1102 that the tracking target is occluded, the determination unit 407 performs control so as not to set the detected subject candidate as the tracking target.
[0098] According to this embodiment, in a scene in which a tracking target subject moves near other similar subjects, erroneous tracking can be suppressed by estimating the positional relationship in the depth direction between the tracking target subject and the other subjects.In addition to the effects of the first embodiment, when it is difficult to identify the difference in distance in the depth direction between subjects, the tracking accuracy can be improved by complementing it with a tracking method using image features.
[0099] <Embodiment 3> In the third embodiment, a method of using information about the operating state of the imaging device 10 together with depth information at the time of imaging will be described. In a single-lens reflex camera, which is an example of the imaging device 10, a sudden change occurs in the operating state of the imaging device 10 when the photographer performs zooming, focusing, framing, etc. during shooting. Such a sudden change in the operating state may affect the accuracy of the depth information obtained from the imaging device 10. Therefore, in this embodiment, a method of performing tracking processing using information about the operating state of the imaging device 10 will be described.
[0100] In this embodiment, the driving state of the focus lens of the photographing lens 201 will be described as an example of the operating state of the imaging device 10. FIG. 15 is a graph showing the time series change in the position of the focus lens of the photographing lens 201 in the optical axis direction. The horizontal axis represents time, and the vertical axis represents the position of the focus lens. Assume that the imaging device 10 starts capturing an image of a subject in tracking mode from time t0. Between time t0 and t1, the position of the focus lens remains almost unchanged, and the lens driving amount is small. Between time t1 and t2, the focus lens position changes significantly, and the lens driving amount is large. As mentioned above, the defocus amount is an amount based on the deviation on the imaging plane (imaging plane 300), and is therefore a value relative to the lens position. For this reason, it is not possible to simply compare the defocus amount at time t2 with the defocus amount before and after a large movement of the focus lens, such as the defocus amount at time t1.
[0101] Therefore, the system control unit 102 may acquire lens drive information related to the drive amount of the focus lens of the photographing lens 201 from the lens control unit 205, and correct the weight w for the chronological relationship estimation process described in the first embodiment according to the lens drive information. In this case, for example, when the lens drive amount is large, the first estimation unit 406 can reduce the influence of unreliable depth information by setting the weight of the frame to a small value.
[0102] For example, when the most recent three frames are used, the weight w is set to w=[0.1, 0.3, 0.6]. In this case, if the lens drive amount is equal to or less than a predetermined value, the first estimation unit 406 may estimate the anteroposterior relationship by taking a weighted average (inner product) of the time-series list of depth differences and the weight w. Here, if the lens drive amount in the previous frame during tracking is greater than a predetermined value, the first estimation unit 406 may set the second value in the list of weight w to be smaller than the original value. This enables more stable anteroposterior relationship estimation. Note that when setting the weight w to be small, it may be multiplied by a predetermined coefficient, or the weight may be set to be smaller as the lens drive amount increases.
[0103] Furthermore, the first estimation unit 406 may adjust the number of elements (number of frames) in the time-series list of defocus amounts in the front-to-back relationship estimation process described in the first embodiment, depending on the lens drive amount. For example, consider a case where the front-to-back relationship between a first subject and a second subject is determined such that the first subject (tracking target) is located at the back if the signs of the depth differences in the most recent M consecutive frames are all positive, and the first subject (tracking target) is located at the front if the signs are all negative. In this case, for example, if the lens drive amount is large, the number of frames M to be referenced is increased. This allows a longer time interval to be used to determine whether the subject is at the back or the front, thereby reducing the risk of erroneous tracking due to incorrect depth information.
[0104] According to the third embodiment described above, even if the operating state of the imaging device changes significantly between frames, stable tracking can be achieved by reducing the influence of unreliable depth information.
[0105] Furthermore, although the present invention has been described in detail based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.
[0106] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0107] The disclosure of each of the above-described embodiments includes the following configurations, methods, and programs. (Configuration 1) an acquisition unit that acquires images captured in time series and depth direction distance information for a plurality of regions of the images; a detection means for detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation means for estimating an occlusion state, which indicates whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image based on time-series data of the distance information; a determining means for determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the result of the occlusion state estimation; An information processing device comprising: (Configuration 2) The information processing device described in Configuration 1 is characterized in that the estimation means estimates a front-to-back relationship in the depth direction between the first subject, which is a subject to be tracked, and the second subject, based on time-series data of the distance information for each of the candidate areas associated with the first subject, which is a subject to be tracked, and the candidate areas associated with a second subject different from the first subject, and determines the occlusion state of the first subject based on the estimation result of the front-to-back relationship. (Configuration 3) the acquiring means acquires frames of a moving image in sequence and acquires the distance information for each frame; the detection means detects the candidate region for each frame; 3. The information processing device according to configuration 2, wherein the estimation means estimates the occlusion state for each frame. (Configuration 4) The information processing device according to configuration 3, wherein the estimation means determines the occlusion state of the current frame based on the estimation result of the context when the occlusion state of the immediately preceding frame indicates occlusion. (Configuration 5) The estimation means determining the occlusion state of the current frame based on the estimation result of the context when at least a part of the candidate region in the immediately preceding or current frame overlaps with another candidate region in the frame; When the candidate area in the previous or current frame does not overlap with other candidate areas in the frame, the candidate area associated with the first subject is determined from the candidate areas in the current frame by performing matching using image features of the candidate areas associated with the first subject and the second subject in the previous frame. 5. The information processing device according to configuration 3 or 4. (Configuration 6) The information processing device described in any one of configurations 2 to 5, wherein the estimation means estimates the context based on the difference in distance information for the candidate areas associated with the first subject and the second subject in a plurality of frames arranged in chronological order including the immediately preceding frame. (Configuration 7) 7. The information processing device according to configuration 6, wherein the estimation means estimates the context when the difference in the distance information has the same sign for a predetermined number of consecutive frames. (Configuration 8) 8. The information processing device according to configuration 6 or 7, wherein the estimation means estimates the context when an average absolute value of the difference in distance information over a predetermined number of frames is equal to or greater than a predetermined value. (Configuration 9) 9. The information processing device according to any one of configurations 1 to 8, wherein the estimation means weights the time series data of the distance information so that the weight decreases with increasing time. (Configuration 10) 7. The information processing device according to configuration 6, wherein the estimation means calculates a moving average of the difference in the distance information. (Configuration 11) The information processing device according to any one of configurations 2 to 10, wherein the estimation means extracts a plurality of candidate areas that overlap with the candidate area associated with the first subject, and determines the occlusion state of the first subject by estimating the front-to-back relationship for each combination of the candidate area associated with the first subject and the extracted plurality of candidate areas. (Configuration 12) The information processing device according to configuration 3, wherein the estimation means performs matching on the candidate area in the current frame using image features of the candidate area associated with the first subject in the immediately preceding frame to determine an occlusion state of the first subject. (Configuration 13) 13. The information processing device according to configuration 12, wherein the estimation means determines that the first subject is not occluded when a matching cost obtained as a result of matching is equal to or less than a threshold value. (Configuration 14) The information processing device according to configuration 12 or 13, characterized in that the estimation means determines that the first subject is occluded when a matching cost obtained as a result of matching is greater than a threshold value and when another candidate area exists in the vicinity of the candidate area associated with the first subject. (Configuration 15) The estimation means If the estimation of the front-to-back relationship can be performed based on the time-series data of the distance information, an occlusion state of the first subject is determined based on the estimation result of the front-to-back relationship; If the anteroposterior relationship cannot be estimated based on the time-series data of the distance information, the candidate area in the current frame is matched with the image features of the candidate area associated with the first subject in the immediately preceding frame, and the occlusion state of the first subject is determined. 15. The information processing device according to any one of configurations 12 to 14. (Configuration 16) 16. The information processing device according to any one of configurations 1 to 15, wherein the estimation means corrects the time-series data of the distance information based on the operating state of an imaging device that captured the image. (Configuration 17) 17. The information processing device according to configuration 16, wherein the estimation means acquires a lens drive amount from the imaging device, and corrects the time-series data of the distance information based on the lens drive amount. (Configuration 18) The imaging device further includes a control unit that controls lens driving so as to focus on the candidate area determined by the determination unit, 18. The information processing device according to any one of configurations 1 to 17, wherein the control means controls the device not to focus on the candidate region when the occlusion state indicates that the candidate region is occluded. (Configuration 19) 19. The information processing device according to any one of configurations 1 to 18, wherein the acquisition means acquires, as the distance information, a defocus amount detected from each focus detection area within an imaging surface. (Configuration 20) an acquisition unit that acquires images captured in time series and depth direction distance information for a plurality of regions of the images; a detection means for detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation means for estimating an occlusion state, which indicates whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image based on time-series data of the distance information; a determining means for determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the result of the occlusion state estimation; An imaging device comprising: (method) an acquisition step of acquiring images captured in time series and depth direction distance information for a plurality of regions of the images; a detection step of detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation step of estimating an occlusion state indicating whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image, based on time-series data of the distance information; a determining step of determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the estimation result of the occlusion state; An information processing method comprising: (program) A program for causing a computer to function as each means of the information processing device according to any one of configurations 1 to 19.
Claims
1. an acquisition unit that acquires images captured in time series and depth direction distance information for a plurality of regions of the images; a detection means for detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation means for estimating an occlusion state, which indicates whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image based on time-series data of the distance information; a determining means for determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the result of the occlusion state estimation; An information processing device comprising:
2. The information processing device according to claim 1, characterized in that the estimation means estimates a front-to-back relationship in the depth direction between the first subject, which is a subject to be tracked, and the second subject, based on time-series data of the distance information for each of the candidate areas associated with the first subject, which is a subject to be tracked, and the candidate areas associated with a second subject different from the first subject, and determines the occlusion state of the first subject based on the estimation result of the front-to-back relationship.
3. the acquiring means acquires frames of a moving image in sequence and acquires the distance information for each frame; the detection means detects the candidate region for each frame; The information processing apparatus according to claim 2 , wherein the estimation means estimates the occlusion state for each frame.
4. 4. The information processing apparatus according to claim 3, wherein the estimation means determines the occlusion state of the current frame based on the estimation result of the context when the occlusion state of the immediately preceding frame indicates occlusion.
5. The estimation means determining the occlusion state of the current frame based on the estimation result of the context when at least a part of the candidate region in the immediately preceding or current frame overlaps with another candidate region in the frame; When the candidate area in the previous or current frame does not overlap with other candidate areas in the frame, the candidate area associated with the first subject is determined from the candidate areas in the current frame by performing matching using image features of the candidate areas associated with the first subject and the second subject in the previous frame.
4. The information processing apparatus according to claim 3,
6. 4. The information processing device according to claim 3, wherein the estimation means estimates the context based on a difference in distance information for the candidate areas associated with the first subject and the second subject in a plurality of frames arranged in chronological order, including the immediately preceding frame.
7. 7. The information processing apparatus according to claim 6, wherein the estimation means estimates the context when the difference in the distance information has the same sign for a predetermined number of consecutive frames.
8. 7. The information processing apparatus according to claim 6, wherein the estimation means estimates the context when an average absolute value of the difference in distance information over a predetermined number of frames is equal to or greater than a predetermined value.
9. 2. The information processing apparatus according to claim 1, wherein said estimation means weights the time series data of said distance information so that the weight decreases with increasing time.
10. 7. The information processing apparatus according to claim 6, wherein the estimation means calculates a moving average of the difference in the distance information.
11. 3. The information processing device according to claim 2, wherein the estimation means extracts a plurality of candidate areas that overlap with the candidate area associated with the first subject, and determines the occlusion state of the first subject by estimating the front-to-back relationship for each combination of the candidate area associated with the first subject and the extracted plurality of candidate areas.
12. The information processing device according to claim 3, wherein the estimation means performs matching of the candidate area in the current frame using image features of the candidate area associated with the first subject in the immediately preceding frame to determine the occlusion state of the first subject.
13. 13. The information processing apparatus according to claim 12, wherein the estimation means determines that the first subject is not occluded when a matching cost obtained as a result of the matching is equal to or less than a threshold value.
14. The information processing device according to claim 12, characterized in that the estimation means determines that the first subject is occluded when a matching cost obtained as a result of matching is greater than a threshold value and when another candidate area exists in the vicinity of the candidate area associated with the first subject.
15. The estimation means If the estimation of the front-to-back relationship can be performed based on the time-series data of the distance information, an occlusion state of the first subject is determined based on the estimation result of the front-to-back relationship; If the anteroposterior relationship cannot be estimated based on the time-series data of the distance information, the candidate area in the current frame is matched with the image features of the candidate area associated with the first subject in the immediately preceding frame, and the occlusion state of the first subject is determined.
13. The information processing apparatus according to claim 12.
16. 2. The information processing apparatus according to claim 1, wherein the estimation means corrects the time-series data of the distance information based on an operating state of an imaging device that captured the image.
17. 17. The information processing apparatus according to claim 16, wherein the estimation means acquires a lens drive amount from the imaging device, and corrects the time series data of the distance information based on the lens drive amount.
18. The imaging device further includes a control unit that controls lens driving so as to focus on the candidate area determined by the determination unit, 2. The information processing apparatus according to claim 1, wherein the control means controls the device so as not to focus on the candidate area when the occlusion state indicates that the candidate area is occluded.
19. 2. The information processing apparatus according to claim 1, wherein the acquisition means acquires, as the distance information, a defocus amount detected from each focus detection area within an image pickup surface.
20. an acquisition unit that acquires images captured in time series and depth direction distance information for a plurality of regions of the images; a detection means for detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation means for estimating an occlusion state, which indicates whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image based on time-series data of the distance information; a determining means for determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the result of the occlusion state estimation; An imaging device comprising:
21. an acquisition step of acquiring images captured in time series and depth direction distance information for a plurality of regions of the images; a detection step of detecting a candidate area of a subject to be tracked from the image based on image features of the image; an estimation step of estimating an occlusion state indicating whether or not a subject to be tracked is occluded by another subject different from the tracking target, for the candidate area detected from the image, based on time-series data of the distance information; a determining step of determining the candidate area of the subject to be tracked from the candidate area detected from the image based on the estimation result of the occlusion state; An information processing method comprising:
22. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 19.
Citation Information
Patent Citations
Information processing apparatus, information processing method, and program
JP2022019339A