Main subject determination device, imaging apparatus, and method for controlling main subject determination device

The imaging device accurately tracks the main subject by detecting and correlating multiple parts within an image, using facial recognition to ensure correct focus on the intended subject, addressing the issue of erroneous tracking in multiple-subject scenarios.

JP2025167322APending Publication Date: 2025-11-07CANON KK
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024071824
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing imaging devices struggle to accurately determine and track the main subject when multiple subjects are present, often focusing on the wrong subject due to erroneous tracking of priority parts.

Method used

The imaging device employs a tracking mechanism that detects and tracks both a first part (e.g., face) and a second part (e.g., upper body) from an image, using a reference image of the main subject to accurately identify the main subject by combining these parts, even when they are occluded or their positions change.

Benefits of technology

This approach enhances the accuracy of tracking the intended subject by utilizing facial recognition and comparing detected features with a reference image, ensuring the main subject is correctly identified and maintained in focus, even amidst multiple subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025167322000001_ABST
    Figure 2025167322000001_ABST
Patent Text Reader

Abstract

To more accurately track a subject intended by a user even when a plurality of different subjects are present in a picked-up image.SOLUTION: An imaging apparatus has: tracking means that detects, from a picked-up image, one or more first portions and one or more second portions and tracks the detected portions; acquisition means that acquires, from the one or more first portions and the one or more second portions, the combination of the first portion and the second portion of the same subject; and detection means that, with the use of a reference image of a main subject, detects the first portion of the main subject from the one or more first portions, and detects, from the one or more second portions, the second portion selected as the portion of the same subject as the first portion of the main subject, as the second portion of the main subject.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an imaging device and a control method for an imaging device. [Background technology]

[0002] When detecting and photographing multiple moving subjects in continuous shooting or video shooting, it is desirable for the imaging device to determine a main subject from the multiple subjects and keep the main subject in focus. Patent Document 1 discloses a method of detecting multiple parts from the subject and tracking the subject based on the results of searching for each part. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-106907 Summary of the Invention [Problem to be solved by the invention]

[0004] In Patent Document 1, priority is given to the results of the search process for a priority part among multiple parts, so if the priority part is mistakenly tracked, the imaging device may determine that a subject other than the user's intention is the main subject, and may focus on the erroneously determined main subject.

[0005] The present invention aims to provide a technology for tracking a subject intended by a user with higher accuracy even when a plurality of different subjects exist in a captured image. [Means for solving the problem]

[0006] The imaging device of the present invention is characterized by having a tracking means for detecting and tracking one or more first parts and one or more second parts from an imaged image, an acquisition means for acquiring combinations of first parts and second parts of the same subject from the one or more first parts and the one or more second parts, and a detection means for detecting the first part of the main subject from the one or more first parts using a reference image of the main subject, and detecting the second part of the main subject from the one or more second parts, which is selected as a part of the subject that is the same as the first part of the main subject, as the second part of the main subject. [Effects of the Invention]

[0007] According to the present invention, even when a plurality of different subjects exist in a captured image, it is possible to track a subject intended by a user with higher accuracy. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of an imaging device. [Figure 2] FIG. 2 is a block diagram illustrating a detailed configuration of a subject detection unit. [Figure 3] 10 is a flowchart illustrating an example of a subject detection process. [Figure 4] 10A and 10B are diagrams illustrating the sizes and center of gravity positions of multiple body parts to be detected. [Figure 5] 10A and 10B are diagrams illustrating changes in the center of gravity position of a detected body part between frames. [Figure 6] 10A and 10B are diagrams illustrating a determination as to whether a plurality of parts are parts of the same subject. [Figure 7] FIG. 10 is a diagram illustrating an example of a tracking process. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. An image capturing apparatus according to this embodiment tracks the face and upper body of a main subject, for example, and can acquire more feature amounts than the face. Even if the camera starts to erroneously track the upper body of the main subject, the camera can use the results of facial recognition of the main subject to correctly track the main subject.

[0010] 1 is a block diagram illustrating an example of the configuration of an imaging device 100. The imaging device 100 is a digital still camera or video camera capable of capturing and recording moving and still images. The components within the imaging device 100 are connected to each other via a bus 160 so that they can communicate with each other. The operation of the imaging device 100 is achieved by a control unit 151 (central processing unit) executing a program to control the components.

[0011] The lens unit 101 (photographing lens) has a fixed first group lens 102, a zoom lens 111, an aperture 103, a fixed third group lens 121, a focus lens 131, a zoom motor 112, an aperture motor 104, and a focus motor 132. The photographing optical system of the image capturing device 100 includes the fixed first group lens 102, the zoom lens 111, the aperture 103, the fixed third group lens 121, and the focus lens 131. Note that although the lens included in the photographing optical system is illustrated as a single lens, each may be composed of multiple lenses. The lens unit 101 may be an interchangeable lens unit (interchangeable lens) that is detachable from the image capturing device 100.

[0012] An aperture control unit 105 controls the operation of an aperture motor 104 that drives the aperture 103, and adjusts the amount of light during shooting by adjusting the aperture diameter of the aperture 103. A zoom control unit 113 controls the operation of a zoom motor 112 that drives a zoom lens 111, and changes the focal length (angle of view) of the lens unit 101.

[0013] The focus control unit 133 acquires the defocus amount and defocus direction of the lens unit 101 based on the phase difference between a pair of focus detection signals (image A and image B) obtained from the image sensor 141. The focus control unit 133 determines the drive amount and drive direction of the focus motor 132 based on the defocus amount and defocus direction. The focus control unit 133 controls the operation of the focus motor 132 based on the determined drive amount and drive direction, thereby moving the focus lens 131 and controlling focus adjustment of the lens unit 101. The focus control unit 133 can achieve automatic focus detection (autofocus, AF) using a phase difference detection method by controlling focus adjustment of the lens unit 101. Note that the focus control unit 133 may also achieve AF using a contrast detection method, which controls focus adjustment of the lens unit 101 by moving the focus lens 131 based on a contrast evaluation value of the image signal obtained from the image sensor 141.

[0014] The subject image formed on the imaging plane of the image sensor 141 by the lens unit 101 is converted into an electrical signal (image signal) by a photoelectric conversion element included in each of a plurality of pixels arranged on the image sensor 141. The image sensor 141 has m pixels arranged in the horizontal direction and n pixels arranged in the vertical direction (m and n are natural numbers) in a matrix. Each pixel has two photoelectric conversion elements (photoelectric conversion regions). The control unit 151 can acquire an image on the imaging plane by adding the outputs of the two photoelectric conversion elements. The control unit 151 can also acquire two images with different parallax (parallax images) by separately processing the outputs of the two photoelectric conversion elements.

[0015] The imaging control unit 143 controls the reading of the image signal from the imaging element 141 in accordance with instructions from the control unit 151. The image signal read from the imaging element 141 is supplied to the signal processing unit 142. The signal processing unit 142 applies signal processing such as noise reduction processing, A / D conversion processing, and automatic gain control processing to the image signal and outputs it to the imaging control unit 143. The imaging control unit 143 stores the image signal (image data) received from the signal processing unit 142 in a RAM (random access memory). The data is stored in the memory 154.

[0016] The image processing unit 152 applies predetermined image processing to the image data stored in the RAM 154. The image processing that the image processing unit 152 applies to the image data includes, but is not limited to, development processes such as white balance adjustment, color interpolation (demosaic) and gamma correction, as well as signal format conversion and scaling. The image processing unit 152 can also generate information on subject brightness to be used for automatic exposure control (AE).

[0017] Information regarding the specific subject region may be supplied from the subject detection unit 161 and used, for example, for white balance adjustment processing. When contrast detection AF is performed, the AF evaluation value may be generated by the image processing unit 152. The image processing unit 152 applies image processing to the image data and stores the resulting image data in the RAM 154.

[0018] When the control unit 151 records the image data stored in the RAM 154 on the recording medium 157, the control unit 151 generates a data file according to the recording format by, for example, adding a predetermined header to the image processing data. The control unit 151 may also compress the amount of information by encoding the image data using the compression / decompression unit 153. The control unit 151 records the generated data file on the recording medium 157, such as a memory card.

[0019] When displaying image data stored in RAM 154, control unit 151 causes image processing unit 152 to scale the image data so that it fits the display size of display unit 150. Control unit 151 writes the scaled image data to an area of ​​RAM 154 used as video memory (VRAM area). Display unit 150 reads the image data for display from the VRAM area of ​​RAM 154 and displays it on a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.

[0020] The imaging device 100 can cause the display unit 150 to function as an electronic viewfinder (EVF) by instantly displaying a captured video on the display unit 150 while in a standby state for capturing still images or while recording a video. When the display unit 150 is made to function as an EVF, the video and frame images contained in the video displayed on the display unit 150 are called live view images or through images. When the imaging device 100 captures a still image, it displays the captured still image on the display unit 150 for a certain period of time so that the user can check it. Image display processing on the display unit 150 is realized under the control of the control unit 151.

[0021] The operation unit 156 includes switches, buttons, keys, a touch panel, an eye-gaze input device, and the like, which are used by the user to input instructions to the imaging device 100. The user's instructions input via the operation unit 156 are notified to the control unit 151 via the bus 160. The control unit 151 controls each unit of the imaging device 100 to realize processing according to the user's instructions.

[0022] The control unit 151 has one or more programmable processors such as a CPU, an MPU, etc. The control unit 151 controls each unit and realizes the functions of the imaging device 100, for example, by reading a program stored in a storage unit 155 into a RAM 154 and executing the program.

[0023] The control unit 151 executes AE processing to automatically determine exposure conditions (shutter speed or accumulation time, aperture value, sensitivity) based on information about the subject brightness. The information about the subject brightness can be acquired, for example, from the image processing unit 152. The control unit 151 can also determine the exposure conditions based on a specific area, such as a person's face.

[0024] The control unit 151 controls exposure by adjusting the electronic shutter speed (accumulation time) and the magnitude of the gain. The control unit 151 notifies the imaging control unit 143 of the determined accumulation time and magnitude of the gain. The imaging control unit 143 controls the operation of the imaging element 141 so that an image is captured according to the notified exposure conditions.

[0025] The power management unit 158 ​​manages the battery 159. The battery 159 supplies power to the entire imaging device 100. The storage unit 155 stores programs executed by the control unit 151, setting values ​​used in executing the programs, GUI data, user setting values, etc. For example, when the user operates the operation unit 156 to instruct a transition from a power-off state to a power-on state, the control unit 151 loads the program stored in the storage unit 155 into a part of the RAM 154, executes it, turns on the power of the imaging device 100, and starts processing.

[0026] The subject detection unit 161 detects a subject to be imaged. The subject detection unit 161 has a function of detecting a part of the subject. The subject detection unit 161 also has a function of detecting a subject stored in RAM 154 between consecutively captured images and a function of tracking the detected part of the subject. Furthermore, the subject detection unit 161 has a function of detecting a first part (e.g., face) and a second part (e.g., upper body) of the subject, and determining whether the detected first part and second part are the same subject.

[0027] The continuously captured images include a moving image. Each continuously captured image corresponds to a frame of the moving image. Although the following description will be given of a subject detection process for each frame of the moving image, this embodiment can be applied to continuously captured images.

[0028] For example, when the subject is a person, the subject detection unit 161 can detect body parts such as the face and torso, track the detected face and torso across multiple frames, and determine whether the tracked face and torso are parts of the same person.

[0029] The results of processing by the subject detection unit 161 are stored in RAM 154 and are used for inter-frame tracking processing, automatic setting of the focus detection area, and the like. The processing by the subject detection unit 161 enables the imaging device 100 to realize a tracking AF function for a specific subject. The imaging device 100 can perform AE processing based on luminance information of the focus detection area, and various types of image processing based on pixel values ​​of the focus detection area. The image processing here includes, for example, gamma correction processing and white balance adjustment processing.

[0030] The control unit 151 may superimpose an index indicating the area of ​​the main subject, which is the subject to be tracked, on the display image. The index indicating the area of ​​the main subject is, for example, a rectangular frame surrounding the area of ​​the main subject.

[0031] 2 is a block diagram illustrating a detailed configuration of subject detection unit 161. Subject detection unit 161 includes a face detection unit 201, an upper body detection unit 202, a face tracking unit 203, an upper body tracking unit 204, a same subject determination unit 205, a face authentication unit 206, and a main subject determination unit 207. Each unit of subject detection unit 161 realizes its respective function under the control of control unit 151.

[0032] The subject detection process executed by subject detection unit 161 will be described with reference to Fig. 3. Fig. 3 is a flowchart illustrating the subject detection process. The control unit 151 controls each unit of imaging device 100 and each unit of subject detection unit 161 to realize the process of each step.

[0033] The subject detection process shown in FIG. 3 begins when the power of the imaging device 100 is turned on and a light is displayed on the display unit 150. The process starts when a live view is displayed and a state is reached in which an instruction to start capturing (recording) a still image or a video can be accepted via the operation unit 156. The subject detection process shown in Fig. 3 is a process that is executed for each captured image (each frame of a video image) that is captured continuously.

[0034] 3 illustrates a subject detection process assuming a scene in which the face and upper body of a person who is a main subject are detected and tracked, and the main subject is occluded by another person passing in front of the main subject. When the person to be tracked is occluded by another person passing in front, subject detection unit 161 may erroneously track either the face or the upper body of the other person in front. Note that the subject detection process shown in FIG. 3 can also be applied to other scenes in which erroneous tracking may occur.

[0035] In step S301, the imaging control unit 143 performs imaging processing by controlling the imaging element 141. The signal processing unit 142 acquires a captured image by A / D converting an image signal from the imaging element 141. The captured image here includes frames of a moving image.

[0036] In step S302, the face detection unit 201 detects a face (corresponding to a first part) of a person (corresponding to a subject) from the captured image acquired in step S301. If there are multiple subjects, the face detection unit 201 detects one or more faces. In step S303, the upper body detection unit 202 detects the upper body (corresponding to a second part) of the person from the captured image acquired in step S301. If there are multiple subjects, the upper body detection unit 202 detects one or more upper bodies.

[0037] In steps S302 and S303, the face detection unit 201 and the upper body detection unit 202 can detect the face and the upper body, respectively, using known methods. For example, the face detection unit 201 can detect the face by performing a feature extraction process on the face of a specific target using Convolutional Neural Networks (hereinafter referred to as CNN). Alternatively, the face detection unit 201 may detect the face of the subject by template matching, by registering an image of the face of the subject to be detected in advance as a template. The upper body detection unit 202 can detect the upper body in the same way as the face detection by the face detection unit 201. The method for detecting each part may be different for each part. Alternatively, a combination of multiple detection methods may be used to detect one part.

[0038] In step S304, face tracking unit 203 tracks the face between frames using the face detected in step S302 and face detection results for past frames captured at different times and stored in RAM 154. If there are multiple subjects, face tracking unit 203 tracks one or more faces detected by face detection unit 201. In step S305, upper body tracking unit 204 tracks the upper body between frames using the upper body detected in step S303 and upper body detection results for past frames captured at different times and stored in RAM 154. If there are multiple subjects, upper body tracking unit 204 tracks one or more upper bodies detected by upper body detection unit 202.

[0039] In steps S304 and S305, the face tracking unit 203 and the upper body tracking unit 204 can track the face and the upper body, respectively, using known methods. For example, the face tracking unit 203 can register faces detected in past frames stored in RAM 154 as templates and perform tracking processing by template matching. The face tracking unit 203 can also compare the positions of faces detected in frame images between frames and acquire faces within a predetermined distance range as tracking results. The upper body tracking unit 204 can track the upper body in the same way as the face tracking unit 203 tracks the face. The method for tracking each part may be different for each part. Furthermore, a combination of multiple tracking methods may be used to track one part.

[0040] In the example shown in step S302, a face is detected as the first part, and in step S303, the upper body is detected as the second part, but the parts to be detected are not limited to these parts. Subject detection unit 161 may perform subject detection processing with the face as the first part and the pupils as the second part. Subject detection unit 161 may also perform subject detection processing with the pupils as the first part and the face as the second part.

[0041] The first and second regions are selected so that they differ in at least one of their center of gravity positions and sizes. With reference to Figures 4(A) and 4(B), the reason why regions with different center of gravity positions or sizes are selected as the first and second regions will be described.

[0042] 4(A) and 4(B) are frames captured at different times. As shown in FIG. 4(A), subject detection unit 161 detects, for example, a face 402 and a head 403 of subject 401. The face 402 and the head 403 have similar center-of-gravity positions and sizes. In addition, an obstacle 405 is present above the head of subject 401.

[0043] In FIG. 4(B), face 402 of subject 401 is hidden by obstacle 405. In this case, head 403, which has approximately the same center of gravity and size as face 402, is also hidden by obstacle 405. Therefore, if the first part is the face and the second part is the head, subject detection unit 161 may lose sight of subject 401.

[0044] On the other hand, upper body 404 detected in the frame of Fig. 4(A) is detected as upper body 406 in the frame of Fig. 4(B) even if face 402 is hidden by obstacle 405. By setting the upper body, which differs from the face, which is the first part, in at least one of center of gravity position and size, as the second part, subject detection unit 161 can track upper body 406 in step S305. By detecting and tracking two parts with different center of gravity positions or sizes, subject detection unit 161 can prevent losing sight of both parts in the same frame and can continuously track subject 401.

[0045] The first and second regions are set so that the change in at least one of the center of gravity position and size between frames is minimized. The reason why the first and second regions are set so that the change in the center of gravity position and size between frames is minimized will be explained with reference to Figures 5(A) and 5(B).

[0046] 5(A) and 5(B) are frames captured at different times. In FIGS. 5(A) and 5(B), unlike in FIGS. 4(A) and 4(B), the upper body of subject 501 is detected as an area including the arms.

[0047] 5(A) and 5(B) show a state in which subject 501 is walking while swinging his arms back and forth. Upper body 502 detected in the frame of FIG. 5(A) shows a state in which subject 501's arms are swung out in front. Upper body 503 detected in the frame of FIG. 5(B) shows a state in which subject 501's arms are positioned at the sides of the body. The change in center of gravity position and size between upper bodies 502 and 503 between frames is greater than when the arms are not included in the upper body, and upper body tracking unit 204 may fail to perform the tracking process in step S305.

[0048] Therefore, it is preferable that the part of subject 501 to be detected in steps S302 and S303 is a part where at least one of the center of gravity position and size changes little between frames. By selecting a part where the center of gravity position and size changes little between frames as the tracking target, the tracking performance in steps S304 and S305 is improved.

[0049] In step S306, the same subject determination unit 205 determines whether the same subject is detected in steps S302 and S304. It is determined whether the detected and tracked face and the upper body detected and tracked in steps S303 and S305 are the same part of the subject.

[0050] In steps S302 and S304, one or more faces (first parts) are detected and tracked. In addition, in steps S303 and S305, one or more upper bodies (second parts) are detected and tracked. Same subject determination unit 205 determines whether or not each combination of one or more tracked faces and one or more tracked upper bodies is a part of the same subject. Same subject determination unit 205 acquires a combination of a face and an upper body that is determined to be a part of the same subject as the same subject.

[0051] The processing of step S306 will be specifically described with reference to Fig. 6. A face 603 of subject 601 and a face 604 of subject 602 are the tracking results of the faces detected in step S302 and tracked in step S304. An upper body 605 of subject 601 and an upper body 606 of subject 602 are the tracking results of the upper bodies detected in step S303 and tracked in step S305.

[0052] Because face 603 is included in upper body 605, same subject determination unit 205 determines that face 603 and upper body 605 are parts of the same subject 601. Similarly, because face 604 is included in upper body 606, same subject determination unit 205 determines that face 604 and upper body 606 are parts of the same subject 602. Same subject determination unit 205 acquires the combination of face 603 and upper body 605 as subject 601, and acquires the combination of face 604 and upper body 606 as subject 602. In this way, same subject determination unit 205 can determine that the first part and the second part are parts of the same subject when the face area, which is the first part, is included in the upper body area, which is the second part, in the captured image.

[0053] The same subject determination unit 205 is not limited to determining that a face included in the upper body is the face of the same subject as the upper body, and may determine whether the face and the upper body are parts of the same subject by other methods. For example, the same subject determination unit 205 may determine whether the face and the upper body are parts of the same subject by using a CNN to execute a feature extraction process with the center of gravity of the face and the center of gravity of the upper body as input. That is, the same subject determination unit 205 can acquire combinations of the face and the upper body of the same subject using a trained model that receives the center of gravity of the face and the center of gravity of the upper body as input and is trained to output whether the face and the upper body are parts of the same subject.

[0054] In step S307, face authentication unit 206 performs face authentication processing by comparing the face detected and tracked in steps S302 and S304 with a reference image of a face stored in advance in RAM 154. The reference image of a face stored in RAM 154 is, for example, an image of the face of a main subject set in advance by the user as a tracking target. Using the reference image of the main subject, face authentication unit 206 can detect the face of the main subject from one or more faces detected in step S302.

[0055] If the first part detected in step S302 is the pupil and the second part detected in step S303 is the face, the reference image stored in RAM 154 may be the pupil and image of the main subject.

[0056] The face authentication unit 206 uses, for example, CNN to acquire feature amounts between a reference image of the main subject's face stored in advance in RAM 154 and the face detected and tracked in steps S302 and S304, and calculates the similarity. The face authentication unit 206 can use cosine similarity as the similarity. The cosine similarity takes a real value between -1 and +1, and the closer to 1, the higher the similarity.

[0057] The reference image stored in RAM 154 is, for example, an image registered by the user. The user can select a face image stored in recording medium 157 such as a memory card via operation unit 156 and register it in RAM 154 as a reference image.

[0058] The reference image may also be an image of a face detected as the face of the main subject in the main subject determination process in step S308. Main subject determination unit 207 registers the image of the face detected as the face of the main subject in RAM 154 as a reference image.

[0059] 3, the first part is the face, and in step S307, facial recognition is performed using a reference image of the face. However, the first part is not limited to the face. The first part may be any part that can identify the main subject to be tracked, and may be, for example, the eyes or the upper body wearing a uniform bearing the individual's sports team number. RAM 154 may store, as a reference image, an image of a part corresponding to the first part of the main subject.

[0060] In step S308, main subject determination unit 207 determines, from the subjects acquired in step S306 (combinations of the face and upper body of the same subject), the subject that includes the face authenticated in step S307 as the main subject. That is, main subject determination unit 207 detects, from the one or more upper bodies detected in step S303, the upper body selected as the same subject part as the face of the main subject as the upper body of the main subject.

[0061] The process of detecting a main subject in a scene where two people cross paths will be described with reference to Figures 7(A) to 7(C). Figure 7(A) shows the (n-2)th frame, Figure 7(B) shows the (n-1)th frame, and Figure 7(C) shows the nth frame that is the target of the main subject detection process.

[0062] The main subject in FIG. 7A is subject 701. The main subject may be a subject set in advance by the user as a tracking target, or may be a subject set by imaging device 100 based on the user's gaze information or the like. Face 703 and upper body 705 are parts of subject 701, and are determined to be parts of the same subject in step S306. Similarly, face 704 and upper body 706 are parts of subject 702, and are determined to be parts of the same subject in step S306. Subject 701, the main subject, faces the user who is taking the picture. Subject 702, the sub-subject, is walking in a direction that will cross in front of subject 701.

[0063] 7(B) shows a state in which subject 701 is temporarily hidden behind subject 702 as subject 702 passes in front of subject 701. In FIG. 7(A), face detection unit 201 detects two faces, face 703 and face 704. In contrast, in FIG. 7(B), face detection unit 201 detects only one face, face 707.

[0064] 7(B), the face tracking unit 203 fails to track the face 703 or the face 704 in step S304. When tracking processing is performed using template matching, the face tracking unit 203 detects the sideways face 707 as a tracking result of the sideways face 704. Furthermore, the face tracking unit 203 determines that it has failed (lost) to track the face 703.

[0065] Similarly, for the upper body, upper body tracking section 204 detects upper body 708 facing sideways as a tracking result of upper body 706 facing sideways. Moreover, upper body tracking section 204 determines that it has failed (lost) tracking upper body 705.

[0066] FIG. 7C shows the state after the subject 702 has passed in front of the subject 701. The positional relationship between the subject 701 and the subject 702 in FIG. 7C is the opposite of the positional relationship in FIG. 7A. It is.

[0067] For example, if subject 701 and subject 702 are wearing clothes of the same color and pattern, similar features will be extracted from upper bodies 711 and 712, and therefore upper body tracking unit 204 may perform erroneous tracking using template matching. Specifically, upper body tracking unit 204 may detect upper body 708 of subject 702 in Fig. 7(B) as a tracking result of upper body 705 of subject 701 in Fig. 7(A). Even if upper body tracking unit 204 detects upper body 708 of subject 702 in Fig. 7(B) as a tracking result of upper body 706 of subject 702 in Fig. 7(A), it may detect upper body 711 of subject 701 as a tracking result of upper body 708 in Fig. 7(C).

[0068] Furthermore, for example, if subject 701 and subject 702 are wearing clothes of different colors or patterns, there is a possibility that upper body tracking unit 204 will erroneously track the upper bodies when performing position-based tracking processing. Because the positional relationship between subject 701 and subject 702 is reversed between Figures 7(A) and 7(C), there is a possibility that upper body tracking unit 204 will detect upper body 712 of subject 702 in Figure 7(C) as a tracking result of upper body 705 of subject 701 in Figure 7(A). Similarly, with regard to faces, there is a possibility that face tracking unit 203 will detect face 710 of subject 702 in Figure 7(C) as a tracking result of face 703 of subject 701 in Figure 7(A).

[0069] The process of detecting a main subject by same subject determination unit 205 based on the results of tracking the face and upper body by template matching in steps S304 and S305 will be described. For example, determination (detection) of a main subject when the face is correctly tracked but the upper body tracking fails will be described. A case will be described in which face 703 of subject 701 facing forward in Fig. 7(A) is tracked as face 709 of subject 701 in Fig. 7(C), and upper body 705 of subject 701 in Fig. 7(A) is mistakenly tracked as upper body 712 of subject 702 in Fig. 7(C).

[0070] In step S306, same subject determination unit 205 determines that face 709 and upper body 711 in the frame of Fig. 7(C) are parts of the same subject, and detects the combination of face 709 and upper body 711 as subject 701. Same subject determination unit 205 also determines that face 710 and upper body 712 are parts of the same subject, and detects the combination of face 710 and upper body 712 as subject 702. Therefore, face 709, which is the tracking result of face 703 of subject 701, and upper body 712, which is the tracking result of upper body 705 of subject 701, are determined to be parts of different subjects in Fig. 7(C).

[0071] In step S308, the main subject determination unit 207 acquires the subject 701 including the face 709 and the subject 702 including the upper body 712 in the frame of Fig. 7(C) as candidates for the main subject. The main subject determination unit 207 can determine and detect the main subject based on the result of the face authentication in step S307.

[0072] The main subject determination unit 207 acquires the similarity (authentication score) between the face tracked in step S304 and the reference image of the face of the main subject, and detects the face with the highest similarity among the tracked faces as the face of the main subject.

[0073] In the example of FIG. 7(A), main subject determination unit 207 calculates authentication scores for faces 709 and 710 using a reference image of the face of subject 701, who is the main subject. Main subject determination unit 207 may use an image of face 703 of subject 701 in FIG. 7(A) as the reference image of the main subject. The authentication score of face 709 will be higher than the authentication score of face 710. Therefore, in step S308, main subject determination unit 207 detects face 709, which has a high authentication score, as the part of the main subject. Main subject determination unit 207 calculates authentication scores for faces 709 and 710 using a reference image of the face of subject 701, which is the main subject, in step S306. The upper body 711 acquired as the part of the photographed body can be detected as the upper body of the main subject.

[0074] In the above embodiment, the imaging device 100 detects the face and upper body of each of multiple subjects from a captured image (frame) to be processed. The imaging device 100 tracks the face and upper body of each of the multiple subjects and detects the main subject based on the tracking results. When tracking of different subjects begins for face tracking and upper body tracking, the imaging device 100 performs authentication processing of the tracked face using a reference image of the main subject's face, thereby enabling the imaging device 100 to accurately determine the main subject to be tracked from among the multiple subjects. This enables the imaging device 100 to detect and track the main subject according to the user's intention in a captured image containing multiple subjects.

[0075] In step S307, if the similarity between one or more faces tracked in step S304 and the reference image of the main subject is all lower than a predetermined threshold, main subject determination unit 207 determines the main subject based on the tracking result of the upper body. Specifically, main subject determination unit 207 detects, as the face of the main subject, a face determined to be the same part of the subject as the upper body tracked as the part of the main subject.

[0076] Even if the reference image of the main subject is not registered in RAM 154 or storage unit 155 in step S307, main subject determination unit 207 can determine the main subject based on the tracking result of the upper body.

[0077] Although the above embodiment has been described with reference to an example in which two parts of a subject are tracked, image capture device 100 may be configured to track three or more parts. Image capture device 100 can accurately detect the main subject by performing authentication processing using a reference image of the main subject for any of the three or more parts.

[0078] This embodiment can also be realized in the following way. That is, a system or device includes a storage medium on which software program code describing procedures for realizing the functions of this embodiment is recorded. A computer (or CPU, MPU, etc.) of the system or device reads and executes the program code stored in the storage medium. The program code read from the storage medium can realize the novel functions according to this embodiment, and the storage medium storing the program code and the program constitute the present invention.

[0079] Examples of storage media for supplying the program code include flexible disks, hard disks, optical disks, magneto-optical disks, etc. The storage medium may also be a CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-R, magnetic tape, a non-volatile memory card, a ROM, etc.

[0080] The functions of this embodiment are realized by making the program code executable by a computer. Furthermore, the functions of this embodiment may be realized by an operating system (OS) running on a computer performing some or all of the actual processing based on instructions in the program code.

[0081] This embodiment may be realized by the following method: First, program code is read from a storage medium and written to a memory provided on a function expansion board inserted into a computer or a function expansion unit connected to the computer. A CPU or the like provided on the function expansion board or function expansion unit performs some or all of the actual processing based on the instructions of the program code written in the memory.

[0082] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). The entire device may be controlled by multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) sharing the processing.

[0083] The above processor is a processor in a broad sense, and includes general-purpose processors and dedicated processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Dedicated processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0084] Although the embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0085] <Other embodiments> The present invention can also be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program, or by a circuit that realizes one or more of the functions.

[0086] The disclosure of this embodiment includes the following configuration, method, program, and medium. (Configuration 1) a tracking means for detecting and tracking one or more first regions and one or more second regions from a captured image; an acquisition means for acquiring a combination of a first region and a second region of the same subject from the one or more first regions and the one or more second regions; a detection means for detecting the first part of the main subject from the one or more first parts using a reference image of the main subject, and detecting the second part selected as the same part of the subject as the first part of the main subject from the one or more second parts as the second part of the main subject; An imaging device comprising: (Configuration 2) The first region and the second region are different in at least one of a center of gravity position and a size. 2. The imaging device according to claim 1, (Configuration 3) The first and second regions are set so that a change in at least one of a center of gravity position and a size between the successively captured images is minimized. 3. The imaging device according to configuration 1 or 2. (Configuration 4) The first part is a person's face. 4. The imaging device according to any one of configurations 1 to 3. (Configuration 5) The reference image is an image registered by a user. 5. The imaging device according to any one of configurations 1 to 4. (Configuration 6) The detection means registers an image of the first region detected as the first region of the main subject as the reference image. 5. The imaging device according to any one of configurations 1 to 4. (Configuration 7) The detection means acquires a similarity between each of the one or more first regions and the reference image, and detects the first region having the highest similarity among the one or more first regions as the first region of the main subject. 7. The imaging device according to any one of configurations 1 to 6. (Configuration 8) When the similarity of the one or more first regions is lower than a predetermined threshold, the detection means detecting the second part of the main subject based on a tracking result of the second part; The first region selected as the same region of the subject as the second region of the main subject is detected as the first region of the main subject. 8. The imaging device according to configuration 7. (Configuration 9) When the reference image is not registered, the detection means detecting the second part of the main subject based on a tracking result of the second part; The first region selected as the same region of the subject as the second region of the main subject is detected as the first region of the main subject. 5. The imaging device according to any one of configurations 1 to 4. (Configuration 10) the first portion is a portion included in the second portion, The acquisition means determines that the first part and the second part are parts of the same subject when the area of ​​the first part is included in the area of ​​the second part in the captured image. 10. The imaging device according to any one of configurations 1 to 9, wherein: (Configuration 11) The acquisition means acquires a combination of the first region and the second region of the same subject using a trained model that has been trained to receive the center of gravity position of the first region and the center of gravity position of the second region as input and output whether the first region and the second region are regions of the same subject. 11. The imaging device according to any one of configurations 1 to 10. (method) Detecting and tracking one or more first regions and one or more second regions from a captured image; obtaining a combination of a first region and a second region of the same subject from the one or more first regions and the one or more second regions; detecting the first region of the main subject from the one or more first regions using a reference image of the main subject, and detecting the second region selected as the same region of the main subject as the first region of the main subject from the one or more second regions as the second region of the main subject; 10. A method for controlling an imaging device, comprising: (program) 12. A program for causing a computer to function as each means of the imaging device according to any one of the first to eleventh aspects. (medium) A computer is caused to function as each means of the imaging device according to any one of the first to eleventh aspects. A computer-readable storage medium that stores a program for [Explanation of symbols]

[0087] 100: imaging device, 151: control unit, 161: subject detection unit, 201: face detection unit, 202: upper body detection unit, 203: face tracking unit, 204: upper body tracking unit, 205: same subject determination unit, 207: main subject determination unit

Claims

1. a tracking means for detecting and tracking one or more first regions and one or more second regions from a captured image; an acquisition means for acquiring a combination of a first region and a second region of the same subject from the one or more first regions and the one or more second regions; a detection means for detecting the first part of the main subject from the one or more first parts using a reference image of the main subject, and detecting the second part selected as the same part of the subject as the first part of the main subject from the one or more second parts as the second part of the main subject; An imaging device comprising:

2. The first portion and the second portion are different in at least one of a center of gravity position and a size.

2. The imaging device according to claim 1.

3. The first and second regions are set so that a change in at least one of a center of gravity position and a size between the successively captured images is minimized.

2. The imaging device according to claim 1.

4. The first part is a person's face.

2. The imaging device according to claim 1.

5. The reference image is an image registered by a user.

2. The imaging device according to claim 1.

6. The detection means registers an image of the first region detected as the first region of the main subject as the reference image.

2. The imaging device according to claim 1.

7. The detection means acquires a similarity between each of the one or more first regions and the reference image, and detects the first region having the highest similarity among the one or more first regions as the first region of the main subject.

2. The imaging device according to claim 1.

8. When the similarity of the one or more first regions is lower than a predetermined threshold, the detection means detecting the second part of the main subject based on a tracking result of the second part; The first region selected as the same region of the subject as the second region of the main subject is detected as the first region of the main subject.

8. The imaging device according to claim 7.

9. When the reference image is not registered, the detection means detecting the second part of the main subject based on a tracking result of the second part; The first region selected as the same region of the subject as the second region of the main subject is detected as the first region of the main subject.

2. The imaging device according to claim 1.

10. the first region is a region included in the second region, The acquisition means acquires a region of the first part in the captured image, the region of the second part being included in the captured image. If the first region and the second region are included in the image, it is determined that the first region and the second region are regions of the same subject.

2. The imaging device according to claim 1.

11. The acquisition means acquires a combination of the first region and the second region of the same subject using a trained model that has been trained to receive the center of gravity position of the first region and the center of gravity position of the second region as input and output whether the first region and the second region are regions of the same subject.

2. The imaging device according to claim 1.

12. Detecting and tracking one or more first regions and one or more second regions from the captured image; obtaining a combination of a first region and a second region of the same subject from the one or more first regions and the one or more second regions; detecting the first part of the main subject from the one or more first parts using a reference image of the main subject, and detecting the second part selected as the same part of the subject as the first part of the main subject from the one or more second parts as the second part of the main subject; 10. A method for controlling an imaging device, comprising:

13. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image processing apparatus

    JP2010263550A

  • Image processing device and image processing method

    JP2012065049A

  • Object detection device

    JP2015121905A

  • Image tracking device

    JP2021184564A

  • Image processing device, control method thereof, imaging apparatus, and program

    JP2022128652A