Image capturing apparatus, method for controlling image capturing apparatus, and storage medium

The imaging device uses posture and gaze detection to accurately determine the user's intended main subject, addressing the issue of erroneous subject selection in existing technologies.

JP2026004896APending Publication Date: 2026-01-15CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024102948
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing imaging technologies may erroneously determine a subject not intended by the user as the main subject during continuous shooting or video shooting.

Method used

The imaging device employs a subject detection system that includes posture detection and gaze detection to determine the user's intended main subject, adjusting determination conditions based on the correlation between detected subject posture and gaze position.

Benefits of technology

This approach allows for accurate determination of the user's desired subject as the main subject, enhancing the imaging device's subject selection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004896000001_ABST
    Figure 2026004896000001_ABST
Patent Text Reader

Abstract

To provide a technique capable of appropriately determining a subject desired by a user of an imaging apparatus as a main subject.SOLUTION: The image capturing apparatus includes an object detection unit configured to detect an object, an orientation detection unit configured to detect an orientation of the object, a line-of-sight detection unit configured to detect a line-of-sight position of a user of the image capturing apparatus, a first determination unit configured to determine a candidate for a main object to which the user pays attention based on the orientation of the object detected by the orientation detection unit, and a second determination unit configured to determine the main object based on the candidate determined by the first determination unit and the line-of-sight position of the user detected by the line-of-sight detection unit. And a second determination unit configured to determine the main subject, wherein the second determination unit changes a condition related to the gaze position used for determining the main subject when there is a correlation between the candidate determined by the first determination unit and the gaze position of the user detected by the gaze detection unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an imaging device, a control method for an imaging device, and a program. [Background technology]

[0002] Patent Document 1 discloses a method for detecting multiple moving subjects in continuous shooting or video shooting, in which an imaging device takes consecutive photographs, by detecting the line of sight of the user of the imaging device and setting priorities for multiple areas on the screen using the results of the user's line of sight detection. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-008076 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method disclosed in Patent Document 1, when a different subject is detected, there is a possibility that a subject not intended by the user may be erroneously determined as the main subject.

[0005] The present invention has been made in view of the above, and has an object to provide a technique for appropriately determining a subject desired by a user of an imaging device as a main subject. [Means for solving the problem]

[0006] The imaging device of the present invention comprises a subject detection means for detecting a subject, a posture detection means for detecting the posture of the subject, a gaze detection means for detecting the gaze position of a user of the imaging device, a first determination means for determining a candidate for a main subject that the user is focusing on based on the posture of the subject detected by the posture detection means, and a second determination means for determining the main subject using the candidate determined by the first determination means and the gaze position of the user detected by the gaze detection means, wherein the second determination means changes the conditions related to the gaze position used to determine the main subject when there is a correlation between the candidate determined by the first determination means and the gaze position of the user detected by the gaze detection means.

[0007] In addition, a control method for an imaging device according to the present invention includes a subject detection step for detecting a subject, a posture detection step for detecting the posture of the subject, a gaze detection step for detecting the gaze position of a user of the imaging device, a first determination step for determining a candidate for a main subject that the user is paying attention to based on the posture of the subject detected by the posture detection step, and a second determination step for determining the main subject using the candidate determined by the first determination step and the gaze position of the user detected by the gaze detection step, wherein the second determination step changes conditions related to the gaze position used to determine the main subject when there is a correlation between the candidate determined by the first determination step and the gaze position of the user detected by the gaze detection step. [Effects of the Invention]

[0008] According to the present invention, a subject desired by a user of an imaging device can be appropriately determined as a main subject. [Brief explanation of the drawings]

[0009] [Figure 1]Block diagram showing the configuration of a camera according to a first embodiment. [Figure 2] FIG. 1 is a diagram showing an example of a pixel array in a camera according to a first embodiment; [Figure 3] 1A and 1B are a plan view and a cross-sectional view of a pixel in a camera according to a first embodiment of the present invention; [Figure 4] FIG. 1 is an explanatory diagram of a pixel structure in a camera according to a first embodiment; [Figure 5] FIG. 1 is an explanatory diagram of pupil division in a camera according to a first embodiment; [Figure 6] FIG. 10 is a diagram showing the relationship between the defocus amount and the image shift amount in the camera according to the first embodiment. [Figure 7] FIG. 1 is a diagram showing a focus detection area in a camera according to a first embodiment. [Figure 8] 1 is a flowchart of AF and image capture processing executed by a camera according to a first embodiment. [Figure 9] 1 is a flowchart of a photographing process executed by a camera according to a first embodiment. [Figure 10] Flowchart of subject tracking AF processing executed by the camera according to the first embodiment [Figure 11A] FIG. 1 is an explanatory diagram of posture information according to the first embodiment; [Figure 11B] FIG. 1 is an explanatory diagram of posture information according to the first embodiment; [Figure 12] FIG. 1 is a diagram showing an example of the structure of a neural network according to the first embodiment. [Figure 13] 1 is a flowchart of a detection and tracking process executed by a camera according to a first embodiment. [Figure 14] 1 is a flowchart of an authentication process executed by a camera according to a first embodiment. [Figure 15] 1 is a flowchart of a gaze detection process executed by a camera according to a first embodiment. [Figure 16] Flowchart of main subject determination processing executed by the camera according to the first embodiment [Figure 17A] FIG. 1 is a conceptual diagram illustrating a main subject determination process performed by a camera according to a first embodiment. [Figure 17B] FIG. 1 is a conceptual diagram illustrating a main subject determination process performed by a camera according to a first embodiment. [Figure 17C] FIG. 1 is a conceptual diagram illustrating a main subject determination process performed by a camera according to a first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the embodiments described below are examples of means for realizing the present invention, and may be modified or changed as appropriate depending on the configuration of the device to which the present invention is applied and various conditions. Furthermore, the embodiments may be combined as appropriate.

[0011] <Embodiment 1> FIG. 1 shows the configuration of a camera 100 as an electronic device or an imaging device equipped with an image processing device according to a first embodiment of the present invention. In FIG. 1, a first lens group 101 is arranged closest to the subject (front side) in the imaging optical system as an image forming optical system, and is held so as to be movable in the direction of the optical axis. An aperture 102 adjusts the light intensity by adjusting its aperture diameter. A second lens group 103 moves in the direction of the optical axis together with the aperture 102, and performs magnification change (zooming) together with the first lens group 101, which moves in the direction of the optical axis.

[0012] The third lens group (focus lens) 105 moves in the optical axis direction to adjust the focus. The optical low-pass filter 106 is an optical element for reducing false colors and moiré in the captured image. The first lens group 101, the aperture 102, the second lens group 103, the third lens group 105, and the optical low-pass filter 106 constitute an imaging optical system.

[0013] The zoom actuator 111 rotates a cam barrel (not shown) around the optical axis, and cams provided on the cam barrel move the first lens group 101 and the second lens group 103 in the optical axis direction to change magnification. The diaphragm actuator 112 drives a plurality of light-shielding blades (not shown) in opening and closing directions to adjust the amount of light from the diaphragm 102. The focus actuator 114 moves the third lens group 105 in the optical axis direction to adjust focus.

[0014] A focus driving circuit 126 as a focus adjustment means drives a focus actuator 114 in response to a focus driving command from the camera CPU 121, and moves the third lens group 105 in the optical axis direction. An aperture driving circuit 128 drives an aperture actuator 112 in response to an aperture driving command from the camera CPU 121. A zoom driving circuit 129 drives the aperture actuator 112 in response to a zoom command from the user. The zoom actuator 111 is driven in response to the zoom operation.

[0015] In this embodiment, it is assumed that the imaging optical system, zoom actuator 111, aperture actuator 112, focus actuator 114, focus drive circuit 126, aperture drive circuit 128, and zoom drive circuit 129 are provided integrally with the camera body. However, an interchangeable lens having the imaging optical system, zoom actuator 111, aperture actuator 112, focus actuator 114, focus drive circuit 126, aperture drive circuit 128, and zoom drive circuit 129 may be detachably attached to the camera body.

[0016] The electronic flash 115 has a light-emitting element such as a xenon tube or LED, and emits light to illuminate the subject. The AF (autofocus) assist light emitter 116 has a light-emitting element such as an LED, and projects an image of a mask with a predetermined aperture pattern onto the subject via a projection lens, thereby improving focus detection performance for dark or low-contrast subjects. The electronic flash control circuit 122 controls the electronic flash 115 to turn on in synchronization with the imaging operation. The assist light drive circuit 123 controls the AF assist light emitter 116 to turn on in synchronization with the focus detection operation.

[0017] The camera CPU 121 is responsible for various controls in the camera 100. The camera CPU 121 has a calculation unit, a ROM (Read Only Memory), a RAM (Random Access Memory), an A / D converter, a D / A converter, a communication interface circuit, etc. The camera CPU 121 drives various circuits in the camera 100 in accordance with a computer program stored in the ROM, and controls a series of processes related to captured images, such as AF processing, imaging processing, image processing, and recording processing. In this embodiment, the camera CPU 121 functions as an image processing device.

[0018] The image sensor 107 is composed of a two-dimensional CMOS (Complementary Metal Oxide Semiconductor) photosensor including multiple pixels and its peripheral circuitry, and is disposed on the imaging plane of the imaging optical system. The image sensor 107 photoelectrically converts the subject image formed by the imaging optical system. The image sensor drive circuit 124 controls the operation of the image sensor 107, and also A / D converts the analog signal generated by the photoelectric conversion and transmits the digital signal to the camera CPU 121.

[0019] The shutter 108 has a focal plane shutter configuration, and drives the focal plane shutter in response to a command from a shutter drive circuit built into the shutter 108 based on an instruction from the camera CPU 121. The image sensor 107 is shielded from light while a signal from the image sensor 107 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened and a photographing light beam is guided to the image sensor 107.

[0020] The image processing circuit 125 applies predetermined image processing to image data stored in the RAM in the camera CPU 121. The image processing applied by the image processing circuit 125 includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. Furthermore, the image processing circuit 125 stores the processed image data, the joint positions of each subject, position and size information of specific objects, center of gravity of the subject, position information of the face and eyes, etc. in the RAM in the camera CPU 121. The result of the determination processing may be used for other image processing (for example, white balance adjustment processing).

[0021] The display 131 as a display means includes a display element such as an LCD (Liquid Crystal Display). The display 131 displays information about the image capture mode of the camera 100, a preview image before image capture, a confirmation image after image capture, an indicator for the focus detection area, and a focused image. etc. The display 131 functions as an electronic viewfinder (EVF) and displays an image projected on a small LCD through the viewfinder. The operation switch group 132 includes a main (power) switch, a release (photography trigger) switch, a zoom operation switch, a photography mode selection switch, etc., and is operated by the user. The flash memory 133 records captured images. The flash memory 133 is detachable from the camera 100.

[0022] The subject detection unit 140, which serves as a subject detection means, performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 140 uses dictionary data corresponding to the type of subject to detect multiple types of subjects. The various types of dictionary data are, for example, data in which the characteristics of the corresponding subjects are registered. The subject detection unit 140 performs subject detection by sequentially switching between dictionary data for each subject. In this embodiment, the dictionary data for each subject is stored in the dictionary data storage unit 141. Therefore, multiple dictionary data are stored in the dictionary data storage unit 141. The camera CPU 121 determines which dictionary data from the multiple dictionary data to use for subject detection based on the subject priority set in advance and the settings of the imaging device. Here, subject detection includes detection of people and detection of organs such as the person's face, eyes, and torso. Detection of objects other than people, such as balls, may also be performed.

[0023] The subject detection unit 140, which serves as a subject detection means, may detect a subject by identifying an individual whose face has been registered in advance (personal authentication). For example, assume that the camera 100 is capable of using a face registration mode for subject authentication. In the face registration mode, feature information indicating the feature amounts of the face area of ​​the detected subject is registered in dictionary data. The subject detection unit 140 extracts the facial feature amounts of the person by detecting organs such as the eyes and mouth through organ detection of the person depicted in the captured image. The subject detection unit 140 then calculates the similarity (reliability) with the facial feature amounts registered in advance in dictionary data. The subject detection unit 140 then determines whether the calculated similarity is equal to or greater than a predetermined threshold, thereby determining whether the face of the person present in the captured image is the face of a person registered in the dictionary data. This allows the subject detection unit 140 to perform personal authentication of the subject using the face registration mode.

[0024] The posture acquisition unit 142, which serves as posture detection means, estimates the posture of each of the multiple subjects detected by the subject detection unit 140, and acquires posture information of the subjects. The content of the posture information to be acquired is determined according to the type of subject. In this example, since the subject is a person, the posture acquisition unit 142 acquires the positions of multiple joints of the person as the subject. Any method may be used to estimate the posture, and for example, the method described in the following document 1 may be used. Details of acquiring posture information of the subjects in this embodiment will be described later. (Reference 1) Cao, Zhe, et al., "Realtime multi-person 2d pose estimation using part affinity fields.", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.

[0025] The dictionary data storage unit 141 stores dictionary data for each object. The object detection unit 140 estimates the position of the object in the image based on the captured image data and the dictionary data stored in the dictionary data storage unit 141. The object detection unit 140 may estimate the position, size, reliability, etc. of the object and output the estimated information. The object detection unit 140 may also output other information. Examples of dictionary data used for object detection include dictionary data for detecting "people" as objects, dictionary data for detecting "animals," dictionary data for detecting "vehicles," and dictionary data for detecting objects such as "balls." Furthermore, dictionary data for detecting "the whole of a person" and dictionary data for detecting "parts of a person" such as a person's face are stored separately in the dictionary data storage unit 141. Good too.

[0026] In this embodiment, the object detection unit 140 estimates the position of an object included in image data using a machine-learned convolutional neural network (CNN). The object detection unit 140 may be realized by a circuit specialized for estimation processing using a graphics processing unit (GPU) or a CNN.

[0027] The machine learning of the CNN may be performed by any method. For example, a predetermined computer such as a server may perform the machine learning of the CNN, and the camera 100 may acquire the trained CNN from the predetermined computer. For example, the CNN of the subject detection unit 140 may be trained by the predetermined computer performing supervised learning using training image data as input and the position of the subject corresponding to the training image data as training data. In this way, a trained CNN is generated. The training of the CNN may be performed by the camera 100 or the image processing device described above.

[0028] Gaze detection unit 143, which serves as a gaze detection means, has a sensor for detecting the user's gaze and photoelectrically converts incident reflected infrared light into an electrical signal. Gaze detection unit 143 detects the user's gaze position from the movement of the user's eyeballs based on the output of the sensor, and outputs gaze detection information to camera CPU 121. Since the user's gaze position also depends on how the user looks into the electronic viewfinder, which is display 131, the gaze position may be adjusted in advance by performing gaze position calibration in camera 100.

[0029] Next, the image array of the image sensor 107 will be described with reference to Fig. 2. Fig. 2 schematically shows the pixel array of the image sensor 107, which is in a range of 4 pixel columns x 4 pixel rows, as viewed from the optical axis direction (z direction).

[0030] Each pixel unit 200 includes four imaging pixels arranged in two rows and two columns. Arranging a large number of pixel units 200 on the image sensor 107 enables photoelectric conversion of a two-dimensional subject image. In FIG. 2, an imaging pixel 200R having a spectral sensitivity of R (red) (hereinafter referred to as an R pixel) is arranged at the upper left of each pixel unit 200, and imaging pixels 200G having a spectral sensitivity of G (green) (hereinafter referred to as G pixel) are arranged at the upper right and lower left. Furthermore, an imaging pixel 200B having a spectral sensitivity of B (blue) (hereinafter referred to as a B pixel) is arranged at the lower right of the pixel unit 200. Each imaging pixel includes a first focus detection pixel 201 and a second focus detection pixel 202, which are divided in the horizontal direction (x direction).

[0031] In the image sensor 107 according to this embodiment, for example, the pixel pitch of the imaging pixels is assumed to be 4 μm, the number of imaging pixels is assumed to be 5575 horizontal (x) columns × 3725 vertical (y) rows = approximately 20.75 million pixels, and the pixel pitch of the focus detection pixels is assumed to be 2 μm, the number of focus detection pixels is assumed to be 11150 horizontal columns × 3725 vertical rows = approximately 41.5 million pixels.

[0032] In this embodiment, we will explain the case where each imaging pixel is divided into two horizontally as shown in Figure 2, but it may also be divided vertically. Also, the image sensor 107 according to this embodiment has a plurality of imaging pixels, each of which includes a first focus detection pixel and a second focus detection pixel, but the imaging pixels and the first and second focus detection pixels may be provided as separate pixels. For example, the first and second focus detection pixels may be discretely arranged among the plurality of imaging pixels.

[0033] FIG. 3A shows one imaging pixel (200R, 200G, 200B) as viewed from the light receiving surface side (+z direction) of the image sensor 107. FIG. 3B shows a cross section of the imaging pixel taken along line aa in FIG. 3A as viewed from the -y direction. As shown in FIG. 3B, one imaging pixel has , one microlens 305 is provided to focus the incident light.

[0034] Each imaging pixel is provided with photoelectric conversion units 301 and 302 that are divided into N parts in the x direction (N is a positive integer; in this embodiment, N=2). The photoelectric conversion units 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively. The centers of gravity of the photoelectric conversion units 301 and 302 are decentered to the -x and +x sides, respectively, with respect to the optical axis of the microlens 305.

[0035] A color filter 306 of any one of R, G, and B is provided between the microlens 305 and the photoelectric conversion units 301 and 302 in each imaging pixel. The spectral transmittance of the color filter may be changed for each photoelectric conversion unit, or the color filter may be omitted.

[0036] Light incident on the imaging pixels from the imaging optical system is collected by the microlens 305, dispersed by the color filter 306, and then received by the photoelectric conversion units 301 and 302 and photoelectrically converted.

[0037] Next, the relationship between the pixel structure shown in Figures 3A and 3B and pupil division will be described using Figure 4. Figure 4 shows a cross section of the imaging pixel taken along line aa in Figure 3A as viewed from the +y direction, and also shows the exit pupil of the imaging optical system. In Figure 4, the x and y directions of the imaging pixel are reversed from the x and y directions in Figure 3B, respectively, to correspond to the coordinate axes of the exit pupil.

[0038] The first pupil region 501, whose center of gravity is decentered in the +x direction, of the exit pupil is a region that is made approximately conjugate with the light receiving surface of the photoelectric conversion unit 301 of the imaging pixel in the -x direction by the microlens 305. The light beam that passes through the first pupil region 501 is received by the photoelectric conversion unit 301, i.e., the first focus detection pixel 201. The second pupil region 502, whose center of gravity is decentered in the -x direction, is a region that is made approximately conjugate with the light receiving surface of the photoelectric conversion unit 302 of the imaging pixel in the +x direction by the microlens 305. The light beam that passes through the second pupil region 502 is received by the photoelectric conversion unit 302, i.e., the second focus detection pixel 202. The pupil region 500 represents the pupil region that can receive light across the entire imaging pixel, including all of the photoelectric conversion units 301 and 302 (the first focus detection pixel 201 and the second focus detection pixel 202).

[0039] FIG. 5 shows pupil division by the image sensor 107. A pair of light beams that pass through the first pupil region 501 and the second pupil region 502, respectively, are incident on each pixel of the image sensor 107 at different angles and are received by the two divided first focus detection pixels 201 and second focus detection pixels 202. In this embodiment, the image sensor 107 generates a first focus detection signal using output signals from the multiple first focus detection pixels 201, and generates a second focus detection signal using output signals from the multiple second focus detection pixels 202. The image sensor 107 also generates an imaging pixel signal by adding the output signals from the first focus detection pixels 201 and the output signals from the second focus detection pixels 202 of the multiple imaging pixels. The image sensor 107 then combines the imaging pixel signals from the multiple imaging pixels to generate an imaging signal for generating an image with a resolution corresponding to the number of effective pixels N.

[0040] Next, the relationship between the defocus amount of the imaging optical system and the phase difference (image shift amount) between the first focus detection signal and the second focus detection signal acquired from the image sensor 107 will be described with reference to FIG. 6. The image sensor 107 is disposed on the imaging plane 600 in the figure, and as described with reference to FIGS. 4 and 5, the exit pupil of the imaging optical system is divided into two, a first pupil region 501 and a second pupil region 502. The defocus amount d is expressed as |d| (absolute value of d), where the distance (magnitude) from the imaging position C of the light beam from the subject (801, 802) to the imaging plane 600 is expressed as a negative sign (d<0) when the imaging position C is closer to the subject than the imaging plane 600. Also, the defocus amount d represents a back-focus state where the imaging position C is on the opposite side of the imaging plane 600 from the subject, with a positive sign (d>0). In an in-focus state where the imaging position C is on the imaging plane 600, d=0. The imaging optical system is in-focus (d=0) with respect to the subject 801, and in a front-focus state (d<0) with respect to the subject 802. The front-focus state (d<0) and the back-focus state (d>0) together are called a defocus state (|d|>0).

[0041] In a front-focus state (d<0), the light beam from the subject 802 that passes through the first pupil region 501 is first focused and then spreads over a width Γ1 centered at the center of gravity G1 of the light beam, forming a blurred image on the image pickup surface 600. Similarly, the light beam from the subject 802 that passes through the second pupil region 502 is first focused and then spreads over a width Γ2 centered at the center of gravity G2 of the light beam, forming a blurred image on the image pickup surface 600. This blurred image is received by each of the first focus detection pixels 201 on the image pickup element 107, and a first focus detection signal is generated. Similarly, this blurred image is received by each of the second focus detection pixels 202 on the image pickup element 107, and a second focus detection signal is generated. In other words, the first focus detection signal is a signal representing an image of the subject 802 that is blurred by the blur width Γ1 at the center of gravity G1 of the light beam on the image pickup surface 600. Similarly, the second focus detection signal is a signal that represents an image of the object 802 at the center of gravity G2 of the light beam on the image pickup surface 600, blurred by the blur width Γ2.

[0042] The blur widths Γ1 and Γ2 of the subject image increase roughly in proportion to the increase in the magnitude |d| of the defocus amount d. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 in the center of gravity positions of the light beams) between the first focus detection signal and the second focus detection signal also increases roughly in proportion to the increase in the magnitude |d| of the defocus amount d. In the back-focus state (d>0), the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state, but a similar proportional relationship exists.

[0043] In this way, the amount of image shift between the first focus detection signal and the second focus detection signal increases as the amount of defocus increases. In this embodiment, focus detection is performed using an image plane phase difference detection method, in which the amount of defocus is calculated from the amount of image shift between the first focus detection signal and the second focus detection signal obtained using the image sensor 107.

[0044] Next, the focus detection area of ​​the image sensor 107 that acquires the first focus detection signal and the second focus detection signal will be described with reference to Fig. 7. In Fig. 7, A(n, m) indicates the nth focus detection area in the x direction and the mth focus detection area in the y direction out of the multiple focus detection areas (three in the x direction and three in the y direction, for a total of nine) set in the effective pixel area 1000 of the image sensor 107. The first focus detection signal and the second focus detection signal are generated from the output signals of the multiple first focus detection pixels 201 and second focus detection pixels 202 included in the focus detection area A(n, m), respectively. I(n, m) indicates an index that displays the position of the focus detection area A(n, m) on the display 131.

[0045] The nine focus detection areas shown in FIG. 7 are merely an example, and the number, positions, and sizes of the focus detection areas are not limited. For example, one or more focus detection areas may be set within a predetermined range centered on a position specified by the user or the position of the subject detected by the subject detector. In this embodiment, the focus detection areas are arranged to obtain focus detection results with higher resolution when acquiring a defocus map, which will be described later. For example, a total of 9,600 focus detection areas can be arranged on the image sensor 107, divided into 120 horizontal and 80 vertical areas.

[0046] Next, the details of the processing executed by the camera 100 in this embodiment will be described. Fig. 8 shows a flowchart of the AF / imaging processing (image processing method) that realizes the AF operation and imaging operation executed by the camera 100 of this embodiment. Specifically, Fig. 8 shows the processing from before imaging to display a live view image on the display 131 of the camera 100 to capturing a still image. The camera CPU 121, which is a computer, executes this process in accordance with a computer program. In the following description, "S" means "step."

[0047] First, in S1, the camera CPU 121 drives the image sensor 107 using the image sensor drive circuit 124 and acquires image data from the image sensor 107. From the acquired image data, the camera CPU 121 acquires first and second focus detection signals from a plurality of first and second focus detection pixels included in each focus detection area shown in FIG. 7. The camera CPU 121 also generates an image signal by adding the first and second focus detection signals from all effective pixels of the image sensor 107, and performs image processing on the image signal (image data) using the image processing circuit 125 to acquire image data. Note that if the image sensor 107 has separate image sensors for the image sensors, the first and second focus detection pixels, and the camera CPU 121 acquires image data by performing interpolation processing on the focus detection pixels.

[0048] Next, in S2, the camera CPU 121 generates a live view image from the image data acquired by the image processing circuit 125 in S1, and displays the generated live view image on the display 131. The live view image is a reduced image matched to the resolution of the display 131, and the user can adjust the image capture composition, exposure conditions, etc. while checking the live view image. Therefore, the camera CPU 121 performs exposure adjustment based on the photometric value obtained from the image data, and displays it on the display 131. The camera CPU 121 achieves exposure adjustment by appropriately adjusting the exposure time, opening and closing the aperture of the shooting lens, and adjusting the gain for the image sensor output.

[0049] Next, in S3, the camera CPU 121 determines whether or not a switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing a release switch included in the operation switch group 132. If Sw1 is not turned on (S3: NO), the camera CPU 121 repeats the determination in S3 to monitor the timing at which Sw1 is turned on. On the other hand, if Sw1 is turned on (S3: YES), the camera CPU 121 proceeds to S400 and performs subject tracking autofocus (AF) processing. In S400, the camera CPU 121 uses the obtained imaging signal and focus detection signal to perform subject area detection, focus detection area setting, predictive AF processing to suppress the influence of a time lag between focus detection processing and image capture processing for recording an image, and the like. Details of each process will be described later.

[0050] Next, the camera CPU 121 proceeds to S5 and determines whether or not the switch Sw2, which instructs the start of an imaging operation, has been turned on by fully pressing the release switch. If Sw2 has not been turned on (S5: NO), the camera CPU 121 returns the process to S3. On the other hand, if Sw2 has been turned on (S5: YES), the camera CPU 121 proceeds to S300 and executes an imaging subroutine. Details of the imaging subroutine will be described later. When the imaging subroutine ends, the camera CPU 121 proceeds to S7.

[0051] In S7, the camera CPU 121 determines whether or not the main switch included in the operation switch group 132 has been turned off. If the main switch has been turned off (S7: YES), the camera CPU 121 ends the processing of this flowchart. If the main switch has not been turned off (S7: NO), the camera CPU 121 returns the processing to S3.

[0052] In this embodiment, after it is detected in S3 that Sw1 is turned on, the camera CPU 121 executes the subject detection process and the AF process, but the timing of these processes is not limited to the above. For example, the camera CPU 121 may execute the subject tracking AF process performed in S400 before Sw1 is turned on. This eliminates the need for the user to perform preparatory actions before capturing images in the camera 100.

[0053] Next, the imaging subroutine executed by the camera CPU 121 in S300 of FIG. 8 will be described with reference to the flowchart shown in FIG.

[0054] In S301, the camera CPU 121 executes exposure control processing to determine imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.). The camera CPU 121 can execute exposure control processing using brightness information acquired from image data of a live view image. The camera CPU 121 then transmits the determined aperture value to the aperture drive circuit 128, which then drives the aperture 102. The camera CPU 121 also transmits the determined shutter speed to the shutter 108, which then opens the focal plane shutter. Furthermore, the camera CPU 121 causes the image sensor drive circuit 124 to accumulate charge in the image sensor 107 during the exposure period.

[0055] Next, in S302, the camera CPU 121 causes the image sensor drive circuit 124 to read out all pixels of the image sensor 107 to capture an image signal for capturing a still image. The camera CPU 121 also causes the image sensor drive circuit 124 to read out one of the first focus detection signal and the second focus detection signal from the focus detection area (focus target area) within the image sensor 107. At this time, the first focus detection signal or the second focus detection signal that is read out is used to detect the focus state of the image during image playback, which will be described later. The other focus detection signal can also be obtained by subtracting one of the first focus detection signal and the second focus detection signal from the image capture signal.

[0056] In S303, the camera CPU 121 performs defective pixel interpolation processing on the imaging data read out and A / D converted in S302 using the image processing circuit 125. The defective pixel interpolation processing in this step can be achieved using well-known technology, and therefore a detailed description of the processing will be omitted.

[0057] In S304, the camera CPU 121 performs image processing and encoding processing to convert the captured image data using the image processing circuit 125. Here, the image processing and encoding processing include demosaic (color interpolation), white balance processing, gamma correction (tone correction), color conversion, and edge enhancement processing for the captured image data after the interpolation processing for defective pixels.

[0058] In S305, the camera CPU 121 records the still image data obtained as the image data by the image processing and encoding processing in S304 and one of the focus detection signals read out in S302 in the flash memory 133 as an image data file.

[0059] In S306, the camera CPU 121 associates the camera characteristic information as characteristic information of the camera 100 with the still image data recorded in S305, and records it in the flash memory 133 and the memory within the camera CPU 121. Here, the camera characteristic information includes, for example, the following information: Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 125 Information about the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107 Information about vignetting of the imaging light beam within the camera 100 Information on the distance from the mounting surface of the imaging optical system of the camera 100 to the imaging element 107 Information about the manufacturing tolerances of the Camera 100

[0060] Information about the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light-receiving sensitivity distribution information) is the sensitivity of the image sensor 107 according to the distance (position) on the optical axis from the image sensor 107. The light sensitivity distribution information may include information about the microlens 305 and the photoelectric conversion units 301 and 302 because the light sensitivity distribution depends on the performance of the microlens 305 and the photoelectric conversion units 301 and 302. The light sensitivity distribution information may also include information about changes in sensitivity with respect to the angle of incidence of light.

[0061] In S307, the camera CPU 121 associates lens characteristic information, which is characteristic information of the imaging optical system, with the still image data recorded in S305 and records it in the flash memory 133 and the memory in the camera CPU 121. Here, the lens characteristic information includes, for example, information about the exit pupil, information about a frame such as a lens barrel that blocks light beams, information about the focal length and F-number at the time of imaging, and information about aberrations of the imaging optical system. Furthermore, the lens characteristic information includes information about manufacturing errors in the imaging optical system and information about the position of the third lens group 105 at the time of imaging (subject distance).

[0062] In S308, the camera CPU 121 records image-related information, which is information related to the still image data, in the flash memory 133 and in a memory within the camera CPU 121. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.

[0063] In S309, the camera CPU 121 displays a preview of the captured image on the display 131. This allows the user to easily check the captured image on the camera 100. When the processing of S309 ends, the camera CPU 121 ends this imaging subroutine and proceeds to S7 in FIG. 8.

[0064] Next, the subroutine of the subject tracking AF process executed by the camera CPU 121 in S400 of FIG. 8 will be described with reference to the flowchart shown in FIG.

[0065] In S401, the camera CPU 121 calculates the amount of image shift between the first focus detection signal and the second focus detection signal obtained in each of the multiple focus detection areas acquired in S2, and calculates the defocus amount for each focus detection area from the image shift amount. As described above, in this embodiment, the group of focus detection results obtained from the focus detection areas in the image sensor 107, which are divided into 120 horizontally and 80 vertically, for a total of 9,600 points, is called a defocus map. In this step, the camera CPU 121 acquires the defocus map.

[0066] In S402, camera CPU 121 performs subject detection processing and tracking processing. The subject detection processing in this step is performed by the above-mentioned subject detection unit 140. In some cases, subject detection is impossible depending on the state of the obtained image. In such cases, subject detection unit 140 performs tracking processing using other means such as template matching to estimate the position of the subject. Details of the subject detection processing and tracking processing will be described later.

[0067] In S403, the camera CPU 121 performs authentication processing of the subject by determining whether the face of the subject depicted in the captured image is the face of a person registered in the dictionary data. The authentication processing in this step is performed by the subject detection unit 140 described above. In this embodiment, the authenticated subject is determined as the main subject in priority over subjects that are candidates for the main subject determined using posture information. For this reason, the authentication processing (S403) is performed before the process of determining candidates for the main subject using posture information (S405), which will be described later.

[0068] In S404, posture information of each subject detected by the subject detection unit 140 is acquired from the joint positions of the subject. FIGS. 11A and 11B are diagrams schematically showing an example of information acquired by the posture acquisition unit 142. FIG. 11A shows the posture of the subject depicted in the captured image to be processed. 11A and 11B show subjects 901 and 902 and a ball 903 that have been caught. Subject 901 is catching ball 903. Subject 901 is an important subject in the captured scene in FIGS. 11A and 11B. In this embodiment, subject detection unit 140 uses posture information of the subjects to determine subjects that are likely to be the subject that the user intends to focus on as candidates for the main subject. On the other hand, subject 902 is a non-main subject. Here, a non-main subject refers to a subject other than the candidate main subject that is likely to be the subject that the user intends to focus on.

[0069] FIG. 11B is a diagram showing an example of the joint positions of subjects 901 and 902, and the position and size of ball 903. In FIG. 11B, each joint 911 represents a joint of subject 901, and each joint 912 represents a joint of subject 902. In FIG. 11B, the positions of the top of the head, neck, shoulders, elbows, wrists, waist, knees, and ankles are shown as joint positions, but the joint positions may be some of these or other positions. Furthermore, information such as axes connecting joints may be used in addition to joint positions, and any information representing the posture of the subject may be used to acquire posture information. The following describes a case where posture information is acquired using joint positions.

[0070] Posture acquisition unit 142 acquires two-dimensional coordinates (x, y) of joints 911 of subject 901 and joints 912 of subject 902 in the image. Here, the unit of two-dimensional coordinates (x, y) is pixels. Center of gravity position 413 indicates the center of gravity position of ball 903, and arrow 914 indicates the size of ball 903 in the image. Subject detection unit 140 acquires the two-dimensional coordinates (x, y) of the center of gravity position of ball 903 in the image and the number of pixels indicating the width of ball 903 in the image.

[0071] In S405, the subject detection unit 140 performs a process of determining candidates for the main subject using the posture information. Here, the subject detection unit 140 is a first determination means for determining candidates for the main subject that the user focuses on based on the posture of the detected subject. First, a method for calculating the reliability (probability) of a subject, which indicates its likelihood of being the main subject, will be described. A case will be described in which the probability that the subject is the main subject of the image is used as the reliability indicating its likelihood of being the main subject (reliability corresponding to the degree of possibility that the subject is the main subject in the captured image to be processed), but a value other than probability may also be used for the reliability. For example, the reciprocal of the distance between the center of gravity of the subject and the center of gravity of an object related to the subject may be used as the reliability.

[0072] <How to calculate probability> We will explain a method for calculating the probability that an object is likely to be the main subject based on the coordinates and size of each joint of the object. Below, we will explain the case where a neural network, which is a machine learning method, is used.

[0073] FIG. 12 is a diagram showing an example of the structure of a neural network. In FIG. 12, 1001 is an input layer, 1002 is a hidden layer, 1003 is an output layer, 1004 is a neuron, 1005 is a neuron, and the connections between the neurons 1004 are shown. For convenience of illustration, only representative neurons and connecting lines are numbered here. The number of neurons 1004 in the input layer 1001 is equal to the dimension of the input data, and the number of neurons in the output layer 1003 is two. This corresponds to a two-class classification problem of determining whether or not an object is likely to be the main subject.

[0074] The line 1005 connecting the i-th neuron 1004 in the input layer 1001 and the j-th neuron 1004 in the hidden layer 1002 has a weight w ij In addition, the value z j is given by the following equations (1) and (2).

number

number

[0075] In formula (1), x i represents the value input to the i-th neuron 1004 of the input layer 1001. The sum is taken over all neurons 1004 of the input layer 1001 that are connected to the j-th neuron. j is called the bias, and is a parameter that controls the likelihood of firing the j-th neuron 1004. The function h defined in equation (2) is an activation function called ReLU (Rectified Linear Unit). It is also possible to use other functions as the activation function, such as a sigmoid function.

[0076] Also, the value y k is given by the following formula:

number

number

[0077] In equation (3), z j represents the value output by the j-th neuron 1004 in the hidden layer 1002, where i, k = 0, 1. 0 corresponds to non-main subject, and 1 corresponds to likelihood of being the main subject. The sum is taken for all neurons in the hidden layer 1002 that are connected to the k-th neuron. The function f defined by equation (4) is called a softmax function, and outputs a probability value that belongs to the k-th class. In this embodiment, f(y1) is used as the probability that an object is likely to be the main subject.

[0078] During training, the coordinates of the person's joints and the coordinates and size of the ball are input to the input layer. Then, all weights and biases are optimized to minimize the loss function that uses the output probability and correct label. Here, the correct label takes two values: "1" if the subject is likely to be the main subject, and "0" if the subject is not the main subject. The loss function L can be the binary cross-entropy shown below.

number

[0079] In equation (5), the subscript m represents the index of the object to be learned. m is the probability value output from the neuron 504 with k=1 in the output layer 1003, and t m is the correct label. In addition to Equation (5), the loss function can also be used to calculate the correct label, such as the mean square error. Any function may be used as long as it can calculate the degree of match with the label. By optimizing based on equation (5), it is possible to determine the weights and biases so that the correct label and the output probability value are closer to each other. The trained weights and bias values ​​are saved in advance in flash memory 133 and are stored in RAM in camera CPU 121 as needed. Multiple types of weights and bias values ​​may be prepared depending on the scene. The trained weights and biases (results of machine learning performed in advance) are used to output the probability value f(y1) based on equations (1) to (4).

[0080] During learning, a state before an important action can be learned as a state highly likely to be the main subject. For example, when a subject throws a ball, the state in which the subject extends his / her hand as he / she throws the ball can be learned as one state highly likely to be the main subject. The reason for adopting this configuration is that it is necessary to accurately control the imaging device when the subject, who is actually the main subject, performs an important action. For example, in camera 100, if the reliability (probability value) indicating the likelihood of the subject being the main subject exceeds a preset threshold, automatic control to record images or videos (recording control) can be initiated, allowing the user to capture important moments without missing them. In this case, information on the state transition of the subject being learned may be used to control the imaging process of camera 100 to obtain information on the typical time until the important action.

[0081] Although a method for calculating probability using a neural network has been described above, other machine learning techniques such as support vector machines and decision trees may be used as long as they can classify an object as likely to be the main subject. Furthermore, other than machine learning, a function that outputs a reliability or probability value based on a certain model may also be employed. It is also possible to use the value of a monotonically decreasing function of the distance between the object and the ball, assuming that the closer the distance between the object and the ball, the greater the reliability indicating the object's likelihood to be the main subject.

[0082] Furthermore, although ball information is also used to calculate the reliability indicating the likelihood of being the main subject, it is also possible to determine candidates for the main subject using only the posture information of the subject. Depending on the type of posture information of the subject (e.g., pass, shoot, etc.), it may be better to use ball information as well, or not. For example, when the subject shoots, the distance between the person and the ball becomes greater, but the user may intend for the subject who took the shot to be the main subject. Therefore, in this case, candidates for the main subject may be determined based only on the posture information of the person who is the subject, regardless of the ball. However, even in this case, ball information may also be used to determine candidates for the main subject depending on the type of posture information of the subject.

[0083] Alternatively, data obtained by performing a predetermined transformation, such as a linear transformation, on the coordinates of each joint of the subject and the coordinates and size of the ball may be used as input data. Furthermore, if the candidate for the main subject frequently switches between two subjects with a defocus difference, the reliability of a subject other than the user's intended main subject may increase. Therefore, for example, if camera 100 detects frequent switching between two subjects based on the time-series data of the reliability of each subject, camera 100 may prevent the switching of the candidate for the main subject by increasing the reliability of one of the subjects (e.g., the subject closest to the subject). Alternatively, camera 100 may determine an area including the two subjects as the area representing the candidate for the main subject.

[0084] As another method, time series data of posture information of the person, the positions of the person and the ball, the defocus amount of each subject, and reliability indicating the likelihood of being the main subject may be used as input data. As another method, the above prediction process may be performed, and the reliability may be calculated using predicted data of the coordinates of the person's joints and the coordinates and size of the ball at the time of exposure of the captured image as input data. Also, the image plane movement speed of the subject and the time series data of the coordinates of each joint position may be used. Depending on the amount of change in the series, it may be possible to switch whether or not to use the data that has undergone prediction processing. This allows the accuracy of the reliability indicating the likelihood of being the main subject to be maintained when the change in the subject's posture is small, and by using the results of the prediction processing when the change in the subject's posture is large, it is possible to detect subjects that are likely to be the main subject at an earlier timing. In the camera 100 according to this embodiment, it is possible to determine candidates for the main subject by calculating the reliability of each of multiple subjects using the above method.

[0085] In S406, the camera CPU 121 performs gaze detection processing. Details of the gaze detection processing in this step will be described later with reference to FIG.

[0086] In S407, camera CPU 121 performs a main subject determination process. Here, camera CPU 121 is a second determination means that determines a main subject using the candidate determined by the first determination means and the user's gaze position detected by the gaze detection means. Specifically, camera CPU 121 performs the main subject determination process using the defocus map acquired in S401, the subject detection result in S402, the authentication result in S403, the main subject candidate determination result in S405, and the gaze detection process result in S406. In this manner, camera CPU 121 determines a main subject from among multiple detected subjects. In this embodiment, the main subject refers to an object to be focused on. In addition, the main subject in this embodiment includes an object on which the user intends to focus and an object (e.g., a ball) to be focused on when there is no object on which the user intends to focus. Details of the main subject determination process in this step will be described later.

[0087] In S408, camera CPU 121 sets a focus detection area using the subject detection area of ​​the subject that is a candidate for the main subject determined in S405. Camera CPU 121 can select a focus detection result that is highly reliable and indicates a subject that is relatively close, based on the results of the focus detection area within the area set as the subject detection area. Furthermore, camera CPU 121 may set the focus detection area by relocating a focus detection area within the area set as the obtained subject detection area, acquiring image data and focus detection signals again, and similarly selecting a focus detection result.

[0088] In S409, the camera CPU 121 acquires the focus detection result of the set focus detection area. Here, the focus detection result can be acquired by selecting a focus detection result that is close to the desired area from the focus detection results included in the defocus map acquired in S401, or by calculating the defocus amount using a focus detection signal corresponding to the set focus detection area. Furthermore, the focus detection area used to calculate the defocus amount is not limited to one, and multiple areas may be arranged around one focus detection area, and the defocus amount may be calculated using these multiple areas.

[0089] In S410, the camera CPU 121 performs predictive AF processing using the defocus amount calculated in S406 and a plurality of defocus amounts that are time-series data of the timings at which focus detection was performed in the past.

[0090] Predictive AF processing is processing that is executed when there is a time lag between the timing of focus detection and the timing of exposure of the captured image. Predictive AF processing is processing that predicts the position of the subject in the optical axis direction at the timing of exposure of the captured image, which is a predetermined time after the timing of focus detection, and performs AF control. Specifically, in the processing of predicting the image plane position of the subject, the camera CPU 121 performs multivariate analysis (for example, the least squares method) using history data of the image plane position of the subject in the past and time, and finds an equation for a prediction curve. Then, the camera CPU 121 substitutes the time of exposure of the captured image into the equation for the found prediction curve, thereby finding the position of the subject. The predicted image plane position is calculated.

[0091] Furthermore, in the predictive AF process, not only the optical axis direction but also three-dimensional position may be predicted. For example, consider a case where the screen is an XY plane, the optical axis direction is the Z direction, and vectors in the XYZ directions are defined in three-dimensional space. In this case, the position of the subject at the timing of exposure for the captured image may be predicted using time-series data of the position in the Z direction from the XY position of the subject obtained in the subject detection and tracking process in S402 and the defocus amount obtained in S407. Furthermore, the position of the subject may be predicted from time-series data of the joint positions of the person who is the subject.

[0092] By predicting the position of the subject as described above, it is possible to estimate the position of each subject even if the ball or person is hidden during shooting, or if some of the person's joint positions become invisible. The subject to be predicted is not only the main subject, but also multiple detected subjects. In this embodiment, predictive AF processing is performed on the person and objects other than the person, such as the ball. Furthermore, the predictive AF processing on the person may use the focus detection results corresponding to each joint position acquired in S404. As described above, it is possible to predict the positions of multiple subjects and the joint positions of the subject.

[0093] In S411, the camera CPU 121 calculates the drive amount of the focus lens using the main subject candidate determination result in S405, the defocus amount obtained in S407, and the predicted AF processing result in S408. Then, the camera CPU 121 performs focus adjustment processing by driving the focus actuator 114 and moving the third lens group 105 in the optical axis direction. Details of the focus adjustment processing in this step will be described later. When the processing of S411 ends, the camera CPU 121 ends the subject tracking AF processing subroutine and proceeds to S5 in FIG. 8.

[0094] Next, the subroutine of the subject detection and tracking process executed by the camera CPU 121 in S402 of FIG. 10 will be described with reference to the flowchart shown in FIG.

[0095] In S2000, the camera CPU 121 sets dictionary data according to the type of subject to be detected from the data detected from the image data acquired in S2. The camera CPU 121 selects dictionary data to be used in this process from multiple dictionary data stored in the dictionary data storage unit 141 based on the preset subject priority and the settings of the imaging device. For example, dictionary data in which subjects are classified, such as "people," "vehicles," "animals," and "balls," is stored. In this embodiment, one or more dictionary data may be selected. When one dictionary data is selected, subjects detectable by one dictionary data are frequently and repeatedly used to detect the subject. On the other hand, when multiple dictionary data are selected, the dictionary data are set sequentially according to the priority of the subject to be detected, so that subjects corresponding to the dictionary data can be detected sequentially.

[0096] In S2001, subject detection unit 140 performs subject detection using the dictionary data set in step S2000, with the image data read in step S2 as the input image. In this embodiment, as an example, dictionary data for "person" and "ball" is set in S2000. At this time, subject detection unit 140 outputs information such as the position, size, and reliability of the detected subject. At this time, camera CPU 121 may display the above information output by subject detection unit 140 on display 131.

[0097] In S2001, the camera CPU 121 detects a plurality of hierarchical regions for one subject from the image data. For example, if a "person" is set as dictionary data, the camera CPU 121 detects a plurality of hierarchical regions for the person's organs, for example, a "face" region, an "eye" region, and so on. The camera CPU 121 detects areas classified based on organs, such as the "whole body" area, and the "eyes" area. While areas related to local organs, such as a person's eyes or face, are areas where the focus and exposure are adjusted as the subject, there is a possibility that these areas may not be detected depending on surrounding obstacles or the direction of the face. Even in such cases, the camera CPU 121 can continue to robustly detect the subject by detecting the whole body area. In this way, in this embodiment, the camera CPU 121 detects the subject using hierarchical areas.

[0098] In S2001, the camera CPU 121 may detect a ball in addition to a person. For example, dictionary data for "ball" is stored in the dictionary data storage unit 141, and after detecting a person, the camera CPU 121 changes the dictionary data to dictionary data for "ball" and detects the ball. Here, ball detection using the dictionary data for "ball" is detection of the entire area of ​​the ball, and the camera CPU 121 outputs the center position and size of the detected ball. In this embodiment, it is assumed that the dictionary data for "ball" is set as the dictionary data, but dictionary data for "object detection" for detecting objects in general, not just balls, may also be used as the dictionary data. In this case, the camera CPU 121 can also detect a ball as a type of object using the dictionary data for "object detection."

[0099] Any object detection method may be used in the object detection in this step. For example, the method described in the following document 2 may be used. In this embodiment, the subject to be detected is a ball, but other specific objects such as a racket may also be used as the detection target. (Reference 2) Redmon, Joseph, et al., "You only look once: Unified, real-time object detection.", Proceedings of the IEEE conference on computer vision and pattern recognition, 2016.

[0100] Next, in S2002, the camera CPU 121 performs a known template matching process using the subject detection area obtained in S2001 as a template to perform subject tracking. Using the multiple images obtained in S2, the camera CPU 121 searches for a similar area in the most recently obtained image using the subject detection area obtained in a previous image as a template. Examples of information used for template matching include brightness information, color histogram information, and feature point information such as the corners and edges of the subject. Any known matching method and template update method may be used. Note that the tracking process performed in S2002 is a process for more reliably performing the subject detection and tracking process of this embodiment by detecting a subject from the most recently obtained image data in an area similar to the previous subject detection data if the subject was not detected in S2001. After completing the process in S2002, the camera CPU 121 ends the subject detection and tracking process subroutine and proceeds to S403 in FIG. 10 .

[0101] Next, the subroutine of the authentication process executed by the camera CPU 121 in S405 of FIG. 10 will be described with reference to the flowchart shown in FIG.

[0102] In S3001, camera CPU 121 sets a priority of the subject for each person in dictionary data in which the faces of multiple people are registered. Since multiple people are registered in the dictionary data, when multiple subjects are recognized using the faces of different people, a priority is set indicating which subject should be prioritized as a candidate for the main subject. Note that camera CPU 121 may automatically change the priority depending on the size or position of the person's face in the captured image, depending on the settings of camera 100, or may set the priority based on a priority set in advance by the user.

[0103] In S3002, the camera CPU 121 extracts features. The CPU 121 extracts facial features of a person by detecting organs such as the eyes and mouth of the person present in the captured image.

[0104] In S3003, the camera CPU 121 calculates the similarity. Specifically, the camera CPU 121 calculates the similarity (second reliability) between the feature amount extracted in S3002 and the feature amount of a face registered in advance in dictionary data by pattern matching or the like.

[0105] In S3004, camera CPU 121 determines candidates for the main subject for the authenticated subject. Specifically, camera CPU 121 determines whether the similarity (second reliability) calculated in S3003 is equal to or greater than a second threshold. Then, camera CPU 121 determines the subject whose similarity is equal to or greater than the second threshold as a candidate for the main subject.

[0106] In S3005, camera CPU 121 determines whether the authentication process is complete. Specifically, camera CPU 121 determines whether the authentication process is complete based on whether a candidate for the main subject has been determined in S3004. If a candidate for the main subject has been determined in S3004, camera CPU 121 determines that the authentication process is complete (S3005: YES), ends the subroutine for the authentication process, and proceeds to S404 in Fig. 10. On the other hand, if a candidate for the main subject has not been determined in S3004, camera CPU 121 determines that the authentication process is not complete (S3005: NO), and returns the process to S3002.

[0107] Next, the subroutine of the line-of-sight detection process executed by the camera CPU 121 in S406 of FIG. 10 will be described with reference to the flowchart shown in FIG.

[0108] In S4001, the camera CPU 121 determines whether or not the user's gaze can be detected. In this step, the camera CPU 121 determines whether or not the user's gaze can be detected, for example, when the user is not looking at the display 131. If the camera CPU 121 determines that the user's gaze can be detected (S4001: YES), the process proceeds to S4002. If the camera CPU 121 determines that the user's gaze cannot be detected (S4001: NO), the camera CPU 121 ends this subroutine.

[0109] In S4002, the camera CPU 121 sets gaze detection conditions, which include a gaze detection time and a number of gaze detection data used to calculate a gaze gaze position, which will be described later.

[0110] In S4003, the camera CPU 121 performs gaze detection. Specifically, the camera CPU 121 records the gaze position of the user at each time as time-series data.

[0111] In S4004, camera CPU 121 determines whether the gaze detection position overlaps with a subject that is a candidate for the main subject determined based on the posture information in S405. If camera CPU 121 determines that the gaze detection position overlaps with a subject that is a candidate for the main subject (S4004: YES), it proceeds to S4005. If camera CPU 121 determines that the gaze detection position does not overlap with a subject that is a candidate for the main subject (S4004: NO), it proceeds to S4006.

[0112] In S4005, the camera CPU 121 changes the gaze detection conditions. By lengthening the gaze detection time included in the gaze detection conditions, the number of time-series data of the gaze detection position increases, and therefore the gaze gaze position described below can be calculated more accurately. However, since the posture of the subject that is a candidate for the main subject determined by the posture information changes over time, the longer the gaze detection time, the greater the posture change, and the candidate for the main subject may be changed to another subject. There is a possibility that this may occur. Therefore, camera CPU 121 sets the gaze detection time of the gaze detection conditions to be shorter than the subject detection time by subject detection unit 140 (for example, the subject detection time using the dictionary data described above or the subject detection time using authentication processing). Instead of or in addition to this, camera CPU 121 changes the gaze detection conditions so as to reduce the amount of time-series data on the gaze detection position. As described above, even when determining candidates for the main subject using posture information that is prone to change over time, by appropriately setting the gaze detection conditions, it is possible to determine an appropriate subject as the main subject using the gaze detection results.

[0113] In S4006, the camera CPU 121 calculates the gaze gaze position. Here, the gaze gaze position is a position obtained by averaging the gaze detection position data of the user acquired in time series. Alternatively, the gaze gaze position may be a position obtained by averaging the gaze detection position data by weighting it in time series (the more recent the time, the heavier the weight of the gaze position data). Furthermore, if the gaze detection position continues to change directionally in time series, an approximation curve may be calculated from the trajectory of the time change of the gaze detection position using the least squares method or the like, and the gaze gaze position at the next time may be calculated using the gaze gaze position vector. As a result, the camera CPU 121 ends the gaze detection processing subroutine and proceeds to S407 in FIG. 10.

[0114] Next, a subroutine of the main subject determination process executed by camera CPU 121 in S407 of Fig. 10 will be described using the flowchart shown in Fig. 16. In this embodiment, there are two methods for detecting a subject: a detection method using a subject detection means (subject detection using dictionary data in S402 or subject detection using authentication processing in S403) and a detection method for determining main subject candidates using posture information. In this case, how the main subject is determined in the main subject determination process in this embodiment will be described below.

[0115] In S5001, the camera CPU 121 acquires gaze gaze information. Specifically, the camera CPU 121 acquires information on the gaze gaze position and / or gaze gaze vector calculated in S4006 in the gaze detection processing of Fig. 15 as gaze gaze information.

[0116] In S5002, the camera CPU 121 determines whether the gaze gaze position identified using the gaze gaze information acquired in S5001 overlaps with a subject that is a candidate for the main subject. If the camera CPU 121 determines that the gaze gaze position overlaps with a subject that is a candidate for the main subject (S5002: YES), the process proceeds to S5005. If the camera CPU 121 determines that the gaze gaze position does not overlap with a subject that is a candidate for the main subject (S5002: NO), the process proceeds to S5003.

[0117] A specific example of the determination process in S5002 will now be described with reference to Fig. 17A. Fig. 17A shows a basketball shooting scene, with a person shooting and a person blocking the shot depicted in the captured image. In Fig. 17A, person 926 is attempting to shoot, person 927 is attempting to block the shot of person 926, and person 926 is attempting to shoot ball 923 into goal 940. In Fig. 17A, people 926 and 927, ball 923, and goal 940 are examples of subjects in the captured image.

[0118] 17A also shows a user's gaze gaze position 1901 calculated by the above-described gaze detection process. Note that the gaze gaze position 1901 does not have to be displayed on the display 131. Person 926 is a subject that is a candidate for the main subject determined by posture information, and person 927 is a subject detected by subject detection means. In FIG. 17A, gaze gaze position 1901 is located at a position overlapping with the face of person 926, so the camera CPU 121 determines that gaze gaze position 1901 overlaps with a subject that is a candidate for the main subject.

[0119] In S5003, the camera CPU 121 determines whether or not a subject that is a candidate for the main subject exists near the gaze position in the captured image. If the camera CPU 121 determines that a subject that is a candidate for the main subject exists near the gaze position in the captured image (S5003: YES), the process proceeds to S5005. If the camera CPU 121 determines that a subject that is a candidate for the main subject does not exist near the gaze position in the captured image (S5003: NO), the process proceeds to S5004.

[0120] Here, a specific example of the determination process in S5003 will be described with reference to FIG. 17B. The captured image shown in FIG. 17B is the same as that in FIG. 17A, depicting a basketball shooting scene. Therefore, the same reference numerals are used to designate the same objects as those in FIG. 17A, and detailed description thereof will be omitted. However, unlike FIG. 17A, in FIG. 17B, the user's gaze position 1901 does not overlap with the person 926 but is located near the person 926. Therefore, in this step, the camera CPU 121 determines that a subject who is a candidate for the main subject is present near the gaze position. Note that the range near the gaze position where it is determined that a subject who is a candidate for the main subject exists in S5003 can be determined using the detection error of the gaze detection position and / or the standard deviation of the time-series data of the gaze detection position. Alternatively, the range near the gaze position where it is determined that a subject who is a candidate for the main subject exists may be determined using other well-known techniques.

[0121] In S5004, the camera CPU 121 determines whether or not a subject that is a candidate for the main subject exists in the moving direction of the gaze position. If the camera CPU 121 determines that a subject that is a candidate for the main subject exists in the moving direction of the gaze position (S5004: YES), the process proceeds to S5005. If the camera CPU 121 determines that a subject that is a candidate for the main subject does not exist in the moving direction of the gaze position (S5004: NO), the process proceeds to S5006.

[0122] Here, a specific example of the determination process in S5004 will be described with reference to FIG. 17C. The captured image shown in FIG. 17B is the same as that in FIG. 17A, showing a basketball shooting scene, and therefore the same objects as those in FIG. 17A are assigned the same reference numerals, and detailed description thereof will be omitted. However, in FIG. 17C, the user's gaze position moves over time toward person 926 in the order of positions 1901a, 1901b, and 1901c. Furthermore, arrow 1902 is a gaze gaze vector calculated by the camera CPU 121 from the time-series gaze gaze positions 1901a, 1901b, and 1901c. Note that the number of gaze gaze positions and the movement time of the gaze gaze positions used by the camera CPU 121 to calculate the gaze gaze vector may be set as appropriate.

[0123] In this step, the camera CPU 121 determines whether or not the gaze gaze vector calculated based on the movement of the gaze gaze position 1901 will abut on a subject that is a candidate for the main subject when extended. Then, based on this determination, the camera CPU 121 determines whether or not a subject that is a candidate for the main subject exists in the movement direction of the gaze gaze position. In the case of FIG. 17C, when the gaze gaze vector 1902 is extended, it abuts on the person 926 who is a candidate for the main subject. Therefore, the camera CPU 121 determines that a subject that is a candidate for the main subject exists in the movement direction of the gaze gaze position.

[0124] Through the above steps S5002, S5003, and S5004, the camera CPU 121 can identify the presence of a subject who is a candidate for the main subject using the gaze position and / or gaze vector. Therefore, in step S5005, the camera CPU 121 determines that there is a correlation between the gaze position and / or gaze vector and the subject who is a candidate for the main subject. Then, the camera CPU 121 determines that the subject who is a candidate for the main subject, in the example of FIGS. 17A to 17C, is the main subject.

[0125] Furthermore, the camera CPU 121 also performs steps S5002, S5003, and S5004. If the presence of a subject who is a candidate for the main subject cannot be identified using the gaze position and / or gaze vector, the process proceeds to S5006. Therefore, in S5006, the camera CPU 121 determines that there is no correlation between the gaze position and / or gaze vector and the subject who is a candidate for the main subject. Then, the camera CPU 121 determines, as the main subject, a subject selected from the subjects detected by the subject detection method without using the gaze information acquired in S5001. Here, the method of determining the main subject may be a method that determines, as the main subject, a subject selected using various conditions, such as the reliability of each subject detection method, a subject closer to the center of the captured image (screen), or a subject that is larger in the captured image (screen).

[0126] Alternatively, the camera CPU 121 may set the subject detected by the subject detection method stored in S5007, which will be described below, as the main subject.

[0127] In S5007, when the camera CPU 121 determines the main subject based on the gaze information through the above process, it records the subject detection method used to detect the main subject. This allows the camera CPU 121 to prioritize the subject detection method that detects a subject that closely matches the gaze position for the same continuously captured scene, thereby determining the main subject as a subject that best matches the user's intention. In this step, the camera CPU 121 records the subject detection method in the flash memory 133 for the same continuously captured scene, and erases the recorded subject detection method when continuous shooting is stopped in the camera 100. When the process of S5007 is completed, the camera CPU 121 ends the gaze detection process subroutine and proceeds to S408 in FIG. 10.

[0128] In this embodiment, person 927 is a person who is registered in advance in the dictionary data. Therefore, person 927 is a subject detected by the authentication process. Furthermore, person 926's shooting posture is identified by the posture information, and the first reliability indicating the likelihood of the person being the main subject is high. As a result, the first reliability is equal to or greater than the first threshold, and the person is determined to be a subject who is a candidate for the main subject. Furthermore, person 927's blocking posture is identified by the posture information, and the first reliability indicating the likelihood of the person being the main subject is high. As a result, the first reliability is equal to or greater than the first threshold, and the person is determined to be a subject who is a candidate for the main subject. That is, in the process of determining a candidate for the main subject using posture information, both person 926 and person 927 are determined to be subjects who are candidates for the main subject. However, in this embodiment, person 927, who is determined to be a subject who is a candidate for the main subject by posture information and who is detected by the authentication process, is determined to be the main subject with priority over person 926, who is not detected by the authentication process.

[0129] While the present invention has been described in detail above based on preferred embodiments, it is not limited to these specific embodiments, and various modifications within the scope of the present invention are also encompassed within the scope of the present invention. Parts of the above-described embodiments may be combined as appropriate. Furthermore, the present invention also encompasses a case in which a software program implementing the functions of the above-described embodiments is supplied to a system or device having a computer capable of executing the program, either directly from a storage medium or via wired or wireless communication, and the program is executed. Therefore, the program code itself supplied to and installed on a computer to implement the functional processing of the present invention also embodies the present invention. In other words, the computer program itself for implementing the functional processing of the present invention is also encompassed within the present invention. In this case, the form of the program is not important, as long as it has the program functionality, such as object code, a program executed by an interpreter, or script data supplied to an OS. Examples of storage media for providing the program include magnetic storage media such as hard disks and magnetic tapes, optical / magneto-optical storage media, and non-volatile semiconductor memory. Furthermore, a method for providing the program includes storing the computer program implementing the present invention on a server on a computer network and transmitting the program to a connected server. Another possible method is for a client computer to download and program a computer program.

[0130] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). The entire device may be controlled by multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) sharing the processing.

[0131] The above processor is a processor in the broad sense, and includes general-purpose processors and dedicated processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Dedicated processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0132] Although the embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0133] (Other embodiments) The present invention can also be realized by a process in which a program that realizes one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in the computer of the system or device read and execute the program, or by a circuit that realizes one or more functions.

[0134] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) An imaging device, a subject detection means for detecting a subject; a posture detection means for detecting the posture of the subject; a gaze detection means for detecting a gaze position of a user of the imaging device; a first determination means for determining a candidate for a main subject that the user focuses on based on the posture of the subject detected by the posture detection means; a second determination means for determining the main subject using the candidate determined by the first determination means and the gaze position of the user detected by the gaze detection means; and The second determining means changes a condition related to the gaze position used to determine the main subject when there is a correlation between the candidate determined by the first determining means and the gaze position of the user detected by the gaze detecting means. An imaging device characterized by: (Configuration 2) 2. The imaging device according to configuration 1, wherein the conditions include a condition regarding a time period for detecting the gaze position by the gaze detection means. (Configuration 3) 3. The imaging device according to configuration 2, wherein the second determining means changes the condition so that the detection time is shortened when the correlation is present. (Configuration 4) The imaging device according to configuration 3, wherein the second determining means changes the condition so that the detection time is shorter than the time it takes for the subject to be detected by the subject detecting means. (Configuration 5) 2. The imaging device according to configuration 1, wherein the conditions include conditions related to time-series data used for detecting the gaze position by the gaze detection means. (Configuration 6) 6. The imaging device according to configuration 5, wherein the second determining means changes the condition so that the amount of time-series data decreases when the correlation exists. (Configuration 7) The imaging device according to any one of configurations 1 to 6, wherein the second determination means determines the candidate as the main subject depending on the degree of coincidence between the gaze position of the user and the position of the candidate determined by the first determination means. (Configuration 8) The imaging device according to any one of configurations 1 to 6, characterized in that the second determination means determines the candidate determined by the first determination means as the main subject when the candidate is located in the direction of movement of the user's gaze position. (method) A control method for an imaging device, comprising: a subject detection step of detecting a subject; a posture detection step of detecting a posture of the subject; a gaze detection step of detecting a gaze position of a user of the imaging device; a first determination step of determining a candidate for a main subject that the user focuses on based on the posture of the subject detected by the posture detection step; a second determination step of determining the main subject using the candidate determined in the first determination step and the gaze position of the user detected in the gaze detection step; and The second determining step changes a condition related to the gaze position used to determine the main subject when there is a correlation between the candidate determined in the first determining step and the gaze position of the user detected in the gaze detecting step. 10. A method for controlling an imaging device, comprising: (program) A program for causing a computer to function as each of the means of the imaging device according to any one of configurations 1 to 8. [Explanation of symbols]

[0135] 1 imaging device, 131 camera control unit, 141 part detection unit, 150 search area setting unit, 151 main part setting unit

Claims

1. An imaging device, a subject detection means for detecting a subject; a posture detection means for detecting the posture of the subject; a gaze detection means for detecting a gaze position of a user of the imaging device; a first determination means for determining a candidate for a main subject that the user focuses on based on the posture of the subject detected by the posture detection means; a second determining means for determining the main subject using the candidate determined by the first determining means and the gaze position of the user detected by the gaze detecting means; and The second determining means changes a condition related to the gaze position used to determine the main subject when there is a correlation between the candidate determined by the first determining means and the gaze position of the user detected by the gaze detecting means. An imaging device characterized by:

2. 2. The imaging device according to claim 1, wherein the conditions include a condition regarding a time period for detecting the gaze position by the gaze detection means.

3. 3. The imaging device according to claim 2, wherein the second determining means changes the condition so that the detection time is shortened when the correlation is present.

4. 4. The imaging device according to claim 3, wherein the second determining means changes the condition so that the detection time is shorter than the time taken by the object detecting means to detect the object.

5. 2. The imaging device according to claim 1, wherein the conditions include a condition related to time-series data used for detecting the gaze position by the gaze detection means.

6. 6. The imaging apparatus according to claim 5, wherein the second determining means changes the condition so that the amount of the time-series data becomes smaller when the correlation exists.

7. 2. The imaging device according to claim 1, wherein the second determining means determines the candidate as the main subject based on the degree of coincidence between the gaze position of the user and the position of the candidate determined by the first determining means.

8. 2. The imaging device according to claim 1, wherein the second determining means determines the candidate determined by the first determining means as the main subject when the candidate is located in a direction in which the user's gaze position is moving.

9. A control method for an imaging device, comprising: a subject detection step of detecting a subject; a posture detection step of detecting a posture of the subject; a gaze detection step of detecting a gaze position of a user of the imaging device; a first determination step of determining a candidate for a main subject that the user pays attention to based on the posture of the subject detected by the posture detection step; a second determination step of determining the main subject using the candidate determined in the first determination step and the gaze position of the user detected in the gaze detection step; and The second determining step changes a condition related to the gaze position used to determine the main subject when there is a correlation between the candidate determined in the first determining step and the gaze position of the user detected in the gaze detecting step.

10. A method for controlling an imaging device, comprising:

10. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Imaging device with function of detecting line of sight

    JP2019008076A