Control apparatus, imaging apparatus, control method, program, and storage medium

The control device enhances focus detection by using pixel pairs in different directions and subject detection to adjust focus areas, ensuring precise focus adjustment based on subject identification, addressing the challenge of suboptimal focus detection in imaging devices.

JP2025180131APending Publication Date: 2025-12-11CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024087264
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing imaging devices struggle with appropriate focus detection due to the lack of a systematic approach to arrange defocus maps relative to subjects, leading to suboptimal focus adjustment.

Method used

A control device that performs focus detection using a pair of pixels arranged in different directions, incorporates subject detection to adjust focus detection areas based on subject identification, and determines focus detection groups accordingly, utilizing a phase difference method for precise focus adjustment.

Benefits of technology

Enables appropriate focus detection tailored to the subject, improving focus accuracy and adaptability in imaging devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025180131000001_ABST
    Figure 2025180131000001_ABST
Patent Text Reader

Abstract

To provide a control apparatus that can perform proper focus detection according to a subject.SOLUTION: A control apparatus (120) includes: focus detection means (129) that performs first focus detection on the basis of a first signal obtained from a pair of pixels arranged on an imaging element (122) in a first direction, and performs second focus detection on the basis of a second signal obtained from a pair of pixels arranged in a second direction different from the first direction; subject detection means (130) that detects a subject on the basis of an image signal obtained from the imaging element; and determination means (125) that determines a first focus detecting area group for the first focus detection and a second focus detecting area group for the second focus detection. The determination means changes at least one of the first focus detecting area group and the second focus detecting area group according to a subject detection result by the subject detection means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a control device, an imaging device, a control method, a program, and a storage medium. [Background technology]

[0002] Patent Document 1 discloses an imaging device that performs focus detection from different focus detection directions using a phase difference detection method with two image sensors. Patent Document 1 also discloses a method for obtaining a defocus amount by combining the defocus amounts from the different focus detection directions into one. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-237215 [Non-patent literature]

[0004] [Non-Patent Document 1] Chen et.al, DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, arXiv, 2016 Summary of the Invention [Problem to be solved by the invention]

[0005] Patent Document 1 does not disclose how to arrange a defocus map, which is a group of defocus amounts for multiple regions, relative to a subject. Therefore, the imaging device disclosed in Patent Document 1 may not be able to perform appropriate focus detection depending on the subject.

[0006] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide a control device that is capable of performing appropriate focus detection in accordance with the subject detection result. [Means for solving the problem]

[0007] A control device as one aspect of the present invention comprises a focus detection means that performs first focus detection based on a first signal obtained from a pair of pixels arranged in a first direction on an image sensor, and performs second focus detection based on a second signal obtained from a pair of pixels arranged in a second direction different from the first direction, a subject detection means that detects a subject based on an image signal obtained from the image sensor, and a determination means that determines a first group of focus detection areas for the first focus detection and a second group of focus detection areas for the second focus detection, and the determination means changes at least one of the first group of focus detection areas or the second group of focus detection areas depending on the subject detection result by the subject detection means.

[0008] Other objects and features of the present invention will be described in the following embodiments. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide a control device that can perform appropriate focus detection in accordance with the subject detection result. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram of an imaging apparatus according to an embodiment of the present invention. [Figure 2(a)] FIG. 2 is a schematic diagram of a pixel array of an image sensor according to the present embodiment. [Figure 2(b)] FIG. 2 is an equivalent circuit diagram of a pixel of the image sensor according to the present embodiment. [Figure 2(c)] FIG. 2 is a diagram illustrating a pixel arrangement of an image sensor according to the present embodiment. [Figure 3] 1A and 1B are a plan view and a cross-sectional view of a pixel according to the present embodiment. [Figure 4] FIG. 2 is an explanatory diagram of pupil division in this embodiment. [Figure 5]FIG. 10 is another explanatory diagram of pupil division in this embodiment. [Figure 6] 10A and 10B are diagrams illustrating the relationship between the image shift amount and the defocus amount in this embodiment. [Figure 7] FIG. 2 is a diagram illustrating the layout of focus detection areas in the present embodiment. [Figure 8] 4 is a flowchart of a live view shooting process in the present embodiment. [Figure 9] 10 is a flowchart of a photography subroutine in the present embodiment. [Figure 10] 10 is a flowchart of subject tracking AF processing in this embodiment. [Figure 11] 4 is a flowchart of subject detection and tracking processing in this embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of a CNN that infers the likelihood of a specific region in this embodiment. [Figure 13] 10 is a flowchart of a flicker determination process according to the present embodiment. [Figure 14] 10A and 10B are diagrams illustrating the influence of flicker on a pair of signals for vertical eye focus detection in this embodiment. [Figure 15] FIG. 10 is a diagram showing a waveform when flicker occurs in this embodiment. [Figure 16] FIG. 10 is another diagram showing a waveform when flicker occurs in this embodiment. [Figure 17] 10 is a flowchart of a defocus amount selection process in the present embodiment. [Figure 18] FIG. 4 is a diagram illustrating a method for setting a defocus map in the present embodiment. [Figure 19] FIG. 10 is another diagram showing a method for setting a defocus map in the present embodiment. [Figure 20] FIG. 10 is a diagram showing a histogram of a defocus map in the present embodiment. [Figure 21] FIG. 10 is a diagram showing a histogram of a defocus map using specific region information in this embodiment. [Figure 22]10 is a flowchart of focus detection processing in this embodiment. [Figure 23] FIG. 4 is a diagram showing the execution order of subject tracking AF processing in this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0012] (imaging system 10) FIG. 1 is a block diagram of an imaging system 10 according to this embodiment. The imaging system 10 is configured to include a camera body (imaging device) 120 serving as a digital camera, and a lens unit (interchangeable lens) 100. The lens unit 100 is detachably attached to the camera body 120 via a mount M, indicated by the dotted line in FIG. 1. This embodiment is also applicable to imaging devices in which the camera body and lens unit are integrally configured. This embodiment is not limited to digital cameras, but can also be applied to other imaging devices such as video cameras.

[0013] The lens unit 100 has an imaging optical system including a first lens group 101, an aperture stop (aperture stop) 102, a second lens group 103, a focus lens (focus lens group) 104 as a focus element, and a drive / control system. The imaging optical system captures light from a subject and forms a subject image (optical image).

[0014] The first lens group 101 is disposed closest to the object and is held movably in the direction of the optical axis OA. The diaphragm 102 adjusts the amount of light by changing its aperture diameter. The diaphragm 102 and the second lens group 103 are movable together in the direction of the optical axis, and perform zooming by moving in conjunction with the first lens group 101. The focus lens 104 is movable in the direction of the optical axis and performs focusing. Focus control (autofocus (AF) control) is performed by controlling the position of the focus lens 104 in accordance with the focus detection result, which will be described later.

[0015] The drive / control system includes a zoom actuator 111, an aperture actuator 112, a focus actuator 113, a zoom drive circuit 114, an aperture drive circuit 115, a focus drive circuit 116, a lens MPU 117, and a lens memory 118. During zooming, the zoom drive circuit 114 drives the zoom actuator 111 to move the first lens group 101 and the second lens group 103 in the optical axis direction. The aperture drive circuit 115 drives the aperture actuator 112 to operate the aperture 102, thereby performing aperture operation and shutter operation.

[0016] During focusing, the focus drive circuit 116 drives the focus actuator 113 to move the focus lens 104 in the optical axis direction. The focus drive circuit 116 functions as a position detection unit that detects the current position of the focus lens 104 (hereinafter referred to as the focus position) through the focus actuator 113.

[0017] The lens MPU 117 is a computer that executes calculations and processing related to the lens unit 100, and controls the zoom drive circuit 114, aperture drive circuit 115, and focus drive circuit 116. The lens MPU 117 is also communicatively connected to a camera MPU (determination means, control means) 125 in the camera body 120 via a communication terminal of the mount M, and exchanges commands and data. For example, the lens MPU 117 notifies the camera MPU 125 of lens information in response to a request from the camera MPU 125. This lens information includes information such as the focus position, the position and diameter of the exit pupil of the imaging optical system in the optical axis direction, and the position and diameter of the lens frame that limits the luminous flux of the exit pupil in the optical axis direction.

[0018] Furthermore, the lens MPU 117 controls the zoom drive circuit 114, the aperture drive circuit 115, and the focus drive circuit 116 in response to requests from the camera MPU 125. The lens memory 118 stores optical information necessary for AF. The camera MPU 125 controls the operation of the lens unit 100 by executing programs stored in the built-in nonvolatile memory and the lens memory 118.

[0019] The camera body 120 has an optical low-pass filter 121, an image sensor 122, an image processing circuit 124, and a drive / control system. The optical low-pass filter 121 is provided to reduce false colors and moire.

[0020] The image sensor 122 is composed of a CMOS (Complementary Metal-Oxide-Semiconductor) sensor and its peripheral circuits. The image sensor 122 photoelectrically converts the subject image (optical image) formed by the imaging optical system and outputs an image signal and a pair of focus detection signals (two image signals). The image sensor 122 has multiple image pixels arranged in m pixels horizontally and n pixels vertically (m and n are integers of 2 or greater). Each image pixel includes a pair of focus detection pixels, as described below, and has a pupil division function that enables focus detection using a phase difference detection method.

[0021] The drive / control system has an image sensor drive circuit 123, a shutter 133, an image processing circuit 124, a camera MPU 125, a display 126, an operation switch (SW) 127, and a memory 128. The drive / control system also has a phase difference AF unit (focus detection means) 129, a subject detection unit (subject detection means) 130, an AE unit 131, and a white balance (WB) adjustment unit 132. In this embodiment, for example, the camera MPU 125, the phase difference AF unit 129, and the subject detection unit 130 configure a control device.

[0022] The image sensor drive circuit 123 controls charge accumulation and signal readout in the image sensor 122, and also A / D converts the image signal and paired focus detection signal output from the image sensor 122 and outputs them to the image processing circuit 124 and camera MPU 125. The image processing circuit 124 performs image processing such as gamma conversion, color interpolation, and compression encoding on the digital image signal from the image sensor drive circuit 123 to generate image data.

[0023] The camera MPU 125 is a computer that executes calculations and processing related to the camera body 120, and controls the image sensor drive circuit 123, the image processing circuit 124, the display 126, the phase difference AF unit 129, the subject detection unit 130, the AE unit 131, and the WB adjustment unit 132. The camera MPU 125 is also communicatively connected to the lens MPU 117 via a communication terminal of the mount M, and exchanges commands and data with the lens MPU 117. For example, the camera MPU 125 requests lens information or optical information from the lens MPU 117, or requests the lens MPU 117 to drive the first lens group 101, the focus lens 104, or the aperture 102. The camera MPU 125 receives the lens information or optical information transmitted from the lens MPU 117.

[0024] The camera MPU 125 has built-in ROM 125a for storing various programs, RAM 125b for storing variables, and EEPROM 125c for storing various parameters. The camera MPU 125 executes various processes including AF processing, which will be described later, in accordance with the programs stored in ROM 125a. The camera MPU 125 generates two-image data from a pair of digital focus detection signals from the image sensor drive circuit 123 and outputs the data to a phase-difference AF unit 129.

[0025] The shutter 133 has a focal plane shutter configuration, and drives the focal plane shutter in response to a command from a shutter drive circuit built into the shutter 133 based on instructions from the camera MPU 125. The image sensor 122 is shielded from light while a signal from the image sensor 122 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened, and a photographing light beam is guided to the image sensor 122.

[0026] The display 126 is configured with an LCD or the like, and displays information about the imaging mode, a preview image before imaging, a confirmation image after imaging, the focus state, etc. The operation switches 127 include a power switch, a release (imaging instruction) switch, a zoom switch, an imaging mode selection switch, etc. The memory 128 is a flash memory that is detachable from the camera body 120, and stores images for recording obtained by imaging.

[0027] The phase-difference AF unit 129 performs focus detection using two-image data generated by the camera MPU 125. The image sensor 122 photoelectrically converts a pair of optical images formed by light beams passing through a pair of different pupil regions (pupil partial regions) of the exit pupil of the imaging optical system, and outputs a pair of focus detection signals. The phase-difference AF unit 129 performs correlation calculations on the two-image data generated by the camera MPU 125 from the pair of focus detection signals to calculate the image shift amount, which is the phase difference between them, and calculates (acquires) a defocus amount as information related to the focus from the image shift amount. The camera MPU 125 calculates the drive amount of the focus lens 104 based on the defocus amount calculated by the phase-difference AF unit 129, and transmits a focus control command including the drive amount to the lens MPU 117. The phase-difference AF unit 129, which serves as focus detection means, also sets the arrangement of the area where focus detection is performed. This will be described in detail later.

[0028] As described above, in this embodiment, image plane phase difference AF is performed using the output of the image sensor 122, without using an AF sensor dedicated to focus detection. In this embodiment, the phase difference AF unit 129 has an acquisition unit 129a that acquires two-image data and a calculation unit 129b that calculates the defocus amount. Note that at least one of the acquisition unit 129a and the calculation unit 129b may be provided in the camera MPU 125.

[0029] The subject detection unit 130 detects subjects based on image signals obtained from the image sensor 122. The subject detection unit 130 also performs subject detection using dictionary data generated by machine learning. In this embodiment, the subject detection unit 130 uses dictionary data for each subject to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 130 performs subject detection by sequentially switching between dictionary data for each subject. The dictionary data for each subject is stored in a dictionary data storage unit (ROM 125a in the camera MPU 125). Therefore, multiple dictionary data are stored in the dictionary data storage unit. The camera MPU 125 determines which dictionary data from the multiple dictionary data to use for subject detection based on pre-set subject priorities and settings of the imaging device.

[0030] The AE unit 131 performs exposure control (AE: Auto Exposure) by performing photometry using image data for AE obtained from the image processing circuit 124. Specifically, the AE unit 131 acquires brightness information of the image data for AE, and calculates the aperture value, shutter speed (shutter time), and ISO sensitivity as imaging conditions from the difference between the exposure amount obtained from this brightness information and a preset exposure amount. Then, AE is performed by controlling the aperture value, shutter speed, and ISO sensitivity to the calculated values.

[0031] The WB adjustment unit 132 calculates the WB of the image data for WB adjustment obtained from the image processing circuit 124, and performs WB adjustment by adjusting the weights of RGB colors according to the difference between the calculated WB and a predetermined appropriate WB.

[0032] Furthermore, the camera MPU 125 can select the image height range for performing phase difference AF, AE, and WB adjustment according to the position and size of the subject detected by the subject detection unit 130.

[0033] (Image sensor 122) 2(a) to 2(c) show pixel arrangements on the imaging surface of an image sensor 122 serving as a two-dimensional CMOS sensor in this embodiment. Fig. 2(a) is a schematic diagram of an example of the overall configuration of the image sensor 122 shown in Fig. 1. The image sensor 122 includes a pixel array section 208, a vertical selection circuit 209, a column circuit 203, and a horizontal selection circuit 204.

[0034] The pixel array unit 208 has a plurality of pixels 205 arranged in a matrix. The output of a vertical selection circuit 209 is input to the pixels 205 via a pixel drive wiring group 207, and pixel signals from the pixels 205 in a row selected by the vertical selection circuit 209 are read out row by row to the column circuit 203 via output signal lines 206. One output signal line 206 can be provided for each pixel column, for each set of pixel columns, or for multiple pixel columns. The column circuit 203 receives signals read out in parallel via the multiple output signal lines 206, performs signal amplification, noise reduction, A / D conversion, and other processing, and stores the processed signals. The horizontal selection circuit 204 sequentially, randomly, or simultaneously selects signals stored in the column circuit 203, and the selected signals are output to the image sensor 122 via horizontal output lines and an output unit (not shown).

[0035] In this way, by sequentially outputting pixel signals of the row selected by the vertical selection circuit 209 to the outside of the image sensor 122 while changing the row selected by the vertical selection circuit 209, a two-dimensional image signal or phase difference signal can be read out from the image sensor 122.

[0036] FIG. 2B is an equivalent circuit diagram of a pixel 205 according to this embodiment. Each pixel 205 has two photodiodes (PDA 211 and PDB 212) that serve as photoelectric conversion units. The PDA 211 performs photoelectric conversion according to the amount of incident light, and the accumulated signal charge is transferred via a transfer switch (TXA) 213 to a floating diffusion (FD) 215 that constitutes a charge accumulation unit. The PDB 212 performs photoelectric conversion and accumulates the signal charge, and the transferred signal charge is transferred via a transfer switch (TXB) 214 to the FD 215. When a reset switch (RES) 216 is turned on, the FD 215 is reset to the voltage of a constant voltage source VDD. The PDA 211 and the PDB 212 can be reset by simultaneously turning on the RES 216, the TXA 213, and the TXB 214.

[0037] When a selection switch (SEL) 217 ​​that selects a pixel is turned on, an amplification transistor (SF) 218 ​​converts the signal charge accumulated in the FD 215 into a voltage, and the converted signal voltage is output from the pixel to an output signal line 206. In addition, the gates of the TXA 213, the TXB 214, the RES 216, and the SEL 217 are each connected to a pixel drive wiring group 207 and controlled by a vertical selection circuit 209.

[0038] In the following description, in this embodiment, the signal charge accumulated in the photoelectric conversion unit is assumed to be electrons, the photoelectric conversion unit is formed of an N-type semiconductor, and is separated by a P-type semiconductor, but the signal charge may also be holes, the photoelectric conversion unit may be formed of a P-type semiconductor, and is separated by an N-type semiconductor.

[0039] Next, we will explain the operation of reading signal charges from the PDA211 and PDB212 after a predetermined charge accumulation time has elapsed after resetting the PDA211 and PDB212 in a pixel having the above-mentioned configuration. First, the SEL217 of the row selected by the vertical selection circuit 209 is turned on, connecting the source of the SF218 to the output signal line 206, and the output signal line 206 enters a state in which a voltage corresponding to the voltage of the FD215 is read out. Next, the RES216 is turned on / off, resetting the potential of the FD215. After that, the system waits until the output signal line 206, which has been subjected to voltage fluctuations in the FD215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as a signal voltage N, where it is processed and held.

[0040] Thereafter, the TXA 213 is turned on / off, and the signal charge accumulated in the PDA 211 is transferred to the FD 215. The voltage of the FD 215 drops by an amount corresponding to the amount of signal charge accumulated in the PDA 211. Thereafter, the system waits until the output signal line 206, which has been subjected to the voltage fluctuation of the FD 215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as signal voltage A, which is then subjected to signal processing and held.

[0041] Thereafter, the TXB 214 is turned on / off, and the signal charge accumulated in the PDB 212 is transferred to the FD 215. The voltage of the FD 215 drops by an amount corresponding to the amount of signal charge accumulated in the PDB 212. Thereafter, the system waits until the output signal line 206, which has been subjected to the voltage fluctuation of the FD 215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as a signal voltage (A+B), which is then subjected to signal processing and held.

[0042] From the difference between the signal voltage N and the signal voltage A thus captured, a signal A corresponding to the amount of signal charge accumulated in the PDA 211 can be obtained. Furthermore, from the difference between the signal voltage A and the signal voltage (A+B), a signal B corresponding to the amount of signal charge accumulated in the PDB 212 can be obtained. This difference calculation may be performed by the column circuit 203, or may be performed after output from the image sensor 122. A phase difference signal can be obtained by using the signals A and B, and an image signal can be obtained by adding the signals A and B together. Alternatively, if the difference calculation is performed after output from the image sensor 122, the image signal may be obtained by taking the difference between the signal voltage N and the signal voltage (A+B).

[0043] Alternatively, the signal voltages N, A, and B may be read out by performing the same drive as that for reading out the signal voltages N and A on the PDB 212 instead of the PDA 211. In this case, the signals A and B obtained from the signal voltages A and B, respectively, can be used as phase difference signals as they are, and an imaging signal can be obtained by adding together the signal voltages A and B, or the signals A and B. In this embodiment, the pixel from which the above-mentioned signal A is obtained is referred to as the first focus detection pixel, and the pixel from which the above-mentioned signal B is obtained is referred to as the second focus detection pixel.

[0044] 2(c) is an array diagram showing imaging pixels in an area of ​​4 columns x 4 rows. One pixel group 200 including 2 columns x 2 rows of imaging pixels includes a pixel 200R with R (red) spectral sensitivity located in the upper left, pixels 200Ga and 200Gb with G (green) spectral sensitivity located in the upper right and lower left, and a pixel 200B with B (blue) spectral sensitivity located in the lower right. Each imaging pixel is composed of a first focus detection pixel 201 and a second focus detection pixel 202.

[0045] In pixels 200R, 200Ga, and 200B, a first focus detection pixel 201 and a second focus detection pixel 202 are arranged in the horizontal direction (first direction), while in pixel 200Gb, the first focus detection pixel 201 and the second focus detection pixel 202 are arranged in the vertical direction (second direction). In this embodiment, the phase-difference AF unit 129 performs first focus detection based on a first signal obtained from a pair of pixels (pixels 200Ga) arranged in the first direction (horizontal direction) in the image sensor 122. The first focus detection is performed using a first focus detection area group including multiple first focus detection areas (ranging frames, ranging areas). The phase-difference AF unit 129 also performs second focus detection based on a second signal obtained from a pair of pixels (pixels 200Gb) arranged in a second direction (vertical direction) different from the first direction. The second focus detection is performed using a second focus detection area group including multiple second focus detection areas.

[0046] Fig. 3(a) is a plan view of pixel 200Ga when viewed from the incident side (+z side) of the image sensor 122. Fig. 3(b) is a cross-sectional view showing the pixel structure when the aa cross section of pixel 200Ga in Fig. 3(a) is viewed from the -y side. In pixel 200Ga, a microlens 305 for collecting incident light is formed on the incident side, and photoelectric conversion units 301 and 302 divided into two in the x direction are formed. The photoelectric conversion units 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively.

[0047] The photoelectric conversion units 301 and 302 may be pin structure photodiodes with an intrinsic layer sandwiched between p-type and n-type layers, or may be pn junction photodiodes without the intrinsic layer. A color filter 306 is formed between the microlens 305 and the photoelectric conversion units 301 and 302. The spectral transmittance of the color filter may be different for each focus detection pixel, or the color filter may be omitted.

[0048] The two light beams incident on the pixel 200Ga from the paired pupil regions are each collected by a microlens 305, dispersed by a color filter 306, and then received by the photoelectric conversion units 301 and 302. In each photoelectric conversion unit, electrons and holes are generated in pairs according to the amount of received light, and after being separated by a depletion layer, the negatively charged electrons are accumulated in the n-type layer. On the other hand, the holes are not shown. The electrons are discharged to the outside of the image sensor 122 through the p-type layer connected to the constant voltage source shown in FIG. 1. The electrons accumulated in the n-type layer of each photoelectric conversion unit are transferred to the capacitance unit (FD) via the transfer gate and converted into a voltage signal.

[0049] FIG. 4 illustrates the relationship between the pixel structure and pupil division shown in FIGS. 3(a) and (b). The lower part of FIG. 4 shows the pixel structure when the aa cross section in FIG. 3(a) is viewed from the +y side, and the upper part shows the pupil plane at pupil distance DS. Note that in FIG. 4, the x-axis and y-axis of the pixel structure are inverted relative to FIG. 3(b) to correspond to the coordinate axes of the pupil plane. The pupil plane corresponds to the entrance pupil position of the image sensor 122. In this embodiment, the microlens position in each pixel is offset (shrunk) from the center of the image sensor 122, so that the entrance pupils of each pixel overlap each other to form the entrance pupil of a single image sensor 122. The pupil distance DS is the distance between the pupil plane and the image sensor, and will be referred to as the sensor pupil distance in the following description.

[0050] As shown in FIG. 4, the first pupil region 501 of the first focus detection pixel 201 is generally conjugate by a microlens to the light receiving surface of the photoelectric conversion unit 301, whose center of gravity is decentered in the -x direction. The first pupil region 501 is a pupil region through which a light beam that can be received by the first focus detection pixel 201 passes. The center of gravity of the first pupil region 501 is decentered on the +X side on the pupil plane. Furthermore, the second pupil region 502 of the second focus detection pixel 202 is generally conjugate by a microlens to the light receiving surface of the photoelectric conversion unit 302, whose center of gravity is decentered in the +x direction. The second pupil region 502 is a pupil region through which a light beam that can be received by the second focus detection pixel 202 passes. The center of gravity of the second pupil region 502 is decentered on the -X side on the pupil plane. The pupil region 500 is a pupil region through which light beams that can be received by the entire pixel 200G, which is a combination of the photoelectric conversion unit 301 and the photoelectric conversion unit 302 (the first focus detection pixel 201 and the second focus detection pixel 202), pass.

[0051] FIG. 5 is another explanatory diagram of pupil division. As shown in FIG. 5, light beams that enter the imaging optical system from the subject (the vertical line on the left side of the figure) and pass through the first pupil region 501 and the second pupil region 502 are incident on each imaging pixel at different angles and are received by the photoelectric conversion units 301 and 302. The pixels 200R, 200Ga, and 200B perform pupil division in the horizontal direction (the x-axis direction in FIG. 4), while the pixel 200Gb performs pupil division in the vertical direction (the y-axis direction in FIG. 4). The imaging pixels, each having a first focus detection pixel and a second focus detection pixel, receive light beams that pass through the first pupil region 501 and the second pupil region 502. A pair of focus detection signals is generated by combining the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of the multiple imaging pixels. Furthermore, an imaging signal with a resolution of the number of effective pixels N (= m × n) is generated by adding together the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of the multiple imaging pixels. Note that one of the pair of focus detection signals may be subtracted from the imaging signal to generate the other focus detection signal.

[0052] In addition, in this embodiment, a first and a second focus detection pixel are provided for each of all imaging pixels of the imaging element 122, but two imaging pixels may be used as the first and second focus detection pixels, or first and second focus detection pixels may be provided for some imaging pixels.

[0053] (Relationship between defocus amount and image shift amount) 6 is a diagram showing the relationship between the defocus amount and the image shift amount of two image data. 800 indicates the imaging plane of the image sensor 122, and the pupil plane of the image sensor 122 is divided into a first pupil region 501 and a second pupil region 502. The defocus amount d is defined as |d|, which is the distance from the imaging position of the subject image (image position) to the imaging plane 800, and a front-focus state in which the image position is located closer to the subject than the imaging plane, is defined as a negative sign (d<0). On the other hand, a back-focus state in which the image position is located on the opposite side of the subject than the imaging plane 800 is defined as a positive sign (d>0). The in-focus state in which the image position is located on the imaging plane 800 is d=0.

[0054] 6, subject 801 shows a focused state (d=0), and subject 802 shows a front-focused state (d<0). The front-focused state (d<0) and the back-focused state (d>0) are combined to form a defocused state (|d|>0).

[0055] In a front-focus state, light beams from the subject 802 that pass through the first pupil region 501 and the second pupil region 502 are focused once and then spread to widths Γ1 and Γ2 centered at the positions G1 and G2 of the centers of gravity of the light beams, forming blurred optical images on the imaging plane 800. These blurred images are received by the first focus detection pixel 201 and the second focus detection pixel 202 in each imaging pixel on the imaging plane 800, which generate a pair of focus detection signals, a first focus detection signal and a second focus detection signal. The first focus detection signal and the second focus detection signal are recorded as blurred images of the subject 802 at the positions G1 and G2 of the centers of gravity on the imaging plane 800, with the blur widths Γ1 and Γ2 spreading across them. The blur widths Γ1 and Γ2 increase approximately in proportion to an increase in the magnitude of the defocus amount d, |d|. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 in the center of gravity positions of the light beams) between the first focus detection signal and the second focus detection signal also increases roughly in proportion to the increase in the magnitude |d| of the defocus amount d. The same is true in the back-focus state (d>0), although the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state.

[0056] In this embodiment, the difference in the centers of gravity of the incident angle distributions in the first pupil region 501 and the second pupil region 502 is referred to as the base line length. The relationship between the defocus amount d and the image shift amount p on the imaging plane 800 is roughly similar to the relationship between the base line length and the sensor pupil distance. Because the magnitude of the image shift amount between the first focus detection signal and the second focus detection signal increases as the magnitude of the defocus amount d increases, the phase difference AF section 129 converts the image shift amount into a defocus amount using a conversion coefficient calculated based on the base line length, based on this relationship.

[0057] In the following description, calculating the defocus amount using a pair of focus detection signals from focus detection pixels that divide the pupil horizontally (horizontally) like pixel 200Ga is referred to as horizontal eye focus detection (first focus detection), and calculating the defocus amount using a pair of focus detection signals from focus detection pixels that divide the pupil vertically (vertically) like pixel 200Gb is referred to as vertical eye focus detection (second focus detection).

[0058] (Layout of focus detection area) Next, with reference to FIG. 7, the focus detection areas of the image sensor 122, which are areas for acquiring paired signal sequences for detecting a phase difference, will be described. In FIG. 7, A(n,m) and B(n,m) indicate the nth focus detection area in the x direction and the mth focus detection area in the y direction among the multiple focus detection areas (three in the x direction and three in the y direction, for a total of nine) set in the effective pixel area 300 of the image sensor 122. A signal sequence of pixel pairs, each with a pupil divided in the horizontal direction, is generated from the multiple pixels included in the focus detection area A(n,m). A signal sequence of pixel pairs, each with a pupil divided in the vertical direction, is generated from the multiple pixels included in the focus detection area B(n,m). I(n,m) indicates an index that displays the position of the focus detection areas A(n,m) and B(n,m) on the display 126. By arranging the focus detection areas in this manner, focus detection can be performed using contrast information corresponding to both the horizontal and vertical directions of the subject at the position of the index I(n,m).

[0059] The nine focus detection areas shown in FIG. 7 are merely examples, and the number, positions, and sizes of the focus detection areas are not limited. For example, one or more focus detection areas may be set within a predetermined range centered on a position specified by the user or the position of the subject detected by the subject detection unit 130. In this embodiment, when acquiring a defocus map (described later), the focus detection areas are arranged so as to obtain focus detection results with higher resolution. For example, a group of focus detection results obtained from sideways eye focus detection areas (first focus detection area group including multiple first focus detection areas) arranged in 17 horizontal and 11 vertical divisions, totaling 187 points, on the image sensor 122 is arranged as the sideways eye defocus map. Also, a group of focus detection results obtained from vertical eye focus detection areas (second focus detection area group including multiple second focus detection areas) arranged in 7 horizontal and 5 vertical divisions, totaling 35 points, is arranged as the vertical eye defocus map. Details of how the focus detection areas for sideways eye focus detection and vertical eye focus detection are arranged relative to the subject will be described later.

[0060] (Image capture processing) 8 is a flowchart showing AF / imaging processing (image processing method) that causes the camera body (imaging device) 120 of this embodiment to perform AF operation and imaging operation. Specifically, it shows processing (live view shooting processing) that causes the camera body 120 to perform operations from before imaging to displaying a live view image on the display 126 to capturing a still image. The camera MPU 125, which is a computer, executes this processing in accordance with a computer program.

[0061] First, in step S1, the camera MPU 125 causes the image sensor drive circuit 123 to drive the image sensor 122 and acquires image data from the image sensor 122. Then, from the acquired image data, the camera MPU 125 acquires first and second focus detection signals from a plurality of first and second focus detection pixels included in each of the focus detection areas shown in FIG. 7. The camera MPU 125 also adds the first and second focus detection signals of all effective pixels of the image sensor 122 to generate an image signal, and causes the image processing circuit 124 to perform image processing on the image signal (image data) to acquire image data. Note that if the image sensor pixels and the first and second focus detection pixels are provided separately, the camera MPU 125 acquires image data by performing interpolation processing on the focus detection pixels.

[0062] Next, in step S2, the camera MPU 125 causes the image processing circuit 124 to generate a live view image from the image data obtained in step S2 and displays it on the display 126. The live view image is a reduced image matched to the resolution of the display 126, and the user can adjust the imaging composition, exposure conditions, etc. while viewing this image. Therefore, the AE unit 131 and camera MPU 125 adjust the exposure based on the photometric value obtained from the image data and display it on the display 126. The exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the photographing lens, and adjusting the gain of the output from the image sensor 122.

[0063] Next, in step S3, the camera MPU 125 determines whether switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing the release switch included in the operation switch 127. If Sw1 is not turned on, the camera MPU 125 repeats the determination of step S3 to monitor the timing at which Sw1 is turned on. On the other hand, if Sw1 is turned on, the camera MPU 125 proceeds to step S400 and performs subject tracking autofocus (AF) processing. Here, it performs predictive AF processing to detect the subject area from the obtained image capture signal and focus detection signal, set the focus detection area, and suppress the influence of the time lag between the focus detection processing and the image capture processing for recording. Details will be described later.

[0064] The camera MPU 125 then proceeds to step S5, where it determines whether or not switch Sw2, which instructs the start of imaging operation, has been turned on by fully pressing the release switch. If Sw2 has not been turned on, the camera MPU 125 returns to step S3. On the other hand, if Sw2 has been turned on, the camera MPU 125 proceeds to step S300, where it executes an imaging subroutine. Details of the imaging subroutine will be described later. When the imaging subroutine ends, the camera proceeds to step S7.

[0065] In step S7, the camera MPU 125 determines whether or not the main switch included in the operation SW 127 has been turned off. If the main switch has been turned off, the camera MPU 125 ends this process, and if the main switch has not been turned off, the process returns to step S3.

[0066] In this embodiment, after it is detected in step S3 that Sw1 is turned on, the subject detection process and AF process are performed, but the timing of performing these processes is not limited to this. By performing the subject tracking AF process in step S400 before Sw1 is turned on, it is possible to eliminate the need for the user to perform preparatory actions before shooting.

[0067] (photography subroutine) Next, the photographing subroutine executed by the camera MPU 125 in step S300 of FIG. 8 will be described with reference to the flowchart shown in FIG.

[0068] First, in step S301, the AE unit 131 performs exposure control processing to determine imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.). This exposure control processing can be performed using brightness information acquired from image data of a live view image. The camera MPU 125 then transmits the determined aperture value to the aperture drive circuit 115 to drive the aperture 102. The camera MPU 125 also transmits the determined shutter speed to the shutter 133 to open the focal plane shutter. The camera MPU 125 also causes the image sensor 122 to accumulate charge during the exposure period via the image sensor drive circuit 123.

[0069] In step S302, after performing the exposure control process, the camera MPU 125 causes the image sensor drive circuit 123 to read out all pixels of the image sensor 122's image sensing signal for capturing a still image. The camera MPU 125 also causes the image sensor drive circuit 123 to read out one of the first and second focus detection signals from the focus detection area (focus target area) within the image sensor 122. By subtracting one of the first and second focus detection signals from the image sensing signal, the other focus detection signal can be obtained.

[0070] Next, in step S303, the camera MPU 125 causes the image processing circuit 124 to perform defective pixel correction processing on the imaging data read out and A / D converted in step S302. Next, in step S304, the camera MPU 125 causes the image processing circuit 124 to perform image processing and encoding processing on the imaging data after the defective pixel correction processing. Image processing includes, for example, demosaic (color interpolation) processing, white balance processing, gamma correction (tone correction) processing, color conversion processing, edge enhancement processing, etc., but is not limited to these. Next, in step S305, the camera MPU 125 records, as an image data file, still image data obtained as a result of the processing in step S304 and one of the focus detection signals read out in step S302 in the memory 128.

[0071] Next, in step S306, camera MPU 125 associates the camera characteristic information as characteristic information of camera body 120 with the still image data recorded in step S305 and records it in memory 128 and in the memory within camera MPU 125. The camera characteristic information includes, for example, the following information:

[0072] Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 124 Information about the light sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 122 Information about vignetting of imaging light beams within the camera body 120 Information on the distance from the mounting surface of the imaging optical system in the camera body 120 to the imaging element 122 ·Information about manufacturing tolerances of the camera body 120.

[0073] Information on the light sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light sensitivity distribution information) is information on the sensitivity of the image sensor 122 according to the distance (position) on the optical axis from the image sensor 122. Since the light sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, it may also be information on these. Furthermore, the light sensitivity distribution information may also be information on changes in sensitivity with respect to the angle of incidence of light.

[0074] Next, in step S307, camera MPU 125 associates lens characteristic information as characteristic information of the imaging optical system with the still image data recorded in step S305 and records it in memory 128 and in a memory within camera MPU 125. The lens characteristic information includes, for example, information on the exit pupil, a frame such as a lens barrel that blocks light beams, the focal length and F-number at the time of imaging, aberration of the imaging optical system, manufacturing error of the imaging optical system, or the position of focus lens 104 at the time of imaging (subject distance).

[0075] Next, in step S308, the camera MPU 125 records image-related information, which is information related to the still image data, in the memory 128 and in a memory within the camera MPU 125. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to focus detection accuracy. Next, in step S309, the camera MPU 125 displays a preview of the captured image on the display 126. This allows the user to easily check the captured image. After the processing of step S309 is completed, the camera MPU 125 ends this image capture subroutine and proceeds to step S7 in FIG. 8.

[0076] (Subroutine for subject tracking AF processing) Next, with reference to Fig. 10, a subject tracking AF processing subroutine executed by the camera MPU 125 in step S400 in Fig. 8 will be described. The chronological order in which steps S401 to S406 in this embodiment are executed will be described later with reference to Fig. 23.

[0077] First, in step S401, the camera MPU 125 and the phase-difference AF unit 129 perform focus detection processing using the first and second focus detection signals obtained from each of the focus detection areas acquired in step S2, as will be described in detail later.

[0078] Next, in step S402, the camera MPU 125 performs subject detection and tracking processing. The subject detection processing is executed by the subject detection unit 130. Subject detection may be impossible depending on the state of the obtained image. In that case, tracking processing is performed using other means such as template matching to estimate the subject position. Details of this will be described later.

[0079] Next, in step S403, camera MPU 125 performs a main subject determination process. The method for determining the main subject is determined according to a priority order based on predetermined criteria. For example, the closer the position of the subject detection area is to the central image height, the higher the priority is set, and when the positions are the same (the distance from the central image height is the same), the larger the size, the higher the priority is set. Also, a configuration may be adopted in which a defocus map is used to select a portion of a particular type of subject (person) that the user often wants to focus on.

[0080] Next, in step S404, the camera MPU 125 and phase difference AF unit 129 determine whether flicker is occurring in each focus detection area (flicker determination). Because the focus detection accuracy of vertical eye focus detection may be reduced due to the influence of flicker, the results of vertical eye focus detection are not used when the influence of flicker is expected to be significant. Details of the flicker detection method and the determination of whether vertical eye focus detection can be used will be described later.

[0081] Next, in step S405, camera MPU 125 and phase-difference AF unit 129 perform a defocus amount selection process. Based on the subject information obtained in step S402 and the flicker determination result obtained in step S404, the camera MPU 125 and phase-difference AF unit 129 select a defocus amount, which is the focus detection result, using the focus detection results obtained from the arranged side-eye defocus map and vertical-eye defocus map. This will be described in detail later.

[0082] Next, in step S406, the camera MPU 125 performs predictive AF processing using the defocus amount obtained in step S405 and multiple defocus amounts, which are time-series data on the timing of past focus detection. This processing is necessary when there is a time lag between the timing of focus detection and the timing of exposure for the captured image. In other words, this processing predicts the position of the subject in the optical axis direction at the timing of exposure for the captured image, which is a predetermined time after the timing of focus detection, and performs AF control.

[0083] The subject's image plane position is predicted by performing multivariate analysis (for example, the least squares method) using historical data of the subject's past image plane positions and times to find an equation for a prediction curve. The predicted image plane position of the subject can be calculated by substituting the time of exposure for the captured image into the equation for the prediction curve thus found. Furthermore, three-dimensional positions may be predicted in addition to the optical axis direction. Consider the case where the screen is represented as XY and the optical axis direction is represented as the Z direction, and the XYZ direction vectors are used. In this case, the subject's position at the time of exposure for the captured image may be predicted from the XY position of the subject obtained in the subject detection and tracking process in step S402 and the time series data of the Z direction position from the defocus amount obtained in step S405.

[0084] Predictions may also be made from time-series data of the joint positions of the person who is the subject. The above predictions make it possible to estimate the positions of the ball or person even if they are hidden during the shot, or if some of the person's joint positions become invisible. Predictions are made not only for the main subject, but also for multiple detected subjects. By performing predictive AF processing on multiple subjects, when the main subject is switched, there is no need to re-store the defocus amount history for the new main subject, and predictive AF can be continued without any time loss.

[0085] In step S406, camera MPU 125 uses the predicted AF processing result to calculate the drive amount of focus lens 104. Then, lens MPU 117 performs focus adjustment processing by driving focus actuator 113 using focus drive circuit 116 in response to the focus drive command from camera MPU 125 and moving focus lens 104 in the optical axis direction. After completing the processing of step S406, camera MPU 125 ends the subroutine of this subject tracking AF processing and proceeds to step S5 in FIG. 8.

[0086] Next, with reference to FIG. 23, the chronological execution order of steps S401 to S406 in FIG. 10 will be described. In this embodiment, the focus detection process in step S401 and the subject tracking process in step S402 are executed simultaneously. Step S401 is executed by the camera MPU 125 and the phase-difference AF unit 129, and step S402 is executed by the subject detection unit 130. Step S402 may be executed after step S401 is completed. The focus detection process in step S401 is executed after step S2201 in FIG. 22 is completed, and step S2202 is executed. In this embodiment, the vertical eye defocus map is calculated after the side-eye defocus map is calculated. This is because, when the image sensor 122 uses the slit rolling method for readout, signals in the side-eye direction are read out first and can be calculated first. However, this embodiment is not limited to this, and the vertical eye defocus map may be calculated before the side-eye defocus map.

[0087] The main subject determination process of step S403 is executed after step S402 is completed. In step S403, a defocus map is used, but in this embodiment, since calculation of the vertical eye defocus map has not been completed, a horizontal eye defocus map is used. Note that step S403 may be executed after step S401 is completed.

[0088] In this embodiment, step S404 is executed after steps S401 and S403 are completed. Also in this embodiment, step S405 is executed after steps S403 and S404 are completed. Also in this embodiment, step S406 is executed after step S405 is completed.

[0089] (Focus detection processing subroutine) Next, the focus detection processing subroutine executed by the camera MPU 125 in step S401 of FIG. 10 will be described with reference to FIG.

[0090] First, in step S2201, the camera MPU 125 sets a focus detection area. In this embodiment, a sideways eye focus detection area is set on the image sensor 122, divided into 17 horizontal and 11 vertical sections, for a total of 187 points. The camera MPU 125 also sets a vertical eye focus detection area on the image sensor 122, divided into 7 horizontal and 5 vertical sections, for a total of 35 points. The center of the focus detection area is set based on either the AF area set via the operation SW 127, the position of the subject detected and tracked in step S402, or the position of the main subject determined in step S403. In this embodiment, the group of focus detection results obtained from the sideways eye focus detection area is referred to as a sideways eye defocus map, and the group of focus detection results obtained from the vertical eye focus detection area is referred to as a vertical eye defocus map.

[0091] A method for setting a defocus map, which is a group of focus detection areas for sideways and vertical eyes, will be described with reference to Figures 18(a) to (h). Figure 18(a) is a diagram showing the subject area detected by the subject detection process described above when the subject is a person. Reference numeral 1801 denotes the upper body detection area (whole detection area, first detection area), 1802 denotes the face detection area (first detection area or second detection area), and 1803 denotes the pupil detection area (local detection area, second detection area).

[0092] The arrangement of the side-eye defocus map, which is a group of side-eye focus detection areas, will now be described. Fig. 18(b) shows the side-eye defocus map when pupils are detected, with 1804 being the side-eye defocus map. The side-eye defocus map is arranged relative to the center of the upper body detection area so as to encompass the subject. This makes it possible to fit the subject within the defocus map even when the subject is moving or when framing with a camera.

[0093] Next, the arrangement of the vertical eye defocus map, which is a group of vertical eye focus detection areas, will be described. Fig. 18(c) shows the vertical eye defocus map during face detection, with 1805 being the vertical eye defocus map. In this embodiment, the description will be made on the premise that the vertical eye defocus map has a smaller area than the horizontal eye defocus map due to limitations on calculation time. In other words, the number of first focus detection areas included in the first group of focus detection areas (horizontal eye defocus map 1804) is greater than the number of second focus detection areas included in the second group of focus detection areas (vertical eye defocus map 1805).

[0094] Because the side-eye defocus map described above can encompass the subject, the vertical-eye defocus map is set based on the area on which the user wants to focus. In the case of a person, the area on which the user wants to focus is often the pupil, so in Fig. 18(c), the vertical-eye defocus map is set with pupil detection area 1803 at the center. This makes it possible to select the defocus amount using both the side-eye defocus map and the vertical-eye defocus map for the area on which the user wants to focus in the defocus amount selection process described later.

[0095] If no pupils are detected, a vertical eye defocus map is set with face detection area 1802 at the center, as shown in Fig. 18(d). If no face is detected, a vertical eye defocus map is set with upper body detection area 1801 at the center, as shown in Fig. 18(e). If no face is detected, a vertical eye defocus map is set with upper body detection area 1801 at the center.

[0096] It is preferable to set the horizontal eye defocus map and the vertical eye defocus map so that the center positions and areas of each focus detection area are similar, which enables focus detection using signals from the same focus detection area, and therefore makes it possible to use both the horizontal eye defocus amount and the vertical eye defocus amount without distinguishing between them in the defocus amount selection process described below.

[0097] 18(f) shows a case where the area of ​​the vertical eye defocus map is reduced and each focus detection area is also reduced. In other words, the density of the second focus detection areas included in the second focus detection area group is higher than the density of the first focus detection areas included in the first focus detection area group. By densely arranging the vertical eye defocus map in the face detection area, it becomes possible to use a larger number of defocus amounts in the defocus amount selection process described below.

[0098] 18(g) shows an example in which the subject is a motorcycle. 1806 is the overall detection area of ​​the motorcycle, and 1807 is the local detection area of ​​the motorcycle helmet. As with people, it is preferable to arrange the side-glance defocus map so that it encompasses the overall detection area.

[0099] Figure 18(h) shows the setting of a vertical eye defocus map during local detection of a motorcycle. The vertical eye defocus map is not placed in the center of the local region 1807, but is placed in an area that encompasses the local detection region and allows the position and size of the horizontal eye defocus map and each focus detection region to be aligned. As described above, this results in defocus amounts that are the results of horizontal eye focus detection and vertical eye focus detection using signals from the same focus detection region. This makes it possible to use the horizontal eye defocus amount and vertical eye defocus amount together without separating them in the defocus amount selection process described below.

[0100] 22, the camera MPU 125 acquires a defocus map. For the focus detection areas set in step S2201, the phase-difference AF unit 129 calculates the amount of image shift between the first focus detection signal and the second focus detection signal obtained in each of the focus detection areas acquired in step S2. The phase-difference AF unit 129 then calculates the amount of defocus and reliability for each focus detection area from the amount of image shift.

[0101] (Subroutine for subject detection and tracking processing) Next, with reference to FIG. 11, the subject detection and tracking process subroutine executed by the camera MPU 125 in step S402 of FIG. 10 will be described.

[0102] First, in step S421, the camera MPU 125 sets dictionary data according to the type of subject to be detected from the image data acquired in step S1. Based on the preset subject priority and the settings of the imaging device, dictionary data to be used in this process is selected from multiple dictionary data stored in the dictionary data storage unit. For example, multiple dictionary data are stored by classifying subjects into categories such as "people," "vehicles," and "animals." In this embodiment, one or more dictionary data may be selected. When one dictionary data is selected, it becomes possible to repeatedly detect subjects that can be detected using one dictionary data at a high frequency. On the other hand, when multiple dictionary data are selected, the dictionary data can be set sequentially according to the priority of the detected subject, allowing subjects to be detected sequentially.

[0103] Next, in step S422, the subject detection unit 130 performs subject detection using the dictionary data set in step S421 and the image data read in step S1 as an input image. At this time, the subject detection unit 130 outputs information such as the position, size, and reliability of the detected subject. At this time, the camera MPU 125 may display the information output by the subject detection unit 130 on the display 126. In step S422, multiple regions of the subject are detected hierarchically from the image data. For example, if "person" or "animal" is set as the dictionary data, multiple organs such as the "whole body" region, the "face" region, and the "eye" region are detected. While local regions such as a person's eyes or face are areas where it is desirable to adjust the focus and exposure as a subject, they may not be detectable due to surrounding obstacles or the orientation of the face. Even in such cases, the subject is detected robustly by performing full-body detection, allowing for hierarchical subject detection. Similarly, when a "vehicle" such as a motorbike is set as dictionary data, the system is configured to hierarchically detect the entire area including the driver and vehicle body, and the helmet (head) as a local area.

[0104] Next, in step S423, camera MPU 125 performs a known template matching process using the subject detection area obtained in step S422 as a template. Using the multiple images obtained in step S1, a similar area is searched for in the most recently obtained image using the subject detection area obtained in the past image as a template. As is well known, any information may be used for template matching, such as brightness information, color histogram information, or feature point information such as corners and edges. Various matching methods and template update methods are possible, and any of these methods may be used. The tracking process performed in step S423 is performed to achieve stable subject detection and tracking when a subject is not detected in step S422 by detecting an area similar to the past subject detection data from the most recently obtained image data.

[0105] Next, in step S424, the subject detection unit 130 performs region division of the detected subject region into specific regions. A specific region is a partial or entire region of the detected subject region. For example, if a person or animal is detected, it may be the region of the person's head, or if a vehicle is detected, it may be the region of the helmet. Unlike subject detection, in which the size and position of the subject are obtained using the size and coordinates of a rectangular region, region division allows the detection result to be obtained as a high-resolution distribution of the specific region. Any method (for example, the method disclosed in Non-Patent Document 1) can be used as the region division method.

[0106] The object detection unit 130 uses a deep-learned CNN to infer the likelihood of a specific region for each pixel region. However, the object detection unit 130 may infer the likelihood of a specific region using a trained model trained by any machine learning algorithm, or may determine the likelihood of a specific region based on a rule base. When a CNN is used to infer the likelihood of a specific region, the CNN performs deep learning using the specific region as a positive example and regions other than the specific region as a negative example. As a result, the CNN outputs the likelihood of a specific region for each pixel region as an inference result.

[0107] (Method of calculating specific areas) 12(a) to 12(c) are diagrams showing an example of a CNN (convolutional neural network) that infers the likelihood of a specific region. FIG. 12(a) shows an example of a subject region of an input image input to the CNN. The subject region 1201 is detected from the image by the subject detection described above. The subject region 1201 includes a face region 1202, which is the detection target of the subject detection. The face region 1202 in FIG. 12(a) includes two occluded regions (occluded regions 1203 and 1204). The occluded region 1203 is a region with no depth difference from the face region, and the occluded region 1204 is a region with a depth difference. The occluded region is also called an occlusion. In this embodiment, the face region 1202 excluding the occluded regions 1203 and 1204 is detected as a specific region.

[0108] FIG. 12(b) shows an example of the definition of specific region information. Each of images 1 to 3 in FIG. 12(b) is divided into a white region and a black region, with the black region indicating a positive example and the white region indicating a negative example. The specific region information obtained by dividing the image of the subject region in FIG. 12(b) is all images assumed to be candidates for training data used when performing deep learning of CNN. Below, we will explain which of the specific region information in FIG. 12(b) is used as training data in this embodiment.

[0109] Image No. 1 in Figure 12(b) shows an example of occlusion information when the image is divided into a subject region (face region) and a region other than the subject, with the subject region being a positive example and regions other than the subject region, such as background and occlusion regions, being negative examples. Image No. 2 in Figure 12(b) shows an example of occlusion information when the image is divided into a foreground occluded region relative to the subject and other regions, with the foreground occluded region being a negative example and regions other than the foreground occluded region relative to the subject being positive examples. Image No. 3 in Figure 12(c) shows an example of occlusion information when the image is divided into a occluded region that causes perspective conflict relative to the subject and other regions, with the occluded region that causes perspective conflict being a negative example and regions other than the occluded region that causes perspective conflict being a positive example.

[0110] As shown in FIG. 12(b) , a person's face in an image has a distinctive visibility pattern with small variance, enabling highly accurate region segmentation. For example, the occlusion information shown in FIG. 12(b) 1 is suitable as training data for the learning process when generating a CNN that detects a person as a subject. From the perspective of detection accuracy, the occlusion information shown in FIG. 12(b) 1 is more suitable than the occlusion information shown in FIG. 12(b) 3. However, an image like the image shown in FIG. 12(b) 3 is suitable as training data for the learning process when generating a CNN that detects occluded areas that cause perspective conflict. A pair of parallax images used in focus detection may be used as training data for the learning process when generating a CNN that detects occluded areas that cause perspective conflict. Furthermore, the occlusion information is not limited to the above example and may be generated based on any method for segmenting an image into an occluded area and an area outside the occluded area. In this embodiment, emphasis is placed on the accuracy of the detection area, and the learning process is performed using the information shown in FIG. 12(b) 1. However, learning may also be performed using other information.

[0111] Fig. 12(c) shows the flow of deep learning of CNN. In this embodiment, an RGB image is used as an input image 1210 for learning. Furthermore, a teacher image 1214 (teacher image of specific region information) as shown in Fig. 12(c) is used as the teacher image. The teacher image 1214 is an image of face region information excluding occlusion information and background information in Fig. 12(b).

[0112] An input image 1210 for training is input to a neural network system 1211 (CNN). The neural network system 1211 can employ, for example, a layer structure in which convolutional layers and pooling layers are alternately stacked between an input layer and an output layer, and a multi-layer structure in which a fully connected layer is connected downstream of the layer structure. A score map indicating the likelihood of a specific region in the input image is output from the output layer 1212 in FIG. 12(c). The score map is output in the form of an output result 1213.

[0113] In CNN deep learning, the error between the output result 1213 and the training image 1214 is calculated as a loss value 1215. The loss value 1215 is calculated using, for example, a method such as cross entropy or squared error. Then, coefficient parameters such as the weight and bias of each node of the neural network system 1211 are adjusted so that the loss value 1215 gradually decreases. By sufficiently performing CNN deep learning using many learning input images 1210, the neural network system 1211 will output a more accurate output result 1213 when an unknown input image is input. In other words, when an unknown input image is input, the neural network system 1211 (CNN) will output, with high accuracy, specific region information obtained by dividing the region into occluded regions and non-occluded regions as the output result 1213. Note that it takes a lot of work to create training data that identifies occluded regions (regions of overlapping objects). For this reason, it is possible to create training data using CG or by using image synthesis, which cuts out and superimposes images of objects.

[0114] As described above, an example has been described in which image 1 in FIG. 12(b) is used as the teacher image 1214, in which the face region, excluding the occluded region and the background region, is the specific region. Here, image 2 in FIG. 12(b) may be used as the teacher image 314, in which a region with no depth difference (a region in the foreground of the subject where the depth difference is less than a predetermined value) is used as the occluded region. Alternatively, image 3 in FIG. 12(b) may be used, in which a region with a depth difference (a region in the foreground of the subject where the depth difference is equal to or greater than a predetermined value) is used as the occluded region. Even if images such as 2 or 3 in FIG. 12(b) are used as the teacher image 1214, the CNN can infer a region that will cause perspective conflict when an unknown input image is input to the CNN.

[0115] Any method other than CNN can be applied to detect a specific region. For example, the detection of a specific region may be realized by a rule-based method. Furthermore, a trained model trained by machine learning using any method other than deep learning CNN may be used to detect a specific region. For example, an occluded region may be detected using a trained model trained by machine learning using any machine learning algorithm such as a support vector machine or logistic regression. This is similar to subject detection.

[0116] In this embodiment, detection of a specific area is performed for all detected subjects, but the amount of calculation can be reduced by performing this after the main subject determination process in step S403 and detecting a specific area only for the main subject.

[0117] When the process of step S424 in FIG. 11 is completed, the camera MPU 125 ends the subject detection and tracking process subroutine, and proceeds to step S404 in FIG.

[0118] (Flicker detection subroutine) Next, the flicker determination subroutine executed by the camera MPU 125 in step S404 of FIG. 10 will be described with reference to FIG.

[0119] First, in step S1301, the camera MPU 125 acquires information (image sensor drive information) related to the drive of the image sensor 122 performed in step S1. In this embodiment, the image sensor 122 selects various drive methods depending on factors such as the brightness of the shooting environment and whether the recorded image is a still image or a video. To read signals within the screen within the time allowed by the frame rate (image sensor drive rate) set based on the brightness of the shooting environment and user settings, the readout rows are thinned out or signals from multiple rows are read simultaneously. In step S1301, information is acquired regarding the drive of the image sensor regarding the result of vertical eye focus detection (image shift amount) that occurs when flicker occurs, which is determined by the number of thinned rows and the number of rows being simultaneously read out. In this embodiment, whether flicker occurs in the shooting environment is determined based on the degree of agreement between the acquired information and the result calculated by the phase-difference AF unit 129 as the image shift amount of vertical eye focus detection. Details will be described later.

[0120] Next, in step S1302, the camera MPU 125 sets a focus detection area for performing flicker detection within the defocus map calculated in step S401 of FIG. 10. In this embodiment, the determination is performed sequentially for each of the 24 areas that make up the vertical eye defocus map. Then, in step S1303, the camera MPU 125 acquires the side-eye focus detection results and the vertical eye focus detection results for the focus detection area set in step S1302 and calculates the difference between them. This process is performed because if the vertical eye focus detection result contains an error due to the effects of flicker, the difference from the side-eye focus detection result may be large.

[0121] Next, in step S1304, the camera MPU 125 acquires image shift amount candidates for vertical eye focus detection. To explain the image shift amount candidates, the correlation calculation for performing focus detection in step S401 will be explained.

[0122] In this embodiment, a pair of signals used for vertical eye focus detection is referred to as the A image signal and the B image signal. The first, second, etc. outputs of the A image signal in each row within the focus detection area are designated A(1), A(2), etc., and similarly, the first, second, etc. outputs of the B image signal are designated B(1), B(2), etc. In this manner, 300 sequentially generated A (B image) signals are concatenated to generate a pair of image signals. In the correlation calculation, the correlation amount is calculated while shifting the relative positions of the pair of image signals, and the shift amount at the position with the highest correlation (the degree of similarity in the shapes of the pair of image signals) is detected as the image shift amount. For example, the correlation amount COR(h) can be calculated using the following equation (1):

[0123]

number

[0124] In equation (1), W1 corresponds to the number of data points within the field of view, and hmax corresponds to the number of shift data points. After calculating the correlation amount COR(h) for each shift amount h, the phase-difference AF unit 129 calculates the shift amount h that maximizes the correlation between the A and B images, i.e., the value of the shift amount h that minimizes the correlation amount COR(h). Note that the shift amount h used to calculate the correlation amount COR(h) is an integer, but when calculating the shift amount h that minimizes the correlation amount COR(h), interpolation or the like is performed to obtain a value in sub-pixel units (real values) in order to improve the accuracy of the defocus amount. Note that in this embodiment, the shift amount at which the sign of the difference value of the correlation amount COR changes is calculated as the shift amount h (in sub-pixel units) that minimizes the correlation amount COR(h). First, the phase difference AF unit 129 calculates a difference value DCOR of the correlation amount according to the following equation (2).

[0125] DCOR(2×h)=COR(h+1)-COR(h-1) …(2) Then, the phase difference AF unit 129 uses the correlation difference value DCOR to calculate a shift amount dh1 at which the sign of the difference amount changes. If the value of h just before the sign of the difference amount changes is h1 and the value of h after the sign changes is h2 (h2 = h1 + 1), the phase difference AF unit 129 calculates the shift amount dh1 according to the following equation (3).

[0126] dh1=(h1+|DCOR1(h1)| / |DCOR1(h1)-DCOR1(h2)|)×2 …(3) In this way, the phase-difference AF unit 129 calculates the shift amount dh1 in sub-pixel units that maximizes the correlation between the images A and B of the first signal, and then completes the process. Note that the method for calculating the shift amount (phase difference) between two one-dimensional image signals is not limited to that described here, and any known method can be used. As a result of the above-described correlation calculation, multiple shift amounts that change the sign of the difference value of the correlation amount COR may be calculated. In normal focus detection, the shift amount that maximizes the difference value is selected, but in step S1304, the multiple calculated shift amounts are acquired as image shift amount candidates. A method for using the image shift amount candidates will be described in detail later.

[0127] Next, in step S1305, the camera MPU 125 determines whether there is a correlation between the image shift amount candidate of step S1304 and the information about the result of vertical eye focus detection (image shift amount) when flicker occurs related to the driving method of the image sensor acquired in step S1301. If the value of the image shift amount candidate acquired in step S1304 or the difference therebetween is close to the image shift amount acquired in step S1301 within a predetermined value, the process proceeds to step S1306; if not close, the process proceeds to step S1308.

[0128] In step S1306, camera MPU 125 determines the magnitude of the difference between the vertical eye and horizontal eye focus detection results acquired in step S1303. If the difference is large, proceed to step S1307; if the difference is small, proceed to step S1308. In step S1307, camera MPU 125 determines that the set focus detection area is affected by flicker because flicker has caused an error in the vertical eye focus detection result. On the other hand, in step S1308, camera MPU 125 determines that the set focus detection area is less affected by flicker on the vertical eye focus detection result.

[0129] After step S1307 or step S1308 is completed, the process proceeds to step S1309, where the camera MPU 125 determines whether flicker detection has been completed for all focus detection areas. If not, the process returns to step S1302, and the above-described processing is repeated. If completed, the process of this subroutine is completed, and the process proceeds to step S405.

[0130] (Impact of image sensor driving method and flicker on vertical eye focus detection) Next, with reference to FIGS. 14(a) to 14(c) through 16(a) to 16(d), a mechanism by which an error occurs in vertical eye focus detection due to flicker depending on the driving method of the image sensor will be described.

[0131] Flicker, which occurs in lighting, digital signage, and the like, is a phenomenon in which blinking occurs repeatedly over time at an invisible frequency. On the other hand, slit rolling image sensors sequentially accumulate and read out signals from each row over time. When a slit rolling image sensor is exposed in an environment where flicker occurs, the accumulation time difference between each row causes the signal from each row to fluctuate due to the influence of flicker. In the present invention, focus detection signals are also sequentially read out from each row. However, since the paired signals used for sideways eye focus detection use signals from the same row, they are affected by flicker to the same extent, and therefore the impact on the focus detection results is small. On the other hand, the paired signals used for longitudinal eye focus detection are affected by flicker-induced blinking within the paired signal sequence because the direction in which the signal sequence is formed coincides with the direction in which the slit rolling readout is performed.

[0132] Figures 14(a) to (c) are explanatory diagrams showing the effect of flicker on a pair of vertical eye focus detection signals. Figure 14(a) shows the passage of time horizontally from left to right, and shows the timing of accumulation and readout of focus detection signals (image A) and image signals (image A+B) for each row of the image sensor on the time axis. As explained with reference to Figure 2(b), the output of signal A and signal A+B for each row is indicated by the accumulation period and readout period shown in the top two lines of Figure 14(a).

[0133] After resetting the PDA211 and PDB212, accumulation of signal A and signal A+B begins, and as soon as accumulation of signal A is completed, the voltage is read out. After readout of signal A is completed, accumulation of signal A+B is completed and the voltage is read out. Similarly, the signal on the second row is read out. The time difference between the accumulation period of signal A on the first row and the accumulation period of signal A on the second row can be considered to be the difference in the centers of the accumulation periods, so the interval is Pa-a. Furthermore, the interval between the accumulation period of signal A+B on the first row and the accumulation period of signal A+B on the second row is Pab-ab. As mentioned above, in an environment where flicker occurs, brightness changes over time, so the signal output of the first and second rows changes as Pa-a and Pab-ab pass. The difference in the accumulation periods of signal A and signal A+B is shown as Pa-ab.

[0134] In an environment where flicker occurs, there is a Pa-ab shift between signal A and signal A+B during their accumulation periods, regardless of the row. Due to the Pa-ab shift, the waveforms of signal A and signal A+B experience an image shift due to the effects of flicker. Due to the difference in the accumulation periods of signal A and signal A+B, the waveform of signal B is shifted horizontally by Pa-ab / Pa-a pixels relative to the waveform of signal A. For example, as shown in Figure 14(a), the accumulation start time for each row is shifted by a time equivalent to the sum of the readout periods of signal A and signal A+B. When the readout periods of signal A and signal A+B are equal, the waveform of signal B is shifted horizontally by Pa-ab / Pa-a pixels = 1 / 4 pixel relative to the waveform of signal A.

[0135] Figure 14(b) shows a case where exposure control for each row is different from that of Figure 14(a), and signals A and B are read out for each row. This shows a case where the start of accumulation for signals A and B in the first row is shifted by the readout period of signal A. As in Figure 14(a), due to the difference in the accumulation periods for signals A and A+B, the waveform of signal B is shifted horizontally by Pa - ab / Pa - a pixels relative to the waveform of signal A. For example, for signal A in the first row, signal B in the first row, signal A in the second row, etc., the start of accumulation is shifted by times corresponding to the readout period of signal A in the first row, the readout period of signal B in the first row, the readout period of signal B in the second row, etc. Furthermore, if the readout periods of signals A and B are equal, the waveform of signal B is shifted horizontally by Pa - ab / Pa - a pixels = 1 / 2 pixels relative to the waveform of signal A.

[0136] FIG. 15(a) shows signals A and B corresponding to the case of FIG. 14(b). The horizontal axis represents the pixel number, and the vertical axis represents the signal output normalized by the maximum value. The undulations in the output for each pixel indicate flickering over time. An enlarged view of a portion is shown in the upper right corner of FIG. 15(a), which shows that the waveforms of signals A and B are slightly shifted. As explained in step S1304, FIG. 15(b) shows the results of calculating the correlation amount. The horizontal axis represents the amount of shift (shift amount) between the positions of signals A and B, and the vertical axis represents the correlation amount, which indicates the magnitude of the correlation. FIG. 15(b) shows that the correlation amount reaches its minimum value when the shift amount is around ±40 pixels and 0 pixel.

[0137] FIG. 15(c) is a diagram showing the calculated difference value DCOR of the correlation amount. The horizontal axis represents the shift amount, and the vertical axis represents the difference in the correlation amount. The shift amounts that intersect with the horizontal axis in an upward sloping manner to the right indicate the vicinity of ±80 pixels and 0 pixel. FIG. 15(d) is an enlarged view of the vicinity of the shift amount of 0 pixel. The candidate image shift amount dh1 in this embodiment indicates -0.5 pixel, which is the intersection with the horizontal axis. Similarly, -80.5 pixels and +79.5 pixels are candidates for the image shift amount.

[0138] The image shift amount candidate of -0.5 pixels is the pixel shift amount that occurs when the readout of FIG. 14(b) is performed in an environment where flicker occurs. In this embodiment, in step S1301, information on the readout method such as FIG. 14(a) or FIG. 14(b) is obtained as information related to driving of the image sensor 122, thereby obtaining the image shift amount that occurs due to flicker. For example, in the case of the driving method of FIG. 14(b), information of -0.5 pixels is obtained.

[0139] On the other hand, as can be seen from Figure 15(a), the image shift amounts of -80.5 pixels and +79.5 pixels are image shift amounts that are offset by -0.5 pixels of the image shift amount caused by the influence of flicker relative to the 80-pixel cycle in which flicker occurs. By canceling out the image shift amount caused by the influence of flicker, it is possible to calculate that the cycle in which flicker occurs is 80 pixels, and the frequency of flicker blinking can be calculated from information on the readout time for each row.

[0140] In step S1305, it is determined whether the image shift amount of -0.5 pixels when flicker occurs is included in the image shift amount candidates obtained in step S1304, based on the read information from the image sensor 122 in FIG. 14(b). If it is included, it is determined that a flicker environment may exist, and the process proceeds to step S1305. In step S1306, to exclude cases where the defocus state of the subject matches the image shift amount detected in a flicker environment, the difference with the sideways eye focus detection result, which is less affected by flicker, is checked. If the difference between the sideways eye focus detection result and the vertical eye focus detection result is small, it is determined that the defocus state of the subject can also be obtained from the vertical eye focus detection result. On the other hand, if the difference is large, it is determined that the vertical eye focus detection result is affected by flicker.

[0141] By performing the determination in step S1306, focus detection using the vertical eye focus detection result can be performed in a wider range of shooting environments, allowing for more accurate focus adjustment. On the other hand, it is also possible to omit the determination in step S1306 and configure the system to minimize the influence of flicker on the vertical eye focus detection result.

[0142] FIG. 14(c) shows a case where multiple rows of the image sensor 122 are simultaneously read out. While FIG. 14(c) shows a case where four rows are simultaneously read out, the number of rows that can be simultaneously read out is not limited to this. Even when multiple rows are simultaneously read out, there is a difference between the readout period of signal A and the readout period of signals A+B. Furthermore, there is a difference in the readout period for each block of rows (four rows per block in FIG. 14(c)). For ease of understanding, FIG. 16(a) shows the waveforms of signals A and B when 10 rows are simultaneously read out. In addition to the effect of flicker, it can be seen that a step occurs every 10 rows. When the above-described correlation calculation is performed on such waveforms, there is a section of shift amount where the change in correlation amount is small, making it impossible to obtain a highly accurate image shift amount, so digital filtering is performed.

[0143] FIG. 16(b) shows the results of the specified filter processing (-4, -11, -21, -28, -28, -17, 0, 17, 28, 28, 21, 11, 4). As with the correlation calculation process described above, FIG. 16(c) shows the correlation amount COR, and FIG. 16(d) shows the correlation amount difference DCOR. It can be seen that the correlation amount difference DCOR slopes upward and intersects with the horizontal axis at approximately -90, -80, -10, 0, +70, and +80 pixels. Here, the -10-pixel shift amount represents the amount of image shift caused by flicker when the image sensor 122 simultaneously reads 10 rows. As with the row-by-row readout described above, step S1301 acquires information about the drive of the image sensor 122 regarding simultaneous 10-row readout, and acquires that the amount of image shift caused by flicker is approximately -10 pixels.

[0144] Thereafter, the determinations made in steps S1305 and S1306 are as described above. Similarly, the flicker blinking frequency can be calculated from the shift amounts of ±80 pixels and 0 pixel. It can also be seen that the image shift amount candidates of -90 pixels and +70 pixels are image shift amounts resulting from the sum of the flicker blinking frequency and the influence of flicker caused by the readout method of the image sensor.

[0145] When multiple rows are read simultaneously, image shift occurs due to two factors: the difference Pa-ab between the readout periods of Signal A and Signal A+B, and the difference Pa-ab due to waveform steps that occur every few rows. If the digital filter processing described above sufficiently reduces the effect of the waveform steps that occur every few rows, the effect of the difference Pa-ab between the readout periods of Signal A and Signal A+B becomes significant. For example, in the case of the four-row simultaneous readout shown in Figure 14(c), if the waveform steps that occur every four rows are eliminated by digital filter processing, an image shift of Pa-ab / Pa-a × 4 pixels = 1 pixel occurs. (This occurs when the accumulation start and readout shift occur every four rows for a time equivalent to the sum of the readout periods of Signal A and Signal A+B, and the readout periods of Signal A and Signal A+B are equal.) On the other hand, in the case of Figure 16, when 10 rows are read simultaneously, the waveform steps that occur every 10 rows are not eliminated by digital filter processing. Therefore, an image shift of -10 pixels is calculated as the image shift candidate.

[0146] In this way, the amount of image shift at which the focus detection result is affected by flicker is calculated in advance by combining the drive information of the image sensor 122 acquired in step S1301 and the digital filter processing used in the correlation calculation, which allows comparison with the image shift amount candidate in step S1304.

[0147] 14(a)-(c) through 16(a)-(d), the effects on the A and B signals and the vertical eye focus detection results in a flicker environment are based on the case where the subject has no contrast and only the effects of flicker occur. In reality, contrast including the defocus state of the subject is superimposed on the A and B signals. Therefore, when the contrast of the subject is low and the brightness difference of the flicker is large, the flicker has a large effect on the vertical eye focus detection results, resulting in a value close to the amount of image shift described above.

[0148] On the other hand, when the contrast of the subject is high or when the difference in brightness due to the flicker is small in a mixed light environment with other flicker-free light sources, the effect of the flicker on the vertical eye focus detection result is small, and a vertical eye focus detection result indicating the defocus state of the subject is obtained. Therefore, in the determination made in step S1305 of Figure 13, it is desirable to make the determination assuming that a certain degree of error will occur in the image shift amount that occurs in a flicker environment due to the readout method of the image sensor 122 and the digital filter. For example, in Figures 15(a) to 15(d), a method can be considered in which a determination of "Yes" is made when a candidate image shift amount for vertical eye focus detection is obtained within the range of -0.5 pixels ±0.25 pixels.

[0149] As described above, in an environment where flicker occurs, errors may occur in the vertical eye focus detection result, but by determining whether or not it can be used according to the driving information of the image sensor, it is possible to avoid using a low-accuracy vertical eye focus detection result, and as a result, high-accuracy focus detection can be performed.

[0150] In this embodiment, the presence or absence of flicker influence is determined for each focus detection area. Flicker may occur due to lighting in the entire shooting environment, or it may occur only in a part of the shooting environment, such as digital signage. By determining the presence or absence of flicker influence for each focus detection area as in this embodiment, more vertical eye focus detection results can be used, enabling more accurate focus detection.

[0151] On the other hand, as mentioned above, the impact of flicker on vertical eye focus detection results varies depending on the contrast of the subject, including defocus. Therefore, using only one focus detection area can result in an incorrect determination. Therefore, a threshold can be set in advance, and if a number of focus detection areas greater than the threshold are determined to be affected by flicker, none of the vertical eye focus detection results can be used. Furthermore, if there is an uneven distribution of focus detection areas affected by flicker, a method can be considered in which vertical eye focus detection areas are not used in only some areas within the shooting range. These methods can more reliably reduce errors due to flicker in vertical eye focus detection results.

[0152] (Defocus amount selection process) Next, the subroutine of the defocus amount selection process (step S405 in FIG. 10) will be described with reference to Figures 17 to 20. Figure 17 is a flowchart showing the defocus amount selection process.

[0153] First, in step S1701, camera MPU 125 acquires subject detection information, namely, subject detection position and size, detected by subject detection unit 130. Next, in step S1702, camera MPU 125 acquires specific region information detected by subject detection unit 130. In this embodiment, the specific region information is a face region excluding occluded regions and background regions. Note that processing using the specific region information will be described later with reference to FIGS. 19(a) to 19(d).

[0154] Next, in step S1703, the camera MPU 125 collects usable focus detection results. Collecting usable focus detection results is a process of collecting defocus amounts, which are focus detection results that can be used in the defocus amount selection process, from the side-eye defocus map and the defocus amounts in the side-eye defocus map. Specifically, whether or not all vertical-eye focus detection results are usable is determined depending on whether the number of focus detection areas determined to be affected by flicker in the flicker determination process shown in FIG. 13 described above is equal to or greater than a predetermined number. The reason why all vertical-eye focus detection results are usable is because if a predetermined number or more are determined to be affected by flicker, there is a high possibility that the vertical-eye focus detection results will contain errors due to the influence of flicker.

[0155] Additionally, when the subject contrast is low, the ISO sensitivity is high, or the exposure is darker than the correct exposure, the error in the focus detection result is large. Therefore, the reliability of the focus detection result may be determined based on the difference in the correlation amount in the correlation calculation process described above, and the result may be determined not to be used as the focus detection result. Furthermore, thinning or adding the rows to be read out depending on the driving method of the image sensor 122 may result in the accuracy of the vertical eye focus detection result being inferior to the accuracy of the horizontal eye focus detection result. Therefore, in modes where imaging is performed using such driving methods, the vertical eye focus detection may be determined not to be used.

[0156] Next, in step S1704, the camera MPU 125 generates a histogram using the defocus amount, which is the focus detection result made available in step S1703. The histogram uses the subject detection information and specific area information to determine which focus detection area's focus detection result to use. As shown in the focus detection area setting process in step S2201 of Figure 22 above, the histogram is generated using a defocus map included in the subject area. Figures 20(a) to 20(f) show histograms of focus detection results.

[0157] A method for creating a histogram using defocus amounts in an upper body detection region will be described with reference to FIGS. 20(a) to 20(c). FIG. 20(a) is a histogram generated from the defocus amounts of a side-eye defocus map within the upper body detection region of the person in FIG. 18(b). FIG. 20(b) is a histogram generated from the defocus amounts of a vertical-eye defocus map within the upper body detection region of the person in FIG. 18(c). FIG. 20(c) is a histogram generated by combining the defocus amounts of the side-eye defocus map and the vertical-eye defocus map within the upper body detection region of the person in FIGS. 18(b) and 18(c). The horizontal axis of the histogram represents the class into which the defocus amount is divided into a certain range, and the vertical axis represents the frequency. In this example, the defocus amount is set such that the positive side is the near side and the negative side is the far side, and the defocus amount of the person's pupil region is 0Fδ.

[0158] In the side-glance histogram in Figure 20(a), the histogram is generated for the entire upper body detection area, and so the area to the left of the upper body below the face is more prevalent, so the maximum value of the histogram frequency is closer to the target. Therefore, if the defocus amount is selected from the range of defocus amounts that results in the maximum value of the histogram frequency, the selected defocus amount will be different from the eye region of the person on whom the user wants to focus.

[0159] In the vertical eye histogram in Figure 20(b), the vertical eye defocus map is placed in the face detection area, and therefore does not include the area on the left side of the upper body below the face, so the frequency of the histogram is maximum in the range around 0Fδ, which is the defocus amount of the pupil area. However, because the number of focus detection areas in the defocus map is small, it may be difficult to extract the point with the maximum frequency under conditions where the defocus amount is likely to fluctuate due to errors.

[0160] Therefore, by generating a histogram that combines horizontal and vertical eye movements, as shown in Figure 20(c), it is possible to generate a histogram using a larger number of defocus amounts. This makes it possible to select a more accurate defocus amount in response to variations in defocus amount or incorrect defocus amounts. However, since this is a histogram of defocus amounts in the upper body detection area, it also includes the area on the left side of the upper body below the face. Therefore, the areas with high frequencies in the defocus amount histogram are high in the ranges of -1Fδ to 0Fδ and 0Fδ to 1Fδ, making it difficult to extract the range of defocus amount where the histogram frequency is maximum. As a result, depending on the variations in defocus amount, the range of defocus amount where the histogram frequency is maximum may fluctuate.

[0161] Therefore, in this embodiment, it is possible to use a vertical eye defocus map in addition to a horizontal eye defocus map. For this reason, a method for creating a histogram using the defocus amount in the face detection area will be described with reference to Figures 20(d) to (f). Let us take an example in which the defocus amount in the pupil area of ​​a person is 0Fδ.

[0162] Fig. 20(d) is a histogram generated from the defocus amount of the side-glance defocus map within the person's face detection area in Fig. 18(b). Because the histogram is generated based on the defocus amount within the face detection area, the frequency of the defocus amount histogram reaches its maximum value in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil area.

[0163] Fig. 20(e) is a histogram generated from the defocus amount of the vertical eye defocus map within the person's face detection area in Fig. 18(d). Because the histogram is generated based on the defocus amount within the face detection area, the frequency of the defocus amount histogram reaches its maximum value in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil area.

[0164] Figure 20(f) is a histogram generated by combining the defocus amounts for side and vertical eyes by combining the histograms in Figures 20(d) and 20(e). The histogram combining the defocus amounts for side and vertical eyes within the face detection area results in a higher frequency in the histogram for defocus amounts in the range from -1Fδ to 0Fδ, which includes the defocus amount in the person's pupil area, than in the histogram for defocus amounts for side or vertical eyes only. This makes it possible to reduce the effects of variations in defocus amounts or defocus amounts that cause perspective conflicts with the background.

[0165] It is desirable to generate a defocus amount histogram using as many defocus amounts as possible in a narrow person detection area. As shown in Figures 20(d) to (f), it is preferable to generate a histogram using the defocus amounts of the sideways eye defocus map and the vertical eye defocus map within the face detection area. However, when the area of ​​the defocus map within the face detection area is small, the number of defocus amount data points is small, so the frequency of the histogram using the defocus amounts is low overall, making it difficult to extract the range of defocus amounts where the frequency is greatest.

[0166] Therefore, when creating a histogram, the necessary number of defocus amount data or person detection area is determined, and it is determined whether the number of defocus amount data or person detection area is equal to or greater than a predetermined value. If the number of defocus amount data or person detection area is less than the predetermined value, the person detection area is expanded so that the number of defocus amount data is equal to or greater than the predetermined value. Furthermore, there is a difference between the area of ​​the side-eye defocus map and the area of ​​the vertical-eye defocus map. For this reason, a defocus amount histogram may be created using, for example, the side-eye defocus map as the person's upper body detection area and the vertical-eye defocus map as the person's face detection area.

[0167] Also, depending on the person detection area, there may be a defocus amount in the side-eye defocus map. However, there may be no defocus amount in the vertical-eye defocus map. For this reason, when there is only a defocus amount in the side-eye defocus map, the number of defocus amount data may be doubled or left as is, and in areas where both side-eye and vertical-eye defocus maps exist, the number of defocus amounts may be left as is or halved for normalization. In this embodiment, the area of ​​the side-eye defocus map is described as being larger than the area of ​​the vertical-eye defocus map, but the area of ​​the vertical-eye defocus map may also be larger than the area of ​​the side-eye defocus map.

[0168] Next, in step S1705, the camera MPU 125 selects a focus detection area using the defocus amount histogram, which is the focus detection result generated in step S1704, and selects the defocus amount, which is the focus detection result for that area. The defocus amount is selected from the range where the frequency of the defocus amount histogram is the maximum value. There are several selection methods. For example, when selecting the defocus amount closest to the defocus amount, which is the predicted AF processing result of step S406, when selecting the defocus amount of a focus detection area that is close to the pupil detection area, which is the detection area of ​​a person, or when selecting a defocus amount on the near side. Alternatively, a defocus amount histogram may be created for each of multiple detection areas (e.g., upper body, face, pupil, etc.), and the defocus amount may be selected from multiple ranges where the frequency of the histograms of multiple detection areas is the maximum value, or from multiple defocus amounts on the near side. Alternatively, the defocus amount may be calculated by averaging the defocus amounts in the range where the frequency of the defocus amount histogram is the maximum value.

[0169] Next, processing using specific region information will be described with reference to FIGS. 19(a) to 19(d). FIGS. 19(a) to 19(d) are diagrams illustrating an example of the arrangement of a defocus map when occlusion occurs. FIG. 19(a) is an image of the moment when an occluded region (an arm) covers a person's face region, and the subject detection information acquired in step S1701 is indicated by a rectangular frame. FIG. 19(b) represents the specific region (in this embodiment, the face region) acquired in step S1702 as a lattice frame, indicating that the portion covered by the arm has not been detected as a specific region (face region). The specific region information (likelihood) acquired in step S1702 may be information that expresses whether or not the region is a specific region as a binary output result of 1 or 0, or may be information that expresses, for example, one byte ranging from 0 to 255, with the larger the value, the higher the likelihood. Here, since the explanation is based on the former, it is assumed that 1 is output for the lattice frame region and 0 is output for other regions such as the arm.

[0170] 19(c) is a diagram in which only areas of the side-eye defocus map that are valid as specific areas are represented by diagonal lines by associating a 3x3 side-eye defocus map with the specific area. The validity can be determined by determining whether the estimated area occupies a certain percentage or more, for example, 50% or more, within each frame of the defocus map. The range within each frame may also be determined based on parameters used in correlation calculations, such as the shift amount used when calculating the defocus amount.

[0171] 19(d) is a diagram in which only areas of the vertical eye defocus map that are valid as specific areas are represented by diagonal lines by associating a 3x3 vertical eye defocus map with specific areas. The determination of whether or not a specific area is valid is the same as in the case of the horizontal eye defocus map, so a description thereof will be omitted.

[0172] 21(a) to (c) are histograms generated from the defocus maps of FIGS. 19(c) and (d). Because the 3×3 defocus map contains occluded areas, if a histogram is generated for the entire area, the influence of the occluded areas makes it easier to detect histogram peaks nearer than the face. However, by generating a histogram for only specific areas as in this embodiment, it is possible to reduce the influence of occluded areas and background areas.

[0173] As described above, by generating a histogram only in a specific region, it is expected that the influence of occluded regions can be prevented. In addition, in this embodiment, a defocus map with a 3x3 frame has been used for explanation, but this is not limited to this, and the number of frames can be freely set to NxM frames (N and M are integers of 2 or more).

[0174] In this embodiment, the camera MPU 125 determines a first group of focus detection areas for the first focus detection and a second group of focus detection areas for the second focus detection. The camera MPU 125 also changes at least one of the first group of focus detection areas or the second group of focus detection areas depending on the subject detection results by the subject detection unit 130 (whether the entire area and local areas have been detected, the center positions of the entire area and local areas, etc.).

[0175] Preferably, the subject detection unit 130 detects a first detection area (global detection area) of the subject and a second detection area (local detection area) smaller than the first detection area. The camera MPU 125 arranges the first focus detection area group so as to include the center of the first detection area, and arranges the second focus detection area group so as to include the center of the second detection area.

[0176] Preferably, when the subject detection section 130 detects both the first detection area and the second detection area of ​​the subject, the camera MPU 125 arranges the first focus detection area group to include the center of the first detection area, and arranges the second focus detection area group to include the center of the second detection area. Furthermore, when the subject detection section 130 detects only the first detection area or the second detection area, the camera MPU 125 arranges the first focus detection area group and the second focus detection area group to include the center of either the detected first detection area or the center of the second detection area.

[0177] Preferably, the camera MPU 125 makes the center of the first group of focus detection areas different from the center of the second group of focus detection areas, depending on the result of the object detection by the object detection section 130.

[0178] In this embodiment, since the area that the user wants to focus on in the case of a person is often the pupil, the camera MPU 125 sets a vertical eye defocus map (second focus detection area group) centered on the pupil detection area 1803 in Fig. 18(c), for example. However, this embodiment is not limited to this.

[0179] The camera MPU 125 may change at least one of the first focus detection area group or the second focus detection area group according to the area designated by the user, along with the subject detection result by the subject detection unit 130. For example, the subject detection unit 130 may detect the first detection area (whole detection area) of the subject, and the camera MPU 125 may arrange the first focus detection area group to include the center of the first detection area, and arrange the second focus detection area group to include the area designated by the user. With this configuration, if the user wants to focus on an area other than the pupil (for example, a person's hands or feet), the arrangement of the second focus detection area group can be changed according to the area designated by the user. Alternatively, the arrangement of the second focus detection area group may be changed according to user settings such as the shooting mode and shooting scene. This allows for more appropriate focus detection according to the user's purpose.

[0180] In this embodiment, the camera MPU 125 may change the first focus detection area group along with the second focus detection area group in accordance with the subject detection result. Alternatively, the camera MPU 125 may change only the first focus detection area group in accordance with the subject detection result. In other words, the camera MPU 125 only needs to be able to change at least one of the first focus detection area group or the second focus detection area group in accordance with the subject detection result so that the centers of the first focus detection area group and the second focus detection area group differ from each other.

[0181] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0182] According to this embodiment, it is possible to appropriately position defocus maps with different focus detection directions relative to the subject and select the defocus amount for the subject. Therefore, according to this embodiment, it is possible to provide a control device, an imaging device, a control method, a program, and a storage medium that are capable of performing appropriate focus detection according to the subject detection result.

[0183] The disclosure of each embodiment includes the following configurations and methods. (Configuration 1) a focus detection means for performing first focus detection based on a first signal obtained from a pair of pixels arranged in a first direction in the image sensor, and for performing second focus detection based on a second signal obtained from a pair of pixels arranged in a second direction different from the first direction; an object detection means for detecting an object based on an image signal obtained from the imaging element; a determination unit that determines a first focus detection area group for the first focus detection and a second focus detection area group for the second focus detection, The control device is characterized in that the determining means changes at least one of the first group of focus detection areas or the second group of focus detection areas in accordance with a result of the subject detection by the subject detection means. (Configuration 2) The control device described in configuration 1, characterized in that the number of first focus detection areas included in the first focus detection area group is greater than the number of second focus detection areas included in the second focus detection area group. (Configuration 3) The control device described in configuration 1 or 2, characterized in that the density of the second focus detection areas included in the second focus detection area group is higher than the density of the first focus detection areas included in the first focus detection area group. (Configuration 4) the subject detection means detects a first detection area of ​​the subject and a second detection area smaller than the first detection area; The determining means The first focus detection area group is arranged so as to include the center of the first detection area; 4. The control device according to any one of configurations 1 to 3, wherein the second focus detection area group is arranged so as to include the center of the second detection area. (Configuration 5) The determining means When the subject detection means detects both a first detection area of ​​the subject and a second detection area smaller than the first detection area, the first focus detection area group is arranged so as to include the center of the first detection area, and the second focus detection area group is arranged so as to include the center of the second detection area, The control device described in configuration 1, characterized in that when the subject detection means detects only one of the first detection area or the second detection area, the first focus detection area group and the second focus detection area group are arranged so as to include the center of the detected first detection area or the center of the detected second detection area. (Configuration 6) 6. The control device according to any one of configurations 1 to 5, wherein the determining unit causes the center of the first group of focus detection areas to differ from the center of the second group of focus detection areas. (Configuration 7) The control device according to any one of configurations 1 to 6, wherein the determination means changes at least one of the first focus detection area group or the second focus detection area group in accordance with an area designated by a user. (Configuration 8) the subject detection means detects a first detection area of ​​the subject, The determining means The first focus detection area group is arranged so as to include the center of the first detection area; 8. The control device according to configuration 7, wherein the second focus detection area group is arranged so as to include the area designated by the user. (Configuration 9) An imaging device comprising the control device according to any one of configurations 1 to 8 and the imaging element. (Configuration 10) 10. The imaging device according to configuration 9, wherein the imaging element has a plurality of pixels that receive light beams that pass through different pupil partial regions of the imaging optical system. (Configuration 11) 11. The imaging device according to claim 10, wherein the plurality of pixels includes the pair of pixels arranged in the first direction and the pair of pixels arranged in the second direction. (Configuration 12) the first direction is a horizontal direction of the imaging device, 12. The imaging device according to claim 11, wherein the second direction is a vertical direction of the imaging device. (Method 1) a focus detection step of performing first focus detection based on a first signal obtained from a pair of pixels arranged in a first direction in the image sensor, and performing second focus detection based on a second signal obtained from a pair of pixels arranged in a second direction different from the first direction; a subject detection step of detecting a subject based on an image signal obtained from the imaging element; a determination step of determining a first focus detection area group for the first focus detection and a second focus detection area group for the second focus detection, A control method characterized in that in the determining step, at least one of the first group of focus detection areas or the second group of focus detection areas is changed depending on the subject detection result in the subject detection step. (Configuration 13) A program that causes a computer to execute the control method described in Method 1. (Configuration 14) A computer-readable storage medium storing the program according to configuration 13.

[0184] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]

[0185] 120 Camera body (control device) 122 Image sensor 125 Camera MPU (Decision Means) 129 Phase difference AF section (focus detection means) 130 Subject detection unit (subject detection means)

Claims

1. a focus detection means (129) that performs first focus detection based on a first signal obtained from a pair of pixels arranged in a first direction in the image sensor (122), and performs second focus detection based on a second signal obtained from a pair of pixels arranged in a second direction different from the first direction; an object detection means (130) for detecting an object based on an image signal obtained from the imaging element; a determination means (125) for determining a first focus detection area group for the first focus detection and a second focus detection area group for the second focus detection, The control device according to claim 1, wherein the determining means changes at least one of the first group of focus detection areas or the second group of focus detection areas in accordance with a result of the subject detection by the subject detection means.

2. 2. The control device according to claim 1, wherein the number of first focus detection areas included in the first focus detection area group is greater than the number of second focus detection areas included in the second focus detection area group.

3. 2. The control device according to claim 1, wherein the density of the second focus detection areas included in the second focus detection area group is higher than the density of the first focus detection areas included in the first focus detection area group.

4. the subject detection means detects a first detection area of ​​the subject and a second detection area smaller than the first detection area, The determining means The first focus detection area group is arranged so as to include the center of the first detection area; 2. The control device according to claim 1, wherein the second focus detection area group is arranged so as to include the center of the second detection area.

5. The determining means When the subject detection means detects both a first detection area of ​​the subject and a second detection area smaller than the first detection area, the first focus detection area group is arranged so as to include the center of the first detection area, and the second focus detection area group is arranged so as to include the center of the second detection area, 2. The control device according to claim 1, wherein, when the subject detection means detects only one of the first detection area or the second detection area, the first focus detection area group and the second focus detection area group are arranged so as to include the center of the detected first detection area or the center of the detected second detection area.

6. 2. The control device according to claim 1, wherein the determining means causes the center of the first group of focus detection areas to differ from the center of the second group of focus detection areas.

7. 2. The control device according to claim 1, wherein the determining means changes at least one of the first group of focus detection areas or the second group of focus detection areas in accordance with an area designated by a user.

8. the subject detection means detects a first detection area of ​​the subject; The determining means The first focus detection area group is arranged so as to include the center of the first detection area; The control device according to claim 7 , wherein the second focus detection area group is arranged so as to include the area designated by the user.

9. An imaging device comprising: the control device according to claim 1; and the imaging element.

10. 10. The imaging apparatus according to claim 9, wherein the imaging element has a plurality of pixels that receive light beams that pass through different pupil partial regions of the imaging optical system.

11. The imaging device according to claim 10 , wherein the plurality of pixels includes the pair of pixels arranged in the first direction and the pair of pixels arranged in the second direction.

12. the first direction is a horizontal direction of the imaging device, The imaging device according to claim 11 , wherein the second direction is a vertical direction of the imaging device.

13. a focus detection step of performing first focus detection based on a first signal obtained from a pair of pixels arranged in a first direction in the image sensor, and performing second focus detection based on a second signal obtained from a pair of pixels arranged in a second direction different from the first direction; a subject detection step of detecting a subject based on an image signal obtained from the imaging element; a determination step of determining a first focus detection area group for the first focus detection and a second focus detection area group for the second focus detection, a control method comprising: changing, in the determining step, at least one of the first group of focus detection areas or the second group of focus detection areas according to the subject detection result in the subject detection step;

14. A program causing a computer to execute the control method according to claim 13.

15. A computer-readable storage medium storing the program according to claim 14.

Citation Information

Patent Citations

  • Imaging apparatus, method for controlling imaging apparatus, and program

    JP2022129931A

  • Electronic instrument, and control method of the same

    JP2023081041A

  • Controller, control method, and control program

    WO2017135276A1

  • Image processing device, image processing method, program, and imaging device

    WO2020195073A1

  • Depth map output device

    JP2011237215A