Control apparatus, image capturing apparatus, control method, and storage medium
The control device addresses the challenge of focusing on a main subject obscured by occlusions by using multiple defocus amounts and specific area information for precise focus adjustment, ensuring accurate focusing despite partial obstructions.
Patent Information
- Application Number
- JP2024104336
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-16
AI Technical Summary
Existing autofocus systems struggle to accurately focus on a main subject when part of the face is obscured by occlusions, such as arms, due to the need for individual tracking of occluded subjects and inappropriate determination of phase difference information.
A control device that acquires multiple defocus amounts and specific area information to select a defocus amount for focus adjustment, using a control unit to accurately focus on the main subject even when partially hidden.
Enables accurate focusing on a main subject despite occlusions by employing a control device that considers multiple defocus amounts and specific area information for precise focus adjustment.
Smart Images

Figure 2026005779000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a control device capable of autofocus (AF). [Background technology]
[0002] In a method of automatically adjusting the focus by simply pointing a camera at a subject such as a person's face, for example, in a sports scene, the face may not be in focus due to the influence of an obscuring subject such as a hand or a stick when the face is obscured. Patent Document 1 discloses a configuration in which a main subject and an object in front of the main subject (obscuring subject) are tracked separately, and focus adjustment is performed at a focus detection point where only the main subject is present. Patent Document 2 also discloses a configuration in which, when the direction of pupil division differs from the direction of the distribution of obscured areas, phase difference information of the obscured areas is excluded to control focus adjustment. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-202875 [Patent Document 2] Japanese Patent Publication No. 2022-125743 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the configuration of Patent Document 1 requires tracking occluded subjects individually, and it is not possible to eliminate the influence of occlusions within the main subject, such as the arms of the main subject. Also, the configuration of Patent Document 2 requires appropriate determination of which phase difference information should be used for focus adjustment from the phase difference information from which the phase difference information of the occluded area has been excluded, otherwise it is not possible to accurately focus on the main subject.
[0005] An object of the present invention is to provide a control device that can accurately focus even when, for example, part of a face is hidden. [Means for solving the problem]
[0006] A control device according to one aspect of the present invention is a control device for controlling focus adjustment, characterized by having an acquisition unit that acquires multiple defocus amounts acquired in multiple focus detection areas and information regarding a specific area detected from a subject area in an imaging area, and a control unit that selects a first defocus amount for controlling focus adjustment in accordance with the multiple defocus amounts and the information regarding the specific area. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide an image generating device that can accurately focus even when, for example, part of a face is hidden. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing the configuration of a camera system according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing a pixel array of an image sensor according to the first embodiment. [Figure 3] FIG. 2 is an explanatory diagram of a pixel according to the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating pupil division according to the first embodiment. [Figure 5] FIG. 4 is another diagram showing pupil division according to the first embodiment. [Figure 6] 10 is a diagram showing the relationship between the amount of image shift and the amount of defocus in the first embodiment. FIG. [Figure 7] FIG. 2 is a layout diagram of focus detection areas according to the first embodiment. [Figure 8] FIG. 2 is a diagram showing an overall flow of live view shooting processing in the first embodiment. [Figure 9] 10 is a flowchart of a photographing subroutine in the first embodiment. [Figure 10] 10 is a flowchart of subject tracking AF processing in the first embodiment. [Figure 11] 4 is a flowchart of a subject detection and tracking process according to the first embodiment. [Figure 12]FIG. 2 is a diagram illustrating an example of a CNN that infers the likelihood of a specific region according to the first embodiment. [Figure 13] 10 is a flowchart of a flicker determination process according to the first embodiment. [Figure 14] 10A and 10B are diagrams for explaining the influence of flicker on a pair of signals for vertical eye focus detection in the first embodiment. [Figure 15] FIG. 10 is a diagram showing waveforms when flicker occurs in Example 1. [Figure 16] FIG. 10 is another diagram showing waveforms when flicker occurs in Example 1. [Figure 17] 10 is a flowchart of a defocus amount selection process according to the first embodiment. [Figure 18] FIG. 4 is a diagram illustrating a method for setting a defocus map in the first embodiment. [Figure 19] FIG. 10 is another diagram showing a method for setting a defocus map according to the first embodiment. [Figure 20] FIG. 10 is a diagram showing a histogram of a defocus map in the first embodiment. [Figure 21] FIG. 10 is a diagram showing a histogram generated from a defocus map using specific region information in the first embodiment. [Figure 22] 10 is a flowchart of focus detection processing according to the first embodiment. [Figure 23] FIG. 4 is a diagram showing the execution order of subject tracking AF processing in the first embodiment. [Figure 24] 10 is a flowchart of a defocus amount selection process according to the second embodiment. [Figure 25] FIG. 10 is a diagram illustrating a recommended direction determination in the second embodiment. [Figure 26] 10 is a flowchart of a selection process of a focus detection area according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same reference numerals are used to designate the same components, and redundant explanations will be omitted. [Example 1] 1 is a block diagram showing the configuration of an imaging system 10 including a camera body (imaging device) 120 of this embodiment. A lens unit (interchangeable lens) 100 is detachably attached to the camera body 120, which is a digital camera, via a mount M indicated by a dotted line in the figure. The camera body 120 may also be one in which an imaging optical system is integrally provided. Furthermore, the camera body 120 is not limited to a digital camera, and may be another imaging device such as a video camera.
[0010] The lens unit 100 includes an imaging optical system that includes a first lens group 101, an aperture 102, a second lens group 103, a focus lens group (hereinafter simply referred to as a focus lens) 104 as a focus element, and a drive / control system. The imaging optical system captures light from a subject and forms an image of the subject.
[0011] The first lens group 101 is disposed closest to the object and is held so as to be movable in the direction of the optical axis OA. The diaphragm 102 adjusts the amount of light by changing its aperture diameter. The diaphragm 102 and the second lens group 103 are movable together in the direction of the optical axis, and perform zooming by moving in conjunction with the first lens group 101. The focus lens 104 performs forcing by moving in the direction of the optical axis. Automatic focus adjustment (AF) is performed by controlling the position of the focus lens 104 in accordance with the focus detection results described below.
[0012] The drive / control system includes a zoom actuator 111 , an aperture actuator 112 , a focus actuator 113 , a zoom drive circuit 114 , an aperture drive circuit 115 , a focus drive circuit 116 , a lens MPU 117 , and a lens memory 118 .
[0013] During zooming, a zoom drive circuit 114 drives a zoom actuator 111 to move the first lens group 101 and the third lens group 103 in the optical axis direction. An aperture drive circuit 115 drives an aperture actuator 112 to operate the aperture 102, thereby performing aperture operation and shutter operation.
[0014] During focusing, the focus drive circuit 116 drives the focus actuator 113 to move the focus lens 104 in the optical axis direction. The focus drive circuit 116 functions as a position detection unit that detects the current position of the focus lens 104 (hereinafter referred to as the focus position) through the focus actuator 113.
[0015] The lens MPU 117 is a computer that executes calculations and processing related to the lens unit 100, and controls the zoom drive circuit 114, the aperture drive circuit 115, and the focus drive circuit 116. The lens MPU 117 is also communicatively connected to a camera MPU 125 in the camera body 120 via a communication terminal of the mount M, and exchanges commands and data with the camera MPU 125. For example, the lens MPU 117 notifies the camera MPU 125 of lens information in response to a request from the camera MPU 125. This lens information includes information such as the focus position, the position and diameter of the exit pupil of the imaging optical system in the optical axis direction, and the position and diameter of the lens frame that limits the luminous flux of the exit pupil in the optical axis direction.
[0016] Furthermore, the lens MPU 117 controls the zoom drive circuit 114, the aperture drive circuit 115, and the focus drive circuit 116 in response to requests from the camera MPU 125. The lens memory 118 stores optical information necessary for AF. The camera MPU 125 controls the lens unit 100 by executing programs stored in the built-in nonvolatile memory and the lens memory 118.
[0017] The camera body 120 has an optical low-pass filter 121, an image sensor 122, and a drive / control system. The optical low-pass filter 121 is provided to reduce false colors and moire.
[0018] The image sensor 122 is composed of a CMOS sensor and its peripheral circuits, and photoelectrically converts the subject image (optical image) formed by the imaging optical system, outputting an image signal and a pair of focus detection signals (two image signals). The image sensor 122 has a plurality of image pixels, m pixels in the horizontal direction and n pixels in the vertical direction perpendicular to the horizontal direction (m and n are integers of 2 or greater). Each image pixel includes a pair of focus detection pixels, as described below, and has a pupil division function that enables focus detection using a phase difference detection method.
[0019] The drive / control system includes an image sensor drive circuit 123, a shutter 133, an image processing circuit 124, and a camera MPU (control unit) 125. The drive / control system also includes a display 126, an operation switch (SW) 127, a memory 128, a phase difference AF unit (focus detection means) 129, a subject detection unit (subject detection means) 130, an AE unit 131, and a white balance (WB) adjustment unit 132. The image sensor drive circuit 123 controls charge accumulation and signal readout in the image sensor 122, and also A / D converts the image signal and the paired focus detection signal output from the image sensor 122 and outputs them to the image processing circuit 124 and the camera MPU 125. The image processing circuit 124 performs image processing such as gamma conversion, color interpolation, and compression encoding on the digital image signal from the image sensor drive circuit 123 to generate image data.
[0020] The camera MPU 125 is a computer that executes calculations and processing related to the camera body 120, and controls the image sensor drive circuit 123, the image processing circuit 124, the display 126, the phase difference AF unit 129, the subject detection unit 130, the AE unit 131, and the WB adjustment unit 132. The camera MPU 125 is also communicatively connected to the lens MPU 117 via a communication terminal of the mount M, and exchanges commands and data with the lens MPU 117. For example, the camera MPU 125 requests lens information and optical information from the lens MPU 117, and requests the lens MPU 117 to drive the first lens group 101, the focus lens 104, and the aperture 102. The camera MPU 125 receives the lens information and optical information transmitted from the lens MPU 117.
[0021] The camera MPU 125 has a built-in ROM 125a for storing various programs, a RAM 125b for storing variables, and an EEPROM 125c for storing various parameters. The camera MPU 125 executes various processes including the AF process described below in accordance with the programs stored in the ROM 125a. The camera MPU 125 generates two-image data from a pair of digital focus detection signals from the image sensor drive circuit 123 and outputs the data to a phase-difference AF unit 129.
[0022] The shutter 133 has a focal plane shutter configuration, and drives the focal plane shutter in response to a command from a shutter drive circuit built into the shutter 133 based on instructions from the camera MPU 125. The image sensor 122 is shielded from light while a signal from the image sensor 122 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened, and a photographing light beam is guided to the image sensor 122.
[0023] The display 126 is configured with an LCD or the like, and displays information about the imaging mode, a preview image before imaging, a confirmation image after imaging, the focus state, etc. The operation switches 127 include a power switch, a release (imaging instruction) switch, a zoom switch, an imaging mode selection switch, etc. The memory 128 is a flash memory that is detachable from the camera body 120, and stores images for recording obtained by imaging.
[0024] The phase-difference AF unit 129 performs focus detection using the two-image data generated by the camera MPU 125. The image sensor 122 photoelectrically converts a pair of optical images formed by light beams that have passed through a pair of different pupil regions of the exit pupil of the imaging optical system, and outputs a pair of focus detection signals. The phase-difference AF unit 129 performs a correlation operation on the two-image data generated by the camera MPU 125 from the pair of focus detection signals, and calculates the amount of image shift, which is the phase difference between them. The phase-difference AF unit 129 also calculates (acquires) a defocus amount as information related to the focus from the calculated image shift amount. The camera MPU 125 calculates the drive amount of the focus lens 104 based on the defocus amount calculated by the phase-difference AF unit 129, and transmits a focus control command including the drive amount to the lens MPU 117.
[0025] As described above, in this embodiment, image plane phase difference AF is performed using the output of the image sensor 122, without using an AF sensor dedicated to focus detection. In this embodiment, the phase difference AF unit 129 has an acquisition unit 129a that acquires two-image data and a calculation unit (calculation means) 129b that calculates the defocus amount. At least one of the acquisition unit 129a and the calculation unit 129b may be provided in the camera MPU 125.
[0026] The subject detection unit 130 performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 130 uses dictionary data for each subject to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 130 performs subject detection by sequentially switching between dictionary data for each subject. The dictionary data for each subject is stored in a dictionary data storage unit (ROM 125a in the camera MPU 125). Therefore, multiple dictionary data are stored in the dictionary data storage unit. The camera MPU 125 determines which dictionary data from the multiple dictionary data to use for subject detection based on the subject priorities set in advance and the settings of the camera body 120.
[0027] The AE unit 131 performs exposure control (AE) by performing photometry using image data for AE obtained from the image processing circuit 124. Specifically, the AE unit 131 acquires brightness information of the image data for AE, and calculates the aperture value, shutter speed (shutter time), and ISO sensitivity as imaging conditions from the difference between the exposure amount obtained from this brightness information and a preset exposure amount. Then, AE is performed by controlling the aperture value, shutter speed, and ISO sensitivity to the calculated values.
[0028] The WB adjustment unit 132 calculates the WB of the image data for WB adjustment obtained from the image processing circuit 124, and performs WB adjustment by adjusting the weights of RGB colors according to the difference between the calculated WB and a predetermined appropriate WB.
[0029] Furthermore, the camera MPU 125 can select the image height range for performing phase difference AF, AE, and WB adjustment according to the position, size, etc. of the subject in the imaging area detected by the subject detection unit 130. (Regarding the image sensor 122) 2A and 2B are diagrams showing the pixel arrangement on the imaging surface of the image sensor 122 serving as a two-dimensional CMOS sensor in this embodiment. Fig. 2A is a diagram showing an example of the overall configuration of the image sensor 122. The image sensor 122 includes a pixel array unit 208, a vertical selection circuit 209, a column circuit 203, and a horizontal selection circuit 204.
[0030] The pixel array unit 208 has a plurality of pixels 205 arranged in a matrix. The output of a vertical selection circuit 209 is input to the pixels 205 via a pixel drive wiring group 207, and pixel signals from the pixels 205 in a row selected by the vertical selection circuit 209 are read out row by row to the column circuit 203 via output signal lines 206. One output signal line 206 may be provided for each pixel column, for each set of multiple pixel columns, or multiple output signal lines 206 may be provided for each pixel column. The column circuit 203 receives signals read out in parallel via the multiple output signal lines 206, performs signal amplification, noise reduction, A / D conversion, and other processing, and stores the processed signals. The horizontal selection circuit 204 sequentially, randomly, or simultaneously selects signals stored in the column circuit 203, and the selected signals are output to the outside of the image sensor 122 via horizontal output lines and an output unit (not shown).
[0031] In this way, by sequentially outputting pixel signals of the row selected by the vertical selection circuit 209 to the outside of the image sensor 122 while changing the row selected by the vertical selection circuit 209, a two-dimensional image signal or phase difference signal can be read out from the image sensor 122.
[0032] FIG. 2(b) is an equivalent circuit diagram of the pixel 205. The pixel 205 has two photodiodes 211 (PDA) and 212 (PDB) that are photoelectric conversion units. The PDA 211 performs photoelectric conversion according to the amount of incident light, and the accumulated signal charge is transferred via a transfer switch (TXA) 213 to a floating diffusion (FD) 215 that constitutes a charge accumulation unit. The PDB 212 performs photoelectric conversion and accumulates the signal charge, and the transferred signal charge is transferred via a transfer switch (TXB) 214 to the FD 215. When a reset switch (RES) 216 is turned on, it resets the FD 215 to the voltage of a constant voltage source VDD. When the RES 216, TXA 213, and TXB 214 are turned on simultaneously, the PDA 211 and PDB 212 can be reset.
[0033] When a selection switch (SEL) 217 that selects a pixel is turned on, an amplification transistor (SF) 218 converts the signal charge accumulated in the FD 215 into a voltage, and the converted signal voltage is output from the pixel to an output signal line 206. In addition, gates of the TXA 213, TXB 214, RES 216, and SEL 217 are each connected to a pixel drive wiring group 207 and controlled by a vertical selection circuit 209.
[0034] In the following description, in this embodiment, the signal charge accumulated in the photoelectric conversion unit is assumed to be electrons, the photoelectric conversion unit is formed of an N-type semiconductor, and is separated by a P-type semiconductor; however, the signal charge may be holes, the photoelectric conversion unit may be formed of a P-type semiconductor, and is separated by an N-type semiconductor.
[0035] Next, we will explain the operation of reading signal charges from the PDA211 and PDB212 after a predetermined charge accumulation time has elapsed after resetting the PDA211 and PDB212 in a pixel having the above-mentioned configuration. First, the SEL217 of the row selected by the vertical selection circuit 209 is turned on, connecting the source of the SF218 to the output signal line 206, and the output signal line 206 enters a state in which a voltage corresponding to the voltage of the FD215 is read out. Next, the RES216 is turned on / off, resetting the potential of the FD215. After that, the system waits until the output signal line 206, which has been subjected to voltage fluctuations in the FD215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as a signal voltage N, where it is processed and held.
[0036] Thereafter, the TXA 213 is turned on / off, and the signal charge accumulated in the PDA 211 is transferred to the FD 215. The voltage of the FD 215 drops by an amount corresponding to the amount of signal charge accumulated in the PDA 211. Thereafter, the system waits until the output signal line 206, which has been subjected to the voltage fluctuation of the FD 215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as signal voltage A, which is then subjected to signal processing and held.
[0037] Thereafter, the TXB 214 is turned on / off, and the signal charge accumulated in the PDB 212 is transferred to the FD 215. The voltage of the FD 215 drops by an amount corresponding to the amount of signal charge accumulated in the PDB 212. Thereafter, the system waits until the output signal line 206, which has been subjected to the voltage fluctuation of the FD 215, settles, and the settled voltage of the output signal line 206 is taken in by the column circuit 203 as a signal voltage (A+B), which is then subjected to signal processing and held.
[0038] From the difference between the signal voltage N and the signal voltage A thus captured, a signal A corresponding to the amount of signal charge accumulated in the PDA 211 can be obtained. Furthermore, from the difference between the signal voltage A and the signal voltage (A+B), a signal B corresponding to the amount of signal charge accumulated in the PDB 212 can be obtained. This difference calculation may be performed by the column circuit 203, or may be performed after output from the image sensor 122. A phase difference signal can be obtained by using the signals A and B, and an image signal can be obtained by adding the signals A and B together. Alternatively, if the difference calculation is performed after output from the image sensor 122, the image signal may be obtained by taking the difference between the signal voltage N and the signal voltage (A+B).
[0039] Alternatively, the signal voltages N, A, and B may be read out by performing the same driving as that for reading out the signal voltages N and A on the PDB 212 instead of the PDA 211. In this case, the signals A and B obtained from the signal voltages A and B, respectively, can be used as phase difference signals as they are, and an imaging signal can be obtained by adding together the signal voltages A and B or the signals A and B.
[0040] In this embodiment, the pixel from which the above-mentioned signal A is obtained is called the first focus detection pixel, and the pixel from which the signal B is obtained is called the second focus detection pixel.
[0041] FIG. 2(c) shows an arrangement of imaging pixels in a range of 4 columns x 4 rows. One pixel group 200 includes imaging pixels arranged in 2 columns x 2 rows. The pixel group 200 includes a pixel 200R located in the upper left and having a spectral sensitivity of R (red), pixels 200Ga and 200Gb located in the upper right and lower left and having a spectral sensitivity of G (green), and a pixel 200B located in the lower right and having a spectral sensitivity of B (blue). Each imaging pixel is composed of a first focus detection pixel 201 and a second focus detection pixel 202. In the pixels 200R, 200Ga, and 200B, the first focus detection pixel 201 and the second focus detection pixel 202 are arranged horizontally, while in the pixel 200Gb, the first focus detection pixel 201 and the second focus detection pixel 202 are arranged vertically.
[0042] FIG. 3 is an explanatory diagram of a pixel. FIG. 3(a) shows pixel 200Ga when viewed from the incident side (+z side) of the image sensor 122. FIG. 3(b) shows the pixel structure when the aa cross section of FIG. 3(a) is viewed from the -y side. In pixel 200Ga, a microlens 305 for collecting incident light is formed on the incident side, and photoelectric conversion units 301 and 302 divided into two in the x direction are formed. The photoelectric conversion units 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively.
[0043] The photoelectric conversion units 301 and 302 may be pin structure photodiodes with an intrinsic layer sandwiched between p-type and n-type layers, or may be pn junction photodiodes without the intrinsic layer. A color filter 306 is formed between the microlens 305 and the photoelectric conversion units 301 and 302. The spectral transmittance of the color filter may be different for each focus detection pixel, or the color filter may be omitted.
[0044] Two light beams incident on pixel 200Ga from the paired pupil regions are each collected by microlens 305, dispersed by color filter 306, and then received by photoelectric conversion units 301 and 302. In each photoelectric conversion unit, electrons and holes are generated in pairs according to the amount of received light, and after being separated by a depletion layer, the negatively charged electrons are accumulated in the n-type layer. Meanwhile, the holes are discharged to the outside of image sensor 122 through a p-type layer connected to a constant voltage source (not shown). The electrons accumulated in the n-type layer of each photoelectric conversion unit are transferred to a capacitance unit (FD) via a transfer gate and converted into a voltage signal.
[0045] FIG. 4 is a diagram showing pupil division. The lower part of FIG. 4 shows the pixel structure when the aa cross section of FIG. 3(a) is viewed from the +y side, and the upper part shows the pupil plane at pupil distance DS. Note that in FIG. 4, the x-axis and y-axis of the pixel structure are inverted relative to FIG. 3(b) in order to correspond to the coordinate axes of the pupil plane. The pupil plane corresponds to the entrance pupil position of the image sensor 122. In this embodiment, the microlens position in each pixel is offset (shrunk) from the center of the image sensor 122, so that the entrance pupils of each pixel overlap each other to form the entrance pupil of a single image sensor 122. The pupil distance DS is the distance between the pupil plane and the image sensor, and will be referred to as the sensor pupil distance in the following description.
[0046] As shown in FIG. 4, the first pupil region 501 of the first focus detection pixel 201 is generally conjugate by a microlens to the light receiving surface of the photoelectric conversion unit 301, whose center of gravity is decentered in the -x direction. The first pupil region 501 is a pupil region through which a light beam that can be received by the first focus detection pixel 201 passes. The center of gravity of the first pupil region 501 is decentered on the +X side on the pupil plane. Furthermore, the second pupil region 502 of the second focus detection pixel 202 is generally conjugate by a microlens to the light receiving surface of the photoelectric conversion unit 302, whose center of gravity is decentered in the +x direction. The second pupil region 502 is a pupil region through which a light beam that can be received by the second focus detection pixel 202 passes. The center of gravity of the second pupil region 502 is decentered on the -X side on the pupil plane. The pupil region 500 is a pupil region through which a light beam that can be received by the entire pixel 200G, which is a combination of the photoelectric conversion units 301 and 302 (the first focus detection pixel 201 and the second focus detection pixel 202), passes.
[0047] As shown in FIG. 5, light beams that enter the imaging optical system from the subject (the vertical line on the left side of the figure) and pass through the first pupil region 501 and the second pupil region 502 are incident on each imaging pixel at different angles and are received by the photoelectric conversion units 301 and 302. The pixels 200R, 200Ga, and 200B perform pupil division in the horizontal direction (x-axis direction in FIG. 4), while the pixel 200Gb performs pupil division in the vertical direction (y-axis direction in FIG. 4). Each imaging pixel, which has a first focus detection pixel and a second focus detection pixel, receives light beams that pass through the first pupil region 501 and the second pupil region 502. A pair of focus detection signals is generated by combining the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of the multiple imaging pixels. In addition, an imaging signal with a resolution of the number of effective pixels N (= m × n) is generated by adding the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of the multiple imaging pixels. Note that one of the pair of focus detection signals may be subtracted from the imaging signal to generate the other focus detection signal.
[0048] In addition, in this embodiment, a first and a second focus detection pixel are provided for each of all imaging pixels of the image sensor 122, but two imaging pixels may be used as the first and second focus detection pixels, or first and second focus detection pixels may be provided for some imaging pixels. (Relationship between defocus amount and image shift amount) 6 is a diagram showing the relationship between the amount of image shift between two image data and the amount of defocus. 800 indicates the imaging plane of the image sensor 122, and the pupil plane of the image sensor 122 is divided into a first pupil region 501 and a second pupil region 502. The amount of defocus d is defined as |d|, which is the distance from the imaging position (image position) of the subject image to the imaging plane 800, with a negative sign (d<0) when the image position is in front focus on the subject side of the imaging plane, and a positive sign (d>0) when the image position is in back focus on the opposite side of the subject from the imaging plane 800. The in-focus state in which the image position is on the imaging plane 800 is d=0.
[0049] 6, subject 801 shows a focused state (d=0), and subject 802 shows a front-focused state (d<0). The front-focused state (d<0) and the back-focused state (d>0) are combined to form a defocused state (|d|>0).
[0050] In a front-focus state, light beams from the subject 802 that pass through the first pupil region 501 and the second pupil region 502 are focused once and then spread to widths Γ1 and Γ2 centered at the center of gravity positions G1 and G2 of the light beams, forming blurred optical images on the imaging plane 800. These blurred images are received by the first focus detection pixel 201 and the second focus detection pixel 202 in each imaging pixel on the imaging plane 800, which generate a pair of focus detection signals, a first focus detection signal and a second focus detection signal. The first focus detection signal and the second focus detection signal are recorded as blurred images of the subject 802 at the center of gravity positions G1 and G2 on the imaging plane 800, with the subject 802 spreading over blur widths Γ1 and Γ2, respectively. The blur widths Γ1 and Γ2 increase approximately in proportion to an increase in the magnitude of the defocus amount d, |d|. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 in the center of gravity positions of the light beams) between the first focus detection signal and the second focus detection signal also increases roughly in proportion to the increase in the magnitude |d| of the defocus amount d. The same is true in the back-focus state (d>0), although the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state.
[0051] In this embodiment, the difference in the centers of gravity of the incident angle distributions in the first pupil region 501 and the second pupil region 502 is referred to as the base line length. The relationship between the defocus amount d and the image shift amount p on the imaging plane 800 is roughly similar to the relationship between the base line length and the sensor pupil distance. Because the magnitude of the image shift amount between the first focus detection signal and the second focus detection signal increases as the magnitude of the defocus amount d increases, the phase difference AF section 129 converts the image shift amount into a defocus amount using a conversion coefficient calculated based on the base line length, based on this relationship.
[0052] In the following description, calculating the defocus amount using a pair of focus detection signals from focus detection pixels that divide the pupil horizontally (horizontally) like pixel 200Ga is referred to as horizontal eye focus detection (first detection), and calculating the defocus amount using a pair of focus detection signals from focus detection pixels that divide the pupil vertically (vertically) like pixel 200Gb is referred to as vertical eye focus detection (second detection). (Focus detection area placement) Next, the focus detection area, which is an area of the image sensor 122 where paired signal sequences for detecting a phase difference are acquired, will be described with reference to FIG. 7. In this embodiment, the camera MPU 125 sets the focus detection area. FIG. 7 is a layout diagram of the focus detection area in this embodiment. A(n,m) and B(n,m) indicate the nth focus detection area in the x direction and the mth focus detection area in the y direction among the multiple focus detection areas (three in the x direction and three in the y direction, for a total of nine) set in the effective pixel area 300 of the image sensor 122. A signal sequence of pixel pairs divided into horizontal pupils is generated from the multiple pixels included in the focus detection area A(n,m). A signal sequence of pixel pairs divided into vertical pupils is generated from the multiple pixels included in the focus detection area B(n,m). I(n,m) indicates an index that displays the position of the focus detection area A(n,m) or B(n,m) on the display 126. By arranging the focus detection areas in this way, focus detection can be performed at the position of the index I(n,m) using contrast information corresponding to both the horizontal and vertical directions of the subject.
[0053] The nine focus detection areas shown in FIG. 7 are merely examples, and the number, position, and size of the focus detection areas are not limited. For example, one or more focus detection areas may be set within a predetermined range centered on a position specified by the user or the subject position detected by the subject detection unit 130. In this embodiment, when acquiring a defocus map (described later), the focus detection areas are arranged to obtain focus detection results with higher resolution. For example, a group of focus detection results obtained from side-eye focus detection areas arranged on the image sensor 122 in 17 horizontal and 11 vertical divisions, totaling 187 points, is used as the side-eye defocus map. Furthermore, a group of focus detection results obtained from vertical-eye focus detection areas arranged in 7 horizontal and 5 vertical divisions, totaling 35 points, is used as the vertical-eye defocus map. The method for arranging the focus detection areas for side-eye focus detection and vertical-eye focus detection relative to the subject will be described in detail below. (Photography processing) 8 is a diagram showing the overall flow of the live view shooting process of this embodiment. Specifically, it shows the process of causing the camera body 120 to perform operations from before capturing a live view image on the display 126 to capturing a still image. The camera MPU 125, which is a computer, executes this process in accordance with a computer program. In the following explanation, "S" means "step."
[0054] In S1, the camera MPU 125 causes the image sensor drive circuit 123 to drive the image sensor 122 and acquires image data from the image sensor 122. The camera MPU 125 then acquires first and second focus detection signals from a plurality of first and second focus detection pixels included in each of the focus detection areas shown in Fig. 7 from the acquired image data. The camera MPU 125 also adds the first and second focus detection signals from all effective pixels of the image sensor 122 to generate an image signal, and causes the image processing circuit 124 to perform image processing on the image signal (image data) to acquire image data. Note that if the image sensor pixels and the first and second focus detection pixels are provided separately, the camera MPU 125 acquires image data by performing interpolation processing on the focus detection pixels.
[0055] In S2, the camera MPU 125 causes the image processing circuit 124 to generate a live view image from the image data obtained in S1 and displays it on the display 126. The live view image is a reduced image matched to the resolution of the display 126, and the user can adjust the image composition, exposure conditions, etc. while viewing this image. Therefore, the AE unit 131 and camera MPU 125 adjust the exposure based on the photometric value obtained from the image data and display it on the display 131. The exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the shooting lens, and adjusting the gain of the image sensor 122 output.
[0056] In S3, the camera MPU 125 determines whether or not a switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing a release switch included in the operation switch 127. If Sw1 is not turned on, the camera MPU 125 repeats the determination in S3 to monitor the timing at which Sw1 will be turned on. On the other hand, if Sw1 is turned on, the camera MPU 125 proceeds to S400 and performs subject tracking autofocus (AF) processing. Here, the camera MPU 125 performs predictive AF processing to detect the subject area from the obtained image capture signal and focus detection signal, set the focus detection area, and reduce the effect of the time lag between the focus detection processing and the image capture processing for recording.
[0057] In S5, the camera MPU 125 determines whether or not the switch Sw2, which instructs the start of an imaging operation, has been turned on by fully pressing the release switch. If Sw2 has not been turned on, the camera MPU 125 returns to S3. On the other hand, if Sw2 has been turned on, the camera MPU 125 proceeds to S300 and executes an imaging subroutine. Details of the imaging subroutine will be described later.
[0058] In S7, the camera MPU 125 determines whether or not the main switch included in the operation SW 127 has been turned off. If the main switch has been turned off, the camera MPU 125 ends this process, and if the main switch has not been turned off, the process returns to S3.
[0059] In this embodiment, the subject detection process and AF process are performed after it is detected that Sw1 is turned on in S3, but the timing of performing these processes is not limited to this. By performing the subject tracking AF process in S400 before Sw1 is turned on, it is possible to eliminate the need for the photographer to take preparatory actions before shooting. (About the shooting subroutine) The photographing subroutine executed by the camera MPU 125 in S300 of Fig. 8 will be described with reference to Fig. 9. Fig. 9 is a flowchart of the photographing subroutine.
[0060] In S301, the AE unit 131 performs exposure control processing to determine imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.) This exposure control processing can be performed using brightness information acquired from image data of a live view image.
[0061] Then, the camera MPU 125 transmits the determined aperture value to the aperture drive circuit 115 to drive the aperture 102. The camera MPU 125 also transmits the determined shutter speed to the shutter 133 to open the focal plane shutter. Furthermore, the camera MPU 125 causes the image sensor 122 to accumulate charge during the exposure period via the image sensor drive circuit 123.
[0062] In S302, the image sensor drive circuit 123 reads out all pixels of the image sensor 122 image signal for capturing a still image. The camera MPU 125 also causes the image sensor drive circuit 123 to read out one of the first and second focus detection signals from the focus detection area (focus target area) within the image sensor 122. By subtracting one of the first and second focus detection signals from the image sensor signal, the other focus detection signal can be obtained.
[0063] In S303, the camera MPU 125 causes the image processing circuit 124 to perform defective pixel correction processing on the imaging data that was read out and A / D converted in S302.
[0064] In S304, the camera MPU 125 causes the image processing circuit 124 to perform image processing and encoding processes such as demosaic (color interpolation), white balance, gamma correction (tone correction), color conversion, and edge enhancement on the image data after the defective pixel correction process.
[0065] In S305, the camera MPU 125 records the still image data obtained as image data by the image processing and encoding processing in S304 and one of the focus detection signals read out in S302 in the memory 128 as an image data file.
[0066] In S306, camera MPU 125 associates the camera characteristic information as characteristic information of camera body 120 with the still image data recorded in S305 and records it in lens memory 118 and memory 128 within camera MPU 125. The camera characteristic information includes, for example, the following information: Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 124 Information about the light sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 122 Information about vignetting of imaging light beams within the camera body 120 Information on the distance from the mounting surface of the imaging optical system in the camera body 120 to the imaging element 122 · Information regarding manufacturing tolerances of the camera body 120.
[0067] Information regarding the light sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light sensitivity distribution information) is information regarding the sensitivity of the image sensor 122 according to the distance (position) on the optical axis from the image sensor 122. This light sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, and therefore may be information regarding these. Furthermore, the light sensitivity distribution information may be information regarding changes in sensitivity with respect to the angle of incidence of light.
[0068] In S307, camera MPU 125 associates lens characteristic information as characteristic information of the imaging optical system with the still image data recorded in S305 and records it in memory 128 and a memory in camera MPU 125. The lens characteristic information may include information about the exit pupil, information about a frame such as a lens barrel that blocks light beams, information about the focal length and F-number at the time of imaging, and information about aberrations of the imaging optical system. The lens characteristic information may also include information about manufacturing errors of the imaging optical system and information about the position of focus lens 104 at the time of imaging (subject distance).
[0069] In S308, the camera MPU 125 records image-related information, which is information related to the still image data, in the memory 128 and in a memory within the camera MPU 125. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.
[0070] In S309, the camera MPU 125 displays a preview of the captured image on the display 126. This allows the user to easily check the captured image.
[0071] When the process of S309 is completed, the camera MPU 125 ends this imaging subroutine and proceeds to S7 in FIG. (Subroutine for subject tracking AF processing) The subject tracking AF processing subroutine executed by the camera MPU 125 in S400 of Fig. 8 will be described with reference to Fig. 10. Fig. 10 is a flowchart of the subject tracking AF processing. The chronological order in which steps S401 to S406 of this flow are executed will be described later with reference to Fig. 23.
[0072] In S401, the camera MPU 125 and the phase difference AF unit 129 perform focus detection processing using the first and second focus detection signals obtained in each of the focus detection areas acquired in S1. Details will be described later.
[0073] In S402, the camera MPU 125 performs subject detection and tracking processing. The subject detection processing is performed by the above-mentioned subject detection unit 130. Since subject detection may be impossible depending on the state of the obtained image, in such cases, tracking processing is performed using other means such as template matching to estimate the position of the subject. Details will be described later.
[0074] In S403, the camera MPU 125 performs a main subject determination process. The method for determining the main subject is determined according to a priority order based on predetermined criteria. For example, the closer the position of the subject detection area is to the central image height, the higher the priority is set, and when the positions are the same (the distance from the central image height is the same), the larger the size, the higher the priority is set. Also, a configuration may be adopted in which a defocus map is used to select a portion of a particular type of subject (person) that the photographer often wants to focus on.
[0075] In S404, the camera MPU 125 and phase difference AF unit 129 perform flicker determination. It is determined whether flicker is occurring in each focus detection area. Since the focus detection accuracy of vertical eye focus detection may decrease due to the influence of flicker, the results of vertical eye focus detection are not used when the influence of flicker is expected to be large. Details of the flicker detection method and the determination of whether vertical eye focus detection can be used will be described later.
[0076] In S405, the camera MPU 125 and the phase difference AF unit 129 perform a defocus amount selection process. Based on the subject information obtained in S402 and the flicker determination result obtained in S404, the camera MPU 125 and the phase difference AF unit 129 select a defocus amount, which is the focus detection result, using the focus detection results obtained from the arranged side-eye defocus map and vertical-eye defocus map. Details will be described later.
[0077] In S406, the camera MPU 125 performs predictive AF processing using the defocus amount obtained in S405 and multiple defocus amounts, which are time-series data on the timing of past focus detection. This processing is necessary when there is a time lag between the timing of focus detection and the timing of exposure for the captured image. Specifically, this processing predicts the position of the subject in the optical axis direction at the timing of exposure for the captured image, which is a predetermined time after the timing of focus detection, and performs AF control. To predict the subject's image plane position, multivariate analysis (e.g., the least squares method) is performed using historical data on the subject's image plane position and time to determine a prediction curve equation. The predicted image plane position of the subject can be calculated by substituting the timing of exposure for the captured image into the obtained prediction curve equation. Furthermore, three-dimensional position may be predicted in addition to the optical axis direction. If the screen is defined as an XY vector and the optical axis direction is defined as the Z direction, the position of the subject at the timing of exposure for the captured image may be predicted from time-series data of the subject's XY position obtained in S402 and the Z direction position based on the defocus amount obtained in S405. Prediction may also be performed from time-series data of the subject's joint positions. This prediction allows for estimation of each position even when the ball or person is hidden or when part of the person's joint position becomes invisible. Prediction of the subject is performed not only for the main subject but also for multiple detected subjects. By performing predictive AF processing on multiple subjects, when the main subject is switched, there is no need to re-accumulate the defocus amount history for the new main subject, and predictive AF can be continued without time loss. In S406, the predictive AF processing result is used to calculate the focus lens drive amount, and focus adjustment processing is performed by driving the focus actuator 113 in accordance with a focus drive command from the camera MPU 125 and moving the focus lens 104 in the optical axis direction.
[0078] When the process of S406 ends, the camera MPU 125 ends the subroutine of the subject tracking AF process and proceeds to S5 in FIG.
[0079] Next, the chronological execution order of S401 to S406 will be described using FIG. 23. FIG. 23 is a diagram showing the execution order of the subject tracking AF process. In this embodiment, the focus detection process of S401 and the subject tracking process of S402 are executed simultaneously. S401 is executed by the camera MPU 125 and the phase difference AF unit 129, and S402 is executed by the subject detection unit 130. S402 may be executed after S401 is completed. As for the focus detection process of S401, S2202 is executed after S2201 is completed. In this embodiment, the vertical eye defocus map is calculated after the side eye defocus map is calculated. Note that the side eye defocus map may be calculated after the vertical eye defocus map is calculated.
[0080] The main subject determination process of S403 is executed after completion of S402. In S403, a defocus map is used, but in this embodiment, since calculation of the vertical eye defocus map has not been completed, a horizontal eye defocus map is used. Note that S403 may be executed after completion of S401.
[0081] In this embodiment, S404 is executed after S401 and S403 are completed.
[0082] In this embodiment, S405 is executed after S404 is completed.
[0083] In this embodiment, S406 is executed after S405 is completed. (Focus detection processing subroutine) The focus detection process subroutine executed by the camera MPU 125 in S401 of Fig. 10 will be described with reference to Fig. 22. Fig. 22 is a flowchart of the focus detection process.
[0084] In S2201, the camera MPU 125 sets a focus detection area. In this embodiment, a sideways eye focus detection area is set on the image sensor 122, divided into 17 horizontal and 11 vertical sections for a total of 187 points. Furthermore, a vertical eye focus detection area is set on the image sensor 122, divided into 7 horizontal and 5 vertical sections for a total of 35 points. The center of the focus detection area is set based on the AF area set via the operation switch 127, the position of the subject detected and tracked in S402, or the position of the main subject determined in S403. Furthermore, the focus detection area may be set only in areas with a high likelihood of being a specific area based on specific area information output by processing described below and acquired in S1702. In this embodiment, the group of focus detection results obtained from the sideways eye focus detection area is referred to as a sideways eye defocus map. Furthermore, the group of focus detection results obtained from the vertical eye focus detection area is referred to as a vertical eye defocus map.
[0085] A method for setting a defocus map, which is a group of focus detection results for sideways and vertical eyes, will be described with reference to Fig. 18. Fig. 18 is a diagram showing a method for setting a defocus map. Fig. 18(a) is a diagram showing the subject area detected by the subject detection process described above when the subject is a person. Reference numeral 1801 denotes the upper body detection area, 1802 denotes the face detection area, and 1803 denotes the pupil detection area.
[0086] The arrangement of the side-eye defocus map, which is a group of side-eye focus detection results, will now be explained. Fig. 18(b) shows the side-eye defocus map when pupils are detected, and 1804 is the side-eye defocus map. The side-eye defocus map is arranged relative to the center of the upper body detection area so that it encompasses the subject. This makes it possible to fit the subject within the defocus map even when the subject is moving or when framing with the camera.
[0087] Next, the arrangement of the vertical eye defocus map, which is a group of vertical eye focus detection results, will be described. Fig. 18(c) shows the vertical eye defocus map when a face is detected, and 1805 is the vertical eye defocus map. In this invention, we will assume that the vertical eye defocus map has a smaller area than the horizontal eye defocus map due to calculation time constraints. Because the aforementioned horizontal eye defocus map can encompass the subject, the vertical eye defocus map is set based on the area on which the photographer wants to focus. In the case of a person, the area on which the photographer wants to focus is often the pupil, so in Fig. 18(c), the vertical eye defocus map is set with pupil detection area 1803 at the center. This makes it possible to select the defocus amount using both the horizontal eye defocus map and the vertical eye defocus map for the area on which the photographer wants to focus in the defocus amount selection process described below.
[0088] If no pupils are detected, a vertical eye defocus map is set with face detection area 1802 at the center, as shown in Fig. 18(d). If no face is detected, a vertical eye defocus map is set with upper body detection area 1801 at the center, as shown in Fig. 18(e). If no face is detected, a vertical eye defocus map is set with upper body detection area 1801 at the center.
[0089] The horizontal eye defocus map and the vertical eye defocus map are set so that the center positions and areas of each focus detection area are the same, which enables focus detection using signals from the same focus detection area, and therefore makes it possible to use both the horizontal eye defocus amount and the vertical eye defocus amount in the defocus amount selection process described below without making any distinction between them.
[0090] 18(f) shows a case where the area of the vertical eye defocus map is made smaller and each focus detection area is made smaller. By densely arranging the vertical eye defocus map in the face detection area, it becomes possible to use more defocus amounts in the defocus amount selection process described later.
[0091] Figure 18(g) shows an example where the subject is a motorcycle. 1806 is the overall detection area of the motorcycle, and 1807 is the local detection area of the motorcycle helmet. As with people, the side-glance defocus map is positioned so as to encompass the overall detection area.
[0092] Figure 18(h) shows the setting of a vertical eye defocus map during local detection of a motorcycle. The vertical eye defocus map is not placed in the center of the local area 1807, but is placed in an area that encompasses the local detection area and that allows the position and size of the horizontal eye defocus map and each focus detection area to be aligned. This results in defocus amounts that are the results of horizontal eye focus detection and vertical eye focus detection using signals from the same focus detection area, as described above. This makes it possible to use the horizontal eye defocus amount and vertical eye defocus amount together without separating them in the defocus amount selection process described below.
[0093] In S2202, the camera MPU 125 acquires a defocus map. For the focus detection areas set in S2201, the phase difference AF unit 129 calculates the amount of image shift between the first and second focus detection signals obtained in each of the multiple focus detection areas acquired in S2, and calculates the amount of defocus and reliability for each focus detection area from the image shift amount. (Subroutine for subject detection and tracking processing) The subject detection and tracking process subroutine executed by the camera MPU 125 in S402 of Fig. 10 will be described with reference to Fig. 11. Fig. 11 is a flowchart of the subject detection and tracking process.
[0094] In S421, the camera MPU 125 sets dictionary data according to the type of subject to be detected from the image data acquired in S1. Based on the preset subject priority and the settings of the imaging device, dictionary data to be used in this process is selected from multiple dictionary data stored in the dictionary data storage unit. For example, multiple dictionary data are stored by classifying subjects into categories such as "people," "vehicles," and "animals." In this embodiment, one or more dictionary data may be selected. When one dictionary data is selected, it becomes possible to repeatedly detect subjects that can be detected using one dictionary data at a high frequency. On the other hand, when multiple dictionary data are selected, the dictionary data can be set sequentially according to the priority of the detected subject, allowing subjects to be detected sequentially.
[0095] In S422, the subject detection unit 130 performs subject detection using the dictionary data set in step S421, using the image data read in step S1 as an input image. At this time, the subject detection unit 130 outputs information such as the position, size, and reliability of the detected subject. At this time, the camera MPU 125 may display the information output by the subject detection unit 130 on the display 126. In S422, multiple regions of the subject are detected hierarchically from the image data. For example, if "person" or "animal" is set as the dictionary data, multiple organs such as the "whole body" region, the "face" region, and the "eye" region are detected. While local regions such as a person's eyes or face are areas where it is desirable to adjust the focus and exposure as a subject, they may not be detectable due to surrounding obstacles or the orientation of the face. Even in such cases, the subject is detected robustly by performing full-body detection, so the subject is detected hierarchically. Similarly, when a "vehicle" such as a motorbike is set as dictionary data, the system is configured to hierarchically detect the entire area including the driver and vehicle body, and the helmet (head) as a local area.
[0096] In S423, the camera MPU 125 performs a known template matching process using the subject detection area obtained in S422 as a template. Using the multiple images obtained in S1, a similar area is searched for in the most recently obtained image using the subject detection area obtained in the past image as a template. As is well known, any information may be used for template matching, such as brightness information, color histogram information, or feature point information such as corners and edges. Various matching methods and template update methods are possible, and any of these methods may be used. The tracking process performed in S423 is performed to achieve stable subject detection and tracking when a subject is not detected in S422 by detecting an area similar to the past subject detection data from the most recently obtained image data.
[0097] In S424, the subject detection unit 130 performs region division for the detected subject region into specific regions. A specific region is a partial or entire region of the detected subject region. For example, if a person or animal is detected, it may be the region of the person's head, or if a vehicle is detected, it may be the region of the helmet. Unlike subject detection, in which the size and position of the subject are obtained using the size and coordinates of a rectangular region, region division allows the detection result to be obtained as a high-resolution distribution of the specific region. Any method (for example, the method described in "Chen et.al, DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, arXiv, 2016") can be applied as a method of region segmentation. The object detection unit 130 uses a deep-learned CNN to infer the likelihood of a specific region for each pixel region. However, the object detection unit 130 may infer the likelihood of a specific region using a trained model machine-learned by any machine learning algorithm, or may determine the likelihood of a specific region based on a rule base. When a CNN is used to infer the likelihood of a specific region, the CNN performs deep learning using the specific region as a positive example and regions other than the specific region as negative examples. As a result, the CNN outputs the likelihood of a specific region for each pixel region as an inference result. (Calculation method for specific areas) FIG. 12 is a diagram showing an example of a CNN that infers the likelihood of a specific region. FIG. 12(A) shows an example of a subject region of an input image input to a CNN. A subject region 1201 is detected from an image by the above-described subject detection. The subject region 1201 includes a face region 1202, which is the detection target of subject detection. The face region 1202 in FIG. 12(A) includes two occluded regions (occluded regions 1203 and 1204). The occluded region 1203 is a region with no depth difference from the face region, and the occluded region 1204 is a region with a depth difference. An occluded region is also called an occlusion. In this embodiment, the face region 1202 excluding the occluded regions 1203 and 1204 is detected as a specific region.
[0098] FIG. 12(B) shows an example of the definition of specific region information. Each of images 1 to 3 in FIG. 12(B) is divided into a white region and a black region, with the black region indicating a positive example and the white region indicating a negative example. The specific region information obtained by dividing the image of the subject region in FIG. 12(B) is all images assumed to be candidates for training data used when performing deep learning of CNN. Below, we will explain which of the specific region information in FIG. 12(B) is used as training data in this embodiment.
[0099] Image No. 1 in Fig. 12(B) shows an example of occlusion information when the image is divided into a subject region (face region) and a region other than the subject, with the subject region being a positive example and regions other than the subject region, such as background and occluded regions, being negative examples. Image No. 2 in Fig. 12(B) shows an example of occlusion information when the image is divided into a foreground occluded region relative to the subject and other regions, with the foreground occluded region being a negative example and regions other than the foreground occluded region relative to the subject being positive examples. Image No. 3 in Fig. 12(C) shows an example of occlusion information when the image is divided into a occluded region that causes perspective conflict relative to the subject and other regions, with the occluded region that causes perspective conflict being a negative example and regions other than the occluded region that causes perspective conflict being a positive example.
[0100] As shown in FIG. 12B, the visibility pattern of a person's face in an image is distinctive, and the pattern variance is small, allowing for highly accurate region segmentation. For example, the occlusion information shown in FIG. 12B is suitable as training data for the learning process when generating a CNN that detects a person as a subject. From the perspective of detection accuracy, the occlusion information shown in FIG. 12B is more suitable than the occlusion information shown in FIG. 12B 3. However, an image like the image shown in FIG. 12B 3 is suitable as training data for the learning process when generating a CNN that detects occluded areas that cause perspective conflict. A pair of parallax images used in focus detection may be used as training data for the learning process when generating a CNN that detects occluded areas that cause perspective conflict. Furthermore, the occlusion information is not limited to the above example and may be generated based on any method for segmenting an area into an occluded area and an area outside the occluded area. In this embodiment, emphasis is placed on the accuracy of the detection area, and the learning process is performed using the information shown in FIG. 12B 1. However, learning may also be performed using other information.
[0101] Fig. 12(C) shows the flow of deep learning of CNN. In this embodiment, an RGB image is used as an input image 1210 for learning. Furthermore, a teacher image 1214 (teacher image of specific region information) as shown in Fig. 12(C) is used as a teacher image. The teacher image 1214 is an image of face region information excluding occlusion information and background information in Fig. 12(B).
[0102] An input image 1210 for training is input to a neural network system 1211 (CNN). The neural network system 1211 can employ, for example, a layer structure in which convolutional layers and pooling layers are alternately stacked between an input layer and an output layer, and a multi-layer structure in which a fully connected layer is connected downstream of the layer structure. A score map indicating the likelihood of a specific region in the input image is output from the output layer 1212 in FIG. 12(C). The score map is output in the form of an output result 1213.
[0103] In CNN deep learning, the error between the output result 1213 and the training image 1214 is calculated as a loss value 1215. The loss value 1215 is calculated using, for example, a method such as cross entropy or squared error. Then, coefficient parameters such as the weight and bias of each node of the neural network system 1211 are adjusted so that the loss value 1215 gradually decreases. By sufficiently performing CNN deep learning using many learning input images 1210, the neural network system 1211 will output a more accurate output result 1213 when an unknown input image is input. In other words, when an unknown input image is input, the neural network system 1211 (CNN) will output, with high accuracy, specific region information obtained by dividing the region into occluded regions and non-occluded regions as the output result 1213. Note that it takes a lot of work to create training data that identifies occluded regions (regions of overlapping objects). For this reason, it is possible to create training data using CG or by using image synthesis, which cuts out and superimposes images of objects.
[0104] As described above, an example has been described in which image No. 1 in Fig. 12(B) is used as the teacher image 1214, with the face region excluding occluded regions and background regions as the specific region. However, even if images such as No. 2 or No. 3 in Fig. 12(B) are used as the teacher image 1214, when an unknown input image is input to the CNN, the CNN can infer regions that cause perspective conflict.
[0105] Any method other than CNN can be applied to detect a specific region. For example, detection of a specific region may be realized by a rule-based method. Furthermore, detection of a specific region may use a trained model trained by machine learning using any method other than deep learning CNN. For example, an occluded region may be detected using a trained model trained by machine learning using any machine learning algorithm such as a support vector machine or logistic regression. This is similar to object detection.
[0106] In this embodiment, detection of a specific region is performed for all detected subjects, but the amount of calculation can be reduced by performing this after the main subject determination process in S403 and detecting a specific region only for the main subject.
[0107] When the processing of S424 is completed, the camera MPU 125 ends the subject detection and tracking processing subroutine and proceeds to S404 in FIG. (Flicker detection subroutine) The flicker determination subroutine executed by the camera MPU 125 in S404 of Fig. 10 will be described with reference to Fig. 13. Fig. 13 is a flowchart of the flicker determination.
[0108] In S1301, the camera MPU 125 acquires information related to the drive of the image sensor 122 performed in S1. In this embodiment, the image sensor 122 selects various drive methods depending on the brightness of the shooting environment, whether the recorded image is a still image or a video, and other factors. To read signals within the screen within the time allowed by the frame rate (image sensor drive rate) set based on the brightness of the shooting environment and the photographer's settings, the readout rows are thinned out or signals from multiple rows are read simultaneously. In S1301, information is acquired regarding the drive of the image sensor regarding the result of vertical eye focus detection (image shift amount) that occurs when flicker occurs, which is determined by the number of thinned rows and the number of rows being simultaneously read out. In the present invention, whether flicker occurs in the shooting environment is determined based on the degree of agreement between the acquired information and the result calculated by the phase-difference AF unit 129 as the image shift amount of vertical eye focus detection. Details will be described later.
[0109] In S1302, a focus detection area for performing flicker determination is set in the defocus map calculated in S401 of Fig. 10. In the present invention, determination is performed sequentially for the 24 areas that make up the vertical eye defocus map.
[0110] In S1303, the camera MPU 125 acquires the side-eye focus detection result and the vertical-eye focus detection result for the focus detection area set in S1302 and calculates the difference between them. This process is performed because if the vertical-eye focus detection result contains an error due to the effect of flicker, the difference from the side-eye focus detection result may be large.
[0111] In S1304, image shift amount candidates in vertical eye focus detection are acquired. In order to explain the image shift amount candidates, the correlation calculation for performing focus detection performed in S401 will be explained.
[0112] In this embodiment, a pair of signals used for vertical eye focus detection is referred to as the A image signal and the B image signal. The first, second, etc. outputs of the A image signal in each row within the focus detection area are designated A(1), A(2), etc., and similarly, the first, second, etc. outputs of the B image signal are designated B(1), B(2), etc. In this manner, 300 sequentially generated A (B image) signals are concatenated to generate a pair of image signals. In the correlation calculation, the correlation amount is calculated while shifting the relative positions of the pair of image signals, and the shift amount at the position with the highest correlation (the degree of similarity in the shapes of the pair of image signals) is detected as the image shift amount. For example, the correlation amount COR(h) can be calculated using the following formula:
[0113]
number
[0114] (1) In equation (1), W1 corresponds to the number of data points within the field of view, and hmax corresponds to the number of shift data points. After determining the correlation amount COR(h) for each shift amount h, the phase difference AF unit 129 determines the shift amount h that maximizes the correlation between the A and B images, i.e., the value of the shift amount h that minimizes the correlation amount COR(h). Note that the shift amount h used to calculate the correlation amount COR(h) is an integer, but when determining the shift amount h that minimizes the correlation amount COR(h), interpolation or the like is performed to determine a value (real value) in sub-pixel units in order to improve the accuracy of the defocus amount.
[0115] In this embodiment, the shift amount at which the sign of the difference value of the correlation amount COR changes is calculated as the shift amount h (in subpixel units) at which the correlation amount COR(h) is minimized. Also, the number of in-field data and the number of shift data used in the calculation may be changed using specific area information, which will be described later, so that the defocus amount of an area with a high likelihood of being a specific area is calculated.
[0116] First, the phase difference AF unit 129 calculates a difference value DCOR of the correlation amount according to the following equation (2).
[0117] DCOR(2×h)=COR(h+1)-COR(h-1) (2) Then, the phase difference AF unit 129 uses the correlation difference value DCOR to calculate a shift amount dh1 at which the sign of the difference amount changes. If the value of h just before the sign of the difference amount changes is h1 and the value of h after the sign changes is h2 (h2 = h1 + 1), the phase difference AF unit 129 calculates the shift amount dh1 according to the following equation (3).
[0118] dh1=(h1+|DCOR1(h1)| / |DCOR1(h1)-DCOR1(h2)|)×2 (3) In this way, the phase difference AF unit 129 calculates the shift amount dh1 in subpixel units that maximizes the correlation between the images A and B of the first signal, and then completes the process. Note that the method for calculating the shift amount (phase difference) between two one-dimensional image signals is not limited to that described here, and any known method can be used. As a result of the above-described correlation calculation, multiple shift amounts that change sign in the difference value of the correlation amount COR may be calculated. In normal focus detection, the shift amount that maximizes the difference value is selected, but in S1304, the multiple calculated shift amounts are acquired as image shift amount candidates. A method for using the image shift amount candidates will be described in detail later.
[0119] In S1305, the camera MPU 125 determines whether there is a correlation between the image shift amount candidate acquired in S1304 and the information about the result of vertical eye focus detection (image shift amount) that occurs when flicker occurs related to the driving method of the image sensor acquired in S1301. If the value of the image shift amount candidate acquired in S1304 or the difference therebetween is close to the image shift amount acquired in S1301 within a predetermined value, proceed to S1306, and if not close, proceed to S1308.
[0120] In S1306, the difference between the vertical and horizontal focus detection results acquired in S1303 is determined. If the difference is large, the process proceeds to S1307, and if the difference is small, the process proceeds to S1308.
[0121] In S1307, the set focus detection area is determined to be affected by flicker because flicker has caused an error in the vertical eye focus detection result.
[0122] In S1308, it is determined that the set focus detection area is less affected by flicker on the vertical eye focus detection result.
[0123] After S1307 or S1308 is completed, the process proceeds to S1309, where it is determined whether flicker detection has been completed for all focus detection areas. If not, the process returns to S1302, and the above-described processing is repeated. If completed, the processing of this subroutine is completed, and the process proceeds to S405. (Impact of image sensor driving method and flicker on vertical eye focus detection) 14 to 16 , a mechanism by which errors in vertical eye focus detection due to flicker occur depending on the driving method of the image sensor 122 will be described. Flicker, which occurs in lighting, digital signage, and the like, is a phenomenon in which blinking occurs repeatedly over time at an invisible frequency. On the other hand, the slit rolling image sensor 122 sequentially accumulates and reads out signals from each row over time. When the slit rolling image sensor 122 is exposed in a flickering environment, the accumulation time difference between each row causes the signal of each row to fluctuate due to the influence of flicker. In this embodiment, focus detection signals are also sequentially read out from each row. However, since the paired signals used for horizontal eye focus detection use signals from the same row, they are affected by flicker to the same extent, and therefore the impact on the focus detection results is small. On the other hand, the paired signals used for vertical eye focus detection are affected by flicker within the paired signal sequence because the direction in which the signal sequence is formed coincides with the direction in which the slit rolling readout is performed.
[0124] FIG. 14 is a diagram illustrating the effect of flicker on a pair of vertical eye focus detection signals. FIG. 14(a) shows time elapsed horizontally from left to right, and the accumulation and readout timing of the focus detection signal (image A) and the image capture signal (image A+B) for each row of the image sensor 122 are shown on the time axis. As explained in FIG. 2(b), the output of signals A and A+B for each row is indicated by the accumulation and readout periods shown in the top two lines of FIG. 14(a). After resetting the PDA 211 and PDB 212, accumulation of signals A and A+B begins, and the voltage of signal A is read out simultaneously with the completion of accumulation. After the readout of signal A is completed, accumulation of signals A+B is completed and the voltage is read out. The signal for the second row is read out in a similar manner. The time difference between the accumulation period of signal A for the first row and the accumulation period of signal A for the second row is considered to be the difference in the centers of the accumulation periods, so the interval is Pa-a. Additionally, the interval between the accumulation period of signal A+B on the first row and the accumulation period of signal A+B on the second row is Pab-ab. As mentioned above, in a flickering environment, brightness changes over time, and the signal output of the first and second rows changes with the passage of time (Pa-a) and Pab-ab). The difference in the accumulation periods of signal A and signal A+B is shown as Pa-ab. In a flickering environment, the accumulation periods of signal A and signal A+B are offset by Pa-ab for every row. Due to the Pa-ab offset, the waveforms of signal A and signal A+B are shifted horizontally by Pa-ab / Pa-a pixels relative to the waveform of signal A. For example, as shown in Figure 14(a), the accumulation start time for each row is shifted by a time equivalent to the sum of the readout periods of signal A and signal A+B. When the readout periods of signal A and signal A+B are equal, the waveform of signal B is shifted horizontally relative to the waveform of signal A by Pa-ab / Pa-a pixels=1 / 4 pixel.
[0125] Figure 14(b) shows a case where exposure control for each row is different from that of Figure 14(a), and signals A and B are read out for each row. This shows a case where the start of accumulation for signals A and B in the first row is shifted by the readout period of signal A. As in Figure 14(a), due to the difference in the accumulation periods for signals A and A+B, the waveform of signal B is shifted horizontally by Pa - ab / Pa - a pixels relative to the waveform of signal A. For example, for signal A in the first row, signal B in the first row, signal A in the second row, etc., the start of accumulation is shifted by times corresponding to the readout period of signal A in the first row, the readout period of signal B in the first row, the readout period of signal B in the second row, etc. Furthermore, if the readout periods of signals A and B are equal, the waveform of signal B is shifted horizontally by Pa - ab / Pa - a pixels = 1 / 2 pixels relative to the waveform of signal A.
[0126] FIG. 15 shows waveforms when flicker occurs. FIG. 15(a) shows signals A and B corresponding to the case of FIG. 14(b). The horizontal axis represents the pixel number, and the vertical axis represents the signal output normalized by the maximum value. The undulations in the output for each pixel indicate the flickering that occurs over time. A partial enlargement is shown in the upper right corner of FIG. 15(a), which shows that the waveforms of signals A and B are slightly shifted. As explained in S1304, FIG. 15(b) shows the results of calculating the correlation amount. The horizontal axis represents the shift amount (shift amount) between signals A and B, and the vertical axis represents the correlation amount, which indicates the magnitude of the correlation. FIG. 15(b) shows that the correlation amount reaches its minimum value when the shift amount is ±40 pixels and 0 pixel. FIG. 15(c) shows the calculated difference value DCOR of the correlation amount. The horizontal axis represents the shift amount, and the vertical axis represents the difference in correlation amount. The shift amounts that intersect the horizontal axis in an upward sloping line to the right indicate that they are in the vicinity of ±80 pixels and 0 pixel. FIG. 14(d) shows an enlarged view of the shift amount near 0 pixel. The image shift amount candidate dh1 in this embodiment indicates -0.5 pixel, which is the intersection with the horizontal axis. Similarly, -80.5 pixels and +79.5 pixels are candidates for the image shift amount.
[0127] The image shift amount candidate, −0.5 pixels, is the pixel shift amount that occurs when the readout shown in FIG. 14B is performed in an environment where flicker is occurring. In this embodiment, in step S1301, information regarding the readout method, such as that shown in FIG. 14A or FIG. 14B, is acquired as information regarding the driving of the image sensor 122, thereby acquiring the image shift amount caused by flicker. For example, in the case of the driving method shown in FIG. 14B, information of −0.5 pixels is acquired. On the other hand, as can be seen from FIG. 15A, the image shift amounts of −80.5 pixels and +79.5 pixels are image shift amounts obtained by offsetting the image shift amount caused by the influence of flicker by −0.5 pixels from the 80-pixel cycle in which flicker occurs. By canceling the image shift amount caused by the influence of flicker, it is possible to calculate that the cycle in which flicker occurs is 80 pixels, and the frequency of flicker blinking can be calculated from information regarding the readout time for each row. In S1305, the image shift amount of −0.5 pixels when flicker occurs is determined from the read information of the image sensor 122 in FIG. 14(b) to determine whether it is included in the image shift amount candidates obtained in S1304. If it is included, it is determined that the environment may be flickering, and the process proceeds to S1305. In S1306, to eliminate cases where the defocus state of the subject coincides with the image shift amount detected in a flickering environment, the process checks the difference with the side-eye focus detection result, which is less affected by flicker. If the difference between the side-eye focus detection result and the vertical-eye focus detection result is small, it is determined that the defocus state of the subject can also be obtained from the vertical-eye focus detection result. On the other hand, if the difference is large, it is determined that the vertical-eye focus detection result is affected by flicker. By performing the determination in S1306, focus detection using the vertical-eye focus detection result can be performed in a wider range of shooting environments, resulting in more accurate focus adjustment. Alternatively, the determination in S1306 can be omitted to minimize the effects of flicker on the vertical-eye focus detection result.
[0128] Returning to FIG. 14, we continue the explanation. FIG. 14(c) shows a case where multiple rows of the image sensor 122 are simultaneously read out. While FIG. 14(c) shows a case where four rows are simultaneously read out, the number of rows simultaneously read out is not limited to this. Even when multiple rows are simultaneously read out, there is a difference between the readout period of signal A and the readout period of signals A+B. Furthermore, there is a difference in the readout period for each block of rows (four rows per block in FIG. 14(c)). FIG. 16 is another diagram showing waveforms when flicker occurs. For ease of understanding, FIG. 16(a) shows the waveforms of signals A and B when 10 rows are simultaneously read out. In addition to the effect of flicker, it can be seen that a step occurs every 10 rows. When the above-described correlation calculation is performed on such waveforms, a section of shift amount where the change in correlation amount is small occurs, making it impossible to obtain a highly accurate image shift amount, so digital filtering is performed. FIG. 16(b) shows the results of the predetermined filter processing (-4, -11, -21, -28, -28, -17, 0, 17, 28, 28, 21, 11, 4). As with the correlation calculation process described above, FIG. 16(c) shows the correlation amount COR, and FIG. 16(d) shows the correlation amount difference DCOR. It can be seen that the correlation amount difference DCOR slopes upward and intersects with the horizontal axis at approximately -90, -80, -10, 0, +70, and +80 pixels. Here, the -10-pixel shift amount is the amount of image shift caused by flicker when the image sensor 122 reads 10 rows simultaneously. As with the row-by-row readout described above, in step S1301, information regarding the driving of the image sensor 122 is acquired, indicating that 10 rows will be read simultaneously. It is also acquired that the amount of image shift caused by flicker is approximately -10 pixels. The subsequent determinations made in S1305 and S1306 are as described above. Similarly, the flicker frequency can be calculated from the shift amounts of ±80 pixels and 0 pixel. It can also be seen that the image shift amount candidates of -90 pixels and +70 pixels are image shift amounts resulting from the sum of the flicker frequency and the effects of flicker caused by the readout method of the image sensor.
[0129] When multiple rows are read out simultaneously, image shift occurs due to two factors: the difference Pa-ab between the readout periods of signal A and signal A+B, and the difference Pa-ab due to waveform steps that occur every few rows. If the digital filter processing described above sufficiently reduces the impact of the waveform steps that occur every few rows, the impact of the difference Pa-ab between the readout periods of signal A and signal A+B becomes significant. For example, in the case of simultaneous readout of four rows shown in Figure 14(c), if the waveform steps that occur every four rows are eliminated by digital filter processing, an image shift of Pa-ab / Pa-a × 4 pixels = 1 pixel occurs. On the other hand, in the case of simultaneous readout of 10 rows shown in Figure 16, the waveform steps that occur every 10 rows are not eliminated by digital filter processing. Therefore, an image shift of -10 pixels is calculated as the image shift candidate.
[0130] In this way, by combining the drive information of the image sensor 122 acquired in S1301 and the digital filter processing used in the correlation calculation, the value of the image shift amount at which the focus detection result is affected by flicker is calculated in advance, and this value can be compared with the candidate image shift amount in S1304.
[0131] Furthermore, the influence of the A signal, B signal, and vertical eye focus detection results in a flicker environment described in Figures 14 to 16 applies when the subject has no contrast and only the effect of flicker occurs. In reality, contrast, including the defocus state of the subject, is superimposed on the A signal and B signal. Therefore, when the subject contrast is low and the flicker contrast is large, the flicker has a large effect on the vertical eye focus detection results, resulting in a value close to the image shift amount described above. On the other hand, when the subject contrast is high or when the flicker contrast is small in a mixed light environment with other flicker-free light sources, the flicker has a small effect on the vertical eye focus detection results, resulting in a vertical eye focus detection result that indicates the defocus state of the subject. Therefore, it is desirable to make the determination performed in S1305 of Figure 13 assuming that a certain amount of error will occur in the image shift amount occurring in a flicker environment due to the readout method of the image sensor 122 and the digital filter. For example, in the case of FIG. 15, a method of determining "Yes" can be considered when a candidate image shift amount for vertical eye focus detection within the range of -0.5 pixels ±0.25 pixels is obtained.
[0132] As described above, in an environment where flicker occurs, errors may occur in the vertical eye focus detection result, but by determining whether or not it can be used according to the driving information of the image sensor, it is possible to avoid using a low-accuracy vertical eye focus detection result, and as a result, high-accuracy focus detection can be performed.
[0133] In this embodiment, the presence or absence of flicker influence was determined for each focus detection area. Flicker may occur due to the lighting in the entire shooting environment, or it may occur only in a part of the shooting environment, such as a digital signage. By determining the presence or absence of flicker influence for each focus detection area as in this embodiment, more vertical eye focus detection results can be used, enabling more accurate focus detection.
[0134] On the other hand, as mentioned above, the impact of flicker on vertical eye focus detection results varies depending on the contrast of the subject, including defocus. Therefore, using only one focus detection area can result in an incorrect determination. Therefore, a threshold can be set in advance, and if a number of focus detection areas greater than the threshold are determined to be affected by flicker, none of the vertical eye focus detection results can be used. Furthermore, if there is an uneven distribution of focus detection areas affected by flicker, it is possible to use only some vertical eye focus detection areas within the shooting range. These methods can more reliably eliminate errors due to flicker in vertical eye focus detection results. (Defocus amount selection process) A subroutine of the defocus amount selection process will be described with reference to Fig. 17 to Fig. 20. Fig. 17 is a flowchart of the defocus amount selection process. Fig. 18 is a diagram showing a method for setting a defocus map. Fig. 19 is another diagram showing a method for setting a defocus map, showing an example of defocus map placement when occlusion occurs. Fig. 20 is a diagram showing a histogram of the defocus map.
[0135] In S1701, the camera MPU 125 acquires the subject detection position and size, which are subject detection information detected by the subject detection unit .
[0136] In S1702, the camera MPU 125 acquires specific region information detected by the subject detection unit 130. In this embodiment, the specific region information is a face region excluding occluded regions and background regions. Processing using the specific region information will be described later.
[0137] In S1703, usable focus detection results are collected. Collecting usable focus detection results is a process of collecting defocus amounts, which are focus detection results that can be used in the defocus amount selection process, from the side-eye defocus map and the defocus amounts in the side-eye defocus map. Specifically, whether or not all vertical-eye focus detection results are usable is determined based on whether or not the number of focus detection areas determined to be affected by flicker in the flicker determination process shown in FIG. 13 above is equal to or greater than a predetermined number. The reason all vertical-eye focus detection results are usable is because if more than a predetermined number are determined to be affected by flicker, it is highly likely that the vertical-eye focus detection results contain errors due to the effects of flicker.
[0138] In addition, when the contrast of the subject is low, the ISO sensitivity is high, or the exposure is darker than the correct exposure, the error in the focus detection result is large. Therefore, the reliability of the focus detection result may be determined based on the difference in the correlation amount in the correlation calculation process described above, and the focus detection result may not be used as the focus detection result. Also, depending on the driving method of the image sensor, thinning or adding the readout rows may result in the accuracy of the focus detection result for vertical eye movements being lower than that of the focus detection result for horizontal eye movements. Therefore, in modes where imaging is performed using such driving methods, vertical eye movement focus detection may not be used.
[0139] In S1704, a histogram is generated using the defocus amount, which is the focus detection result made available in S1703. The histogram is generated using the subject detection information and specific area information to determine which focus detection area's focus detection result should be used. As shown in the focus detection area setting process in S2201 in Figure 22 above, the histogram is generated using the defocus map included in the subject area.
[0140] A method for creating a histogram using the defocus amount in the upper body detection region will be described with reference to FIGS. 20(a) to 20(c).
[0141] Figure 20(a) is a histogram generated from the defocus amount of the side-eye defocus map within the upper body detection region of the person in Figure 18(b). Figure 20(b) is a histogram generated from the defocus amount of the vertical eye defocus map within the upper body detection region of the person in Figure 18(c). Figure 20(c) is a histogram generated by combining the defocus amounts of the side-eye defocus map and the vertical eye defocus map within the upper body detection region of the person in Figures 18(b) and 18(c). The horizontal axis of the histogram represents the defocus amount divided into certain ranges, and the vertical axis represents frequency. The defocus amount is expressed as the near side on the positive side and the far side on the negative side, with the defocus amount of the person's pupil region set to 0Fδ. The side-eye histogram in Figure 20(a) was generated for the entire upper body detection region, so the area to the left of the upper body below the face is more prevalent, resulting in the maximum frequency of the histogram being on the near side. Therefore, if a defocus amount is selected from the range of defocus amounts that maximizes the histogram frequency, the defocus amount selected will differ from the pupil area of the person the photographer wants to focus on. In the vertical eye histogram in Figure 20(b), the vertical eye defocus map is located in the face detection area, so the area to the left of the upper body below the face is not included. This results in the histogram frequency being maximized in the range around 0Fδ, which is the defocus amount for the pupil area. However, because the number of focus detection areas in the defocus map is small, it may be difficult to extract the point with the maximum frequency under conditions where the defocus amount is prone to fluctuating due to errors. Therefore, by generating a histogram that combines horizontal and vertical eye histograms as in Figure 20(c), it is possible to generate a histogram using a larger number of defocus amounts. This allows for a more accurate defocus amount to be selected in response to variations in defocus amount or incorrect defocus amounts. However, because this is a histogram of defocus amounts in the upper body detection area, it also includes the area to the left of the upper body below the face. In this case, the frequency of the defocus amount histogram is high in the ranges of -1Fδ to 0Fδ and 0Fδ to 1Fδ, making it difficult to extract the range of the defocus amount where the frequency of the histogram is maximum.Therefore, depending on the variation in the defocus amount, the range of the defocus amount in which the frequency of the histogram is maximum may vary.
[0142] Therefore, in this embodiment, since it is possible to use a vertical eye defocus map in addition to a horizontal eye defocus map, a method for creating a histogram using the defocus amount in the face detection area will be explained using Figures 20(d) to 20(f). Let us assume that the defocus amount in the pupil area of a person is 0Fδ as an example.
[0143] Fig. 20(d) is a histogram generated from the defocus amount of the side-glance defocus map within the person's face detection area in Fig. 18(d). Because the histogram is generated based on the defocus amount within the face detection area, the frequency of the defocus amount histogram reaches its maximum value in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil area.
[0144] Fig. 20(e) is a histogram generated from the defocus amount of the vertical eye defocus map within the person's face detection area in Fig. 18(d). Because the histogram is generated based on the defocus amount within the face detection area, the frequency of the defocus amount histogram reaches its maximum value in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil area.
[0145] Figure 20(f) is a histogram generated by combining the defocus amounts for side and vertical eyes by combining the histograms in Figures 20(d) and 20(e). The histogram combining the defocus amounts for side and vertical eyes within the face detection area results in a higher frequency in the histogram for defocus amounts in the range from -1Fδ to 0Fδ, which includes the defocus amount in the person's pupil area, than in the histogram for only side or vertical eyes. This makes it possible to reduce the influence of variations in defocus amounts or defocus amounts that cause perspective conflicts with the background.
[0146] It is desirable to generate a defocus amount histogram using a larger number of defocus amounts within a narrow person detection area. As shown in Figures 20(d) to 20(f), it is desirable to generate a histogram using the defocus amounts of the side-eye defocus map and the vertical-eye defocus map within the face detection area. However, when the defocus map area within the face detection area is small, the number of defocus amount data is small, resulting in a low overall frequency of the histogram using the defocus amount, making it difficult to extract the range of defocus amounts with the highest frequency. Therefore, when creating a histogram, the necessary number of defocus amount data or person detection area is determined, and it is determined whether the number of defocus amount data or person detection area is equal to or greater than a predetermined value. If it is less than the predetermined value, the person detection area is expanded so that the number of defocus amount data is equal to or greater than the predetermined value. Furthermore, since there is a difference between the area of the side-eye defocus map and the area of the vertical-eye defocus map, a defocus amount histogram may be generated, for example, using the side-eye defocus map as the person's upper body detection area and the vertical-eye defocus map as the person's face detection area.
[0147] Furthermore, depending on the person detection area, there may be a defocus amount in the side-eye defocus map but no defocus amount in the vertical-eye defocus map. Therefore, when there is only a defocus amount in the side-eye defocus map, the number of defocus amounts may be doubled or left as is, and in areas where both side-eye defocus maps and vertical-eye defocus maps exist, the number of defocus amounts may be left as is or halved for normalization. In this embodiment, the area of the side-eye defocus map is described as being larger than the area of the vertical-eye defocus map, but the area of the vertical-eye defocus map may also be larger than the area of the side-eye defocus map.
[0148] In S1705, a focus detection area is selected using the defocus amount histogram, which is the focus detection result generated in S1704, and the defocus amount, which is the focus detection result for that area, is selected. The defocus amount is selected from the range where the frequency of the defocus amount histogram is the maximum value. There are multiple selection methods, such as selecting the defocus amount closest to the defocus amount, which is the predicted AF processing result of S406, selecting the defocus amount of a focus detection area that is close in position to the pupil detection area, which is the detection area for a person, or selecting a defocus amount on the near side. Alternatively, a defocus amount histogram may be created for each of multiple detection areas (e.g., upper body, face, pupil, etc.), and the defocus amount may be selected from the near side, or from a range where the frequency of the histograms of multiple detection areas is the maximum value.
[0149] Alternatively, the defocus amount may be calculated by averaging the defocus amounts in the range where the frequency of the defocus amount histogram is the maximum.
[0150] Next, processing using specific region information will be described with reference to FIG. 19. FIG. 19(a) shows an image of the moment when a person's face region is covered by an occluded region (an arm), and the subject detection information acquired in S1701 is indicated by a rectangular frame. FIG. 19(b) shows the specific region (in this embodiment, the face region) acquired in S1702 as a lattice frame, indicating that the portion covered by the arm has not been detected as a specific region (face region). The specific region information (likelihood) acquired in S1702 may be information that expresses whether or not the region is a specific region with a binary output result of 1 or 0, or may be information expressed in one byte, for example, from 0 to 255, where the larger the value, the higher the likelihood. Here, we will assume the former, and assume that the lattice frame region is output as 1, and other regions, such as the arm, are output as 0. FIG. 19(c) shows a diagram in which only regions of the side-eye defocus map that are valid as specific regions are indicated by diagonal lines by associating the specific regions with a 3×3 side-eye defocus map. The validity of a defocus map can be determined by determining whether the estimated region occupies a certain percentage or more, for example, 50% or more, within each frame of the defocus map. The range of each frame may be determined based on parameters used in correlation calculations, such as the shift amount used to calculate the defocus amount. FIG. 19(d) shows a diagram in which only the areas of the vertical eye defocus map that are valid as specific regions are represented by diagonal lines by associating the 3×3 vertical eye defocus map with the specific region. The determination of validity is similar to that for the horizontal eye defocus map, and therefore will not be described here. FIG. 21 shows histograms generated from the defocus maps of FIGS. 19(c) and 19(d). The 3×3 defocus map includes an occluded region. Therefore, if a histogram were generated for all regions, the influence of the occluded region would result in histogram peaks being more likely to be detected nearer to the face. However, by generating a histogram for only the specific region as in this embodiment, the influence of the occluded region, background region, etc. can be eliminated.
[0151] As described above, by generating a histogram only for a specific region, it is expected that the influence of occluded regions can be prevented. Also, in this embodiment, a defocus map with a 3x3 frame has been used for explanation, but the number of frames can be freely set to NxM frames (N and M are integers of 2 or more). Furthermore, as mentioned above, it is also possible to set a threshold value to determine whether a specific region is valid or not, and generate a histogram only for valid regions (regions above the threshold value). [Example 2] In this embodiment, the recommended direction is determined using the specific region information, and the result is used to perform the defocus amount selection process.
[0152] The configuration of the camera system of this embodiment is the same as that of embodiment 1, but the defocus amount selection process is partially different. Here, the following description will focus on the differences from embodiment 1 in the defocus amount selection process.
[0153] FIG. 24 is a flowchart of the defocus amount selection process of this embodiment. S2401 and S2402 are the same as in the first embodiment, so their explanation will be omitted. In S2403, the camera MPU 125 determines the recommended direction using the specific area information generated by the subject detection unit 130. Details will be explained using FIG. 25. FIG. 25 is a diagram showing the recommended direction determination. FIG. 25(a) is an example of an image in which a horizontal occluding subject is present relative to the center of the face. FIG. 25(b) is an example in which a 3×3 focus detection area is set relative to the center of the face, and sideways eye defocus and vertical eye defocus are detected from each focus detection area. FIG. 25(c) is a diagram showing a 9×9 specific area for FIG. 25(a) obtained in S2402. In Figure 25(c), one focus detection area corresponds to a 3 x 3 specific area, for example, focus detection area 2501 corresponds to the specific area in 2502, and focus detection area 2503 corresponds to the specific area in 2504. The nine horizontal frames in the center where a horizontal occluding subject exists and the areas around the four corners are not specific areas (likelihood is 0).
[0154] To calculate the recommended direction, first the specific area information contained within the focus detection area is projected horizontally and vertically. Using Figures 25(b) and 25(c) as examples, when projections are taken at 2501 and 2502, the horizontal projection is 1 / 1 / 1 from left to right, and the vertical projection is 1 / 1 / 1 from top to bottom. When projections are taken at 2503 and 2504, the horizontal projection is 0.6 / 0.6 / 0.6 from left to right, and the vertical projection is 1 / 0 / 1 from top to bottom.
[0155] Next, the maximum / minimum value differences for the horizontal and vertical projections, MM_V and MM_H, are obtained. When obtained in 2501 and 2502, MM_V is 0 and MM_H is 0. When obtained in 2503 and 2504, MM_V is 0 and MM_H is 1. Finally, MM_H - MM_V is calculated, and this value becomes the recommended vertical / horizontal judgment value HV_Judge. If the recommended judgment value is a positive value, the defocus amount detected in the horizontal direction (sideways eye defocus) is recommended. If it is a negative value, the defocus amount detected in the vertical direction (vertical eye defocus) is recommended. Because 2501 and 2502 are 0 and 2503 and 2504 are 1, it can be seen that there is no recommended direction in the focus detection area of 2501, and sideways eye defocus is recommended in the focus detection area of 2503. In terms of the image, there is also an occluding subject in the horizontal direction, so in the focus detection areas of the three central horizontal frames, including 2503, when vertical eye defocus is used, a defocus result is calculated that is due to perspective conflict with the occluding subject. It is clear that there is a high possibility that the amount of defocus for the face will not be detected correctly, but that this effect is likely to be small when using horizontal eye defocus.
[0156] As described above, the recommended vertical / horizontal determination value can be calculated using specific area information, and the recommended orientation can be determined. While an example in which a 3x3 specific area corresponds to one focus detection area has been described here, the corresponding specific area can be NxM (N and M are integers greater than or equal to 2). Furthermore, while it has been explained that a positive recommended vertical / horizontal determination value indicates a horizontal orientation, and a negative recommended vertical orientation, as previously mentioned, the likelihood of a specific area can also be information expressed in one byte ranging from 0 to 255. In this case, since the recommended vertical / horizontal determination value can take values from -255 to 255, it is also possible to set a threshold value and determine that the recommended orientation is valid only when the value exceeds that threshold.
[0157] S2404 and S2405 are similar to those in the first embodiment, and therefore will not be described here. However, in generating the histogram in S2405, the histogram may be generated without using the defocus amount in a direction not recommended in this embodiment, such as the vertical eye defocus of 2503 in the example of FIG. 25.
[0158] The selection of the focus detection area in S2406 will be described with reference to Fig. 26. Fig. 26 is a flowchart of the selection process of the focus detection area.
[0159] In S2601, a focus detection area is selected using the histogram generated in S2405. Details of the process are the same as in S1705 in the first embodiment, and are therefore omitted. In S2602, it is determined whether selection of a focus detection area using a histogram was possible. Here, cases in which selection using a histogram is not possible include, for example, when there are no areas with a sufficiently high likelihood of being a specific area, and therefore a histogram cannot be generated for the specific area. Also, a threshold value for the histogram frequency is set, and there are few areas with a sufficiently high likelihood of being a specific area, and the threshold is not exceeded. If selection is possible, the process proceeds to S2606, where the focus detection area is determined and processing ends. If selection is not possible, the process proceeds to S2603. In S2603, it is determined whether a frame with a recommended direction exists within the focus detection area. If it exists, the process proceeds to S2604, where a focus detection area is selected using the recommended direction. If it does not exist, the process proceeds to S2605, where a focus detection area is selected. The selection of the focus detection area in S2605 is a process for selecting a focus detection area that does not use specific area information or a recommended direction, and is not a feature of this embodiment, so a description thereof will be omitted. The method for selecting a focus detection area using a recommended direction in S2604 prioritizes the defocus amount of the focus detection area that is closest in position to a person detection area, such as an pupil detection area, when there are multiple frames with recommended directions within the focus detection area. When the distances are equal, the frame with the highest likelihood of specific area information is prioritized for focus detection area selection; when the likelihoods are equal, the frame with the highest reliability of the defocus calculation result is prioritized.
[0160] As described above, by determining the recommended direction and selecting the defocus amount using the determination result, it is expected that the influence of the occluded area can be prevented. [Other Examples] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0161] The disclosure of this embodiment includes the following configurations and methods. (Configuration 1) A control device for controlling focus adjustment, an acquisition unit that acquires a plurality of defocus amounts acquired in a plurality of focus detection areas and information about a specific area detected from a subject area in an imaging area; a control unit that selects a first defocus amount for controlling the focus adjustment in accordance with the plurality of defocus amounts and information about the specific region. (Configuration 2) The control device according to configuration 1, wherein the control unit creates a histogram using the plurality of defocus amounts and information about the specific region, and selects the first defocus amount using the histogram. (Configuration 3) The control device according to configuration 2, wherein the control unit creates the histogram using defocus amounts for the plurality of focus detection areas, each of which includes valid information among information about the specific area. (Configuration 4) The control device according to configuration 2 or 3, wherein the control unit creates the histogram using the defocus amount for each of the plurality of focus detection areas, each of which has a proportion of a specific area equal to or greater than a threshold value. (Configuration 5) The control device described in any one of configurations 1 to 4, characterized in that the control unit associates information about the specific area with the multiple focus detection areas based on the number of shift data when acquiring the multiple defocus amounts. (Configuration 6) The control device described in any one of configurations 1 to 5 is characterized in that the control unit sets at least one of the positions of the plurality of focus detection areas and the number of shift data when acquiring the plurality of defocus amounts so that a defocus amount is acquired for a focus detection area where the proportion of a specific area is equal to or greater than a threshold. (Configuration 7) The control device described in any one of configurations 1 to 6, characterized in that the control unit determines, based on information about the specific area, whether it is better to acquire the multiple defocus amounts in a first direction included in the specific area or in a second direction perpendicular to the first direction. (Configuration 8) The control device according to configuration 7, wherein the control unit selects the first defocus amount according to the plurality of defocus amounts acquired in one of the first direction and the second direction and information about the specific region. (Configuration 9) 9. The imaging device according to configuration 8, wherein the control unit does not use the plurality of defocus amounts acquired in the other of the first direction and the second direction. (Configuration 10) A configuration control device according to any one of configurations 1 to 9; An imaging device comprising: an imaging element. (Method 1) A control method for controlling focus adjustment, comprising: acquiring information on a plurality of defocus amounts acquired in a plurality of focus detection areas and a specific area detected from a subject area in an imaging area; selecting a defocus amount for controlling the focus adjustment in accordance with the plurality of defocus amounts and information relating to the specific region. (Configuration 11) A program that causes a computer to execute the control method described in Method 1.
[0162] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. [Explanation of symbols]
[0163] 125 Camera MPU (acquisition unit, control unit, control device)
Claims
1. A control device for controlling focus adjustment, an acquisition unit that acquires a plurality of defocus amounts acquired in a plurality of focus detection areas and information about a specific area detected from a subject area in an imaging area; a control unit that selects a first defocus amount for controlling the focus adjustment in accordance with the plurality of defocus amounts and information about the specific region.
2. 2. The control device according to claim 1, wherein the control unit creates a histogram using the plurality of defocus amounts and information about the specific region, and selects the first defocus amount using the histogram.
3. 3. The control device according to claim 2, wherein the control unit creates the histogram using defocus amounts for the plurality of focus detection areas, each of which includes valid information from among information about the specific area.
4. 3. The control device according to claim 2, wherein the control unit creates the histogram using defocus amounts for the plurality of focus detection areas, each of which has a ratio of a specific area equal to or greater than a threshold value.
5. 3. The control device according to claim 1, wherein the control unit associates information about the specific area with the plurality of focus detection areas based on the number of shift data when the plurality of defocus amounts are acquired.
6. The control device according to claim 1 or 2, characterized in that the control unit sets at least one of the positions of the plurality of focus detection areas and the number of shift data when acquiring the plurality of defocus amounts so that a defocus amount is acquired for a focus detection area where the proportion of a specific area is equal to or greater than a threshold value.
7. The control device according to claim 1 or 2, characterized in that the control unit determines, based on information about the specific area, whether it is better to obtain the multiple defocus amounts in a first direction included in the specific area or in a second direction perpendicular to the first direction.
8. 8. The control device according to claim 7, wherein the control unit selects the first defocus amount according to the plurality of defocus amounts acquired in one of the first direction and the second direction and information about the specific region.
9. The imaging device according to claim 8 , wherein the control unit does not use the plurality of defocus amounts acquired in the other of the first direction and the second direction.
10. The control device according to claim 1 or 2; An imaging device comprising: an imaging element.
11. A control method for controlling focus adjustment, comprising: acquiring information on a plurality of defocus amounts acquired in a plurality of focus detection areas and a specific area detected from a subject area in an imaging area; selecting a defocus amount for controlling the focus adjustment in accordance with the plurality of defocus amounts and information relating to the specific region.
12. A program causing a computer to execute the control method according to claim 11.
Citation Information
Patent Citations
Imaging device
JP2019008075A
Information processing device, control method thereof, and program
JP2021068932A
Electronic device, method of controlling the same, and program
JP2022125743A
Subject tracking device
JP2014202875A