Imaging device

JP7927789B2Active Publication Date: 2026-10-01CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024095603
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2026-10-01
Estimated Expiration
2044-06-13

AI Technical Summary

Benefits of technology

【0009】 本発明によれば、高精度な焦点検出を実現することができる撮像装置を提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927789000002
    Figure 0007927789000002
  • Figure 0007927789000003
    Figure 0007927789000003
  • Figure 0007927789000004
    Figure 0007927789000004
Patent Text Reader

Abstract

To provide an imaging apparatus capable of achieving high accuracy focus detection.SOLUTION: An imaging apparatus 120 comprises a plurality of focus detection pixels arranged side by side in a row direction and a column direction and receiving a light flux transmitting through different pupil partial areas of an imaging optical system. The imaging apparatus further comprises: an image pickup device 122 sequentially starting exposing and reading for each row of the pixels or for each of plural rows of blocks; focus detection means for performing first focus detection using an output signal from a focus detection pixel with row direction as pupil division direction and second focus detection using an output signal from a focus detection pixel with column direction as pupil division direction; and determination means 125 for determining whether to use the result of the second focus detection according to drive information of the image pickup device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image pickup apparatus capable of autofocus (AF). [Background Art]

[0002] As an image pickup apparatus that performs focus detection by a phase difference detection method using an image pickup element such as a CMOS sensor, Patent Document 1 discloses a configuration that performs focus detection in a plurality of mutually different focus detection directions. In an image pickup apparatus having an image pickup element such as a CMOS sensor, a rolling electronic shutter that reads out an image screen line by line is used. For this reason, when a live view mode is operated in such an image pickup apparatus, there arises a problem that horizontal stripe-shaped flicker (line flicker) occurs within an image pickup screen. Flicker occurs when displaying or recording a moving image of a subject under a light source directly lit by a commercial power source such as under fluorescent lamp illumination, depending on the accumulation time of the image pickup element, the frame frequency of the image pickup element, and the AC lighting frequency of the fluorescent lamp. By setting the accumulation time of the image pickup element to an integer multiple of the lighting cycle of the fluorescent lamp, the exposure amount for each line can be made uniform, and the influence of flicker can be reduced.

[0003] Patent Document 2 discloses a configuration in which focus detection control is changed depending on whether the influence of flicker is reduced or not, in order to improve focus detection accuracy in a situation where flicker occurs. In the configuration of Patent Document 2, utilization is made of the fact that the magnitude of the influence of flicker differs depending on the pupil division direction of the focus detection signal, and different directions are used accordingly.

[0004] Furthermore, in recent years, higher-frequency flicker occurs in digital signage and the like compared to that in fluorescent lamps and the like. Patent Document 3 discloses a configuration for detecting high-frequency flicker. [Prior Art Documents] [Patent Documents]

[0005] [Patent Document 1] Japanese Unexamined Patent Publication No. 2011-237215 [Patent Document 2] Japanese Patent Publication No. 2010-263568 [Patent Document 3] Japanese Patent Publication No. 2022-129925 [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] In the configuration described in Patent Document 2, the focus detection means used is changed to take into account the influence of flicker on the entire shooting environment. However, since the flicker that occurs in modern digital signage etc. occurs only in a part of the shooting environment (shooting range), if the focus detection area is outside the range of the digital signage, its use is unnecessarily restricted. Furthermore, Patent Document 3 does not disclose a method for detecting flicker using a focus detection signal.

[0007] The present invention aims to provide an imaging device capable of achieving highly accurate focus detection. [Means for solving the problem]

[0008] An imaging device as one aspect of the present invention includes an image sensor having a plurality of focus detection pixels arranged in a row direction and a column direction, which receive light beams passing through different pupil regions of the imaging optical system, and which sequentially starts exposure and reads out each row of pixels or each block of multiple rows; a focus detection means that performs a first focus detection using output signals from focus detection pixels with the row direction as the pupil division direction, and a second focus detection using output signals from focus detection pixels with the column direction as the pupil division direction; and a determination means that determines whether to use the result of the second focus detection according to the drive information of the image sensor. The determination means determines that the result of the second focus detection will not be used if the difference between the result of the second focus detection and the amount of image shift based on the drive information is within a predetermined value. It is characterized by the following. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide an imaging device that can achieve highly accurate focus detection. [Brief explanation of the drawing]

[0010] [Figure 1] 1 is a block diagram showing the configuration of a camera system according to an embodiment of the present invention. [Figure 2] It is a diagram showing a pixel array of an image sensor. [Figure 3] It is an explanatory diagram of a pixel. [Figure 4] It is a diagram showing pupil division. [Figure 5] It is another diagram showing pupil division. [Figure 6] It is a diagram showing the relationship between an image shift amount and a defocus amount. [Figure 7] It is an arrangement diagram of focus detection areas. [Figure 8] It is a diagram showing the overall flow of live view imaging processing. [Figure 9] It is a flowchart of an imaging subroutine. [Figure 10] It is a flowchart of subject tracking AF processing. [Figure 11] It is a flowchart of subject detection and tracking processing. [Figure 12] It is a diagram showing an example of a CNN that infers the likelihood of a specific region. [Figure 13] It is a flowchart of flicker determination. [Figure 14] It is a diagram for explaining the influence that flicker blinking exerts on a pair of signals in vertical direction focus detection. [Figure 15] It is a diagram showing a waveform when flicker occurs. [Figure 16] It is another diagram showing a waveform when flicker occurs. [Figure 17] It is a flowchart of defocus amount selection processing. [Figure 18] It is a diagram showing a setting method of a defocus map. [Figure 19] It is another diagram showing a setting method of a defocus map. [Figure 20] It is a diagram showing a histogram of a defocus map. [Figure 21]FIG. 11 is a diagram showing a histogram generated from a defocus map using specific area information. [Figure 22] FIG. 12 is a flowchart of focus detection processing. [Figure 23] FIG. 13 is a diagram showing an execution order of subject tracking AF processing. DETAILED DESCRIPTION OF EMBODIMENTS

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each drawing, the same reference numerals are assigned to the same members, and overlapping descriptions are omitted.

[0012] FIG. 1 is a block diagram showing a configuration of an imaging system 10 including a camera body (imaging apparatus) 120 according to an embodiment of the present invention. A lens unit (interchangeable lens) 100 is detachably attached to the camera body 120 as a digital camera via a mount M indicated by a dotted line in the figure. Note that the camera body 120 may be one provided with an imaging optical system integrally. Further, the camera body 120 is not limited to a digital camera, and may be another imaging apparatus such as a video camera.

[0013] The lens unit 100 includes a first lens group 101, an aperture 102, a second lens group 103, and a focus lens group (hereinafter simply referred to as a focus lens) 104 as a focus element, which constitute an imaging optical system, and a drive / control system. The imaging optical system takes in light from a subject to form a subject image.

[0014] The first lens group 101 is disposed closest to the object side and is held movably in an optical axis direction in which an optical axis OA extends. The aperture 102 adjusts the amount of light by changing its opening diameter. The aperture 102 and the second lens group 103 are integrally movable in the optical axis direction, and perform zooming by moving in conjunction with the first lens group 101. The focus lens 104 performs focusing by moving in the optical axis direction. Automatic focus adjustment (AF) is performed by controlling the position of the focus lens 104 in accordance with a focus detection result described later.

[0015] The drive / control system includes a zoom actuator 111, an aperture actuator 112, a focus actuator 113, a zoom drive circuit 114, an aperture drive circuit 115, a focus drive circuit 116, a lens MPU 117, and a lens memory 118.

[0016] The zoom drive circuit 114 drives the zoom actuator 111 during zooming to move the first lens group 101 and the third lens group 103 in the direction of the optical axis. The aperture drive circuit 115 drives the aperture actuator 112 to operate the aperture 102, thereby performing aperture operation and shutter operation.

[0017] The focus drive circuit 116 drives the focus actuator 113 during focusing to move the focus lens 104 in the direction of the optical axis. The focus drive circuit 116 also functions as a position detection unit that detects the current position of the focus lens 104 (hereinafter referred to as the focus position) through the focus actuator 113.

[0018] The lens MPU 117 is a computer that performs calculations and processing related to the lens unit 100, and controls the zoom drive circuit 114, aperture drive circuit 115, and focus drive circuit 116. The lens MPU 117 is also connected to the camera MPU 125 in the camera body 120 via the communication terminal of the mount M, and exchanges commands and data with the camera MPU 125. For example, the lens MPU 117 notifies the camera MPU 125 of lens information in response to a request from the camera MPU 125. This lens information includes information such as the focus position, the position and diameter of the exit pupil of the imaging optical system in the optical axis direction, and the position and diameter of the lens frame that limits the light beam of the exit pupil in the optical axis direction.

[0019] Furthermore, the lens MPU 117 controls the zoom drive circuit 114, aperture drive circuit 115, and focus drive circuit 116 in response to requests from the camera MPU 125. The lens memory 118 stores the optical information necessary for autofocus. The camera MPU 125 controls the lens unit 100 by executing programs stored in its built-in non-volatile memory and the lens memory 118.

[0020] The camera body 120 includes an optical low-pass filter 121, an image sensor 122, and a drive / control system. The optical low-pass filter 121 is provided to reduce false colors and moiré patterns.

[0021] The image sensor 122 consists of a CMOS sensor and its peripheral circuits, and converts the subject image (optical image) formed by the imaging optical system into an image by photoelectric conversion, outputting an imaging signal and a pair of focus detection signals (two-image signals). The image sensor 122 has multiple imaging pixels arranged in a horizontal direction of m pixels and a vertical direction of n pixels (where m and n are integers of 2 or more) that are orthogonal to the horizontal direction. Each imaging pixel includes a pair of focus detection pixels as described later, and has a pupil division function that enables focus detection using a phase difference detection method.

[0022] The drive / control system includes an image sensor drive circuit 123, a shutter 133, an image processing circuit 124, and a camera MPU (determination means) 125. The drive / control system also includes a display 126, an operation switch (SW) 127, a memory 128, a phase-difference AF unit (focus detection means) 129, a subject detection unit 130, an AE unit 131, and a white balance (WB) adjustment unit 132. The image sensor drive circuit 123 controls charge accumulation and signal readout in the image sensor 122, and performs A / D conversion on the imaging signal and the pair of focus detection signals output from the image sensor 122 and outputs them to the image processing circuit 124 and the camera MPU 125. The image processing circuit 124 performs image processing such as gamma conversion, color interpolation, and compression encoding on the digital imaging signal from the image sensor drive circuit 123 to generate image data.

[0023] The camera MPU 125 is a computer that performs calculations and processing related to the camera body 120, and controls the image sensor drive circuit 123, image processing circuit 124, display 126, phase-detection AF unit 129, subject detection unit 130, AE unit 131, and WB adjustment unit 132. The camera MPU 125 is also connected to the lens MPU 117 via the communication terminal of the mount M, and exchanges commands and data with the lens MPU 117. For example, the camera MPU 125 requests lens information and optical information from the lens MPU 117, and requests the drive of the first lens group 101, the focus lens 104, and the aperture 102. The camera MPU 125 receives lens information and optical information transmitted from the lens MPU 117.

[0024] The camera MPU 125 incorporates a ROM 125a for storing various programs, a RAM 125b for storing variables, and an EEPROM 125c for storing various parameters. The camera MPU 125 executes various processes, including the AF process described later, according to the program stored in the ROM 125a. The camera MPU 125 generates two-image data from a pair of digital focus detection signals from the image sensor drive circuit 123 and outputs it to the phase-detection AF unit 129.

[0025] The shutter 133 has a focal-plane shutter configuration and drives the focal-plane shutter based on instructions from the camera MPU 125 and commands from the shutter drive circuit built into the shutter 133. While reading the signal from the image sensor 122, the shutter blocks light from the image sensor 122. When exposure is taking place, the focal-plane shutter opens and the photographic light beam is directed to the image sensor 122.

[0026] The display unit 126 consists of an LCD or the like and displays information related to the imaging mode, a preview image before imaging, a confirmation image after imaging, and the focus status. The operation SW127 includes a power switch, a release (imaging instruction) switch, a zoom switch, and an imaging mode selection switch. The memory 128 is a flash memory that can be attached to the camera body 120 and records the images obtained by imaging.

[0027] The phase-difference AF unit 129 performs focus detection using two image data generated by the camera MPU 125. The image sensor 122 performs photoelectric conversion of a pair of optical images formed by light beams passing through different pairs of pupil regions of the exit pupil of the imaging optical system and outputs a pair of focus detection signals. The phase-difference AF unit 129 performs correlation calculation on the two image data generated by the camera MPU 125 from the pair of focus detection signals to calculate the image shift amount, which is the phase difference between them. The phase-difference AF unit 129 also calculates (acquires) a defocus amount as information about the focus from the calculated image shift amount. The camera MPU 125 calculates the drive amount of the focus lens 104 based on the defocus amount calculated by the phase-difference AF unit 129 and transmits a focus control command including this drive amount to the lens MPU 117.

[0028] Thus, in this embodiment, instead of using an AF sensor dedicated to focus detection, image plane phase-detection AF is performed using the output of the image sensor 122. In this embodiment, the phase-detection AF unit 129 has an acquisition unit 129a that acquires two image data and a calculation unit (calculation means) 129b that calculates the amount of defocus. At least one of the acquisition unit 129a and the calculation unit 129b may be provided in the camera MPU 125.

[0029] The subject detection unit 130 performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 130 uses dictionary data for each subject in order to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 130 performs subject detection by sequentially switching between the dictionary data for each subject. The dictionary data for each subject is stored in the dictionary data storage unit (ROM 125a in the camera MPU 125). Therefore, multiple dictionary data are stored in the dictionary data storage unit. The camera MPU 125 determines which dictionary data to use for subject detection from the multiple dictionary data based on the priority of subjects set in advance and the settings of the camera body 120.

[0030] The AE unit 131 performs exposure control (AE) by performing photometering using AE image data obtained from the image processing circuit 124. Specifically, the AE unit 131 acquires brightness information from the AE image data and calculates the aperture value, shutter speed (shutter time in seconds), and ISO sensitivity as imaging conditions from the difference between the exposure amount obtained from this brightness information and a preset exposure amount. Then, it performs AE by controlling the aperture value, shutter speed, and ISO sensitivity to match the calculated values.

[0031] The WB adjustment unit 132 calculates the white balance (WB) of the WB adjustment image data obtained from the image processing circuit 124, and performs WB adjustment by adjusting the weights of the RGB colors according to the difference between the calculated WB and a predetermined appropriate WB.

[0032] Furthermore, the camera MPU 125 can select an image height range for performing phase-detection AF, AE, and WB adjustment according to the position and size of the subject in the imaging area detected by the subject detection unit 130. (Regarding image sensor 122) Figure 2 shows the pixel arrangement on the imaging surface of the image sensor 122 as a two-dimensional CMOS sensor. Figure 2(a) is a schematic diagram showing an example of the overall configuration of the image sensor 122. The image sensor 122 includes a pixel array section 208, a vertical selection circuit 209, a column circuit 203, and a horizontal selection circuit 204.

[0033] Multiple pixels 205 are arranged in a matrix in the pixel array section 208. The output of the vertical selection circuit 209 is input to the pixels 205 via the pixel drive wiring group 207, and the pixel signals of the pixels 205 in the row selected by the vertical selection circuit 209 are read out row by row to the column circuit 203 via the output signal line 206. One output signal line 206 can be provided for each pixel column, or for each of multiple pixel columns, or multiple output signal lines 206 can be provided for each pixel column. The column circuit 203 receives the signals read out in parallel via the multiple output signal lines 206, performs processing such as signal amplification, noise reduction, and A / D conversion, and holds the processed signals. The horizontal selection circuit 204 sequentially, randomly, or simultaneously selects the signals held in the column circuit 203, and the selected signals are output outside the image sensor 122 via a horizontal output line and output section (not shown).

[0034] By sequentially performing the operation of outputting the pixel signals of the rows selected by the vertical selection circuit 209 to the outside of the image sensor 122, while changing the rows selected by the vertical selection circuit 209, it is possible to read out a two-dimensional imaging signal or a phase difference signal from the image sensor 122.

[0035] Figure 2(b) is an equivalent circuit diagram of pixel 205. Pixel 205 has two photodiodes 211 (PDA) and 212 (PDB), which are photoelectric conversion units. Depending on the amount of incident light, the signal charge converted photoelectrically by PDA 211 and stored is transferred via transfer switch (TXA) 213 to the floating diffusion unit (FD) 215, which constitutes the charge storage unit. The signal charge converted photoelectrically by PDB 212 and stored is transferred via transfer switch (TXB) 214 to FD 215. When the reset switch (RES) 216 is turned on, it resets FD 215 to the voltage of the constant voltage source VDD. Also, by turning on RES 216 and TXA 213 and TXB 214 simultaneously, PDA 211 and PDB 212 can be reset.

[0036] When the pixel selection switch (SEL) 217 ​​is turned ON, the amplification transistor (SF) 218 ​​converts the signal charge stored in FD 215 into a voltage, and the converted signal voltage is output from the pixel to the output signal line 206. In addition, the gates of TXA 213, TXB 214, RES 216, and SEL 217 are connected to the pixel drive wiring group 207 and controlled by the vertical selection circuit 209.

[0037] In the following description, in this embodiment, the signal charge accumulated in the photoelectric conversion unit is assumed to be electrons, the photoelectric conversion unit is formed from an N-type semiconductor, and the separation is performed using a P-type semiconductor. However, the signal charge may be assumed to be a hole, the photoelectric conversion unit may be formed from a P-type semiconductor, and the separation may be performed using an N-type semiconductor.

[0038] Next, we will describe the operation of reading the signal charge from PDA211 and PDB212 after resetting PDA211 and PDB212 and after a predetermined charge accumulation time has elapsed, in a pixel having the above-described configuration. First, when the SEL217 of the row selected by the vertical selection circuit 209 is turned on and the source of SF218 and the output signal line 206 are connected, the output signal line 206 enters a state where a voltage corresponding to the voltage of FD215 is read out. Subsequently, RES216 is turned on / off and the potential of FD215 is reset. After that, the system waits until the output signal line 206, which has been affected by the voltage fluctuation of FD215, stabilizes, and the voltage of the stabilized output signal line 206 is acquired as the signal voltage N by the column circuit 203, and the signal is processed and held.

[0039] Subsequently, TXA213 is switched on / off, and the signal charge stored in PDA211 is transferred to FD215. The voltage of FD215 decreases by an amount corresponding to the amount of signal charge stored in PDA211. Then, the system waits until the output signal line 206, which has been affected by the voltage fluctuation of FD215, stabilizes. The voltage of the stabilized output signal line 206 is then taken as signal voltage A by the column circuit 203, processed, and held.

[0040] Subsequently, TXB214 is switched on / off, and the signal charge stored in PDB212 is transferred to FD215. The voltage of FD215 decreases by an amount corresponding to the amount of signal charge stored in PDB212. Then, the circuit waits until the output signal line 206, which has been affected by the voltage fluctuation of FD215, stabilizes. The voltage of the stabilized output signal line 206 is then taken as the signal voltage (A+B) by the column circuit 203, processed, and held.

[0041] In this way, the difference between the acquired signal voltage N and signal voltage A allows us to obtain signal A corresponding to the amount of signal charge stored in PDA211. Furthermore, the difference between signal voltage A and signal voltage (A+B) allows us to obtain signal B corresponding to the amount of signal charge stored in PDB212. This difference calculation may be performed in the column circuit 203 or after output from the image sensor 122. A phase difference signal can be obtained using signals A and B respectively, and the imaging signal can be obtained by adding signals A and B together. Alternatively, if the difference calculation is performed after output from the image sensor 122, the imaging signal may be obtained by taking the difference between signal voltage N and signal voltage (A+B).

[0042] Alternatively, the same drive used to read out signal voltages N and A may be applied to the PDB212 instead of the PDA211 to read out signal voltages N, A, and B, respectively. In this case, signals A and B obtained from signal voltages A and B, respectively, can be used directly as phase difference signals, and an imaging signal can be obtained by adding signal voltages A and B, or signal A and B.

[0043] In this embodiment, the pixel from which signal A is obtained is referred to as the first focus detection pixel, and the pixel from which signal B is obtained is referred to as the second focus detection pixel.

[0044] Figure 2(c) shows the arrangement of imaging pixels in a 4x4 grid. One pixel group 200 contains imaging pixels in a 2x2 grid. Pixel group 200 includes pixel 200R, located in the upper left and having a spectral sensitivity of R (red); pixels 200Ga and 200Gb, located in the upper right and lower left and having a spectral sensitivity of G (green); and pixel 200B, located in the lower right and having a spectral sensitivity of B (blue). Each imaging pixel consists of a first focus detection pixel 201 and a second focus detection pixel 202. In pixels 200R, 200Ga, and 200B, the first focus detection pixel 201 and the second focus detection pixel 202 are arranged horizontally (in the row direction), while in pixel 200Gb, the first focus detection pixel 201 and the second focus detection pixel 202 are arranged vertically (in the column direction).

[0045] Figure 3 is an explanatory diagram of a pixel. Figure 3(a) shows pixel 200Ga as viewed from the incident side (+z side) of the image sensor 122. Figure 3(b) shows the pixel structure as viewed from the -y side of cross-section aa in Figure 3(a). In pixel 200Ga, a microlens 305 for focusing incident light is formed on the incident side, and two photoelectric conversion sections 301 and 302 are formed, divided in the x direction. The photoelectric conversion sections 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively.

[0046] The photoelectric conversion units 301 and 302 may be pin-structured photodiodes with an intrinsic layer sandwiched between a p-type layer and an n-type layer, or they may be pn-junction photodiodes without an intrinsic layer. A color filter 306 is formed between the microlens 305 and the photoelectric conversion units 301 and 302. The spectral transmittance of the color filter may be changed for each focus detection pixel, or the color filter may be omitted.

[0047] Two beams of light, incident on the 200Ga pixel from a pair of pupil regions, are each focused by a microlens 305, spectrally separated by a color filter 306, and then received by photoelectric conversion units 301 and 302. In each photoelectric conversion unit, electrons and holes are generated in pairs according to the amount of light received. After separation in the depletion layer, the negatively charged electrons are accumulated in the n-type layer. Meanwhile, the holes are discharged outside the image sensor 122 through a p-type layer connected to a constant voltage source (not shown). The electrons accumulated in the n-type layer of each photoelectric conversion unit are transferred to the capacitance unit (FD) via a transfer gate and converted into a voltage signal.

[0048] Figure 4 shows the pupil division. The lower part of Figure 4 shows the pixel structure when the cross-section aa in Figure 3(a) is viewed from the +y side, and the upper part shows the pupil plane at pupil distance DS. Note that in Figure 4, the x and y axes of the pixel structure are inverted compared to Figure 3(b) in order to correspond with the coordinate axes of the pupil plane. The pupil plane corresponds to the entrance pupil position of the image sensor 122. In this embodiment, the entrance pupil of one image sensor 122 is formed by offsetting (shrinking) the microlens position of each pixel from the center of the image sensor 122 so that the entrance pupils of each pixel overlap each other. The pupil distance DS is the distance between the pupil plane and the imaging plane, and will be referred to as the sensor pupil distance in the following explanation.

[0049] As shown in Figure 4, the first pupil region (first pupil portion region) 501 of the first focus detection pixel 201 is roughly conjugate to the light-receiving surface of the photoelectric conversion unit 301, whose centroid is eccentric in the -x direction, by a microlens. The first pupil region 501 is the pupil region through which the light beam that can be received by the first focus detection pixel 201 passes. The centroid of the first pupil region 501 is eccentric to the +X side on the pupil surface. Similarly, the second pupil region (second pupil portion region) 502 of the second focus detection pixel 202 is roughly conjugate to the light-receiving surface of the photoelectric conversion unit 302, whose centroid is eccentric in the +x direction, by a microlens. The second pupil region 502 is the pupil region through which the light beam that can be received by the second focus detection pixel 202 passes. The centroid of the second pupil region 502 is eccentric to the -X side on the pupil surface. The pupil region 500 is the pupil region through which the light beam that can be received by the entire 200G of pixels, which are the photoelectric conversion units 301 and 302 (first focus detection pixel 201 and second focus detection pixel 202), passes.

[0050] As shown in Figure 5, the light beam that enters the imaging optical system from the subject (vertical line on the left in the figure) and passes through the first pupil region 501 and the second pupil region 502 respectively is incident on each imaging pixel at different angles and received by the photoelectric conversion units 301 and 302. Pixels 200R, 200Ga, and 200B perform pupil division in the horizontal direction (x-axis direction in Figure 4), while pixel 200Gb performs pupil division in the vertical direction (y-axis direction in Figure 4). Each imaging pixel, having a first focus detection pixel and a second focus detection pixel, receives the light beam passing through the first pupil region 501 and the second pupil region 502. A pair of focus detection signals is generated by combining the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of multiple imaging pixels. Furthermore, an imaging signal with a resolution of effective pixels N (=m × n) is generated by adding the output signals of the first focus detection pixel 201 and the second focus detection pixel 202 of multiple imaging pixels. Alternatively, the other focus detection signal may be generated by subtracting one of the paired focus detection signals from the imaging signal.

[0051] Furthermore, in this embodiment, first and second focus detection pixels are provided for each of the imaging pixels of the image sensor 122, but two imaging pixels may be used as the first and second focus detection pixels, or some of the imaging pixels may be provided with first and second focus detection pixels. (Regarding the relationship between the amount of defocus and the amount of image displacement) Figure 6 shows the relationship between the amount of image shift and the amount of defocus in two image data. 800 represents the imaging surface of the image sensor 122, and the pupil surface of the image sensor 122 is divided into two parts: the first pupil region 501 and the second pupil region 502. The amount of defocus d is defined as |d| being the magnitude of the distance from the imaging position of the subject image to the imaging surface 800, with a negative sign (d<0) indicating a front-focus state where the image position is on the subject side of the imaging surface, and a positive sign (d>0) indicating a back-focus state where the image position is on the opposite side of the imaging surface 800 from the subject. The in-focus state where the image position is on the imaging surface 800 is d=0.

[0052] In Figure 6, subject 801 is in focus (d=0), while subject 802 is front-focused (d<0). The front-focused state (d<0) and the back-focused state (d>0) together constitute a defocused state (|d|>0).

[0053] In the front-focused state, the light beam from the subject 802 that has passed through the first pupil region 501 and the second pupil region 502, respectively, is focused once, then spreads out with widths Γ1 and Γ2 centered on the centroid positions G1 and G2 of the light beam, forming a blurred optical image on the image sensor 800. These blurred images are received by the first focus detection pixel 201 and the second focus detection pixel 202 of each image sensor on the image sensor 800, thereby generating a pair of focus detection signals: a first focus detection signal and a second focus detection signal. The first and second focus detection signals are recorded as blurred images of the subject 802 spread out with blur widths Γ1 and Γ2 at the centroid positions G1 and G2 on the image sensor 800, respectively. The blur widths Γ1 and Γ2 increase approximately proportionally to the increase in the magnitude of the defocus amount d |d|. Similarly, the magnitude of the image shift amount p (= difference in the centroid position of the light beam G1-G2) between the first focus detection signal and the second focus detection signal, |p|, increases roughly in proportion to the increase in the magnitude of the defocus amount d, |d|. The same applies to the back-focused state (d>0), although the direction of the image shift between the first and second focus detection signals is opposite to that of the front-focused state.

[0054] In this embodiment, the difference in the centroids of the incident angle distributions in the first pupil region 501 and the second pupil region 502 is called the baseline length. The relationship between the defocus amount d and the image shift amount p on the imaging surface 800 is generally similar to the relationship between the baseline length and the sensor pupil distance. As the magnitude of the defocus amount d increases, the magnitude of the image shift amount p between the first focus detection signal and the second focus detection signal increases. Therefore, the phase difference AF unit 129 converts the image shift amount p to the defocus amount d using a conversion coefficient calculated based on the baseline length from this relationship.

[0055] In the following explanation, calculating the amount of defocus using a pair of focus detection signals from a focus detection pixel that divides the pupil horizontally (lateral direction), such as a 200Ga pixel, is called horizontal focus detection (first focus detection). Similarly, calculating the amount of defocus using a pair of focus detection signals from a focus detection pixel that divides the pupil vertically (vertical direction), such as a 200Gb pixel, is called vertical focus detection (second focus detection). (Regarding the placement of the focus detection area) Next, the focus detection region of the image sensor 122, which is the region that acquires pairs of signal sequences for detecting phase difference, will be explained using Figure 7. In this embodiment, the camera MPU 125 sets the focus detection region. Figure 7 is a diagram showing the arrangement of the focus detection region in this embodiment. A(n,m) and B(n,m) indicate the nth focus detection region in the x direction and the mth focus detection region in the y direction, out of a total of nine focus detection regions (three each in the x and y directions) set in the effective pixel region 300 of the image sensor 122. A signal sequence of pixel pairs divided horizontally is generated from multiple pixels contained in focus detection region A(n,m). A signal sequence of pixel pairs divided vertically is generated from multiple pixels contained in focus detection region B(n,m). I(n,m) indicates an index that displays the position of focus detection region A(n,m) and B(n,m) on the display unit 126. By arranging the focus detection area in this way, focus detection can be performed at the position of the I(n,m) index using contrast information corresponding to both the horizontal and vertical directions of the subject.

[0056] The nine focus detection regions shown in Figure 7 are merely examples, and the number, position, and size of the focus detection regions are not limited. For example, one or more regions may be set as focus detection regions within a predetermined range centered on a position specified by the user or the position of the subject detected by the subject detection unit 130. In this embodiment, the focus detection regions are arranged to obtain focus detection results with higher resolution when acquiring the defocus map, which will be described later. For example, the group of focus detection results obtained from horizontal focus detection regions arranged on the image sensor 122 at a total of 187 points (17 horizontal divisions and 11 vertical divisions) is used as the horizontal defocus map. Similarly, the group of focus detection results obtained from vertical focus detection regions arranged at a total of 35 points (7 horizontal divisions and 5 vertical divisions) is used as the vertical defocus map. Further details on the arrangement of the focus detection regions for horizontal focus detection and vertical focus detection relative to the subject will be described later. (Photo processing) Figure 8 shows the overall flow of the live view shooting process. Specifically, it shows the process of having the camera body 120 perform actions from before image capture to still image capture, displaying the live view image on the display unit 126. The camera MPU 125, which is a computer, executes this process according to the computer program. In the following explanation, S means step.

[0057] In S1, the camera MPU 125 drives the image sensor 122 using the image sensor drive circuit 123 and acquires imaging data from the image sensor 122. Subsequently, the camera MPU 125 acquires first and second focus detection signals from multiple first and second focus detection pixels included in each of the focus detection regions shown in Figure 7 from the acquired imaging data. The camera MPU 125 also generates an imaging signal by adding the first and second focus detection signals of all effective pixels of the image sensor 122, and has the image processing circuit 124 perform image processing on the imaging signal (imaging data) to acquire image data. If the imaging pixels and the first and second focus detection pixels are provided separately, the camera MPU 125 performs interpolation processing on the focus detection pixels to acquire image data.

[0058] In S2, the camera MPU 125 causes the image processing circuit 124 to generate a live view image from the image data obtained in S1, and displays this on the display unit 126. The live view image is a scaled-down image matched to the resolution of the display unit 126, allowing the user to adjust the imaging composition, exposure conditions, etc., while viewing it. Therefore, the AE unit 131 and the camera MPU 125 adjust the exposure based on the photometric values ​​obtained from the image data and display it on the display unit 131. Exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the photographic lens, and adjusting the gain for the output of the image sensor 122.

[0059] In S3, the camera MPU125 determines whether the switch Sw1, which instructs the start of the image preparation operation, has been turned on by half-pressing the release switch included in the operation SW127. If Sw1 is not turned on, the camera MPU125 repeats the determination in S3 to monitor the timing when Sw1 will be turned on. On the other hand, if Sw1 is turned on, the camera MPU125 proceeds to S400 and performs subject tracking autofocus (AF) processing. Here, it performs subject area detection from the obtained image capture signal and focus detection signal, sets the focus detection area, and performs predictive AF processing to suppress the effect of the time lag between the focus detection processing and the image capture processing of the recorded image.

[0060] In S5, the camera MPU125 determines whether the switch Sw2, which instructs the start of the imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera MPU125 returns to S3. On the other hand, if Sw2 is turned on, the process proceeds to S300 and the imaging subroutine is executed. Details of the imaging subroutine will be described later.

[0061] In S7, the camera MPU125 determines whether the main switch included in the control SW127 is turned off or not. If the main switch is turned off, the camera MPU125 terminates this process; otherwise, it returns to S3.

[0062] In this embodiment, subject detection processing and AF processing are performed after Sw1 is detected as being ON in S3, but the timing of these processes is not limited to this. By performing the subject tracking AF processing in S400 before Sw1 is turned ON, it is possible to eliminate the need for the photographer to perform preparatory actions before shooting. (Regarding the shooting subroutine) Referring to Figure 9, the shooting subroutine executed by the camera MPU 125 in S300 of Figure 8 will be described. Figure 9 is a flowchart of the shooting subroutine.

[0063] In step S301, the AE unit 131 performs exposure control processing and determines the imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.). This exposure control processing can be performed using brightness information obtained from the image data of the live view image.

[0064] The camera MPU 125 then transmits the determined aperture value to the aperture drive circuit 115 to drive the aperture 102. The camera MPU 125 also transmits the determined shutter speed to the shutter 133 to open the focal-plane shutter. Furthermore, the camera MPU 125 causes the image sensor 122 to accumulate charge during the exposure period via the image sensor drive circuit 123.

[0065] In step S302, the image sensor drive circuit 123 is instructed to read out all pixels of the imaging signal from the image sensor 122 for still image capture. The camera MPU 125 also instructs the image sensor drive circuit 123 to read out one of the first and second focus detection signals from the focus detection area (focus target area) within the image sensor 122. The other focus detection signal can be obtained by subtracting one of the first and second focus detection signals from the imaging signal.

[0066] In S303, the camera MPU 125 instructs the image processing circuit 124 to perform defective pixel correction processing on the image data read out in S302 and converted by A / D.

[0067] In S304, the camera MPU 125 instructs the image processing circuit 124 to perform image processing and encoding processes such as demosaicing (color interpolation), white balance processing, gamma correction (gradation correction), color conversion, and edge enhancement on the captured data after defective pixel correction processing.

[0068] In S305, the camera MPU 125 records the still image data obtained as image data through the image processing and encoding process in S304, along with one of the focus detection signals read out in S302, as an image data file in the memory 128.

[0069] In S306, the camera MPU 125 records camera characteristic information, which is characteristic information of the camera body 120, in the lens memory 118 and the memory 128 within the camera MPU 125, corresponding to the still image data recorded in S305. The camera characteristic information includes, for example, the following information. • Imaging conditions (aperture value, shutter speed, ISO sensitivity, etc.) • Information regarding image processing performed by image processing circuit 124 • Information regarding the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 122. • Information regarding vignetting of the imaging light beam within the camera body 120 • Information on the distance from the mounting surface of the imaging optical system in the camera body 120 to the image sensor 122. • Information regarding the manufacturing tolerance of the camera body (120 units).

[0070] Information regarding the light sensitivity distribution of imaging pixels and focus detection pixels (hereinafter simply referred to as light sensitivity distribution information) is information regarding the sensitivity of the image sensor 122 according to the distance (position) on the optical axis from the image sensor 122. Since this light sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, it may also be information regarding these. Furthermore, the light sensitivity distribution information may also be information regarding the change in sensitivity with respect to the angle of incidence of light.

[0071] In S307, the camera MPU 125 records lens characteristic information as characteristic information of the imaging optical system in memory 128 and the memory within the camera MPU 125, corresponding to the still image data recorded in S305. The lens characteristic information may include information about the exit pupil, information about the frame such as the lens barrel that deflects the light beam, information about the focal length and F number at the time of imaging, and information about aberrations of the imaging optical system. The lens characteristic information may also include information about the manufacturing tolerance of the imaging optical system and information about the position (subject distance) of the focus lens 104 at the time of imaging.

[0072] In S308, the camera MPU 125 records image-related information, which is information related to still image data, in the memory 128 and in the memory within the camera MPU 125. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.

[0073] In S309, the camera MPU 125 displays a preview of the captured image on the display unit 126. This allows the user to easily check the captured image.

[0074] Once processing in S309 is complete, the camera MPU125 terminates this imaging subroutine and proceeds to S7 in Figure 8. (Subroutine for subject tracking AF processing) Referring to Figure 10, the subject-tracking AF processing subroutine executed by the camera MPU 125 in S400 of Figure 8 will be explained. Figure 10 is a flowchart of the subject-tracking AF processing. The order in which steps S401 to S406 of this flow are executed will be explained later using Figure 23.

[0075] In S401, the camera MPU 125 and the phase-detection AF unit 129 perform focus detection processing using the first and second focus detection signals obtained in each of the multiple focus detection regions acquired in S1. Details will be described later.

[0076] In S402, the camera MPU 125 performs subject detection and tracking. Subject detection is performed by the subject detection unit 130 described above. Subject detection may not be possible depending on the state of the obtained image; in such cases, tracking is performed using other means such as template matching to estimate the position of the subject. Details will be described later.

[0077] In S403, the camera MPU125 performs primary subject determination processing. The primary subject is determined according to a priority order based on predetermined criteria. For example, the closer the subject detection area is to the central image height, the higher the priority is set. If the positions are the same (the same distance from the central image height), the larger the size, the higher the priority is set. Alternatively, a configuration may be adopted that uses a defocus map to select the part of a particular type of subject (person) that the photographer often wants to focus on.

[0078] In S404, the camera MPU 125 and the phase-detection AF unit 129 perform flicker detection. They determine whether flicker is occurring within each focus detection area. Vertical focus detection may be affected by flicker, so the results of vertical focus detection are not used when the effect of flicker is expected to be significant. Details of the flicker detection method and the determination of whether or not to use vertical focus detection will be described later.

[0079] In S405, the camera MPU 125 and the phase-detection AF unit 129 perform defocus amount selection processing. Based on the subject information obtained in S402 and the flicker detection result obtained in S404, the focus detection result, which is the defocus amount, is selected using the positioned horizontal and vertical defocus maps. Details will be described later.

[0080] In S406, the camera MPU125 performs predictive AF processing using the defocus amount obtained in S405 and multiple defocus amounts, which are time-series data of past focus detection timings. This processing is necessary when there is a time lag between the timing of focus detection and the timing of image exposure. Specifically, it is a process that predicts the position of the subject in the optical axis direction at the time of image exposure, which is a predetermined time after the timing of focus detection, and performs AF control. The prediction of the subject's image plane position is performed by multivariate analysis (e.g., least squares method) using historical data of the subject's image plane position and time, and the equation of the prediction curve is obtained. By substituting the time of image exposure into the obtained prediction curve equation, the predicted image plane position of the subject can be calculated. Furthermore, the position may be predicted not only in the optical axis direction but also in three dimensions. If the screen is considered as an XYZ vector with the XY direction and the optical axis direction as the Z direction, the position of the subject at the time of exposure of the captured image may be predicted from the time-series data of the subject's XY position obtained in S402 and its Z-direction position based on the defocus amount obtained in S405. Furthermore, prediction may also be made from time-series data of the joint positions of the subject, which is a person. With the above prediction, even if the ball or person is hidden or part of the joint positions of the person become invisible, the positions can be estimated by prediction. Prediction is performed not only on the main subject but also on multiple detected subjects. By performing predictive AF processing on multiple subjects, when the main subject is switched, there is no need to re-store the history of the defocus amount of the new main subject, and predictive AF can be continued without time loss. In S406, the amount of drive of the focus lens is calculated using the predictive AF processing result, and the focus actuator 113 is driven in accordance with the focus drive command from the camera MPU 125, and the focus adjustment processing is performed by moving the focus lens 104 in the optical axis direction.

[0081] Once the processing in S406 is complete, the camera MPU125 terminates the subject-tracking AF processing subroutine and proceeds to S5 in Figure 8.

[0082] Next, the execution order of S401 to S406 will be explained using Figure 23. Figure 23 is a diagram showing the execution order of the subject tracking AF process. In this embodiment, the focus detection process in S401 and the subject tracking process in S402 are executed simultaneously. S401 is executed by the camera MPU 125 and the phase-difference AF unit 129, and S402 is executed by the subject detection unit 130. S402 may be executed after S401 is completed. The focus detection process in S401 is performed after S2201 is completed, followed by S2202. In this embodiment, the horizontal defocus map is calculated first, and then the vertical defocus map is calculated. Alternatively, the vertical defocus map may be calculated first, and then the horizontal defocus map may be calculated.

[0083] The main subject determination process in S403 is executed after the completion of S402. In S403, a defocus map is used, but in this embodiment, since the calculation of the vertical defocus map has not been completed, the horizontal defocus map is used. Note that S403 may be executed after the completion of S401.

[0084] In this embodiment, S404 is executed after S401 and S403 are completed.

[0085] In this embodiment, S405 is executed after the completion of S404.

[0086] In this embodiment, S406 is executed after the completion of S405. (Subroutine for focus detection processing) Referring to Figure 22, the focus detection processing subroutine executed by the camera MPU 125 in S401 of Figure 10 will be described. Figure 22 is a flowchart of the focus detection process.

[0087] In S2201, the camera MPU 125 sets the focus detection area. In this embodiment, a total of 187 horizontal focus detection points are set on the image sensor 122, divided into 17 horizontal and 11 vertical sections. Furthermore, a total of 35 vertical focus detection points are set on the image sensor 122, divided into 7 horizontal and 5 vertical sections. The center of the focus detection area is set based on one of the following: the AF area set via the operation switch 127, the position of the subject detected and tracked in S402, or the position of the main subject determined in S403. In this embodiment, the group of focus detection results obtained from the horizontal focus detection area is called the horizontal defocus map. The group of focus detection results obtained from the vertical focus detection area is called the vertical defocus map.

[0088] The method for setting the defocus map, which is a set of focus detection results for horizontal and vertical eyes, will be explained using Figure 18. Figure 18 is a diagram showing the method for setting the defocus map. Figure 18(a) shows the subject area detected by the subject detection process described above, when the subject is a person. 1801 shows the upper body detection area, 1802 shows the face detection area, and 1803 shows the pupil detection area.

[0089] This section explains the arrangement of the side-eye defocus map, which is a set of focus detection results for side-eye detection. Figure 18(b) shows the side-eye defocus map during pupil detection, and 1804 is the side-eye defocus map. The side-eye defocus map is positioned relative to the center of the upper body detection area so as to encompass the subject. This makes it possible to keep the subject within the defocus map even when the subject is moving or when framing with the camera.

[0090] Next, we will explain the arrangement of the vertical eye defocus map, which is the group of vertical eye focus detection results. Figure 18(c) shows the vertical eye defocus map during face detection, and 1805 is the vertical eye defocus map. In this invention, we will explain under the premise that the vertical eye defocus map will be a smaller area than the horizontal eye defocus map due to computation time constraints. As the aforementioned horizontal eye defocus map can encompass the subject, the vertical eye defocus map is set based on the area that the photographer wants to focus on. In the case of a person, the area that the photographer wants to focus on is often the pupil, so in Figure 18(c), the vertical eye defocus map is set with the pupil detection area 1803 as the center. This allows the photographer to select the amount of defocus using both the horizontal eye defocus map and the vertical eye defocus map in the defocus amount selection process described later, in the area that the photographer wants to focus on.

[0091] If pupils are not detected, a vertical eye defocus map is set centered on the face detection region 1802, as shown in Figure 18(d). If a face is not detected, a vertical eye defocus map is set centered on the upper body detection region 1801, as shown in Figure 18(e). If a face is not detected, a vertical eye defocus map is set centered on the upper body detection region 1801.

[0092] The horizontal and vertical defocus maps are set so that the center positions and areas of each focus detection region are the same. This enables focus detection using signals from the same focus detection region, making it possible to use the horizontal and vertical defocus amounts together without distinction in the defocus amount selection process described later.

[0093] Figure 18(f) shows the case where the area of ​​the vertical eye defocus map is reduced, and each focus detection area is reduced. By densely arranging the vertical eye defocus map in the face detection area, it becomes possible to perform the defocus amount selection process described later using a larger amount of defocus.

[0094] Figure 18(g) shows an example where the subject is a motorcycle. 1806 is the overall detection area of ​​the motorcycle, and 1807 is the local detection area of ​​the motorcycle's helmet. Similar to the human subject, the side-view defocus map is positioned to encompass the overall detection area.

[0095] Figure 18(h) shows the setting of the vertical defocus map during local detection of a motorcycle. The vertical defocus map is not placed at the center of the local region 1807, but rather in a region that can match the position size of the horizontal defocus map and each focus detection region, and that also encompasses the local detection region. As a result, as mentioned above, the defocus amount is the result of both the horizontal focus detection result and the vertical focus detection result using the signal of the same focus detection region. This makes it possible to use the horizontal defocus amount and the vertical defocus amount together without separating them in the defocus amount selection process described later.

[0096] In S2202, the camera MPU 125 acquires a defocus map. For the focus detection region set in S2201, the phase-detection AF unit 129 calculates the amount of image shift between the first and second focus detection signals obtained in each of the multiple focus detection regions acquired in S2, and calculates the amount of defocus and reliability for each focus detection region from the amount of image shift. (Subroutine for subject detection and tracking) Referring to Figure 11, the subject detection and tracking subroutine executed by the camera MPU 125 in S402 of Figure 10 will be described. Figure 11 is a flowchart of the subject detection and tracking process.

[0097] In S421, the camera MPU 125 sets dictionary data according to the type of subject to be detected from the image data acquired in S1. Based on the pre-set subject priority and imaging device settings, it selects the dictionary data to be used in this process from multiple dictionary data stored in the dictionary data storage unit. For example, subjects are classified and stored as multiple dictionary data such as "people," "vehicles," and "animals." In this embodiment, one or more dictionary data may be selected. If one is selected, it becomes possible to repeatedly detect subjects that can be detected by one dictionary data at a high frequency. On the other hand, if multiple dictionary data are selected, subjects can be detected sequentially by setting the dictionary data sequentially according to the priority of the subjects to be detected.

[0098] In S422, the subject detection unit 130 uses the image data read in step S1 as the input image and performs subject detection using the dictionary data set in step S421. At this time, the subject detection unit 130 outputs information such as the position, size, and confidence level of the detected subject. At this time, the camera MPU 125 may display the above information output by the subject detection unit 130 on the display unit 126. In S422, multiple regions of the subject are detected hierarchically from the image data. For example, if "person" or "animal" is set as the dictionary data, multiple organs such as the "whole body" region, the "face" region, and the "eyes" region are detected. Local areas such as the eyes and face of a person are areas where the focus and exposure state should be adjusted as a subject, but they may not be detectable due to surrounding obstacles or the orientation of the face. Even in such cases, the system is configured to detect the subject hierarchically in order to continue robustly detecting the subject by performing whole-body detection. Similarly, if "vehicle" such as a motorcycle is set as dictionary data, the system is configured to detect the driver, the entire vehicle, and the helmet (head) as a local area in a hierarchical manner.

[0099] In S423, the camera MPU125 performs known template matching processing using the subject detection area obtained in S422 as a template. Using multiple images obtained in S1, it searches for similar areas in the most recently obtained image, using the subject detection area obtained in past images as a template. As is well known, any of the following information can be used for template matching: brightness information, color histogram information, feature point information such as corners and edges, etc. Various matching methods and template update methods can be considered, and any of them can be used. The tracking processing performed in S423 is done to achieve stable subject detection and tracking processing by detecting areas similar to past subject detection data from the most recently obtained image data when no subject was detected in S422.

[0100] In S424, the subject detection unit 130 performs region division on the detected subject area, specifically focusing on a particular region. A specific region is a part or the entirety of the detected subject area. For example, when detecting a person or animal, it might be the area of ​​the person's head, or when detecting a vehicle, it might be the area of ​​the helmet. Unlike subject detection, where the size and position of the subject are obtained from the size and coordinates of a rectangular area, region division allows the detection result to be obtained as a high-resolution distribution of the specific region. Any method can be applied to segment the image region (for example, the method described in "Chen et.al, DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, arXiv, 2016"). The subject detection unit 130 uses a deeply trained CNN to infer the likelihood of a specific region for each pixel region. However, the subject detection unit 130 may also infer the likelihood of a specific region using a trained model trained with any machine learning algorithm, or it may determine the likelihood of a specific region based on a rule base. When a CNN is used to infer the likelihood of a specific region, the CNN performs deep learning using the specific region as a positive example and regions other than the specific region as negative examples. As a result, the CNN outputs the likelihood of the specific region for each pixel region as the inference result. (Calculation method for specific areas) Figure 12 shows an example of a CNN that infers the likelihood of a specific region. Figure 12(A) shows an example of a subject region in an input image input to the CNN. The subject region 1201 is detected from the image by the subject detection described above. The subject region 1201 includes the face region 1202, which is the target of the subject detection. The face region 1202 in Figure 12(A) includes two occlusion regions (occlusion regions 1203 and 1204). Occlusion region 1203 is a region with no depth difference from the face region, and occlusion region 1204 is a region with a depth difference. Occlusion regions are also called occlusions. In this embodiment, the face region 1202 excluding occlusion regions 1203 and 1204 is detected as a specific region.

[0101] Figure 12(B) shows an example of a specific region information definition. Images 1 through 3 in Figure 12(B) are divided into white and black regions, with the black region representing a positive example and the white region representing a negative example. In Figure 12(B), the specific region information obtained by image segmentation of the subject region image is an image intended to represent a candidate for training data used when performing deep learning of a CNN. Below, we will explain which of the specific region information in Figure 12(B) is used as training data in this embodiment.

[0102] Image 1 in Figure 12(B) shows an example of occlusion information when the area is divided into a subject area (face area) and an area other than the subject, with the subject area being the positive example and the area other than the subject area, such as the background and occluding areas, being the negative example. Image 2 in Figure 12(B) shows an example of occlusion information when the area is divided into a foreground occluding area relative to the subject and an area other than the foreground occluding area, with the foreground occluding area being the negative example and the area other than the foreground occluding area relative to the subject being the positive example. Image 3 in Figure 12(C) shows an example of occlusion information when the area is divided into an occluding area that causes perspective conflict with the subject and an area other than the foreground occluding area that causes perspective conflict, with the occluding area that causes perspective conflict being the negative example and the area other than the occluding area that causes perspective conflict being the positive example.

[0103] As shown in Figure 12(B), number 1, the faces of people in the image have distinctive visibility patterns, and because the pattern variance is small, it is possible to divide the region with high accuracy. For example, the occlusion information shown in Figure 12(B), number 1 is suitable as training data in the training process when generating a CNN to detect people as subjects. From the viewpoint of detection accuracy, the occlusion information shown in Figure 12(B), number 1 is more suitable than the occlusion information shown in Figure 12(B), number 3. However, for training data in the training process when generating a CNN to detect occlusion regions that cause near-field competition, an image like the one shown in Figure 12(B), number 3 is suitable. A pair of disparity images used in focus detection may be used as training data in the training process when generating a CNN to detect occlusion regions that cause near-field competition. In addition, the occlusion information is not limited to the above example and may be generated based on any method of dividing the region into an occluded region and a region outside the occluded region. In this embodiment, the accuracy of the detection region is emphasized, and the training process is performed using the information shown in Figure 12(B), number 1, but training may be performed using other information.

[0104] Figure 12(C) shows the flow of deep learning in a CNN. In this embodiment, an RGB image is used as the input image 1210 for training. In addition, a training image 1214 (a training image for specific region information) as shown in Figure 12(C) is used as the training image. The training image 1214 is an image of the face region information excluding the occlusion information and background information from Figure 12(B).

[0105] The training input image 1210 is input to the neural network system 1211 (CNN). The neural network system 1211 can employ, for example, a layer structure in which convolutional layers and pooling layers are alternately stacked between the input layer and the output layer, or a multilayer structure in which a fully connected layer is connected after the layer structure. The output layer 1212 in Figure 12(C) outputs a score map showing the likelihood of a specific region of the input image. This score map is output in the format of output result 1213.

[0106] In CNN deep learning, the error between the output result 1213 and the training image 1214 is calculated as the loss value 1215. The loss value 1215 is calculated using methods such as cross-entropy or squared error. The coefficient parameters such as weights and biases of each node in the neural network system 1211 are adjusted so that the loss value 1215 gradually decreases. After sufficient deep learning of the CNN using many training input images 1210, the neural network system 1211 will be able to output a more accurate output result 1213 when an unknown input image is input. In other words, when an unknown input image is input, the neural network system 1211 (CNN) will be able to output specific region information, obtained by dividing the image into occluded and non-occluded regions, as the output result 1213 with high accuracy. Note that creating training data that identifies occluded regions (regions of overlapping objects) requires a lot of work. Therefore, it is conceivable to create training data using computer graphics (CG) or by creating training data using image synthesis, which involves cutting out and superimposing object images.

[0107] As described above, we explained an example where image 11 in Figure 12(B), where the face region excluding occluded and background regions is defined as the training image 1214, is used. However, even if images like 2 or 3 in Figure 12(B) are used as training images 1214, the CNN can infer the regions that will cause near-far competition when an unknown input image is input to it.

[0108] For detecting specific regions, any method other than a CNN can be applied. For example, detecting specific regions may be achieved using a rule-based method. Furthermore, for detecting specific regions, a pre-trained model trained using any method other than a deep-trained CNN may be used. For example, occluded regions may be detected using a pre-trained model trained using any machine learning algorithm such as a support vector machine or logistic regression. This is similar to subject detection.

[0109] In this embodiment, detection of a specific region is performed on all detected subjects, but by performing it after the main subject determination process in S403, and detecting the specific region only on the main subject, the amount of computation can be reduced.

[0110] Once processing in S424 is complete, the camera MPU125 terminates the subject detection and tracking subroutine and proceeds to S404 in Figure 11. (Flicker detection subroutine) Referring to Figure 13, the flicker detection subroutine executed by the camera MPU 125 in S404 of Figure 10 will be described. Figure 13 is a flowchart of the flicker detection process.

[0111] In S1301, the camera MPU 125 acquires information regarding the driving of the image sensor 122, which was performed in S1. The image sensor 122 in this embodiment selects various driving methods depending on the brightness of the shooting environment, whether the recorded image is a still image or a video, etc. In order to read the signals in the screen within the time allowed by the frame rate (image sensor driving rate) set according to the brightness of the shooting environment and the settings of the photographer, the system either thins out the rows to be read or reads the signals of multiple rows simultaneously. In S1301, regarding the driving of the image sensor, the system acquires information regarding the result of vertical focus detection (image shift amount) that occurs when flicker occurs, which is determined from the number of rows thinned out and the number of rows being read simultaneously. In this invention, the degree of agreement between the acquired information and the result calculated by the phase-difference AF unit 129 as the actual image shift amount of vertical focus detection is used to determine whether flicker is occurring in the shooting environment. Details will be described later.

[0112] In S1302, a focus detection region for flicker detection is set within the defocus map calculated in S401 of Figure 10. In this invention, the 24 regions constituting the vertical defocus map are sequentially judged.

[0113] In S1303, the camera MPU125 acquires the horizontal and vertical focus detection results for the focus detection area set in S1302 and calculates the difference between them. This process is performed because if the vertical focus detection result contains errors due to flicker, the difference between it and the horizontal focus detection result may be large.

[0114] In S1304, candidate image displacement amounts for vertical focus detection are obtained. To explain the candidate image displacement amounts, the correlation calculation performed in S401 for focus detection will be described.

[0115] In this embodiment, the pair of signals used for vertical focus detection are referred to as the A image signal and the B image signal. The first, second, etc. outputs of the A image signal for each row within the focus detection region are designated as A(1), A(2), etc., and similarly, the first, second, etc. outputs of the B image signal are designated as B(1), B(2), etc. 300 of these sequentially generated A (B) signals are concatenated to produce a pair of image signals. In the correlation calculation, the correlation amount is calculated by shifting the relative positions of the pair of image signals, and the amount of shift at the position with the highest correlation (highest degree of agreement in the shapes of the pair of image signals) is detected as the image shift amount. For example, the correlation amount COR(h) can be calculated using the following formula.

[0116]

number

[0117] (1) In equation (1), W1 corresponds to the number of data points within the field of view, and hmax corresponds to the number of shift data points. The phase-difference AF unit 129 calculates the correlation amount COR(h) for each shift amount h, and then finds the shift amount h that results in the highest correlation between image A and image B, i.e., the shift amount h that minimizes the correlation amount COR(h). Note that the shift amount h used when calculating the correlation amount COR(h) is an integer, but when finding the shift amount h that minimizes the correlation amount COR(h), interpolation processing is performed to obtain a value (real value) in subpixel units in order to improve the accuracy of the defocus amount.

[0118] In this embodiment, the shift amount that changes the sign of the difference value of the correlation amount COR is calculated as the shift amount h (in subpixel units) that minimizes the correlation amount COR(h).

[0119] First, the phase difference AF unit 129 calculates the difference value DCOR of the correlation amount according to the following equation (2).

[0120] DCOR(2×h)=COR(h+1)-COR(h-1) (2) Then, the phase difference AF unit 129 uses the difference value DCOR of the correlation quantity to determine the shift amount dh1 at which the sign of the difference quantity changes. If the value of h immediately before the sign of the difference quantity changes is h1, and the value of h after the sign has changed is h2 (h2 = h1 + 1), the phase difference AF unit 129 calculates the shift amount dh1 according to the following equation (3).

[0121] dh1=(h1+|DCOR1(h1)| / |DCOR1(h1)-DCOR1(h2)|)×2 (3) As described above, the phase-difference AF unit 129 calculates the shift amount dh1 in sub-pixel units that maximizes the correlation between the A image and the B image of the first signal, and completes the processing. Note that the method for calculating the shift amount (phase difference) of the two one-dimensional image signals is not limited to the method described here, but any known method can be used. As a result of the correlation calculation described above, multiple shift amounts may be calculated in which the sign of the difference value of the correlation amount COR changes. In normal focus detection, the shift amount with the largest difference value is selected to perform focus detection, but in S1304, the multiple calculated shift amounts are acquired as image shift amount candidates. Details on how to use the image shift amount candidates will be described later.

[0122] In S1305, the camera MPU 125 determines whether there is a relationship between the image displacement amount candidate acquired in S1304 and the information obtained in S1301 regarding the vertical focus detection result (image displacement amount) that occurs when flicker occurs due to the image sensor driving method. If the value of the image displacement amount candidate acquired in S1304, or the difference thereof, is close to the image displacement amount obtained in S1301 within a predetermined value, the process proceeds to S1306; otherwise, it proceeds to S1308. The camera MPU 125 may also determine whether the number of image displacement amount candidates acquired in S1304 that are related to the information obtained in S1301 regarding the vertical focus detection result (image displacement amount) that occurs when flicker occurs due to the image sensor driving method is above a threshold. If it is above the threshold, the process proceeds to S1306.

[0123] In S1306, the difference between the vertical and horizontal focus detection results obtained in S1303 is determined. If the difference is large, the process proceeds to S1307; otherwise, it proceeds to S1308.

[0124] In S1307, the set focus detection area is affected by flicker, causing an error in the vertical focus detection result, and therefore it is determined that flicker is present.

[0125] In S1308, the set focus detection area is determined to have little influence from flicker on the vertical focus detection result.

[0126] After completing S1307 or S1308, the process proceeds to S1309 to determine if flicker detection has been completed in the entire focus detection area. If not, the process returns to S1302 and the above process is repeated. If completed, the processing of this subroutine is completed and the process proceeds to S405. (The driving method of the image sensor and the effect of flicker on vertical focus detection) Referring to Figures 14 to 16, the mechanism by which the driving method of the image sensor 122 causes errors in vertical focus detection due to flicker will be explained. Flicker, which occurs in lighting and digital signage, is a phenomenon in which light flashes repeatedly at an invisible frequency over time. On the other hand, the slit-rolling image sensor 122 sequentially accumulates and reads out the signals of each row over time. When sequential exposure is performed with the slit-rolling image sensor 122 in an environment where flicker occurs, the difference in the accumulation time of each row causes an increase or decrease in the signal of each row due to the effect of the flicker. In this embodiment, the focus detection signal is also read out sequentially for each row, but since the pair of signals used for horizontal focus detection use the signals of the same row, they are affected by the flicker to the same extent, and therefore the effect on the focus detection result is small. On the other hand, the pair of signals used for vertical focus detection are affected by the flicker within the pair of signal sequences because the direction in which the signal sequence is formed and the direction in which the slit-rolling sequence is read out coincide.

[0127] Figure 14 illustrates the effect of flicker on the pair of vertical focus detection signals. Figure 14(a) shows that time progresses horizontally from left to right, and on the time axis, it shows the timing of accumulation and readout of the focus detection signal (image A) and imaging signal (image A+B) for each row of the image sensor 122. As explained in Figure 2(b), the upper two rows of the diagram in Figure 14(a) show the accumulation period and readout period, respectively, as signals A and A+B are output for each row. After resetting PDA211 and PDB212, accumulation of signals A and A+B begins, and the voltage of signal A is read out simultaneously with the completion of accumulation. After the readout of signal A is complete, the accumulation of signal A+B is completed, and the voltage is read out. Similarly, the signal for the second row is read out. The time difference between the accumulation period of signal A for the first row and the accumulation period of signal A for the second row is considered to be the difference in the centers of the accumulation periods, so the interval is Pa-a. Furthermore, the interval between the storage period of signal A+B in the first row and the storage period of signal A+B in the second row is Pa-ab. As mentioned above, in an environment where flicker occurs, the brightness changes over time, so the signal output of the first and second rows changes as Pa-a and Pa-ab time elapse. The difference in storage periods between signal A and signal A+B is shown as Pa-ab. In an environment where flicker occurs, the storage periods of signal A and signal A+B are shifted by Pa-ab in every row. Due to the Pa-ab shift, the waveforms of signal A and signal A+B are shifted by an amount due to the flickering effect. Due to the difference in storage periods between signal A and signal A+B, the waveform of signal B is shifted horizontally relative to the waveform of signal A by Pa-ab / Pa-a pixels. For example, as shown in Figure 14(a), the start time of storage for each row is shifted by a time equivalent to the sum of the readout periods of signal A and signal A+B. If the readout periods for signal A and signal A+B are equal, the waveform of signal B will be shifted horizontally relative to the waveform of signal A by Pa-ab / Pa-a pixels = 1 / 4 pixels.

[0128] Figure 14(b) shows a case where the exposure control for each row is different from that of Figure 14(a), and signals A and B are read out in each row. It shows a case where the start of accumulation of signals A and B in the first row is shifted by the readout period of signal A. Similar to Figure 14(a), due to the difference in the accumulation periods of signals A and A+B, the waveform of signal B is shifted laterally relative to the waveform of signal A by Pa-ab / Pa-a pixels. For example, for signal A in the first row, signal B in the first row, signal A in the second row, etc., the start time of accumulation is shifted by time corresponding to the readout period of signal A in the first row, the readout period of signal B in the first row, the readout period of signal B in the second row, etc. Furthermore, if the readout periods of signals A and B are equal, the waveform of signal B is shifted laterally relative to the waveform of signal A by Pa-ab / Pa-a pixels = 1 / 2 pixels.

[0129] Figure 15 shows the waveform when flicker occurs. Figure 15(a) shows signals A and B corresponding to the case in Figure 14(b). The horizontal axis represents the pixel number, and the vertical axis represents the signal output normalized by the maximum value. The undulation of the output for each pixel indicates the flickering over time. A magnified view of a part of Figure 15(a) is shown in the upper right, where it can be seen that the waveforms of signals A and B are slightly out of sync. As explained in S1304, Figure 15(b) shows the result of calculating the correlation amount. The horizontal axis represents the amount of positional shift between signals A and B, and the vertical axis represents the correlation amount, which indicates the magnitude of the correlation. In Figure 15(b), it can be seen that the correlation amount takes a minimum value around ±40 pixels and 0 pixels. Figure 15(c) shows the DCOR, which is the difference value of the correlation amount. The horizontal axis represents the shift amount, and the vertical axis represents the difference in the correlation amount. The upward-sloping shift amount that intersects the horizontal axis indicates that it is near ±80 pixels and 0 pixels. Figure 15(d) is a magnified view of the area near 0 pixels in terms of shift amount. In this embodiment, candidate image displacement amount dh1 is -0.5 pixels, which is the intersection with the horizontal axis. Similarly, -80.5 pixels and +79.5 pixels are candidates for image displacement amounts.

[0130] The candidate for image displacement, -0.5 pixels, is the amount of pixel displacement that occurs when reading out as shown in Figure 14(b) in an environment where flicker is present. In this embodiment, in S1301, information regarding the reading method, such as that shown in Figures 14(a) and 14(b), is acquired as information regarding the driving of the image sensor 122, thereby obtaining the amount of image displacement caused by flicker. For example, in the case of the driving method shown in Figure 14(b), information of -0.5 pixels is acquired. On the other hand, the image displacement amounts of -80.5 pixels and +79.5 pixels are, as can be seen from Figure 15(a), the amount of image displacement that is offset by the amount of image displacement caused by flicker, which is -0.5 pixels, relative to the period of 80 pixels in which flicker occurs. By canceling out the amount of image displacement caused by flicker, it is possible to calculate that the period in which flicker occurs is 80 pixels, and the frequency of the flicker blinking can be calculated from the information on the reading time for each row. In S1305, the system determines whether the image shift amount of -0.5 pixels, which occurs when flicker occurs, is included in the image shift amount candidates obtained in S1304, based on the readout information from the image sensor 122 in Figure 14(b). If it is included, the system proceeds to S1306, assuming that a flicker environment is possible. In S1306, to eliminate cases where the defocus state of the subject matches the image shift amount detected in a flicker environment, the system checks the difference with the horizontal focus detection result, which is less affected by flicker. If the difference between the horizontal focus detection result and the vertical focus detection result is small, the system determines that the defocus state of the subject is also obtained from the vertical focus detection result. On the other hand, if the difference is large, the system determines that the vertical focus detection result is affected by flicker. Performing the determination in S1306 allows for focus detection using the vertical focus detection result in a wider range of shooting environments, enabling more accurate focus adjustment. Alternatively, the determination in S1306 can be omitted, and the system can be configured to minimize the flicker effect on the vertical focus detection result.

[0131] Let's return to Figure 14 and continue the explanation. Figure 14(c) shows the case where multiple rows of the image sensor 122 are read out simultaneously. Figure 14(c) shows the case where four rows are read out simultaneously, but the number of rows read out simultaneously is not limited to this. Even when multiple rows are read out simultaneously, there is a difference between the readout period of signal A and the readout period of signal A+B, and furthermore, there is a difference in the readout period at the row block level (in Figure 14(c), four rows make up one block). Figure 16 is another diagram showing the waveform when flicker occurs. For clarity, Figure 16(a) shows the waveforms of signal A and signal B when 10 rows are read out simultaneously. In addition to the effect of the flicker, it can be seen that steps occur every 10 rows. If the correlation calculation described above is performed on such waveforms, intervals of shift amount with small changes in the correlation amount will occur, and a highly accurate image shift amount cannot be obtained, so digital filtering is applied. Figure 16(b) shows the results of applying a predetermined filter (-4, -11, -21, -28, -28, -17, 0, 17, 28, 28, 21, 11, 4). Similar to the correlation calculation process described above, Figure 16(c) shows the correlation amount COR, and Figure 16(d) shows the difference DCOR of the correlation amounts. It can be seen that the difference DCOR of the correlation amounts is upward sloping and intersects the horizontal axis around the pixels at -90, -80, -10, 0, +70, and +80. Here, the shift amount of -10 pixels is the amount of image displacement caused by the flicker effect when reading out 10 rows simultaneously from the image sensor 122. Similar to the case of reading out one row at a time as described above, in S1301, information regarding the driving of the image sensor 122 is obtained that 10 rows will be read out simultaneously, and at that time, it is obtained that the amount of image displacement caused by the flicker effect is approximately -10 pixels. The subsequent determinations in S1305 and S1306 are as described above. Similarly, the flicker frequency can be calculated from the shift amounts of ±80 pixels and 0 pixels. It can also be seen that the candidate image displacement amounts for -90 pixels and +70 pixels are the sum of the flicker frequency and the effect of flicker caused by the image sensor readout method.

[0132] When reading multiple rows simultaneously, there are two factors contributing to image shift: the difference in reading periods between signal A and signal A+B (Pa-ab), and the step difference in waveforms that occurs every few rows. If the effect of the step difference in waveforms that occurs every few rows is sufficiently reduced by the digital filtering process described above, then the effect of the former, the difference in reading periods between signal A and signal A+B (Pa-ab), becomes larger. For example, in the case of simultaneous reading of 4 rows in Figure 14(c), if the step difference in waveforms that occurs every 4 rows is eliminated by the digital filtering process, then an image shift of Pa-ab / Pa-a × 4 pixels = 1 pixel occurs. On the other hand, in the case of simultaneous reading of 10 rows in Figure 16, the step difference in waveforms that occurs every 10 rows is not eliminated by the digital filtering process. Therefore, an image shift of -10 pixels is calculated as a candidate for the amount of image shift.

[0133] The amount of image shift caused by waveform steps that occur for each of the multiple rows read out simultaneously as described above is affected by adding the signals of the multiple rows after they have been read out from the image sensor 122. For example, in the case of simultaneous readout of 10 rows, if the signals of 2 rows are added after reading, and the length of the signal sequence (number of signals) is compressed to 1 / 2 before correlation calculation, then naturally, the candidate for image shift will be an image shift of ±5 pixels (in the example explained above, it will be -5 pixels). Therefore, the determination should be made based on the drive information of the image sensor 122 as well as the content of the subsequent signal processing.

[0134] In this way, by combining the drive information of the image sensor 122 acquired in S1301 with the digital filter processing used in the correlation calculation, the value of the image shift amount affected by flicker in the focus detection result can be calculated in advance and compared with the image shift amount candidate in S1304.

[0135] Furthermore, the effects of the A and B signals and the vertical focus detection results under a flicker environment, as explained in Figures 14 to 16, are based on the case where the subject has no contrast and only the effect of flickering occurs. In reality, the contrast, including the defocus state of the subject, is superimposed on the A and B signals. Therefore, when the contrast of the subject is low and the difference in brightness of the flicker is large, the flicker has a large impact on the vertical focus detection results, producing values ​​close to the image shift amount described above. On the other hand, when the contrast of the subject is high, or when the difference in brightness of the flicker is small in a mixed light environment with other flicker-free light sources, the impact of the flicker on the vertical focus detection results becomes small, and a vertical focus detection result indicating the defocus state of the subject is obtained. Therefore, in the judgment performed in S1305 in Figure 13, it is desirable to assume that there will be some error in the amount of image shift that occurs under a flicker environment due to the way the image sensor 122 is read out and the digital filter. For example, in the case of Figure 15, one possible method is to determine "Yes" if a candidate for vertical focus detection image displacement within the range of -0.5 pixels ± 0.25 pixels is obtained.

[0136] As described above, in environments where flicker occurs, vertical focus detection results may be inaccurate. However, by determining whether or not to use the image sensor based on its drive information, it is possible to avoid using low-accuracy vertical focus detection results. As a result, highly accurate focus detection can be achieved.

[0137] In this embodiment, the presence or absence of flicker effect was determined for each focus detection area. Flicker may occur due to the lighting of the entire shooting environment, but it may also occur only in a part of the shooting environment, such as a digital signage screen. By determining the presence or absence of flicker effect for each focus detection area, as in this embodiment, more vertical focus detection results can be used, enabling more accurate focus detection.

[0138] On the other hand, as mentioned above, the impact of flicker on the vertical focus detection result varies depending on the contrast of the subject, including defocus areas. Therefore, using only one focus detection area may lead to incorrect judgments. For this reason, one approach is to pre-set a threshold and, if flicker is detected in more focus detection areas than the threshold, to completely disregard the vertical focus detection results. Alternatively, if there is an uneven distribution of focus detection areas affected by flicker, one could consider disabling the vertical focus detection area in only a portion of the shooting range. These methods can more reliably eliminate errors caused by flicker included in the vertical focus detection results.

[0139] In this embodiment, the effect of flicker was determined by correlation calculation, but the determination method is not limited to the method described above. Focusing on the number of lines read simultaneously, the brightness of the signal stream may be added in units of the number of lines read simultaneously and compared with each other. If the difference in signal amount is greater than or equal to a predetermined value, it may be determined that a flicker environment exists. Alternatively, the presence or absence of flicker may be determined by comparing it with different driving states of the image sensor 122. Different numbers of lines read simultaneously and different readout speeds result in different effects of flicker on the signal. This difference may be used to determine if flicker occurs. (Defocus amount selection process) The subroutine for the defocus amount selection process will be explained with reference to Figures 17 to 20. Figure 17 is a flowchart of the defocus amount selection process. Figure 18 is a diagram showing how to set up the defocus map. Figure 19 is another diagram showing how to set up the defocus map, and shows an example of the placement of the defocus map when occlusion is present. Figure 20 is a diagram showing the histogram of the defocus map.

[0140] In S1701, the camera MPU 125 acquires subject detection information, which is the subject detection position and size detected by the subject detection unit 130.

[0141] In S1702, the camera MPU 125 acquires specific area information detected by the subject detection unit 130. In this embodiment, the specific area information is the face area excluding occluded areas and background areas. Processing using the specific area information will be described later.

[0142] In S1703, usable focus detection results are collected. The collection of usable focus detection results involves collecting defocus amounts, which are focus detection results that can be used as defocus amount selection results, from the horizontal defocus map and the defocus amount of the horizontal defocus map. Specifically, it is decided whether to make all vertical focus detection results usable based on whether the number of focus detection regions determined to have a flicker effect in the flicker detection process in Figure 13 above is greater than or equal to a predetermined number. All vertical focus detection results are used because if more than a predetermined number of areas have a flicker effect, there is a high possibility that the vertical focus detection results contain errors due to the flicker.

[0143] Furthermore, when the contrast of the subject is low, the ISO sensitivity is high, or the exposure is darker than the correct exposure, the error in the focus detection result is large. For this reason, the reliability of the focus detection result can be determined from the difference in the correlation amount in the correlation calculation process mentioned above, and it may be decided not to use it as the focus detection result. In addition, depending on the driving method of the image sensor, the accuracy of the vertical focus detection result may be lower than that of the horizontal focus detection result due to the decimation or addition of rows read out. For this reason, in modes where imaging is performed using such a driving method, it may be decided not to use vertical focus detection.

[0144] In S1704, a histogram is generated using the defocus amount, which is the focus detection result made available in S1703. The histogram is generated using subject detection information and specific region information to determine which focus detection region's focus detection result to use. As shown in the focus detection area setting process of S2201 in Figure 22 above, the histogram is generated using the defocus map encompassed by the subject region.

[0145] Using Figures 20(a) to 20(c), we will explain how to create a histogram using the defocus amount in the upper body detection region.

[0146] Figure 20(a) is a histogram generated from the defocus amount of the horizontal eye defocus map within the upper body detection area of ​​the person in Figure 18(b). Figure 20(b) is a histogram generated from the defocus amount of the vertical eye defocus map within the upper body detection area of ​​the person in Figure 18(c). Figure 20(c) is a histogram generated by combining the defocus amounts of the horizontal eye defocus map and the vertical eye defocus map within the upper body detection area of ​​the person in Figures 18(b) and 18(c). The horizontal axis of the histogram represents classes that divide the defocus amount into fixed ranges, and the vertical axis represents the frequency. The defocus amount is defined as near on the + side and far on the - side, and we assume that the defocus amount of the pupil area of ​​the person is 0Fδ. In the horizontal eye histogram of Figure 20(a), the histogram is generated for the entire upper body detection area, so the area on the left side of the upper body below the face is included, and the maximum frequency of the histogram is on the near side. Therefore, if the defocus amount is selected from the range of defocus amounts that yields the maximum frequency in the histogram, the selected defocus amount will differ from the pupil area of ​​the person the photographer wants to focus on. In the vertical eye histogram in Figure 20(b), the vertical eye defocus map is placed in the face detection area, so it does not include the left side of the upper body below the face, resulting in the maximum frequency of the histogram in the range around 0Fδ, which is the defocus amount for the pupil area. However, because the number of focus detection areas in the defocus map is small, it may be difficult to extract the point where the frequency is maximum under conditions where the defocus amount is prone to fluctuation due to errors. Therefore, by generating a histogram that combines horizontal and vertical eyes as shown in Figure 20(c), it becomes possible to generate a histogram using a larger number of defocus amounts. This makes it possible to select the defocus amount more accurately in the event of variability or incorrect defocus amounts. However, since this is a histogram of defocus amounts in the upper body detection area, it also includes the left side of the upper body below the face. In this case, the histogram of defocus amounts shows high frequencies in the ranges of -1Fδ to 0Fδ and 0Fδ to 1Fδ, making it difficult to extract the range of defocus amounts where the histogram frequency is at its maximum.Therefore, depending on the variability of the defocus amount, the range of defocus amounts at which the histogram frequency is maximized may change.

[0147] Using Figures 20(d) to 20(f), we will explain how to create a histogram using the defocus amount, with the face detection region being the defocus amount selection region. We will use an example where the defocus amount of the pupil region of the person is 0Fδ.

[0148] Figure 20(d) is a histogram generated from the defocus amounts of the side-eye defocus map within the face detection region of the person in Figure 18(d). Since the histogram is created based on the defocus amount within the face detection region, the frequency of the defocus amount histogram is maximum in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil region.

[0149] Figure 20(e) is a histogram generated from the defocus amounts of the vertical eye defocus map within the face detection region of the person in Figure 18(d). Since the histogram is created based on the defocus amount within the face detection region, the frequency of the defocus amount histogram is maximum in the range from -1Fδ to 0Fδ, which includes the defocus amount of the person's pupil region.

[0150] Figure 20(f) is a histogram generated by combining the histograms in Figure 20(d) and Figure 20(e) to show the defocus amounts for both horizontal and vertical eyes. By combining the defocus amounts for horizontal and vertical eyes within the face detection area, the frequency of the histogram for defocus amounts in the range of -1Fδ to 0Fδ, which includes the defocus amount in the pupil area of ​​the person, becomes larger than the histogram for defocus amounts for only horizontal or vertical eyes. This makes it less susceptible to variations in defocus amounts and defocus amounts that compete with the background in terms of distance.

[0151] A histogram of defocus amounts should ideally be able to generate a histogram using a larger number of defocus amounts within a narrow human detection area. As shown in Figures 20(d) to 20(f), it is desirable to generate a histogram using the defocus amounts of the horizontal-eye defocus map and the vertical-eye defocus map within the face detection area. However, if the area of ​​the defocus map within the face detection area is small, the number of defocus amount data points is small, and therefore the frequency of the histogram using defocus amounts is small overall, making it difficult to extract the range of defocus amounts with the highest frequency. Therefore, when creating a histogram, determine the required number of defocus amount data points or the human detection area, and check whether the number of defocus amount data points or the human detection area is greater than or equal to a predetermined value. If it is less than the predetermined value, expand the human detection area so that the number of defocus amount data points is greater than or equal to the predetermined value. Also, since there is a difference between the area of ​​the horizontal-eye defocus map and the area of ​​the vertical-eye defocus map, for example, the horizontal-eye defocus map may be used as the upper body detection area of ​​the person, and the vertical-eye defocus map as the face detection area of ​​the person, and the histogram of defocus amounts may be generated.

[0152] Furthermore, depending on the detection area of ​​a person, there may be defocus in the horizontal defocus map but not in the vertical defocus map. In such cases, if only the horizontal defocus map has defocus, the number of defocus amounts may be doubled or left as is, and in areas where both the horizontal and vertical defocus maps are present, the number of defocus amounts may be left as is or halved, thus normalizing the results. In this embodiment, the area of ​​the horizontal defocus map was described as being larger than the area of ​​the vertical defocus map, but the area of ​​the vertical defocus map may also be larger than the area of ​​the horizontal defocus map.

[0153] In S1705, a focus detection region is selected using the histogram of defocus amounts generated in S1704, and the defocus amount of that region is selected. The defocus amount is selected from the range where the frequency of the defocus amount histogram is maximum. There are several selection methods, such as selecting the defocus amount closest to the defocus amount of the predicted AF processing result in S406, selecting the defocus amount of the focus detection region that is geographically close to the pupil detection region (a person detection region), or selecting a defocus amount on the near side. Alternatively, histograms of defocus amounts can be created for each of the multiple detection regions (e.g., upper body, face, pupils, etc.), and the defocus amount can be selected from the range where the frequency of the histograms of the multiple detection regions is maximum in multiple ranges, or from the near side of multiple defocus amounts.

[0154] Alternatively, the amount of defocus may be calculated by averaging the defocus amounts within the range where the frequency of the defocus amount histogram is maximum.

[0155] Next, processing using specific region information will be explained with reference to Figure 19. Figure 19(a) is an image of the moment when an occluding region (arm) covers the face region of a person, and the subject detection information acquired in S1701 is shown by a rectangular frame. Figure 19(b) shows the specific region (face region in this embodiment) acquired in S1702 represented in a grid frame, indicating that the part covered by the arm is not detected as a specific region (face region). The specific region information (likelihood) acquired in S1702 may be information that expresses whether or not it is a specific region with a binary output result of 1 or 0, or it may be information that expresses the likelihood as higher as the numerical value, for example, 0 to 255 in 1 byte. Here, we will explain assuming the former, so the grid frame region is output as 1, and other regions such as the arm are output as 0. Figure 19(c) is a diagram in which only the regions that are valid as specific regions in the horizontal defocus map are represented by diagonal lines by associating the horizontal defocus map with a 3x3 frame horizontal defocus map and the specific region. One way to determine whether it is effective or not is to check if the proportion of the estimated area within each frame of the defocus map is above a certain level, for example, 50% or more. The range of each frame may be determined based on the parameters used in the correlation calculation, such as the shift amount used when calculating the defocus amount. Figure 19(d) is a diagram in which only the areas that are effective as specific areas within the vertical defocus map are represented by diagonal lines by associating the vertical defocus map of 3x3 frames with specific areas. The determination of whether it is effective or not is the same as in the case of the horizontal defocus map, so the explanation is omitted. Figure 21 is a histogram generated from the defocus maps of Figure 19(c) and Figure 19(d). The 3x3 defocus map includes occluded areas. Therefore, if a histogram is generated for all areas, the histogram peak is more likely to be detected closer than the face due to the influence of the occluded areas. However, by generating a histogram only for specific areas as in this embodiment, the influence of occluded areas and background areas can be excluded.

[0156] As described above, generating a histogram only in a specific region can be expected to prevent the influence of occluded regions. Furthermore, although this embodiment was explained using a 3x3 frame defocus map, the number of frames can be freely set to NxM frames (where N and M are integers of 2 or more). [Other examples] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0157] This embodiment includes the following configuration. (Composition 1) An image sensor having multiple focus detection pixels arranged in rows and columns, each receiving a light beam passing through different pupil regions of the imaging optical system, and sequentially starting exposure and reading out each row of pixels or each block of multiple rows, A focus detection means that performs a first focus detection using an output signal from a focus detection pixel with the row direction as the pupil division direction, and a second focus detection using an output signal from a focus detection pixel with the column direction as the pupil division direction. An imaging apparatus characterized by having determination means for determining whether to use the result of the second focus detection according to the drive information of the image sensor. (Configuration 2) The imaging apparatus according to configuration 1, characterized in that the determination means determines not to use the result of the second focus detection if the difference between the result of the second focus detection and the amount of image shift based on the drive information is within a predetermined value. (Composition 3) The imaging apparatus according to configuration 2, characterized in that the predetermined value is changed according to the driving method of the image sensor. (Composition 4) The imaging apparatus according to any one of configurations 1 to 3, characterized in that the determination means determines not to use the second focus detection results if the number of second focus detection results related to the image displacement amount based on the drive information is greater than or equal to a threshold. (Composition 5) The focus detection means performs the second focus detection in the plurality of focus detection regions, The imaging apparatus according to any one of configurations 1 to 4, characterized in that the determination means determines whether to use the results of the second focus detection in accordance with the drive information and the results of a plurality of second focus detections.

[0158] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. [Explanation of Symbols]

[0159] 120 Camera body (imaging device) 122 Image sensor 125 Camera MPU (Determination Method) 129 Phase-detection AF unit (focus detection means)

Claims

1. An image sensor having multiple focus detection pixels arranged in rows and columns, each receiving a light beam passing through different pupil regions of the imaging optical system, and sequentially starting exposure and reading out each row of pixels or each block of multiple rows, A focus detection means that performs a first focus detection using an output signal from a focus detection pixel with the row direction as the pupil division direction, and a second focus detection using an output signal from a focus detection pixel with the column direction as the pupil division direction. The system includes a determination means for determining whether to use the second focus detection result according to the drive information of the image sensor, The imaging apparatus is characterized in that the determination means determines not to use the result of the second focus detection if the difference between the result of the second focus detection and the amount of image shift based on the drive information is within a predetermined value.

2. The imaging apparatus according to claim 1, characterized in that the predetermined value is changed according to the driving method of the image sensor.

3. The imaging apparatus according to claim 1 or 2, characterized in that the determination means determines not to use the second focus detection results if the number of second focus detection results related to the information on the amount of image displacement based on the drive information is greater than or equal to a threshold.

4. The focus detection means performs the second focus detection in a plurality of focus detection regions. The imaging apparatus according to claim 1 or 2, characterized in that the determination means determines whether to use the results of the second focus detection in accordance with the drive information and the results of a plurality of second focus detections.

Citation Information

Patent Citations

  • Imaging device

    JP2010263568A

  • Depth map output device

    JP2011237215A

  • Imaging apparatus, flicker detection method and program

    JP2017069741A

  • Focus adjustment device

    JP2019091093A

  • Imaging apparatus, method for controlling the same, and program

    JP2022129925A