Focus detection device, imaging device, focus detection method, and program

The focus detection device and method address inaccuracies in defocus detection by interpolating defocus amounts based on depth values, improving autofocus and depth estimation accuracy in imaging devices.

JP2026061028APending Publication Date: 2026-04-09NIKON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing focus detection methods, such as image plane phase difference detection and TOF techniques, face challenges in accurately calculating defocus amounts for all pixels in an image, particularly due to noise and errors in defocus detection, which can lead to inaccuracies in autofocus and depth estimation.

Method used

A focus detection device and method that utilizes an image acquisition unit to detect a first defocus amount for a predetermined number of pixels, generates a depth map, and interpolates a second defocus amount for any pixel based on the correlation between the first defocus amount and depth value, using a predetermined function to enhance accuracy.

Benefits of technology

Improves the accuracy of defocus amount calculation by reducing noise-related errors, enabling precise autofocus and depth estimation, thereby enhancing imaging device performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026061028000001_ABST
    Figure 2026061028000001_ABST
Patent Text Reader

Abstract

The present invention provides a focus detection device, an imaging device, a focus detector, and a program. [Solution] The focus detection device comprises an image acquisition unit 150 that acquires an image of a subject, a defocus acquisition unit that acquires a first defocus amount for each of a predetermined number of pixels in the subject image which is an image of the subject in the image, and an interpolation unit 154 that calculates a second defocus amount for any pixel in the image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a focus detection device, an imaging device, focus detection, and a program. [Background technology]

[0002] Regarding image plane phase difference detection systems, techniques for calculating the amount of defocus for multiple points on an image are known. Techniques using TOF (Time of Flight) or techniques for learning depth maps and adding them to images are also known. In learning depth maps, results detected beforehand by TOF or stereo ranging are used as training values. On the other hand, a technique has been proposed to directly infer depth maps using the degree of blur in the image as the amount of defocus (Non-Patent Literature 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] “Multi-task Learning for Monocular Depth and Defocus Estimations with Real Images,” [online], Renzhi He, Hualin Hong, Boya Fu, Fei Liu, [Accessed March 17, 2023], Internet <URL: https: / / arxiv.org / ftp / arxiv / papers / 2208 / 2208.09848.pdf> [Overview of the project]

[0004] One aspect of the present invention is a focus detection device comprising: an image acquisition unit that acquires an image of a subject; a defocus acquisition unit that acquires a first defocus amount for each of a predetermined number of pixels in the subject image, which is an image of the subject within the image; and an interpolation unit that calculates a second defocus amount for any pixel of the image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels.

[0005] One aspect of the present invention is an imaging device comprising the above-described focus detection device and an imaging unit that outputs the captured image.

[0006] One aspect of the present invention is a focus detection method comprising: an image acquisition step of acquiring an image in which a subject has been captured; a defocus acquisition step of acquiring a first defocus amount for each of a predetermined number of pixels in the subject image, which is an image of the subject within the image; and an interpolation step of calculating a second defocus amount for any pixel of the image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels.

[0007] One aspect of the present invention is a program that causes a computer to perform the following steps: an image acquisition step of acquiring an image in which a subject has been captured; a defocus acquisition step of acquiring a first defocus amount for each of a predetermined number of pixels in the subject image, which is an image of the subject within the image; and an interpolation step of calculating a second defocus amount for any pixel of the image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example of the configuration of the imaging device 1 according to the first embodiment. [Figure 2] This figure shows an example of the functional configuration of the imaging control unit 40 according to the first embodiment. [Figure 3]This figure shows an example of an image A1 captured according to the first embodiment. [Figure 4] This figure shows an example of a depth map C1 according to the first embodiment. [Figure 5] This figure shows an example of the flow of the defocus map generation process according to the first embodiment of this embodiment. [Figure 6] This figure shows an example of the relationship between the first defocus amount and the depth value according to the first embodiment. [Figure 7] This figure shows an example of the functional configuration of the imaging control unit 40a according to the second embodiment. [Figure 8] This figure shows an example of the flow of the defocus map generation process according to the second embodiment of this embodiment. [Figure 9] This figure shows an example of the slope calculation process flow according to the second embodiment. [Figure 10] This figure shows an example of the difference in depth between the nose and eyes of a person's face according to the second embodiment. [Figure 11] This figure shows an example of the ratio of the size of the subject on the imaging surface to the size of the imaging surface according to the second embodiment. [Figure 12] This figure shows an example of a graph that illustrates the relative defocus difference with respect to the depth value of the face portion according to the second embodiment, with respect to the eyes. [Figure 13] This figure shows an example of the flow of the defocus map generation process according to a modified example of the second embodiment. [Figure 14] This figure shows an example of the functional configuration of the imaging control unit 40b according to the third embodiment. [Figure 15] This figure shows an example of the flow of the defocus map generation process according to the third embodiment. [Figure 16] This figure shows an example of the outline of the defocus map generation process according to the third embodiment. [Figure 17] This figure shows an example of the functional configuration of the imaging control unit 40c according to the first modified example of the third embodiment. [Figure 18]FIG. is a diagram showing an example of the flow of defocus map generation processing according to the first modification of the third embodiment. [Figure 19] FIG. is a diagram showing an example of the outline of defocus map generation processing according to the first modification of the third embodiment. [Figure 20] FIG. is a diagram showing an example of the functional configuration of the imaging control unit 40d according to the second modification of the third embodiment. [Figure 21] FIG. is a diagram showing an example of the flow of defocus map generation processing according to the second modification of the third embodiment. [Figure 22] FIG. is a diagram showing an example of the outline of defocus map generation processing according to the second modification of the third embodiment.

BEST MODE FOR CARRYING OUT THE INVENTION

[0009] (First Embodiment) Hereinafter, the first embodiment will be described in detail with reference to the drawings. FIG. 1 is a diagram showing an example of the configuration of an imaging apparatus 1 according to the present embodiment. The imaging apparatus 1 includes a lens 10, an aperture mechanism 11, an imaging unit 12, a TG (timing generator) 13, an analog front-end unit (hereinafter referred to as "AFE") 14, an image processing unit 15, a RAM (Random Access Memory) 16, a recording I / F (recording interface) 17, a display unit 18, an operation unit 19, a ROM (Read Only Memory) 21, a TOF (Time of Flight) sensor 25, a bus 22, and a control unit 30. The image processing unit 15, the RAM 16, the recording I / F 17, the display unit 18, the ROM 21, the TOF sensor 25, and the control unit 30 are connected to each other via the bus 22. The operation unit 19 is connected to the control unit 30.

[0010] The imaging unit 12, the TG 13, and the AFE 14 are collectively referred to as a solid-state imaging device 100. The control unit 30 and the image processing unit 15 are collectively referred to as an imaging control unit 40.

[0011] Lens 10 is composed of multiple lens groups, including a focusing lens and a zoom lens. In Figure 1, lens 10 is shown as a single lens. This lens 10 is controlled by a lens drive device (not shown).

[0012] The aperture mechanism 11 adjusts the amount of light incident on the imaging unit 12. This adjustment is performed by an aperture drive unit (not shown) in accordance with instructions from the control unit 30.

[0013] The imaging unit 12 has multiple pixels arranged in rows and columns, and generates analog pixel signals of an image by photoelectric conversion of the subject image formed on its imaging surface. In other words, the imaging unit 12 captures a subject image and generates pixel data of the captured image. In this embodiment, for example, an XY address type CMOS (Complementary Metal-Oxide Semiconductor) is used for the imaging unit 12. Multiple photodetectors constituting pixels arranged in a matrix are provided on the light-receiving surface of this imaging unit 12. The imaging unit 12 outputs the analog pixel signals generated by these photodetectors as pixel data to the AFE 14 under control by the TG 13 based on instructions from the control unit 30. The exposure time of the imaging unit 12 is controlled for each row (line) based on the drive signal from the TG 13.

[0014] The AFE14 is an analog front-end circuit that performs signal processing on the analog pixel signals generated by the imaging unit 12. The AFE14 performs gain adjustment of the analog pixel signals and A / D (analog-to-digital) conversion of the analog pixel signals. The AFE14 outputs the converted digital pixel signals as pixel data to the image processing unit 15.

[0015] The TG (Timing Generator) 13 (Control Unit) supplies drive signals to the imaging unit 12 and AFE 14 respectively, according to instructions from the control unit 30, and controls the drive timing of both based on the supplied drive signals. The TG 13 supplies drive signals to read out the pixel values ​​based on the exposure time for each line of the imaging unit 12 as information used for imaging control. Here, imaging control refers to, for example, autofocus control (hereinafter referred to as AF control), automatic exposure control (hereinafter referred to as AE control), and automatic white balance control (hereinafter referred to as AWB control). In addition, imaging control may include control to display a live view image and a through image on the display unit 18 for confirmation of the subject.

[0016] For example, TG13 supplies a drive signal to read out pixel values ​​based on the exposure time for each line read from the imaging unit 12 as information for calculating an evaluation value in autofocus control (AF control). Also, for example, TG13 supplies a drive signal to read out pixel values ​​based on the exposure time for each line read from the imaging unit 12 as information for calculating an evaluation value in automatic exposure control (AE control). Also, for example, TG13 supplies a drive signal to read out pixel values ​​based on the exposure time for each line read from the imaging unit 12 as information for calculating an evaluation value in auto white balance control (AWB control). Also, for example, TG13 supplies a drive signal to read out pixel values ​​based on a predetermined exposure time among the exposure times for each line as a display image. TG13 supplies a drive signal to output the pixel values ​​read from the imaging unit 12 as pixel data to the image processing unit 15 via AFE14.

[0017] The display unit 18 is, for example, a liquid crystal display, which displays various images according to instructions from the control unit 30. The display unit 18 displays, for example, image data captured by the imaging unit 12, and operation screens, etc.

[0018] The operation unit 19 includes, for example, a release button, a directional pad, a command dial, a touch panel, and other operation keys, and supplies operation input signals to the control unit 30 according to the user's operation. For example, the operation unit 19 gives an imaging instruction to the control unit 30 when the user fully presses the release button.

[0019] ROM21 stores control programs for controlling the imaging device 1. For example, ROM21 pre-stores sequence programs for imaging operations and correction processing, which will be described later. RAM16 stores image data captured by the imaging unit 12, as well as various setting information for the imaging process.

[0020] The recording interface 17 is connected to a removable recording medium 23, such as a card memory, and is used to write, read, or erase image data to this recording medium 23. The recording medium 23 is a storage unit that is detachably connected to the imaging device 1, and stores (records) image data processed by the image processing unit 15, for example.

[0021] The TOF sensor 25 is an image sensor used in the known TOF method. The TOF method is a technique that detects the distance to a subject based on the time it takes for the light pulses (illumination light) to be reflected by the subject and returned to the TOF sensor from a light source (not shown).

[0022] Bus 22 is connected to the image processing unit 15, RAM 16, recording interface 17, display unit 18, ROM 21, and control unit 30, and transfers image data, control signals, etc., output from each unit.

[0023] The image processing unit 15 stores the image data output by the AFE 14 in the RAM 16. The image processing unit 15 also performs various image processing operations, such as calculating the defocus amount, which will be described later. For example, the image processing unit 15 calculates the AF evaluation value, AE evaluation value, and AWB evaluation value. Based on the calculated AWB evaluation value, the image processing unit 15 performs auto white balance processing. The image processing unit 15 also outputs the calculated AF evaluation value and AE evaluation value to the control unit 30 via the bus 22. The detailed configuration of the image processing unit 15 will be described later.

[0024] The control unit 30 includes, for example, a processor such as a CPU (Central Processing Unit) and performs overall control of the imaging device 1. The control unit 30 controls each part of the imaging device 1 by executing a control program pre-stored in the ROM 21. For example, when the control unit 30 receives an imaging instruction via the operation unit 19, it stores the image data obtained via the imaging unit 12 and AFE 14 as an captured image in the recording medium 23.

[0025] Furthermore, the control unit 30 performs AF control based on, for example, the defocus amount calculated by the image processing unit 15. Specifically, the control unit 30 calculates the defocus amount based on the AF evaluation value and controls the position of the focus lens of the lens 10 based on the result of this distance measurement process. The AF evaluation value is, for example, a waveform based on the output of the AF pixels. Furthermore, the control unit 30 performs AE control based on, for example, the AE evaluation value calculated by the image processing unit 15. Specifically, the control unit 30 changes the sensitivity and exposure time of the solid-state image sensor 100 and controls the aperture mechanism 11 based on the AE evaluation value.

[0026] Furthermore, the control unit 30 includes an AF control unit 31 and an AE control unit 32. The AF control unit 31 performs the AF control described above based on the AF evaluation value calculated by the image processing unit 15. The AE control unit 32 performs the above-described AE control based on the AE evaluation value calculated by the image processing unit 15.

[0027] As described above, in this embodiment, the imaging control unit 40, which corresponds to the control unit 30 and the image processing unit 15, performs the above-described imaging control based on the pixel values ​​output from the solid-state image sensor 100.

[0028] Figure 2 shows an example of the functional configuration of the imaging control unit 40 according to this embodiment. The imaging control unit 40 comprises an image processing unit 15 and a control unit 30.

[0029] The image processing unit 15 performs a process to calculate the amount of defocus. This process involves interpolating and detecting the amount of defocus around the subject. The image processing unit 15 includes an image acquisition unit 150, a defocus detection unit 151, a depth map generation unit 152, an estimation unit 153, and an interpolation unit 154.

[0030] The image acquisition unit 150 acquires an image A1 of the subject captured from the solid-state image sensor 100. Image A1 is the image data output by the AFE 14.

[0031] The defocus detection unit 151 detects a first defocus amount for each of a predetermined number of pixels contained in the subject image B1. The subject image B1 is an image of the subject in the captured image A1. The first defocus amount is the amount of defocus for each of a predetermined number of pixels detected by the defocus detection unit 151. The defocus detection unit 151 detects the first defocus amount for each of a predetermined number of pixels based, for example, on the image plane phase-difference AF method. Note that detecting the first defocus amount is one example of obtaining the first defocus amount.

[0032] The number of pixels from which the defocus amount is detected by the defocus detection unit 151 is less than the total number of pixels in the captured image A1. A map showing the first defocus amount for each of a predetermined number of pixels in the subject image B1 is referred to as the coarse defocus map E1. The predetermined number of pixels in the subject image B1 are pixels from which the desired portion of the subject for which the defocus amount is calculated is captured. For example, the predetermined number of pixels in the subject image B1 are image plane phase-difference pixels.

[0033] Figure 3 shows image A10 as an example of captured image A1. Image A10 contains subject image B10 as an example of subject image B1. In Figure 3, the subject is, for example, a person's face. In Figure 3, autofocus is used to focus on the eyes of the face. In this case, the desired parts for calculating the amount of defocus are, for example, the forehead, cheeks, and nose.

[0034] Figure 3 shows pixels P1, P2, and P3 as an example of a predetermined number of pixels included in the subject image B1. Pixels P1, P2, and P3 are pixels in the regions where the forehead, cheeks, and nose were imaged, respectively. Near each of pixels P1, P2, and P3, the value of the first defocus amount detected by the defocus detection unit 151 is shown. The first defocus amounts for each of pixels P1, P2, and P3 are defocus amounts with eye height as the reference (0). In this embodiment, the subject may not be a person's face, but any object.

[0035] The depth map generation unit 152 generates a depth map C1 that shows the depth value for each pixel of the subject image B1 in the captured image A1. The depth map generation unit 152 generates the depth map C1 based on, for example, TOF technology. The depth map generation unit 152 generates the depth map C1 based on the distance to the subject detected by the TOF sensor 25.

[0036] Figure 4 shows depth map C10 as an example of depth map C1. In depth map C10, the depth value for each pixel is shown by the intensity of the grayscale. In addition, as shown in Figure 3, the values ​​of the first defocus amount are shown for pixels P1, P2, and P3 in Figure 4.

[0037] The estimation unit 153 estimates a predetermined function that shows the relationship between the first defocus amount and the depth value. In this embodiment, this predetermined function is denoted as function D1.

[0038] The interpolation unit 154 calculates a second defocus amount for any pixel in the subject image B1 within the captured image A1 based on the correlation between the first defocus amount and the depth value for each of a predetermined number of pixels. The second defocus amount is the defocus amount for each arbitrary pixel of the subject image B1 calculated by the interpolation unit 154. A map showing the second defocus amount for any pixel of the subject image B1 is referred to as the defocus map F1. The interpolation unit 154 may also calculate the defocus amount for any pixel of the captured image A1 as the second defocus amount. In this case, the map showing the second defocus amount for any pixel of the captured image A1 may be referred to as the defocus map F1. In this embodiment, the interpolation unit 154 calculates a second defocus amount based on the function D1 from the depth value indicated by the depth map C1. The depth value is obtained based on the depth map C1.

[0039] The control unit 30 includes a focus adjustment unit 300. The focus adjustment unit 300 performs autofocus based on the second defocus amount. The focus adjustment unit 300 includes the AF control unit 31 shown in Figure 1. The focus adjustment unit 300 may also perform tracking based on the second defocus amount. The focus adjustment unit 300 may also perform autofocus based on the defocus map F1. The focus adjustment unit 300 may also perform tracking based on the defocus map F1.

[0040] The memory unit 160 stores various types of information. The memory unit 160 includes the RAM 16 shown in Figure 1. The memory unit 160 stores, for example, depth map information C11 and function information D11. The depth map information C11 is information indicating the depth map C1 generated by the depth map generation unit 152. The function information D11 is information indicating the function D1.

[0041] Referring now to Figure 5, the defocus map generation process, which is the process by which the image processing unit 15 generates the defocus map F1, will be described. Figure 5 is a diagram showing an example of the flow of the defocus map generation process according to this embodiment. The defocus map generation process is started by the image processing unit 15 when AF control is performed by the AF control unit 31.

[0042] Step S10: The image acquisition unit 150 acquires an image A1 of the subject from the solid-state image sensor 100.

[0043] Step S20: The defocus detection unit 151 generates a coarse defocus map E1. Here, the defocus detection unit 151 detects a first defocus amount for each of a predetermined number of pixels contained in the subject image B1 within the captured image A1.

[0044] Step S30: The depth map generation unit 152 generates a depth map C1 that shows the depth value for each pixel of the subject image B1 in the captured image A1.

[0045] Step S40: The estimation unit 153 estimates a predetermined function (i.e., function D1) that shows the relationship between the first defocus amount and the depth value. The estimation unit 153 estimates the coefficients of function D1 based on a predetermined regression. In this embodiment, function D1 is, for example, a linear function as shown in equation (1). The coefficients of function D1 estimated by the estimation unit 153 are the slope and intercept of function D1. The predetermined regression is, for example, a linear regression.

[0046]

number

[0047] The estimation unit 153 estimates the coefficients of function D1 based, for example, on the least squares method. The estimation unit 153 estimates the coefficients of function D1 based on equation (2).

[0048]

number

[0049] In equation (2), xi represents the depth value obtained for a predetermined number of pixels based on the depth map C1. yi represents the first defocus amount detected for a predetermined number of pixels. In equation (2), the values ​​of the slope A and intercept B are calculated as variables that minimize the value of L, based on the least squares method.

[0050] Figure 6 shows the relationship between the first defocus amount and the depth value for each of a predetermined number of pixels for which the first defocus amount is detected by the defocus detection unit 151. The predetermined number of pixels are pixels P1, P2, and P3 shown in Figure 4. The depth value for each of the predetermined number of pixels is obtained based on the depth map C10. In the example shown in Figure 6, the slope and intercept of function D1 are estimated based on the least squares method using three points Q1, Q2, and Q3, which correspond to pixels P1, P2, and P3, respectively. In other words, the slope and intercept of function D1 are estimated based on linear fitting.

[0051] Returning to Figure 5, we will continue the explanation of the defocus map generation process. Step S50: The interpolation unit 154 generates a defocus map F1. Here, the interpolation unit 154 calculates a second defocus amount based on the function D1 from the depth value shown in the depth map C1. The slope and intercept of the function D1 are estimated in step S40. The interpolation unit 154 performs the process of substituting the depth value shown in the depth map C1 into the function D1 to calculate the second defocus amount for a predetermined number of pixels included in the subject image B1. The interpolation unit 154 repeats this process for all pixels included in the subject image B1. As a result of this process, the estimation unit 153 generates a defocus map F1.

[0052] Step S60; The interpolation unit 154 outputs the generated defocus map F1 to the focus adjustment unit 300. Alternatively, the interpolation unit 154 may output the calculated second defocus amount to the focus adjustment unit 300 instead of outputting the defocus map F1. In this embodiment, the interpolation unit 154 outputs the second defocus amount as the defocus map F1 to the focus adjustment unit 300. With this, the image processing unit 15 terminates the defocus map generation process.

[0053] The order in which the processes in step S20 and step S30 are executed is not limited to the order described above. The process in step S30 may be executed before the process in step S20, or the processes in step S20 and step S30 may be executed in parallel. In other words, the process of detecting the first defocus amount and the process of generating the depth map C1 may be executed in any order. Also, the process of detecting the first defocus amount and the process of generating the depth map C1 may be executed in parallel.

[0054] (Second embodiment) A second embodiment of the present invention will be described in detail below with reference to the drawings. In the first embodiment described above, the second defocus amount was calculated by calculating the slope and intercept of a linear function based on linear regression. On the other hand, the detection of the first defocus amount is susceptible to noise, and the error in the detected first defocus amount may be large. Therefore, in the method for calculating the second defocus amount of the first embodiment, the accuracy of calculating the second defocus amount may not be sufficiently high unless the first defocus amount is detected for a sufficient number of pixels. In other words, overfitting may occur in linear regression. Therefore, in this embodiment, a case is described in which the slope of a predetermined function is calculated based on the magnification of the captured image, and only the intercept of the predetermined function is estimated based on a predetermined regression. In this embodiment, the subject is the face of a person or an animal.

[0055] In this embodiment, the imaging control unit is referred to as the imaging control unit 40a, and the image processing unit is referred to as the image processing unit 15a. Note that components identical to those in the first embodiment described above are denoted by the same reference numerals, and descriptions of identical components and operations may be omitted.

[0056] Figure 7 shows an example of the functional configuration of the imaging control unit 40a according to this embodiment. The imaging control unit 40a comprises an image processing unit 15a and a control unit 30.

[0057] The image processing unit 15a includes an image acquisition unit 150, a defocus detection unit 151, a depth map generation unit 152, an estimation unit 153a, an interpolation unit 154, and a magnification estimation unit 155a.

[0058] The estimation unit 153a calculates the slope of a predetermined function based on the magnification G1 of the captured image A1, and estimates the intercept of the predetermined function based on a predetermined regression.

[0059] The magnification estimation unit 155a estimates the magnification G1 of the captured image A1.

[0060] The storage unit 160 stores depth map information C11, function information D11, magnification information G11, and slope calculation parameter information H11. Magnification information G11 is information indicating the magnification G1 estimated by the magnification estimation unit 155a. Slope calculation parameter information H11 is information indicating the slope calculation parameter H1. Slope calculation parameter information H11 may be stored in the storage unit 160 in advance, or the result of the slope calculation parameter H1 being estimated by the estimation unit 153a or the magnification estimation unit 155a may be stored in the storage unit 160.

[0061] The tilt calculation parameter H1 is, for example, a parameter that indicates the actual size of the subject used by the magnification estimation unit 155a to estimate the magnification G1, and a parameter that indicates the actual size of the depth difference between the first part and the second part of the subject. In this embodiment, as an example, the subject is a person's face.

[0062] For example, the first part and the second part are predetermined parts included in a person's face. The first part is, for example, an eye or a mouth, and the second part is, for example, a nose or an ear. In this embodiment, as an example, the first part is an eye and the second part is a nose. The first part may be an eye and the second part is an ear. The first part may be a mouth and the second part is a nose. The first part may be a mouth and the second part is an ear. The first part may be an eye and the second part is a mouth. The first part may be a nose and the second part is an ear. Regardless of which part the first and second parts are, the processing is the same as the processing described below for the case where the first part is an eye and the second part is a nose.

[0063] Referring now to Figure 8, the defocus map generation process, which is the process by which the image processing unit 15a generates the defocus map F1, will be described. Figure 8 is a diagram showing an example of the flow of the defocus map generation process according to this embodiment.

[0064] Note that the processes from step S110 to step S130, and from step S160 to step S170, are the same as the processes from step S10 to step S30, and from step S50 to step S60 in Figure 5, so their explanation is omitted. The processes in steps S140 and S150 are executed in place of step S40 in Figure 5.

[0065] Step S140: The estimation unit 153a calculates the slope of a predetermined function (i.e., function D1) based on the magnification G1 of the captured image A1. Details of the process for calculating the slope of the predetermined function will be described later.

[0066] Step S150: The estimation unit 153a estimates the intercept of a predetermined function (i.e., function D1). The estimation unit 153a estimates the intercept of function D1 based on a predetermined regression, for example. The predetermined regression is linear regression as an example. The estimation unit 153a estimates the intercept of function D1 based, for example, on the least squares method. The estimation unit 153a estimates the intercept of function D1 based on equation (3).

[0067]

number

[0068] In equation (3), xi represents the depth value obtained based on the depth map C1 for a predetermined number of pixels. yi represents the first defocus amount detected for a predetermined number of pixels. In equation (3), the slope A β This is the slope of function D1 calculated in step S140. In equation (2) above, the slope A was a variable, but in equation (3) it is the slope A β It is substituted and treated as a constant. In equation (3), the value of the intercept B is calculated as the variable that minimizes the value of L, based on the least squares method.

[0069] The slope of function D1 was calculated in step S140, and the intercept of function D1 was estimated in step S150, thus the coefficients of function D1 were estimated. Subsequently, the image processing unit 15a executes the process of step S160. With this, the image processing unit 15a terminates the defocus map generation process.

[0070] Referring now to Figure 9, the slope calculation process, in which the estimation unit 153a calculates the slope of a predetermined function, will be described. Figure 9 is a diagram showing an example of the flow of the slope calculation process according to this embodiment. The slope calculation process shown in Figure 9 is executed as the process of step S140 shown in Figure 8.

[0071] Step S210: The estimation unit 153a obtains the tilt calculation parameter H1. The estimation unit 153a obtains the actual size of the person's face and the actual size of the depth difference between the person's nose and eyes as the tilt calculation parameter H1.

[0072] First, let's explain how to obtain the actual size of a person's face. For example, the actual size of the face is the vertical size of the face, for example, the average size from the eyebrows to the chin of a person's face. The actual size of the face is a preset size. The average vertical size of a person's face is preset based on, for example, literature. The actual size of the face is stored in the storage unit 160 in advance as tilt calculation parameter information H11. When a preset size is used as the actual size of the face, a uniform value is used, which may result in a large error in the actual size of the face as a subject captured in the subject image B1. On the other hand, when a preset size is used as the actual size of the face, the processing load is lighter compared to when estimation processing is performed by the image processing unit 15a, as in the other examples described below.

[0073] As another example, the actual size of the face may be a size selected from a set of predefined sizes according to the person's attributes. In this case, the estimation unit 153a estimates the person's attributes from the subject image B1 based on machine learning. The person's attributes include, for example, age and gender. In the machine learning used by the estimation unit 153a for estimation, pairs of facial features and attributes are used as training data.

[0074] The estimation unit 153a obtains the actual face size from an attribute-specific average value table based on the estimated attributes. The attribute-specific average value table is, for example, a two-dimensional tabular data consisting of rows and columns that store the average actual face size for each attribute. The attribute-specific average value table is included in the slope calculation parameter information H11 and is pre-stored in the storage unit 160.

[0075] As another example, the actual size of the face may be a size predetermined based on a specific person. In that case, the estimation unit 153a determines the face of the person as the subject based on the subject image B1. For example, the estimation unit 153a determines identification information (e.g., an identification number) that indicates a specific person by determining the face of the person based on image recognition technology. Based on the determined identification information, the estimation unit 153a obtains the actual size of the face of the specific person from the person information table. The person information table is, for example, a two-dimensional tabular data consisting of rows and columns in which the actual size of the face is stored for each piece of identification information that indicates a specific person. The person information table is included in the slope calculation parameter information H11 and is pre-stored in the storage unit 160.

[0076] As another example, the actual size of the face may be an estimated size based on the facial shape characteristics. In this case, the estimation unit 153a extracts facial shape characteristics based on the measurement results of the distances of each part of the face by TOF. The estimation unit 153a estimates the actual size of the face based on the extracted facial shape characteristics. The measurement results used by the depth map generation unit 152 to generate the depth map C1 may be used as the measurement results of the distances of each part of the face by TOF.

[0077] The actual size of the face may be obtained by combining several of the methods described above. For example, the estimation unit 153a may estimate the actual size of the face based on the characteristics of the face shape, estimate the person's attributes, and correct the size estimated based on the characteristics of the face shape using an attribute-specific average value table based on the attributes. By correcting using an attribute-specific average value table based on the attributes, the accuracy of the size estimated based on the characteristics of the face shape can be improved.

[0078] Next, we will explain how to obtain the actual size of the depth difference between the nose and eyes on a person's face. For example, the actual size of the depth difference between the nose and eyes of a person's face is an estimated size based on the facial shape features. In this case, the estimation unit 153a extracts facial shape features based on the measurement results of the distances of each part of the face by TOF. The estimation unit 153a estimates the actual size of the depth difference between the nose and eyes based on the extracted facial shape features. Note that the measurement results used by the depth map generation unit 152 to generate the depth map C1 may be used as the distance measurement results for each part of the face by TOF.

[0079] Figure 10 shows an example of the depth difference between the nose and eyes of a person's face. In Figure 10, the depth difference d1 between the eye X1 and the nose X2 is shown.

[0080] As another example, the actual size of the depth difference between the nose and eyes of a person's face may be a preset size. A preset size is, for example, the average size of the depth difference between the nose and eyes of a person's face. The average size of the depth difference between the nose and eyes of a person's face is preset based on, for example, literature. The actual size of the depth difference between the nose and eyes is stored in the storage unit 160 in advance as tilt calculation parameter information H11.

[0081] As another example, the actual size of the depth difference between the nose and eyes of a person's face may be a size selected from a plurality of sizes pre-set according to the person's attributes. In this case, the estimation unit 153a estimates the person's attributes based on the subject image B1. The person's attributes include, for example, age and gender. Based on the estimated attributes, the estimation unit 153a obtains the actual size of the depth difference between the nose and eyes from the depth difference table. The depth difference table is, for example, a two-dimensional tabular data consisting of rows and columns that store the average value of the actual depth difference between the nose and eyes for each attribute. The depth difference table is included in the slope calculation parameter information H11 and is pre-stored in the storage unit 160.

[0082] As another example, the actual size of the depth difference between the nose and eyes of a person's face may be a size predetermined based on a specific person. In that case, the estimation unit 153a determines the face of a person as the subject based on the subject image B1. For example, the estimation unit 153a determines identification information (e.g., an identification number) that indicates a specific person by determining the face of a person based on image recognition technology. The person information table described above stores the actual size of the depth difference between the nose and eyes for each piece of identification information that indicates a specific person. Based on the determined identification information, the estimation unit 153a obtains the actual size of the depth difference between the nose and eyes of the specific person's face from the person information table.

[0083] The actual size of the depth difference between the nose and eyes of a person's face may be obtained by combining several of the methods described above. For example, the estimation unit 153a may estimate the actual size of the depth difference between the nose and eyes of a person's face based on the facial shape features, estimate the person's attributes, and correct the size estimated based on the facial shape features using a depth difference table based on the attributes. By correcting using a depth difference table based on the attributes, the accuracy of the size estimated based on the facial shape features can be improved.

[0084] Step S220: The estimation unit 153a causes the magnification estimation unit 155a to estimate the magnification G1 of the captured image A1.

[0085] The magnification estimation unit 155a estimates the magnification G1 of the captured image A1 based, for example, on the size of the subject on the imaging plane and the actual size of the subject. The magnification estimation unit 155a estimates the magnification G1 as the value obtained by dividing the size of the subject on the imaging plane by the actual size of the subject. The size of the subject on the imaging plane is the vertical size of the face on the imaging plane. The size of the subject on the imaging plane is calculated from the ratio of the size of the imaging plane to the size of the subject. The size of the imaging plane is the vertical size of the imaging plane. For example, in the imaging plane A11 shown in Figure 11, the vertical size of the imaging plane A11 is 24 mm. The ratio of the size of the subject can be calculated as the ratio of the size of the subject image B1 to the size of the captured image A1. In Figure 11, the subject image B11 is an example of subject image B1. In Figure 11, the ratio of the size of the subject image B11 to the size of the captured image A1, i.e., the ratio of the size of the subject, is 16%. Therefore, the size of the subject on the imaging plane is calculated as 3.84 mm by multiplying the vertical size of imaging plane A11 (24 mm) by the proportion of the subject's size (16%). On the other hand, the actual size of the subject is obtained in the same way as the method described above for obtaining the actual size of a person's face. The actual size of the face is, for example, 150 mm.

[0086] As another example, the magnification estimation unit 155a may estimate the magnification G1 based on the optical system conditions. Here, magnification is generally expressed as the ratio of the focal length to the imaging distance. Therefore, the magnification estimation unit 155a estimates the value obtained by dividing the focal length by the imaging distance as the magnification G1. In this case, the parameter indicating the actual size of the subject used by the magnification estimation unit 155a to estimate the magnification G1 from the tilt calculation parameter information H11 does not need to be stored in the storage unit 160 in advance, nor does this parameter need to be estimated by the estimation unit 153a or the magnification estimation unit 155a.

[0087] The magnification estimation unit 155a estimates the magnification G1 as the ratio of the size of the subject on the imaging plane to the actual size of the subject. In the example shown in Figure 11, the magnification G1 is approximately 1 / 39. Note that the values ​​used in the example in Figure 11 are just examples.

[0088] The magnification estimation unit 155a may also estimate the magnification G1 based on the magnification estimated based on the optical system conditions and the magnification estimated based on the size of the subject on the imaging plane and the actual size of the subject. For example, the average of both magnifications may be used as the magnification G1. By estimating the magnification G1 based on both magnifications, the accuracy of the magnification G1 estimation can be improved.

[0089] Step S230: The estimation unit 153a calculates the slope of a predetermined function based on the magnification G1 calculated by the magnification estimation unit 155a. Here, the estimation unit 153a calculates the slope of the predetermined function by dividing the difference in the amount of defocus between the first part and the second part by the difference between the depth value of the first part and the depth value of the second part shown by the depth map C1. The difference in the amount of defocus between the first part and the second part is calculated based on the actual size of the difference in depth between the first part and the second part of the subject and the magnification G1 of the captured image A1.

[0090] Here, we will explain in detail how the estimation unit 153a calculates the slope of a predetermined function. The estimation unit 153a calculates the vertical magnification from the magnification G1. The vertical magnification is generally calculated as the square of the magnification. The vertical magnification is the ratio of the difference in the amount of defocus between the first and second parts of the subject to the actual size of the difference in depth between the first and second parts of the subject. The estimation unit 153a calculates the difference in the amount of defocus between the eyes and the nose based on the vertical magnification. Specifically, the estimation unit 153a calculates the difference in the amount of defocus between the eyes and the nose based on equation (4).

[0091]

number

[0092] In equation (4), dZ represents the difference in the amount of defocus between the eye and the nose. β represents the magnification G1. Therefore, β 2 This indicates the vertical magnification. The actual size of the depth difference between the eyes and nose is obtained in the same way as the method described above for obtaining the actual size of the depth difference between the eyes and nose. If the actual difference in depth between the eyes and nose is, for example, 40 mm, and the magnification G1 is 1 / 39, the difference in the amount of defocus between the eyes and nose can be calculated as 0.026 mm from equation (4).

[0093] The estimation unit 153a calculates the slope of a predetermined function based on equation (5) from the difference in the amount of defocus between the eyes and the nose, and the difference between the depth value of the eyes and the depth value of the nose. In other words, the estimation unit 153a calculates the slope of a predetermined function by dividing the difference in the amount of defocus between the eyes and the nose by the difference between the depth value of the eyes and the depth value of the nose.

[0094]

number

[0095] In equation (5), A β The function's slope is shown. Xe and Xn represent the depth values ​​of the eyes and nose, respectively. The difference between the eye depth value and the nose depth value (Xn-Xe) is obtained from depth map C1.

[0096] The method used by the estimation unit 153a to calculate the slope of a predetermined function corresponds to calculating the slope of the straight line shown in Figure 12. The straight line shown in Figure 12 is a graph showing the relative defocus difference with respect to the depth value of the face, with respect to the eyes. The relative defocus difference with respect to the eyes is the difference in the amount of defocus between the eyes and the face. In Figure 12, point Q11 represents the relative defocus difference with respect to the depth value of the eyes (i.e., zero). Point Q12 represents the relative defocus difference with respect to the depth value of the nose (dZ), with respect to the eyes.

[0097] Furthermore, to calculate the slope of the predetermined function, three parts of the face (for example, the eyes, nose, and ears) may be used. For example, the estimation unit 153a may calculate the slope of the predetermined function based on linear fitting or the like, using the depth value of each of the three parts and the relative defocus difference with respect to a predetermined part (for example, the eyes). With this, the estimation unit 153a completes the slope calculation process.

[0098] (Modified version of the second embodiment) In the second embodiment described above, the case in which the slope of a predetermined function is calculated based on the magnification of the captured image, and only the intercept of the predetermined function is estimated based on a predetermined regression, was explained. In this modified example, the case in which the slope of the predetermined function is calculated based on the magnification of the captured image, while both the slope and intercept of the predetermined function are estimated based on a predetermined regression, is explained. In this modified example, the same reference numerals as in the second embodiment described above are used, and descriptions of the same configuration and operation may be omitted.

[0099] Figure 13 shows an example of the flow of the defocus map generation process related to this modified example. Note that the processes from step S210 to step S240, and from step S260 to step S270, are the same as the processes from step S110 to step S140, and from step S160 to step S170 in Figure 8, so their explanation will be omitted.

[0100] Step S250: The estimation unit 153a estimates a predetermined function (i.e., function D1) by estimating the slope and intercept of the predetermined function based on a predetermined regression, under the constraint that the slope of the predetermined function is the slope based on the magnification G1 of the captured image A1.

[0101] Here, the estimation unit 153a estimates the slope and intercept of a predetermined function (i.e., function D1). The estimation unit 153a estimates the slope and intercept of function D1 based on a predetermined regression, for example. The predetermined regression is, as an example, a ridge regression. The estimation unit 153a estimates the intercept of the function D1 based on the formula (6).

[0102]

Equation

[0103] [[ID=...]] In the formula (6), x i represents the depth value obtained based on the depth map C1 for a predetermined number of pixels. y i represents the first defocus amount detected for a predetermined number of pixels. In the formula (6), the slope A β is the slope of the function D1 calculated in step S140. In the formula (6), based on the least squares method, the value of the slope A as a variable is the slope A β close to. Under the constraint, the values of the slope A and the intercept B are calculated as variables that minimize the value of the objective function (the first term on the right side of the formula (6)). The constant λ is the slope A calculated in step S140 β is set according to how much the value of is trusted. When the confidence level of the value of the slope A calculated in step S140 β is high, the value of the constant λ is increased, and when the confidence level is low, the value of the constant λ is decreased. The constant λ is preferably a value corresponding to the reciprocal of the variance when the distribution of the slope A is assumed to be a distribution centered on the slope A β .

[0104] In this embodiment, an example in the case where the subject is a human face has been described, but it is not limited to this. The subject may be an animal face. Even when the subject is an animal face, the slope of the predetermined function is calculated in the same manner as in this embodiment by the slope calculation process described above.

[0105] (Third Embodiment) Hereinafter, the third embodiment of the present invention will be described in detail with reference to the drawings. In the first embodiment described above, the case in which the second defocus amount is calculated from the depth map C1 and the coarse defocus map E1 was explained. In the second embodiment described above, the case in which the second defocus amount is calculated from the depth map C1, the coarse defocus map E1, and the magnification G1 of the captured image A1 was explained.

[0106] On the other hand, there are cases where the error in the calculated magnification G1 of the captured image A1 is large. The information that defines the magnification G1 is optical system information (focal length, shooting conditions, etc.), or the size of the subject on the imaging plane and the actual size of the subject. Therefore, in this embodiment, we will explain a case in which the second defocus amount is calculated by performing a process equivalent to calculating the magnification G1 from the information that defines the magnification G1, based on machine learning.

[0107] In this embodiment, the imaging control unit is referred to as the imaging control unit 40b, and the image processing unit is referred to as the image processing unit 15b. Note that components identical to those in the above-described embodiments are denoted by the same reference numerals, and descriptions of identical components and operations may be omitted.

[0108] Figure 14 shows an example of the functional configuration of the imaging control unit 40b according to this embodiment. The imaging control unit 40b comprises an image processing unit 15b and a control unit 30.

[0109] The image processing unit 15b comprises an image acquisition unit 150, a defocus detection unit 151, and an interpolation unit 154b.

[0110] The interpolation unit 154b calculates a third defocus amount for each pixel of the subject image B2 based on the pixel values ​​of the subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane, using the trained model J1. The subject image B2 is the image of the subject in the captured image A1. For example, the subject image B2 is an image that, vertically, extends from the eyebrows to below the chin, and horizontally, extends from the left end of the left eyebrow to the right end of the right eyebrow.

[0111] The trained model J1 is a model that has learned the relationship between the pixel values ​​of the subject image B2, the vertical and horizontal dimensions of the subject on the imaging plane, and a third defocus amount for each pixel of the subject image B2, for which the reference value is undetermined. Here, vertical and horizontal dimensions refer to the vertical and horizontal dimensions. In this embodiment, the trained model J1 is, as an example, a neural network based on deep learning.

[0112] The third defocus amount is a defocus amount for which the reference value is undetermined. The reference value is a value corresponding to the reference depth value. Therefore, in the third defocus amount output from the trained model J1, the reference value corresponding to the depth value is undetermined. In other words, the depth value is output as a relative value.

[0113] The trained model J1 is trained to calculate a depth map C2 based on the pixel values ​​of the subject image B2. The subject is, for example, a human or animal face. The trained model J1 also infers attributes such as the subject's age and gender from the pixel values ​​of the subject image B2 and the subject's length and width on the imaging plane, and infers the actual size of the subject's face. The trained model J1 is also trained to calculate a magnification G1 from the estimated actual size of the face and the subject's length and width on the imaging plane. In other words, the trained model J1 is trained to calculate a magnification G1 based on the characteristics of the subject. Finally, the trained model J1 is trained to calculate a third defocus amount for each pixel of the subject image B2 based on the calculated magnification G1 and depth map C2.

[0114] Here, we will explain how to generate the pre-trained model J1. The model used to train the pre-trained model J1 will be called model J01. The training data consists of pairs of pixel values ​​from the training subject image and the vertical and horizontal dimensions of the subject on the training imaging plane. Tens of thousands of these pairs are used as the training dataset.

[0115] The training subject image is, for example, an image of an arbitrary region containing a subject at a fixed aspect ratio. An image of an arbitrary region containing a subject at a fixed aspect ratio is input as the training subject image. The vertical and horizontal dimensions of the subject on the training imaging surface can be obtained, for example, from the output result of subject detection based on computer vision (CV) in the imaging device (the number of pixels in the vertical and horizontal dimensions of the subject). At this time, the interpolation unit 154b can determine the vertical and horizontal dimensions of the subject on the imaging surface by calculating the pixel pitch, which is the distance between pixels on the imaging surface, and the reciprocal of the reduction ratio of the image on which subject detection is performed, and multiplying each of these by the output result of subject detection. The pixel pitch and reduction ratio may be those that have been previously stored in the storage unit 160.

[0116] Furthermore, if an image of an arbitrary region containing a subject at a fixed aspect ratio is input as a training subject image, the vertical and horizontal dimensions of the image of the arbitrary region containing the subject at a fixed aspect ratio on the imaging plane may be input instead of the vertical and horizontal dimensions of the subject on the training imaging plane. Specifically, the vertical and horizontal dimensions of the image of the arbitrary region containing the subject at a fixed aspect ratio on the imaging plane can be obtained by multiplying the vertical and horizontal dimensions (number of pixels vertically and horizontally) of the image by the pixel pitch and the reciprocal of the reduction ratio. In this case, the trained model J1 will learn the process of calculating the vertical and horizontal dimensions of the subject on the imaging plane when the vertical and horizontal dimensions of the image of the arbitrary region containing the subject at a fixed aspect ratio are input.

[0117] Furthermore, the training subject image is, for example, an image of a region containing a subject at an arbitrary ratio. An image of a region containing a subject at an arbitrary ratio is input as the training subject image. The vertical and horizontal dimensions of the training subject on the imaging plane can be obtained from the output result of subject detection (CV) in the imaging device (the number of pixels in the vertical and horizontal directions of the subject). The vertical and horizontal dimensions of the subject on the imaging plane can be obtained by multiplying the output result of subject detection by the pixel pitch and the reciprocal of the reduction ratio.

[0118] Furthermore, if an image of a region containing a subject at an arbitrary ratio is input as a training subject image, the vertical and horizontal dimensions of the image of the region containing the subject at an arbitrary ratio on the imaging plane may be input instead of the vertical and horizontal dimensions of the subject on the training imaging plane. Specifically, the vertical and horizontal dimensions of the image of the region containing the subject at an arbitrary ratio on the imaging plane can be obtained by multiplying the vertical and horizontal dimensions (number of pixels vertically and horizontally) of the image of the arbitrary region containing the subject at a fixed ratio by the pixel pitch and the reciprocal of the reduction ratio. In this case, the trained model J1 will learn to calculate the vertical and horizontal dimensions of the subject on the imaging plane when the vertical and horizontal dimensions of the image of the region containing the subject at an arbitrary ratio are input.

[0119] If training is performed using an image of an arbitrary region containing the subject at a fixed aspect ratio as the training subject image, the trained model J1 will be input with the image of the arbitrary region containing the subject at a fixed aspect ratio as the subject image B2. Similarly, if training is performed using an image of a region containing the subject at an arbitrary aspect ratio, the trained model J1 will be input with the image of the region containing the subject at an arbitrary aspect ratio as the subject image B2.

[0120] If training is performed using the output of subject detection (CV) as the length and width dimensions of the subject on the training imaging plane, the trained model J1 will receive the output of subject detection (CV) as the length and width dimensions of the subject on the imaging plane. If training is performed using the length and width dimensions of an arbitrary region containing the subject at a fixed aspect ratio on the image imaging plane instead of the length and width dimensions of the subject on the training imaging plane, the trained model J1 will receive the length and width dimensions of the image of the arbitrary region containing the subject at a fixed aspect ratio as the input for the length and width dimensions of the subject on the imaging plane. Similarly, if training is performed using the length and width dimensions of the image of the region containing the subject at an arbitrary aspect ratio on the imaging plane, the trained model J1 will receive the length and width dimensions of the image of the region containing the subject at an arbitrary aspect ratio as the input for the length and width dimensions of the subject on the imaging plane.

[0121] Model J01 receives the pixel values ​​of the training subject image and the vertical and horizontal dimensions of the subject on the training image sensor as input. Model J01 outputs a depth map C0 corresponding to each pixel position in the training subject image, and a defocus amount with an undefined reference value corresponding to each pixel position in the training subject image. A trained model J1 is generated by repeatedly training Model J01 using the training dataset so that its loss function is minimized. Depth map C0 is a depth map that shows the depth value for each pixel in the training subject image.

[0122] The loss function is a function based on the difference between the value calculated by model J01 (predicted value) and the correct value. The loss function can be calculated, for example, by the squared error obtained by squaring the difference between the predicted value and the correct value, but it is not limited to this, and other functions such as the mean square error may also be used. Below, as an example, we will explain the case where the sum of the first quantity and the second quantity is used as the loss function.

[0123] The first quantity is the sum of the squared differences at each of the following locations: the difference between the value of the first defocus amount (the first defocus amount included in the coarse defocus map E1) obtained when the training subject image was captured and its average value, and the difference between the value of the defocus amount at the pixel position corresponding to the first defocus amount and its average value among the undefined reference defocus amounts output by model J01. Equation (7) is the equation that represents the first quantity.

[0124]

number

[0125] Here, the quantities represented by each letter in equation (7) are as follows: Loss1: First amount n: A natural number indicating the number of pixels from which the first defocus amount was obtained. def 1n :nth first defocus amount def 2n : Defocus amount with an undetermined reference value corresponding to the nth first defocus amount def 1ave :Average of the first defocus amount def 2ave : The average of the undefined defocus amount output by Model J01.

[0126] When calculating the second quantity, the variance and mean of the defocus map, which is output by Model J01 with an undetermined reference value for defocus, are adjusted to match the variance and mean of the depth map C0 generated by Model J01. The second quantity is the sum of the squares of the difference between the pixel-by-pixel defocus amount shown by the defocus map, which is output by Model J01 with adjusted variance and mean, and the pixel-by-pixel depth value shown by the depth map C0 generated by Model J01. Note that the variance and mean of a map are the variance and mean of the pixel-by-pixel values ​​shown by that map over all pixels included in that map. For example, the variance and mean of the defocus map are the variance and mean of the pixel-by-pixel defocus amount shown by that defocus map over all pixels included in that defocus map. The variance and mean of the depth map C0 are the variance and mean of the pixel-by-pixel depth value shown by that depth map C0 over all pixels included in that depth map C0.

[0127] For each training dataset, the trained model J1 is generated by repeatedly updating model J01 so that the sum of the first and second quantities is minimized.

[0128] As described above, the trained model J1 is trained so that for each predetermined number of pixels in the subject image B2, the difference between the third defocus amount output by the trained model J1 and the first defocus amount detected for each predetermined number of pixels matches, and the variance and mean of the third defocus amount output by the trained model J1 and the depth value for each pixel of the training subject image match. Furthermore, the loss function of model J01 may be evaluated using only the first quantity, or only the second quantity. The loss function may also be any other loss function.

[0129] Furthermore, the interpolation unit 154b calculates a second defocus amount for any pixel in the subject image B2 based on the results of a predetermined regression or averaging based on the first defocus amount detected for each of a predetermined number of pixels in the subject image B2 and the third defocus amount for each of the predetermined number of pixels among the third defocus amounts calculated by the trained model J1, which have an undetermined reference value. The interpolation unit 154b obtains a map showing the second defocus amount for any pixel in the subject image B2 as the defocus map F1. The interpolation unit 154b may also calculate the defocus amount for any pixel in the captured image A1 as the second defocus amount. In this case, the map showing the second defocus amount for any pixel in the captured image A1 may be the defocus map F1.

[0130] The memory unit 160 stores depth map information C11 and trained model information J11. The trained model information J11 is information indicating the trained model J1. The trained model J1 is generated by a device separate from the imaging device 1. The trained model J1 generated by the device separate from the imaging device 1 is pre-stored in the memory unit 160 as trained model information J11.

[0131] The trained model J1 may be stored on an external server instead of the memory unit 160. The external server is, for example, a cloud server. When the AF control unit 31 executes AF control and the image processing unit 15 starts the defocus map generation process, the imaging device 1 acquires the trained model J1 from the external server. Alternatively, the imaging device 1 may acquire the trained model J1 from the external server immediately after the power is turned on.

[0132] Furthermore, if the trained model J1 is stored on an external server, the imaging device 1 communicates with the external server via wireless communication. In this case, the imaging device 1 includes hardware for communication via a wireless network.

[0133] Here, referring to Figures 15 and 16, the defocus map generation process, which is the process by which the image processing unit 15b generates the defocus map F1, will be described. Figure 15 is a diagram showing an example of the flow of the defocus map generation process according to this embodiment. Figure 16 is a diagram showing an example of the overview of the defocus map generation process according to this embodiment.

[0134] Note that the processes in steps S310, S320, and S350 are the same as the processes in steps S10, S20, and S60 in Figure 5, so their explanations will be omitted.

[0135] Step S330: The interpolation unit 154b calculates a third defocus amount for each pixel of the subject image B2 based on the pixel values ​​of the subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane, using the trained model J1.

[0136] The interpolation unit 154b inputs the pixel values ​​of the subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane to the trained model J1. Here, the pixel values ​​of the subject image B1 are the values ​​of all pixels included in the region of the imaging image A1 that contains the subject image B2. The shape of this region is, for example, a rectangle. The subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane that are input to the trained model J1 are as described above.

[0137] The trained model J1, upon receiving the pixel values ​​of the subject image B2 and the size of the subject on the imaging plane, outputs a third defocus amount. A map showing the third defocus amount for each pixel in the subject image B2 is referred to as the undefined reference value defocus map F10. The interpolation unit 154b generates the undefined reference value defocus map F10 based on the third defocus amount output from the trained model J1.

[0138] Step S340: The interpolation unit 154b generates a defocus map F1 based on the results of a predetermined regression based on the coarse defocus map E1 and the defocus map F10 with an undetermined reference value. Here, the interpolation unit 154b calculates a second defocus amount for any pixel in the captured image A1 or subject image B2 based on the results of a predetermined regression based on the first defocus amount detected for each of a predetermined number of pixels in the subject image B2 and the third defocus amount for each of the calculated third defocus amounts. The process in step S340 corresponds to the process of calculating a reference value for the third defocus amount.

[0139] The predetermined regression is the least squares method or ridge regression, etc. The interpolation unit 154b calculates a reference value for the third defocus amount for each pixel in the coarse defocus map E1 that contains the first defocus amount, such that the difference between the first defocus amount and the third defocus amount shown in the undefined reference value defocus map F10 is minimized. The interpolation unit 154b may generate a defocus map F1 based on the result of averaging based on the coarse defocus map E1 and the defocus map F10 with an undetermined reference value.

[0140] With this, the image processing unit 15b terminates the defocus map generation process.

[0141] (First modified example of the third embodiment) In the third embodiment described above, only a portion of the information necessary for generating the defocus map F1 is input to the trained model, and a defocus map F10 with an undetermined reference value is obtained as the output from the trained model J1, after which the reference value is determined based on a predetermined regression. In this modification, the information necessary for generating the defocus map F1 is input to the trained model, and the defocus map F1 is directly inferred as the output from the trained model.

[0142] The imaging control unit in this modified example is referred to as the imaging control unit 40c, and the image processing unit is referred to as the image processing unit 15c. In this modified example, the same reference numerals as those used in the third embodiment described above are used, and descriptions of the same configuration and operation may be omitted.

[0143] Figure 17 shows an example of the functional configuration of the imaging control unit 40c according to this embodiment. The imaging control unit 40c comprises an image processing unit 15c and a control unit 30.

[0144] The image processing unit 15c comprises an image acquisition unit 150, a defocus detection unit 151, and an interpolation unit 154c.

[0145] The interpolation unit 154b calculates a second defocus amount for each pixel of the subject image B2 based on the pixel values ​​of the subject image B2 and the size of the subject on the imaging plane, using the trained model J2.

[0146] The trained model J2 is a model in which the relationship between the pixel values ​​of the training subject image in which the training subject was captured, the size of the subject on the training imaging plane, and the amount of defocus detected for each predetermined number of pixels in the training subject image, and the amount of defocus for each pixel of the training subject image has been learned.

[0147] The trained model J2 is trained in the same way as the trained model J1 described above. Specifically, the trained model J2 is trained so that the loss function is minimized by using the sum of squared differences at each position between the value of the first defocus amount (the first defocus amount included in the coarse defocus map E1) obtained when the training subject image was captured and its mean value, and the difference between the value of the defocus amount at the pixel position corresponding to the first defocus amount and its mean value among the undetermined reference defocus amounts output by model J02 during the training stage of the trained model J2. Furthermore, the trained model J2 is adjusted so that the variance and mean of the defocus map based on the defocus amount output by model J02 match the variance and mean of the depth map C0 generated by model J02. The loss function is minimized by using the sum of squared differences between the pixel-specific defocus amount output by model J02 and the pixel-specific depth value shown in the depth map C0 after these adjustments.

[0148] Therefore, the trained model J2 is trained so that, for each predetermined number of pixels in the training subject image, the difference between the third defocus amount output by the trained model J2 and the first defocus amount detected for each of those predetermined pixels matches, and the variance and mean of the third defocus amount output by the trained model J2 and the depth value for each pixel in the training subject image match.

[0149] In the pre-trained model J2, similar to the pre-trained model J1, the process of calculating a depth map C2 based on the pixel values ​​of the training subject image is learned. Also, similar to the pre-trained model J1, the pre-trained model J2 is learned to calculate a magnification G1 from the pixel values ​​of the subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane, and to calculate a third defocus amount for each pixel of the subject image based on the calculated magnification G1 and depth map C2. Furthermore, the pre-trained model J2 is learned to perform a predetermined regression, such as the process in step S340, from the first defocus amount obtained when the subject image B2 was captured and the defocus map based on the third defocus amount, calculate a reference value, and calculate a defocus map F1 for each pixel of the subject image B2.

[0150] The memory unit 160 stores depth map information C11 and trained model information J12. The trained model information J12 is information indicating the trained model J2. The trained model J2 is generated by a device separate from the imaging device 1. The memory unit 160 pre-stores the trained model J2 generated by the device separate from the imaging device 1.

[0151] Here, referring to Figures 18 and 19, the defocus map generation process, which is the process by which the image processing unit 15c generates the defocus map F1, will be explained. Figure 18 is a diagram showing an example of the flow of the defocus map generation process according to this modified example. Figure 19 is a diagram showing an example of the overview of the defocus map generation process according to this modified example.

[0152] Note that the processes in steps S410, S420, and S440 are the same as the processes in steps S310, S320, and S350 in Figure 15, so their explanation will be omitted.

[0153] Step S430: The interpolation unit 154b generates a defocus map F1 from the subject image B2, the size of the subject on the imaging plane, and the coarse defocus map E1 based on the trained model J2. Here, the interpolation unit 154b calculates a second defocus amount for the pixels of the subject image B2 from the pixel values ​​of the subject image B1, the size of the subject on the imaging plane, and the coarse defocus map E1 based on the trained model J2. The subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane input to the trained model J2 are as described above for the trained model J1.

[0154] With this, the image processing unit 15c terminates the defocus map generation process.

[0155] (Second modified example of the third embodiment) This modified example describes a case where the coefficients for the transformation from a depth map estimated by a trained model to a defocus map are estimated by the trained model.

[0156] The imaging control unit in this modified example is referred to as the imaging control unit 40d, and the image processing unit is referred to as the image processing unit 15d. In this modified example, the same reference numerals as those used in the third embodiment described above are used, and descriptions of the same configuration and operation may be omitted.

[0157] Figure 20 shows an example of the functional configuration of the imaging control unit 40d according to this embodiment. The imaging control unit 40d comprises an image processing unit 15d and a control unit 30.

[0158] The image processing unit 15d comprises an image acquisition unit 150, a defocus detection unit 151, and an interpolation unit 154d.

[0159] The interpolation unit 154d generates a normalized surface contour map K1 based on the trained model J3. The normalized surface contour map K1 is obtained by generating a depth map C1 that shows the depth value for each pixel of the subject image B2 in the captured image A1, and then normalizing the generated depth map C1 based on the mean and variance of the depth values. As an example, the process of generating the depth map C1 that shows the depth value for each pixel of the subject image B2 may use a deep learning neural network with predetermined weights.

[0160] The interpolation unit 154d calculates the coefficients of a predetermined function from the subject image B2, the size of the subject on the imaging plane, and the first defocus amount (coarse defocus map) based on the trained model J3.

[0161] The trained model J3 is a model in which the relationship between the pixel values ​​of the subject image B2, the size of the subject on the imaging plane, the first defocus amount (coarse defocus map) detected for each predetermined number of pixels in the subject image B2, and the defocus amount for each arbitrary pixel in the subject image B2 has been learned. One example of a trained model J3 is an autoencoder.

[0162] Model J03 is a model in the training stage for training the pre-trained model J3. Model J03 is inputted with the pixel values ​​of the training subject image, the vertical and horizontal dimensions of the subject on the training imaging plane, and the amount of defocus detected for each of a predetermined number of pixels in the training subject image. Model J03 calculates coefficients based on the pixel values ​​of the training subject image, the vertical and horizontal dimensions of the subject on the training imaging plane, and the amount of defocus detected for each of the predetermined number of pixels in the training subject image. Model J03 also calculates a normalized surface contour map K1 based on the training subject image. Model J03 calculates the amount of defocus from the depth value of each point in the normalized surface contour map K1 based on a predetermined function that uses these coefficients as coefficients. The pre-trained model J3 is trained to minimize the loss function, which is the sum of the squares of the differences between the amount of defocus calculated from the depth value of each point in the normalized surface contour map K1 and the amount of defocus detected for each of the predetermined number of pixels in the training subject image.

[0163] The trained model J3, like the trained model J1, is trained to calculate the magnification G1 from the pixel values ​​of the subject image B2 and the vertical and horizontal dimensions of the subject on the imaging plane. Here, as an example of a predetermined function that shows the relationship between the magnification G1, the amount of defocus detected for each predetermined number of pixels in the subject image B2, and the normalized depth map C1, there is the linear function shown in equation (1) above. In this modified example, the coefficient A in equation (1) represents the square of the magnification G1 (i.e., the vertical magnification). Also, x represents the depth value of the normalized depth map C1. The trained model J3 is trained to calculate the intercept B, which is the coefficient in equation (1), based on the magnification G1 and the amount of defocus detected for each predetermined number of pixels in the subject image B2. The trained model J3 is trained to calculate a second defocus amount based on the magnification G1, the amount of defocus detected for each predetermined number of pixels in the training subject image, and the normalized unevenness shape map K1. In this modified example, the predetermined function is a function that represents a linear transformation. The coefficients of a given function are its slope and intercept. Note that the slope of a function representing a linear transformation corresponds to a scale transformation, and the intercept of that function corresponds to a shift in the reference value.

[0164] The interpolation unit 154d calculates a second defocus amount based on a predetermined function from the depth values ​​of each point in the normalized unevenness shape map K1.

[0165] The memory unit 160 stores depth map information C11 and trained model information J13. The trained model information J13 is information indicating the trained model J3. The trained model J3 is generated by a device separate from the imaging device 1. The memory unit 160 pre-stores the trained model J3 generated by the device separate from the imaging device 1.

[0166] Here, referring to Figures 21 and 22, the defocus map generation process, which is the process by which the image processing unit 15d generates the defocus map F1, will be explained. Figure 21 is a diagram showing an example of the flow of the defocus map generation process according to this modified example. Figure 22 is a diagram showing an example of the overview of the defocus map generation process according to this modified example.

[0167] Note that the processes in steps S510, S520, and S560 are the same as the processes in steps S310, S320, and S350 in Figure 15, so their explanation will be omitted.

[0168] Step S530: Generate a normalized surface contour map K1. Here, the trained model J3 generates a depth map C1 from the input subject image B2. When pixels of the subject image B2 are input to the trained model J3, it outputs a depth map C1 showing the depth value for each pixel of the subject image B2.

[0169] The trained model J3 generates a normalized surface contour map K1 by normalizing the generated depth map C1 based on the mean and variance of the depth value for each pixel shown in the depth map C1. For example, the normalized surface contour map K1 is normalized so that the variance of the pixel values ​​is 1 and the mean of the pixel values ​​is 0.

[0170] Step S540: The interpolation unit 154d calculates the coefficients of a predetermined function from the subject image B2, the size of the subject on the imaging plane, and the first defocus amount (coarse defocus map) based on the trained model J3.

[0171] Step S550: The interpolation unit 154d generates a defocus map F1 from the normalized unevenness shape map K1 based on a predetermined function. Here, the interpolation unit 154d calculates a second defocus amount from the depth value of each point in the normalized unevenness shape map K1 based on a predetermined function.

[0172] With this, the image processing unit 15d terminates the defocus map generation process.

[0173] In the third embodiment and its various modifications described above, an example was given where the trained models J1, J2, and J3 are deep learning neural networks, but the invention is not limited to this. Other machine learning models besides deep learning may be used as these trained models.

[0174] In the embodiments and modifications described above, an example was given in which the interpolation unit (interpolation unit 154, interpolation unit 154b, interpolation unit 154c, and interpolation unit 154d) generates a defocus map F1 showing a second defocus amount for any pixel of the subject image B1 or subject image B2, but the invention is not limited to this example. The interpolation unit (interpolation unit 154, interpolation unit 154b, interpolation unit 154c, and interpolation unit 154d) only needs to calculate the second defocus amount for a number of pixels greater than the number of pixels included in the coarse defocus map E1. In other words, the number of pixels for which the second defocus amount is calculated by the interpolation unit (interpolation unit 154, interpolation unit 154b, interpolation unit 154c, and interpolation unit 154d) only needs to be greater than the number of pixels for which the first defocus amount is detected by the defocus detection unit 151.

[0175] As described above, each embodiment of the focus detection device (in this embodiment, the imaging device 1) comprises an image acquisition unit 150, a defocus detection unit 151, and an interpolation unit (interpolation unit 154, interpolation unit 154b, interpolation unit 154c, and interpolation unit 154d). The image acquisition unit 150 acquires an image A1 in which the subject has been captured. The defocus acquisition unit (in this embodiment, the defocus detection unit 151) acquires a first defocus amount for each of a predetermined number of pixels in the subject image B1, which is an image of the subject in the captured image A1 (in this embodiment, the first defocus amount is detected). The interpolation units (interpolation units 154, 154b, 154c, and 154d) calculate a second defocus amount for any pixel of the subject image B1 based on the correlation between the first defocus amount and the depth value for each of a predetermined number of pixels.

[0176] With this configuration, the focus detection device (in this embodiment, the imaging device 1) according to each embodiment can calculate a second defocus amount for any pixel of the subject image B1 or subject image B2, so that the second defocus amount can be calculated for a greater number of pixels than the number of pixels for which the first defocus amount was obtained. In other words, a defocus map F1 with a higher density of pixels can be calculated than the coarse defocus map E1.

[0177] Conventionally, autofocus technology, specifically image plane phase-difference detection systems, detected the amount of defocus at a predetermined number of detection points on the imaging plane. However, conventional image plane phase-difference detection systems could not calculate the amount of defocus with sufficient density. This was due to hardware limitations, such as the number of lines in the phase-difference image, or computational limitations. If the amount of defocus cannot be calculated with sufficient density, it may not be possible to detect the amount of defocus corresponding to small targets such as pupils.

[0178] On the other hand, known methods for obtaining depth values ​​from images include techniques for inferring the actual depth value from a depth map, and techniques for directly inferring the amount of defocus from the degree of blur in the image. While depth maps can increase spatial density, two challenges remain: the need to use distance information on the subject side as training data, and the requirement of information not present in the image for conversion to the amount of defocus. This information not present in the image includes the reference value for the amount of defocus (the point where the amount of defocus becomes zero) and the magnification at the time of shooting. On the other hand, techniques that directly infer the amount of defocus from the degree of blur in an image require a high enough resolution to distinguish the blur, making it difficult to implement with lightweight computation. Furthermore, creating the training data used for inference is difficult, and the accuracy of the inference is not sufficient.

[0179] A depth map, which infers actual dimensions based on paraxial theory, can be converted into a relative defocus difference by multiplying it by the vertical magnification when the defocus difference due to unevenness is not large. Therefore, when "defocus amount = A × depth value + B", it is sufficient to accurately calculate A (corresponding to the vertical magnification) and B (corresponding to the defocus reference value). Accurately calculating A, which is the coefficient corresponding to the vertical magnification, and B, which is the coefficient corresponding to the defocus reference value, is the core technology of the imaging device 1 according to this embodiment.

[0180] Furthermore, some of the image processing units 15, 15a, 15b, 15c, and 15d in the above-described embodiments, for example, the image acquisition unit 150, the defocus detection unit 151, the depth map generation unit 152, the estimation unit 153, 153a, the interpolation units 154, 154b, 154c, 154d, the magnification estimation unit 155a, the normalized unevenness map generation unit 156d, and the estimation unit 157d, may be implemented using a computer. In that case, the program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Herein, "computer system" refers to the computer system built into the image processing units 15, 15a, 15b, 15c, and 15d, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Furthermore, "computer-readable recording media" may include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or communication lines such as telephone lines, as well as those that hold programs for a certain period of time, such as volatile memory within a computer system that acts as a server or client in such cases. In addition, the above-mentioned program may be for the purpose of realizing some of the functions described above, and may also be a program that can realize the above-mentioned functions in combination with a program already recorded in the computer system. Furthermore, some or all of the image processing units 15, 15a, 15b, 15c, and 15d in the above-described embodiments may be implemented as integrated circuits such as LSIs (Large Scale Integration). Each functional block of the image processing units 15, 15a, 15b, 15c, and 15d may be individually implemented as a processor, or some or all of them may be integrated into a single processor. In addition, the method of implementing the integrated circuit is not limited to LSIs; it may also be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit implementation technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit using that technology may be used.

[0181] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention. [Explanation of Symbols]

[0182] 1...Imaging device, 150...Image acquisition unit, 151...Defocus detection unit, 154, 154b, 154c, 154d...Interpolation unit, A1...Captured image, B1...Subject image

Claims

1. An image acquisition unit that acquires an image of the subject, A defocus acquisition unit that acquires a first defocus amount for each of a predetermined number of pixels included in the subject image, which is an image of the subject in the captured image, An interpolation unit that calculates a second defocus amount for any pixel of the captured image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels, A focus detection device equipped with the following features.

2. The system includes a depth map generation unit that generates a depth map showing the depth value for each pixel of the subject image in the captured image, The aforementioned depth value is obtained based on the depth map. The focus detection device according to claim 1.

3. The system further includes an estimation unit that estimates a predetermined function showing the relationship between the first defocus amount and the depth value, The interpolation unit calculates the second defocus amount based on the predetermined function from the depth value shown in the depth map. The focus detection device according to claim 2.

4. The estimation unit calculates the slope of the predetermined function based on the magnification of the captured image and estimates the intercept of the predetermined function based on a predetermined regression. The focus detection device according to claim 3.

5. The defocus acquisition unit acquires the first defocus amount based on the difference in the amount of defocus between the first part and the second part of the subject, which is calculated based on the difference in depth between the first part and the second part and the magnification. The focus detection device according to claim 4.

6. The estimation unit calculates the slope of the predetermined function based on the difference in the amount of defocus between the first part and the second part, which is calculated based on the difference in depth between the first part and the second part and the magnification. The focus detection device according to claim 5.

7. The system further includes a magnification estimation unit that estimates the magnification based on the size of the subject on the imaging surface and the actual size of the subject. The focus detection device according to any one of claims 4 to 6.

8. The interpolation unit calculates the second defocus amount based on the trained model, The aforementioned trained model has learned the relationship between the pixel values ​​of the training subject image in which the training subject is captured, the size of the subject on the imaging plane, and a third defocus amount for each pixel of the training subject image, for which the reference value is undetermined. The focus detection device according to claim 1.

9. The interpolation unit calculates the second defocus amount based on the trained model, The pre-trained model has learned the relationship between the pixel values ​​of the training subject image in which the training subject is captured, the size of the training subject on the imaging plane, and the amount of defocus detected for each of a predetermined number of pixels in the training subject image, and the amount of defocus for each pixel of the training subject image. The focus detection device according to claim 1.

10. The trained model is trained to generate a depth map from the captured image showing the depth value for each pixel of the training subject image, and to generate a normalized surface map by normalizing the generated depth map based on the mean and variance of the depth values. The focus detection device according to claim 9.

11. The number of pixels from which the second defocus amount is calculated by the interpolation unit is greater than the number of pixels from which the first defocus amount is acquired by the defocus acquisition unit. A focus detection device according to any one of claims 1 to 10.

12. The focus detection device according to claim 1, The imaging unit outputs the aforementioned captured image, An imaging device equipped with the following features.

13. The system further includes a focus adjustment unit that performs autofocus or tracking based on the second defocus amount. The imaging apparatus according to claim 12.

14. Image acquisition step: Acquire an image of the subject captured in the image, A defocus acquisition step in which a first defocus amount is acquired for each of a predetermined number of pixels included in the subject image, which is an image of the subject in the captured image, An interpolation step of calculating a second defocus amount for any pixel of the captured image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels, A focus detection method having the following characteristics.

15. On the computer, Image acquisition step: Acquire an image of the subject captured in the image, A defocus acquisition step in which a first defocus amount is acquired for each of a predetermined number of pixels included in the subject image, which is an image of the subject in the captured image, An interpolation step of calculating a second defocus amount for any pixel of the captured image based on the correlation between the first defocus amount and the depth value for each of the predetermined number of pixels, A program to execute.