Information processing apparatus, imaging apparatus, control method, and program
The information processing device stabilizes focus adjustment by determining target parameters based on subject area distribution and depth range inference, addressing the challenge of maintaining focus on specific subject areas despite changing depth ranges and obstructions.
Patent Information
- Application Number
- JP2024075409
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-11-19
AI Technical Summary
Existing focus adjustment technologies struggle to maintain optimal focus on specific subject areas, such as faces, when subjects are moving and the depth range changes rapidly, especially when parts of the subject like arms obscure the face, leading to inaccurate focus adjustment.
An information processing device that determines target parameters for focus adjustment by acquiring subject area distribution and focus detection results, inferring the depth range, and switching focus adjustment methods based on this inferred depth range.
Enables stable focus adjustment regardless of changes in subject state, ensuring that specific subject areas remain in focus even when obscured by moving objects.
Smart Images

Figure 2025170647000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an imaging device, a control method, and a program, and in particular to a focus adjustment technique. [Background technology]
[0002] There is a technology that performs focus adjustment to focus on a subject based on the defocus amount and subject distance in a predetermined focus detection area of a captured image. In such focus adjustment technology, if the subject is temporarily blocked by an object passing between the subject and the imaging device, focus adjustment may be performed to focus on the object in front (the blocking object). Patent Document 1 discloses that focus adjustment is performed by excluding from the focus target an area that indicates a subject distance that is closer than a predetermined amount relative to the average subject distance corresponding to a plurality of focus detection areas. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-137760 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in scenes where the depth range in which the subjects are distributed changes from moment to moment, particularly when the subjects are moving, it can be difficult to adjust the focus so that a specific part of the subject (for example, the face area) remains in focus.
[0005] For example, FIG. 7 shows a subject (athlete) spinning in a figure skating competition scene. Each image in FIG. 7 is captured at a different time, with the capture timing changing over time in the order of (a) → (b) → (c) → (d) → (e) → (f) → (g). In the example shown, the subject's head 701 is obscured by the left arm 702 in FIGS. 7(c) to 7(e), and the manner of obscuration differs at each capture timing. In such a scene, assume that multiple focus detection areas are set in the detected face area to maintain focus on the subject's face area. In this case, because the head 701 and the left arm 702 are the same subject, the change in defocus amount in the focus detection areas distributed at their boundary can be perceived as continuous. In other words, unlike an obscuring object that is located close to the subject and separated from it, the left arm 702 can be perceived as a part of the subject integrated with the head 701. Therefore, even if the average value of the subject distances of multiple focus detection areas derived including the left arm 702 is adopted as in Patent Document 1, it may not be possible to achieve focus adjustment that maintains optimal focus on the subject's facial area.
[0006] The present invention has been made in consideration of the above-mentioned problems, and has as its object to provide an information processing device, an imaging device, a control method, and a program that perform stable focus adjustment regardless of changes in the state of the subject. [Means for solving the problem]
[0007] In order to achieve the above-mentioned object, the information processing device of the present invention is an information processing device that determines target parameters to be used for focus adjustment when photographing a subject, and includes a first acquisition means that acquires information regarding the distribution of subject areas in an imaging angle of view, a second acquisition means that acquires focus detection results of multiple focus detection areas provided for the imaging angle of view, an inference means that infers a depth range in which the subject exists based on the information regarding the distribution of the subject areas and the focus detection results of the multiple focus detection areas, and a determination means that determines target parameters based on the inferred depth range, and is characterized in that the determination means switches the method of determining the target parameters depending on the depth range. [Effects of the Invention]
[0008] With this configuration, the present invention makes it possible to perform stable focus adjustment regardless of changes in the state of the subject. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram illustrating a hardware configuration of an imaging device 10 according to an embodiment and a modification of the present invention. [Figure 2] 1 is a diagram illustrating a detailed configuration of an image sensor 122 according to an embodiment and a modification of the present invention; [Figure 3] 1A and 1B are plan and cross-sectional views of pixels of an image sensor 122 according to an embodiment and a modification of the present invention; [Figure 4] FIG. 1 is a diagram for explaining the correspondence between a pixel structure and a pupil plane according to an embodiment and a modification of the present invention; [Figure 5] FIG. 1 is a diagram for explaining the correspondence between pixel structures and pupil division according to an embodiment and a modification of the present invention; [Figure 6] FIG. 10 is a diagram illustrating the relationship between the defocus amount and the image shift amount according to the embodiment and the modified example of the present invention. [Figure 7] FIG. 10 is a diagram illustrating a captured image relating to a figure skating competition scene. [Figure 8] FIG. 1 is a diagram illustrating a focus detection area according to an embodiment and a modification of the present invention; [Figure 9] FIG. 1 is a block diagram showing an example of the functional configuration of a subject detection unit 130 according to an embodiment and a modification of the present invention. [Figure 10] FIG. 1 is a block diagram showing an example of the functional configuration of an inference unit 132 according to an embodiment and a modification of the present invention. [Figure 11] FIG. 1 is a diagram illustrating a subject map according to an embodiment and a modification of the present invention. [Figure 12] A block diagram illustrating a learning device 1200 that builds a trained model. [Figure 13] FIG. 10 is a diagram illustrating an extracted defocus amount according to an embodiment and a modification of the present invention. [Figure 14] FIG. 10 is a diagram illustrating an example of the time transition of the width of the inference range according to the first embodiment of the present invention; [Figure 15] 1 is a flowchart illustrating a photographing process executed by the camera body 120 according to an embodiment and a modification of the present invention. [Figure 16] 1 is a flowchart illustrating a control process executed by the camera body 120 according to the first embodiment of the present invention. [Figure 17] FIG. 10 is a diagram illustrating the time transition of the amount of change over time of the width of the inference range according to the second embodiment of the present invention; [Figure 18] 10 is a flowchart illustrating a control process executed by the camera body 120 according to the second embodiment of the present invention. [Figure 19] FIG. 10 is a diagram illustrating an example of a change over time in focus position based on an inference range according to a third embodiment of the present invention. [Figure 20] 10 is a flowchart illustrating a control process executed by the camera body 120 according to the third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] [Embodiment 1] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0011] In the embodiment described below, the present invention is applied to an image capturing apparatus that captures an image for recording by driving a focus lens according to the focus state of a subject, as an example of an information processing apparatus. However, the present invention can be applied to any device that can determine target parameters that serve as a basis for focus adjustment when capturing an image of a subject.
[0012] Hardware Configuration of Imaging Device 10 Fig. 1 is a block diagram illustrating the hardware configuration of an imaging device 10 according to this embodiment. The imaging device 10 illustrated in Fig. 1 is a digital single-lens camera with an interchangeable lens. The imaging device 10 is a camera system having a lens unit 100 (interchangeable lens) and a camera body 120. The lens unit 100 is detachably attached to the camera body 120 via a mount M indicated by a dashed line in Fig. 1.
[0013] 1, the image capture device 10 is described as being configured with an interchangeable lens, but it will be readily understood that the present invention can also be realized with an image capture device configured with an integrated lens unit 100 and camera body 120. Alternatively, the image capture device 10 can include an image capture device such as a video camera.
[0014] Lens unit 100 is a photographing lens (image pickup optical system) that forms a subject image on an image pickup element 122 of camera body 120, which will be described later. In the example shown in the figure, lens unit 100 includes a first lens group 101, an aperture 102, a second lens group 103, a focus lens group (hereinafter simply referred to as a "focus lens") 104, and a drive / control system.
[0015] The first lens group 101 is disposed at the tip of the lens unit 100 and is held so as to be able to move back and forth in the optical axis direction OA. The diaphragm 102 adjusts the amount of light during shooting by adjusting its aperture diameter, and also functions as a shutter for adjusting the exposure time during still image shooting. The diaphragm 102 and the second lens group 103 are movable together in the optical axis direction OA, and realize a zoom function in conjunction with the movement of the first lens group 101 forward and backward. The focus lens 104 is movable in the optical axis direction OA, and changes the subject distance (focusing distance) at which the lens unit 100 focuses depending on its position. In other words, by controlling the position of the focus lens 104 in the optical axis direction OA, focus adjustment (focus control) to adjust the focusing distance of the lens unit 100 is possible.
[0016] In the illustrated example, the drive / control system of the lens unit 100 mainly includes actuators and drive circuits for individually driving three types of lenses: a first lens group 101, a second lens group 103, an aperture 102, and a focus lens 104. A zoom drive circuit 114 uses a zoom actuator 111 to drive the first lens group 101 and the second lens group 103 in the optical axis direction OA, thereby controlling the imaging angle of view of the imaging optical system of the lens unit 100 (realizing zoom operation). An aperture drive circuit 115 uses an aperture actuator 112 to drive the aperture 102, thereby controlling the aperture diameter and opening / closing operation of the aperture 102. A focus drive circuit 116 uses a focus actuator 113 to drive the focus lens 104 in the optical axis direction OA, thereby controlling the focal length of the imaging optical system of the lens unit 100 (performing focus control). The focus drive circuit 116 also functions as a position detector that uses the focus actuator 113 to detect the current position (lens position) of the focus lens 104.
[0017] The lens MPU (processor) 117 performs all calculations and controls related to the lens unit 100 and controls the zoom drive circuit 114, aperture drive circuit 115, and focus drive circuit 116. The lens MPU 117 also connects to the camera MPU 125 via a mount M to send and receive commands and data. For example, the lens MPU 117 detects the position of the focus lens 104 and notifies the camera MPU 125 of lens position information in response to a request. The lens position information includes information such as the position of the focus lens 104 in the optical axis direction OA, the position and diameter of the exit pupil in the optical axis direction OA when the imaging optical system is not moving, and the position and diameter of the lens frame that limits the light beam of the exit pupil in the optical axis direction OA. The lens MPU 117 also controls the zoom drive circuit 114, aperture drive circuit 115, and focus drive circuit 116 in response to a request from the camera MPU 125. The lens memory 118 stores optical information necessary for automatic focus adjustment (AF control). The camera MPU 125 controls the operation of the lens unit 100 by reading out a program stored in, for example, an internal nonvolatile memory or the lens memory 118, and expanding and executing the program in a volatile memory (not shown).
[0018] The camera body 120 has an optical low-pass filter 121, an image sensor 122, and a drive / control system. The optical low-pass filter 121 and the image sensor 122 function as an imaging section that photoelectrically converts an object image (optical image) formed via the lens unit 100 and outputs image data. In this embodiment, the image sensor 122 photoelectrically converts an object image formed via the imaging optical system and outputs an imaging signal and a focus detection signal as image data. In the following description, the first lens group 101, the aperture 102, the second lens group 103, the focus lens 104, and the optical low-pass filter 121 may be referred to as an imaging optical system.
[0019] The optical low-pass filter 121 reduces the occurrence of false colors and moiré in captured images. The image sensor 122 is composed of, for example, a CMOS image sensor and its peripheral circuits. The image sensor 122 has m horizontal pixels and n vertical pixels (m and n are integers of 2 or greater) arranged as photoelectric conversion elements. In this embodiment, the image sensor 122 also serves as a focus detection element, has a pupil-splitting function, and includes pupil-splitting pixels capable of phase-difference detection focus detection (phase-difference AF) using image data (image signals).
[0020] In the illustrated example, various types of hardware are provided as a drive / control system for camera body 120. In the example of Fig. 1, each piece of hardware is provided independently as a different circuit or processor, but at least some of these may be realized by the camera MPU 125 or the like executing a program related to the corresponding processing.
[0021] The image sensor drive circuit 123 controls the operation of the image sensor 122. The image sensor drive circuit 123 A / D converts the image signal (image data) output from the image sensor 122 and transmits it to the camera MPU 125. The image processing circuit 124 performs general image processing performed in digital cameras, such as gamma conversion, color interpolation processing, and compression encoding processing, on the image signal output from the image sensor 122. The image processing circuit 124 also generates signals for phase difference AF, AE, and subject detection.
[0022] In this embodiment, the image processing circuit 124 is described as generating a signal for phase-difference AF, a signal for AE, and a signal for subject detection, but the implementation of the present invention is not limited to this. For example, the signal for AE and the signal for subject detection may be generated as a common signal. Furthermore, the combination of signals that form a common signal is not limited to this.
[0023] The camera MPU 125 (processor, control device) performs all calculations and control related to the camera body 120. That is, the camera MPU 125 controls the image sensor drive circuit 123, image processing circuit 124, display 126, operation switch group 127, memory 128, phase difference AF unit 129, subject detection unit 130, AE unit 131, inference unit 132, and focus adjustment unit 133. The camera MPU 125 is connected to the lens MPU 117 via a signal line of the mount M to send and receive commands and data. The camera MPU 125 issues requests to the lens MPU 117 to acquire the lens position and to drive the lens by a predetermined drive amount. The camera MPU 125 also issues requests to acquire optical information specific to the lens unit 100 and transmits the requests to the lens MPU 117.
[0024] The camera MPU 125 has built-in ROM 125a that stores a program for controlling the operation of the camera body 120, RAM 125b (camera memory) that stores variables, and EEPROM 125c that stores various parameters. The camera MPU 125 also reads out the program stored in ROM 125a, expands it into RAM 125b, and executes it to perform processes such as focus detection. For example, in the focus detection process, a known correlation calculation process is performed using a pair of image signals obtained by photoelectrically converting optical images formed by light beams that have passed through different pupil regions (pupil partial regions) of the imaging optical system.
[0025] The display 126 is a display device such as an LCD, etc. The display 126 displays information about the shooting mode of the imaging device 10, a preview image before shooting and a confirmation image after shooting, an image showing the in-focus state during focus detection, etc.
[0026] The operation switch group 127 is a variety of user input interfaces provided on the camera body 120. The operation switch group 127 can include, for example, a power switch, a release (photography trigger) switch, a zoom operation switch, a photography mode selection switch, and the like.
[0027] The memory 128 is, for example, a removable flash memory (device), and records images for recording obtained by shooting.
[0028] The phase-difference AF unit 129 performs focus detection processing using a phase-difference detection method based on a phase-difference AF signal (focus detection signal) obtained from the image sensor 122 and the image processing circuit 124. More specifically, the image processing circuit 124 generates a pair of image data formed by light beams passing through a pair of pupil regions of the imaging optical system as focus detection signals, and the phase-difference AF unit 129 derives a focus deviation amount (defocus amount) based on the image deviation amount between the image data. As described above, the phase-difference AF unit 129 of this embodiment is configured to be able to perform phase-difference AF (image-surface phase-difference AF) based on the output of the image sensor 122 without using a dedicated AF sensor. As will be described in detail later, in the imaging device 10 of this embodiment, multiple focus detection areas are set for the imaging angle of view, and the image processing circuit 124 outputs information on the defocus amount (focus detection result) derived for each area as focus detection information.
[0029] The subject detection unit 130 performs subject detection processing to detect a predetermined type of subject captured in the imaging field of view on the subject detection signal (imaging signal, image signal, or captured image) generated by the image processing circuit 124. The subject detection processing can acquire information, for example, on the type and state of the subject, and the position and size of the area (subject detection area) occupied by each part of the subject in the imaging field of view. In other words, the subject detection unit 130 acquires information (hereinafter referred to as subject detection information) on the distribution of subject areas in the imaging field of view.
[0030] The AE unit 131 measures the light of the subject using AE signals obtained from the image sensor 122 and the image processing circuit 124, and performs exposure adjustment processing to set appropriate shooting conditions based on the light metering results. Specifically, the AE unit 131 measures the light using the AE signals and derives the amount of exposure for the subject at the currently set aperture value, shutter speed, and ISO sensitivity. The AE unit 131 then derives appropriate aperture value, shutter speed, and ISO sensitivity to be set during shooting from the difference between the derived exposure amount and a predetermined appropriate exposure amount, and applies them as new shooting conditions to perform exposure adjustment.
[0031] The inference unit 132 receives the imaging signal used for subject detection, the subject detection information, and the focus detection information, and infers the depth range in which the subject captured in the imaging field of view exists. In this embodiment, the inference unit 132 infers the depth range (optical axis direction) in which the subject exists using a range of defocus amounts. That is, the inference unit 132 can infer, for an imaging signal, the distance range in the depth direction in which the subject captured in the imaging signal is distributed, using the range of defocus amounts in the imaging signal.
[0032] The focus adjustment unit 133 determines the position of the focus lens 104 to be set when photographing a subject. Focus adjustment for photographing is achieved by controlling the movement of the focus lens 104 to the position determined by the focus adjustment unit 133. Although details will be described later, when determining the position of the focus lens 104, the focus adjustment unit 133 refers to the focus detection information output by the phase difference AF unit 129 and information on the value range of the defocus amount inferred by the inferrer 132.
[0033] In this way, the imaging device 10 of this embodiment is configured to be able to perform a combination of phase difference AF, photometry (exposure adjustment), and subject detection.
[0034] <Configuration of the image sensor 122> The detailed configuration of the image sensor 122 of this embodiment will be described below. Fig. 2 is a schematic diagram illustrating an example of the arrangement of image sensing pixels (and focus detection pixels) of the image sensor 122. Fig. 2 shows a 4 column x 4 row pixel (image sensing pixel) array of the two-dimensional CMOS sensor (image sensor 122).
[0035] In this embodiment, color filters of the same pattern are applied to each pixel array (pixel group 200) of 2 columns by 2 rows in the image sensor 122. In the example shown in the figure, the color filters are configured so that the upper left pixel 200R of the pixel group 200 has a high spectral sensitivity to red (R), the upper right and lower left pixels 200G have a high spectral sensitivity to green (G), and the lower right pixel 200B has a high spectral sensitivity to blue (B). As shown in the figure, one image capturing pixel is horizontally divided into two by two focus detection pixels (a first focus detection pixel 201 and a second focus detection pixel 202), resulting in a pixel array of 8 columns by 4 rows for focus detection pixels.
[0036] 2 is repeated on the surface of the image sensor 122. In one aspect, the image sensor 122 can have an imaging pixel period P of 4 μm and a pixel count N of 5,575 columns horizontally and 3,725 rows vertically, or approximately 20.75 million pixels. In this case, the image sensor 122 has a focus detection pixel column period PAF of 2 μm and a focus detection pixel count NAF of 11,150 columns horizontally and 3,725 rows vertically, or approximately 41.5 million pixels.
[0037] Figure 3(a) shows a plan view of one pixel 200G of the imaging element 122 as viewed from the light receiving surface side (+z side) of the imaging element 122, and Figure 3(b) shows a cross-sectional view of the aa section of Figure 3(a) as viewed from the -y side.
[0038] 3(b), one pixel 200G has a microlens 305 formed on the light-receiving side to collect incident light, and is provided with a photoelectric conversion unit 301 and a photoelectric conversion unit 302 that are divided into N-H divisions (two divisions) in the x direction and N-V divisions (one division) in the y direction. The photoelectric conversion unit 301 and the photoelectric conversion unit 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively.
[0039] The photoelectric conversion units 301 and 302 may be pin structure photodiodes in which an intrinsic layer is sandwiched between a p-type layer 300 and an n-type layer, or may be pn junction photodiodes by omitting the intrinsic layer as needed. In each pixel, a color filter 306 is formed between the microlens 305 and the photoelectric conversion units 301 and 302. Furthermore, the spectral transmittance of the color filter may be changed for each subpixel, or the color filter may be omitted as needed.
[0040] Light incident on pixel 200G shown in FIG. 3(b) is collected by microlens 305, dispersed by color filter 306, and then received by photoelectric conversion unit 301 and photoelectric conversion unit 302. In photoelectric conversion unit 301 and photoelectric conversion unit 302, electron-hole pairs are generated according to the amount of received light, and after being separated by a depletion layer, negatively charged electrons are accumulated in the n-type layer, while holes are discharged to the outside of the image sensor through p-type layer 300 connected to a constant voltage source (not shown). The electrons accumulated in the n-type layers of photoelectric conversion unit 301 and photoelectric conversion unit 302 are transferred to a capacitance unit (FD) via a transfer gate and converted into a voltage signal.
[0041] Fig. 4 shows a schematic diagram of the correspondence between the pixel structure of this embodiment shown in Fig. 3 and pupil division. Fig. 4 shows a cross-sectional view of the aa cross section of the pixel structure of Fig. 3(a) viewed from the +y side, and the pupil plane (pupil distance DS) of the image sensor 122. In Fig. 4, the x-axis and y-axis of the cross-sectional view are reversed with respect to Fig. 3(b) in order to correspond to the coordinate axes of the pupil plane of the image sensor 122.
[0042] In Fig. 4, the first partial pupil region 501 is generally conjugate with the light receiving surface of the photoelectric conversion unit 301, whose center of gravity is decentered in the -x direction, via a microlens, and represents the pupil region 500 that can receive light at the first focus detection pixel 201. The center of gravity of the first partial pupil region 501 is decentered on the +X side on the pupil plane. In Fig. 4, the second partial pupil region 502 of the second focus detection pixel 202 is generally conjugate with the light receiving surface of the photoelectric conversion unit 302, whose center of gravity is decentered in the +x direction, via a microlens, and represents the pupil region 500 that can receive light at the second focus detection pixel 202. The center of gravity of the second partial pupil region 502 of the second focus detection pixel 202 is decentered on the -X side on the pupil plane. In addition, in FIG. 4, a pupil region 500 is a pupil region that can receive light from the entire pixel 200G when all the photoelectric conversion units 301 and 302 (the first focus detection pixel 201 and the second focus detection pixel 202) are combined.
[0043] Image plane phase-difference AF is affected by diffraction because it uses microlenses on the image sensor 122 to divide the pupil. In Figure 4, the pupil distance to the pupil plane of the image sensor 122 is several tens of mm, while the diameter of the microlenses is several microns. This results in an aperture value of several tens of thousands, which causes diffraction blurring on the order of several tens of mm. Therefore, the image on the light receiving surface of the photoelectric conversion unit does not show a clear pupil region or pupil subregion, but rather shows light receiving sensitivity characteristics (incident angle distribution of light receiving rate).
[0044] 5 is a schematic diagram showing the correspondence between the image sensor 122 and pupil division in this embodiment. Light beams that pass through different pupil partial regions, the first pupil partial region 501 and the second pupil partial region 502, are incident on each pixel of the image sensor at different angles, and are received by the first focus detection pixel 201 and the second focus detection pixel 202, which are divided into 2×1 regions. Note that in this embodiment, the pupil region is described as being divided into two in the horizontal direction, as described above, but the direction of pupil division is not limited to this and can also include the vertical direction.
[0045] The image sensor 122 of this embodiment has an array of imaging pixels, each having a first focus detection pixel 201 and a second focus detection pixel 202. The first focus detection pixel 201 receives a light beam that passes through a first pupil partial region 501 of the imaging optical system. The second focus detection pixel 202 receives a light beam that passes through a second pupil partial region 502 of the imaging optical system that is different from the first pupil partial region 501. The imaging pixel also receives a light beam that passes through a pupil region that is the combined first pupil partial region 501 and second pupil partial region 502 of the imaging optical system.
[0046] In the image sensor 122 of this embodiment, each imaging pixel is described as being composed of a first focus detection pixel 201 and a second focus detection pixel 202. However, the imaging pixel, the first focus detection pixel 201, and the second focus detection pixel 202 may be configured as separate pixels, and the first focus detection pixel 201 and the second focus detection pixel 202 may be configured to be partially arranged in a portion of the imaging pixel array.
[0047] In this embodiment, the light receiving signals of the first focus detection pixels 201 of each pixel of the image sensor 122 are collected to generate a first focus signal, and the light receiving signals of the second focus detection pixels 202 of each pixel are collected to generate a second focus signal, and focus detection is performed using these signals. Furthermore, by adding the signals of the first focus detection pixels 201 and the second focus detection pixels 202 for each pixel of the image sensor 122, an image signal (captured image) with a resolution of N effective pixels is generated. Note that the method of generating each signal is not limited to the aspect shown in this embodiment, and other methods may be used, for example, the second focus detection signal may be generated from the difference between the image signal and the first focus signal.
[0048] <Relationship between defocus amount and image shift amount> The relationship between the amount of image shift in the first and second focus detection signals acquired by the image sensor 122 and the amount of defocus of the subject will be described below with reference to FIG. 6. In FIG. 6, the image sensor 122 (not shown) is disposed on an image sensor plane 600, and similarly to FIGS. 4 and 5, the pupil plane of the image sensor 122 is divided into a first pupil partial region 501 and a second pupil partial region 502. The magnitude of the defocus amount d is the distance from the image sensor plane to the image sensor plane, denoted by |d|. A front-focus state in which the image sensor plane is closer to the subject than the image sensor plane is defined as a negative sign (d<0). A back-focus state in which the image sensor plane is on the opposite side of the subject than the image sensor plane is defined as a positive sign (d>0). A focused state in which the image sensor plane is on the image sensor plane (focus position) is defined as d=0. In FIG. 6, the image sensor 601 is in a focused state (d=0), and the image sensor 602 is in a front-focus state (d<0). The front focus state (d<0) and the back focus state (d>0) are sometimes referred to together as the defocus state (|d|>0).
[0049] In a front-focus state (d<0), a light beam from the subject 602 that passes through the first pupil partial region 501 (second pupil partial region 502) is first focused and then spreads to a width Γ1 (Γ2) around the center of gravity G1 (G2) of the light beam, forming a blurred image on the imaging surface 600. The blurred image is received by the first focus detection pixels 201 (second focus detection pixels 202) that constitute the pixels arrayed on the image sensor and output as a first focus detection signal (second focus detection signal). Therefore, the first focus detection signal (second focus detection signal) records an image of the subject 602 at the center of gravity G1 (G2) on the imaging surface 600, with the subject 602 blurred to a width Γ1 (Γ2). The blur width Γ1 (Γ2) of the subject image increases roughly proportionally with an increase in the magnitude of the defocus amount d, |d|. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 in the center of gravity positions of the light beams) of the subject image between the first focus detection signal and the second focus detection signal also increases roughly proportionally as the magnitude |d| of the defocus amount d increases. In the back-focus state (d>0), the direction of the image shift of the subject image between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state, but a similar trend is shown.
[0050] <<Generation of focus detection results>> Since the image sensor 122 can obtain a focus detection signal in this manner, the phase-difference AF unit 129 generates focus adjustment information based on this signal. In this embodiment, as shown in FIG. 8, multiple focus detection areas are set to cover the entire imaging angle of view. In the example shown in the figure, the focus detection areas are set in a grid pattern with 18 columns in the horizontal direction and 17 rows in the vertical direction of the imaging angle of view. The phase-difference AF unit 129 derives a defocus amount for each of these focus detection areas and generates two-dimensional information (a defocus map) that stores the defocus amounts in correspondence with the arrangement of the focus detection areas. In other words, when the focus detection areas are distributed as shown in FIG. 8, an 18-column by 17-row defocus map is generated as the focus detection result. Each pixel in the defocus map indicates the defocus amount for the focus detection area corresponding to the focus detection signal.
[0051] The phase difference AF unit 129 also derives a reliability that evaluates the reliability of the derived defocus amount. Generally, the correlation calculation performed when deriving the defocus amount can be performed with higher accuracy as the amount of signal included in the spatial frequency band being evaluated increases. For example, a more accurate defocus amount can be derived for a signal with high contrast or a signal with many high-frequency components.
[0052] In this embodiment, the phase-difference AF unit 129 derives the reliability of each focus detection area based on the signal amounts of two types of focus detection signals used during focus detection for that focus detection area and their correlation value. For example, the phase-difference AF unit 129 derives a higher reliability value the greater the signal amount. The correlation value can be the degree of change in the correlation amount at the position where the correlation is highest in the correlation calculation, or the sum of absolute values of the differences between the signal used for focus detection and adjacent signals. The greater the degree of change in the correlation amount and the greater the sum of absolute values, the more accurately the defocus amount can be derived, and therefore the phase-difference AF unit 129 derives a higher reliability value.
[0053] The phase difference AF unit 129 derives, for example, reliability as a three-level value (low reliability "0," medium reliability "1," and high reliability "2"). Similar to the defocus map, the phase difference AF unit 129 also derives two-dimensional information (a reliability map) that stores the reliability derived for each of these focus detection areas in correspondence with the arrangement of the focus detection areas. That is, similar to the defocus map, the reliability map is composed of 18 columns by 17 rows.
[0054] <<Functional Configuration of the Subject Detection Unit 130>> Next, an example of the functional configuration of the subject detection unit 130 of this embodiment will be described with reference to the block diagram of FIG.
[0055] As described above, the subject detection unit 130 performs subject detection processing on the image signal (a signal for subject detection) and generates subject detection information. The subject detection information includes information (the position and size of the area) indicating which area within the imaging angle of view contains which part of the subject. There are various methods for recognizing a subject captured within the imaging angle of view from the image signal and identifying which part of the subject it is. As one aspect of such a method, the subject detection unit 130 of this embodiment employs a method that uses dictionary data in which the characteristics of each part of a predetermined type of subject are registered, and estimates, from the input image signal, the area where the part exhibiting the corresponding characteristic is likely to be distributed.
[0056] Dictionary data used to detect the subject region is stored in dictionary storage unit 905. The subject detection unit 130 of this embodiment is configured to be able to detect multiple types of subjects, such as people, vehicles, and animals, and dictionary data is provided for each of these. Furthermore, since the subject detection unit 130 further detects the distribution positions of the body parts that make up each type of subject, dictionary data is also provided for each body part. Therefore, for each type of detectable subject, dictionary data for detecting the region where the subject is distributed from the image signal input to the subject detection unit 130 and dictionary data for detecting the regions of the body parts that make up the subject are stored in dictionary storage unit 905.
[0057] Here, since the area of a part that constitutes a subject is basically part of the area of the subject, in the following explanation, the area of the part will be referred to as a "local area" and the area of the subject will be referred to as the "whole area" in comparison. In other words, the dictionary data stored in the dictionary storage unit 905 includes dictionary data for the whole area and dictionary data for the local area.
[0058] Although the term "whole area" is used in this specification as opposed to a "local area," this does not necessarily mean that the detection of an object by the object detection unit 130 is limited to detecting an area capturing the entire object. The detection of an object may vary depending on the registration mode of the object's features, such as the upper body of a person, the fuselage of an airplane, or the lead car of a train. In contrast, the detection of an object's parts differs in that it detects an area that is even smaller than the object's area.
[0059] The dictionary data stored in the dictionary storage unit 905 is used when the detection unit 902, which will be described later, detects the corresponding object region. Which dictionary data is to be used for detecting the object region is selected by the dictionary selection unit 904. In the object detection unit 130 of this embodiment, in order to detect the entire region and local regions related to a specified object, the dictionary storage unit 905 switches between multiple dictionary data for the input image signal.
[0060] The dictionary data may be selected based on, for example, a selection input of a subject type by the user or a history of detection results. In the subject detection unit 130 of this embodiment, information on the detection results performed by the detection unit 902 on the image signal is stored in the history storage unit 903. The information on the detection results may include, for example, information on the dictionary data used for detection, the number of detections, and the position and size of the area in which the subject (or part) is detected, associated with identification information that uniquely identifies the image signal to be detected.
[0061] As described above, in the subject detection unit 130 of this embodiment, the subject detection process is performed by switching between the entire region and the local region, and therefore, in order to make the process more efficient, the generation unit 901 generates image data to be used for subject detection in the detection unit 902. The generation unit 901 generates image data to be input to the detection unit 902 from the image signal input to the subject detection unit 130, depending on the detection target of the detection unit 902, i.e., the dictionary data selected by the dictionary selection unit 904. In other words, the generation unit 901 generates image data for entire region detection when the detection unit 902 detects the entire region, and generates image data for the local region when the detection unit 902 detects the local region.
[0062] The detection unit 902 executes subject detection processing on the input image data based on dictionary data selected by the dictionary selection unit 904. The subject detection processing outputs a detection result that estimates the area in the image data where the subject to be detected is distributed. In one aspect, the detection result includes information on the position and size of the area in the image data where the image of the subject to be detected appears, and the reliability of the detection result. In this embodiment, the detection unit 902 is configured as a machine-learned convolutional neural network (CNN). More specifically, the detection unit 902 is configured so that the mode of the CNN changes based on weight information defined in the selected dictionary data, for example, and is configured to be able to detect different types of subjects and parts of subjects according to the dictionary data.
[0063] Therefore, in this embodiment, the dictionary data stored in the dictionary storage unit 905 is generated by machine learning. That is, each dictionary data is generated by supervised learning in which learning image data for a subject to be detected in the dictionary data is input and information on the position and size of an area in which an image of the same type of subject (object) appears in the image data is used as training data (annotation).
[0064] The CNN may be, for example, a network in which a fully connected layer and an output layer are connected to a layer structure in which convolutional layers and pooling layers are alternately stacked. In this case, for example, backpropagation or the like may be applied to the CNN training. The CNN may also be a neocognitron CNN, which includes a feature detection layer (S layer) and a feature integration layer (C layer). In this case, a training method called "Add-if Silent" may be applied to the CNN training.
[0065] The dictionary data may be generated by machine learning in an external device such as a server and acquired by camera body 120 from that device, or may be generated by machine learning in camera body 120.
[0066] The detection unit 902 may be realized by, for example, a circuit specialized for estimation processing using a graphics processing unit (GPU) or a CNN. In the latter case, for example, a field programmable gate array (FPGA) configured to allow weight parameters to be changed can be used. Alternatively, the detection unit 902 can use any trained model other than a trained CNN. The detection unit 902 may be, for example, a trained model generated by machine learning such as a support vector machine or a decision tree.
[0067] Although the present embodiment describes an embodiment using dictionary data generated by machine learning, a configuration in which dictionary data generated by a rule base is also possible is also possible. Here, the dictionary data generated by the rule base is configured by storing, for example, images of objects to be detected or feature quantities specific to the objects, determined by a designer. The objects can be detected by comparing the images or feature quantities of the dictionary data with the images or feature quantities of the image data. Rule-based dictionary data is less complicated than models defined by trained models, and therefore has a smaller data volume. Therefore, object detection using rule-based dictionary data can shorten the processing time to obtain detection results and reduce the processing load compared to when only trained models are used.
[0068] In addition, it goes without saying that the object detection in the object detection unit 130 can employ any object detection method that does not use a trained model.
[0069] Overview of focus adjustment An outline of focus adjustment control (control of the position of the focus lens) performed for photographing a subject in the imaging device 10 of this embodiment will be described below.
[0070] In the imaging device 10 of this embodiment, focus adjustment control is performed so that an image focused on a specific part of a subject is captured. To facilitate understanding of the invention, the following description will focus on an aspect in which focus adjustment control is performed so that an image focused primarily on the face (head) and torso of a person who is a figure skating competition scene, such as that shown in FIG. 7, is captured. That is, in the aspect exemplified below, the subject detection unit 130 outputs subject detection information including information on the position and size of an area in the imaging signal in which an image of the person appears and an area in which images of the person's head and torso appear. In the following description, the person's head and torso may be referred to as a "target subject," meaning that they are the subject to be focused on.
[0071] The defocus amount for the area of the person's head and torso can be obtained from the subject detection information output by subject detection section 130 and the defocus map output by phase difference AF section 129. On the other hand, when the condition of the player changes from moment to moment as described using FIG. 7, other parts of the body, such as left arm 702, may enter the area of head 701. Therefore, even if focus adjustment control is performed based on the defocus amount for that area, there is a possibility that head 701 will not be brought into focus.
[0072] For this reason, the imaging device 10 of this embodiment is configured to infer the state of the target subject using the inference unit 132, and to switch the method of determining the defocus amount that is used as the control reference (target) in focus adjustment control based on the inference result. In other words, whether or not focus adjustment control based on the defocus amount of the head and torso area derived from the focus detection signal works appropriately varies depending on the state of the player, and the phase difference AF unit 129 changes the control processing depending on the state of the player inferred by the inference unit 132.
[0073] <Functional configuration of the inference unit 132> The functional configuration of the inference unit 132 of this embodiment will now be described with reference to the block diagram of FIG. 10. As shown in the figure, the inference unit 132 has an inference unit 1002 as a functional configuration for inferring the state of a target subject. The inference unit 1002 will be described as being configured with a machine-learned CNN (inference model) similar to the detection unit 902 of the subject detection unit 130. Like the detection unit 902, the inference unit 1002 can use, for example, a graphics processing unit (GPU) or a circuit specialized for inference using a CNN. In addition, the inference unit 1002 can also use various models, such as a Vision Transformer (ViT) or a Support Vector Machine (SVM) combined with a feature extractor.
[0074] In the illustrated example, the inference unit 1002 repeatedly performs convolution operations in the convolution layer and pooling in the pooling layer on the input data as appropriate. The inference unit 1002 then performs global average pooling (GAP) to reduce the data. The inference unit 1002 then inputs the GAP-processed data to a multilayer perceptron (MLP). The inference unit 1002 then processes any hidden layer, and outputs the inference result via the output layer. The weight parameters for each layer are acquired, for example, when the camera body 120 is shipped from the factory or when a firmware update is performed, and stored in the storage unit 1004. The inference in the inference unit 1002 is performed based on these weight parameters.
[0075] The inference unit 1002 of this embodiment is configured to infer the depth range in which the target subject is distributed as the state of the target subject for input subject-related data. Specifically, the inference unit 1002 receives an image signal, object detection information related to the image signal, a defocus map, and a reliability map as input, and infers the depth range in which the target subject captured in the image signal is located. In this embodiment, the inference unit 1002 infers the depth range in which the target subject is located as a range of defocus amounts corresponding to the nearest and farthest ends of the image signal. Hereinafter, the defocus range in which the target subject is located, as inferred by the inference unit 1002, will be referred to as the "inference range." Furthermore, the region in the image signal (image capture angle of view) in which the image of the target subject is distributed will be referred to as the "target region." Therefore, the inference unit 1002 derives information on the defocus amount of images purely related to the target subject among the images distributed in the target region as the inference range.
[0076] For this reason, the input unit 1001 generates input data having multiple channels by integrating the imaging signal, the subject detection information related to the imaging signal, the defocus map, and the reliability map input to the inference unit 132, and inputs the generated input data to the inference unit 1002. At this time, the input unit 1001 can generate input data by converting the information on the head region and the information on the torso region included in the subject detection information into a subject map in which the pixel value of the target region is set to "1" and the pixel values of other regions are set to "0", based on the information on the head region and the information on the torso region included in the subject detection information.
[0077] The subject map may be configured, for example, in a mode where the subject detection information has regions as shown in FIG. 11(a) as subject detection results, with the pixel value of region 1111 shown in FIG. 11(b) being set to "1." The example of FIG. 11(a) includes region 1101 where the entire subject is detected, region 1102 where the subject's head is detected, and region 1103 where the subject's torso is detected. In this case, region 1111 is set to a rectangular region that encompasses regions 1102 and 1103. In the mode of FIG. 11(b), region 1111 is defined in a mode where it circumscribes regions 1102 and 1103.
[0078] When generating input data, the input unit 1001 may perform upsampling or downsampling on at least some of the maps (images) in order to make the sizes (resolution, number of pixels) of the maps of multiple channels uniform.
[0079] When the inference unit 1002 obtains an inference result, the output unit 1003 outputs the inference result as output data in a predetermined format to the camera MPU 125. In this embodiment, the output data includes information on the range of defocus amounts for the target region. The output unit 1003 associates meta information, such as identification information of the imaging signal of the input data used for the inference, with the output data and outputs it.
[0080] The inference unit 1002 that infers the range of the defocus amount for such a target subject can be generated by, for example, a learning device 1200 as shown in Fig. 12. The learning device 1200 may be an external device (such as a PC or a server) different from the image capture device 10, or may be included in the image capture device 10.
[0081] The learning device 1200 is configured to perform machine learning of the depth direction range in which the head and torso of a subject are actually distributed in a learning image (learning image) for the learning target image. To this end, the acquisition unit 1202 of the learning device 1200 acquires ground truth information indicating the range of defocus amounts for a target region in the learning image as learning data 1201. In addition, the learning data 1201 includes the learning image, a defocus map corresponding to the learning image, a reliability map, and subject detection information.
[0082] The correct answer information is determined, for example, based on the distribution of defocus amounts of pixels remaining after excluding pixels in the focus detection region that include background and foreground obstructions and parts other than the head and torso of the subject from a defocus map corresponding to the training image. That is, the minimum and maximum defocus amounts of the remaining pixels indicate the defocus amounts corresponding to the farthest and nearest ends of the distribution of the head and torso of the subject in the depth direction, respectively. The selection of such excluded pixels during training may be performed, for example, by a person visually checking the training image.
[0083] The inference unit 1203 is configured with a CNN, similar to the inference unit 1002 of the inference device 132. The inference unit 1203 is configured so that weight parameters of the network used in machine learning can be changed. In the embodiment of FIG. 12 , the weight parameters are stored in the storage unit 1205, and the inference unit 1203 reads out the weight parameters when performing inference during learning. Based on the weight parameters, the inference unit 1203 infers the range of defocus amounts for the target region from the training images, defocus maps, confidence maps, and subject detection information in the training data 1201.
[0084] The loss calculation unit 1204 derives the difference between the inference result of the inference unit 1203 and the correct information in the training data 1201 as a loss. The update unit 1206 updates the weight parameters in the storage unit 1205 so as to reduce the loss derived by the loss calculation unit 1204. The updated weight parameters are used the next time the inference unit 1203 learns. By repeating this type of learning for multiple types of training images, the inference unit 1203 reduces the loss and becomes a machine learning model that can accurately infer the range of defocus amounts related to the target area of the training images. The weight parameters updated by the update unit 1206 may be output, for example, when a learning convergence condition is satisfied, and supplied to the camera body 120 or the like to be stored in the storage unit 1004. In other words, by storing the weight parameters updated through such machine learning in the storage unit 1004, the inference unit 1002 can also derive highly accurate inference results for input data.
[0085] <Method of determining target defocus amount> The focus adjustment control process performed in the imaging device 10 moves the focus lens 104 so that the defocus amount of the image of the target subject becomes 0. That is, in the control process executed by the focus adjustment unit 133, the defocus amount related to the image of the target subject in the obtained imaging signal is determined as the target of focus adjustment, and focus adjustment is performed so that this becomes 0. More specifically, the focus adjustment unit 133 determines the target defocus amount based on the distribution of defocus amounts included in the target area of the defocus map.
[0086] On the other hand, as described above, depending on the state of the person (subject), the target area may also include objects other than the target subject (other parts of the person, obstructions in the foreground, etc.) For this reason, the focus adjustment unit 133 first performs processing to extract the defocus amount included in the inference range from the defocus amounts included in the target area.
[0087] For example, when the frequency distribution of the defocus amount in the target region is as shown in Fig. 13, the focus adjustment unit 133 extracts the defocus amount included in the inference range 1301. In the frequency distribution shown in Fig. 13, the horizontal axis represents the defocus amount, and a larger defocus amount corresponds to an object on the closer side (closer to the imaging device 10).
[0088] By this extraction, it is possible to limit the defocus amount referred to in determining the target defocus amount to the inference range inferred as the distribution range of the target subject by the inference unit 1002. In other words, it is possible to exclude, by extraction, defocus amounts derived in relation to objects other than the target subject and defocus amounts in which errors have occurred due to the presence of such objects from among the defocus amounts distributed in the target region of the defocus map.
[0089] The focus adjustment unit 133 then determines a target defocus amount based on the defocus amount extracted in this manner (hereinafter referred to as the extracted defocus amount). At this time, the focus adjustment unit 133 switches the method for determining the target defocus amount depending on the width of the inference range (the width of the value range of the defocus amount). The width of the inference range indicates the extent to which the target subject extends in the depth direction, and therefore the position to which the focus lens should be moved for shooting changes depending on that extent.
[0090] Here, if the width of the inference range is less than the threshold (narrow), it can be assumed that the target subject does not extend in the depth direction, and therefore the focus adjustment unit 133 determines the value corresponding to the nearest end of the extracted defocus amounts as the target defocus amount. In other words, if the width of the inference range is less than the threshold, by performing focus adjustment control with the defocus amount corresponding to the nearest end as the target defocus amount, it can be expected that the target subject can be photographed in a state where it is suitably in focus from the nearest end to the farthest end.
[0091] On the other hand, if the width of the inference range is equal to or greater than the threshold (greater than or equal to the threshold), the target subject is deemed to extend in the depth direction. In this state, if focus adjustment control is performed using the defocus amount corresponding to the nearest end as the target defocus amount, as in the case where the width is below the threshold, the portion of the target subject that is farther from the image capture device 10 may not be included in the depth of field. Also, in a situation where an athlete spins as shown in FIG. 7, a sudden change in the posture of the target subject occurs between time T(a) and time T(g), as shown in FIG. 14, and the depthwise extent of the target subject fluctuates. The alphabet in parentheses around time T corresponds to the corresponding imaging signal in FIG. 7. For example, time T(c) corresponds to the acquisition timing of the imaging signal in FIG. 7(c), and time T(f) corresponds to the acquisition timing of the imaging signal in FIG. 7(f). The example in FIG. 14 indicates that the width of the inference range exceeds the threshold Dw during the period from time T(c) to time T(f). Therefore, if the defocus amount corresponding to the nearest end is set as the target defocus amount under circumstances in which such fluctuations occur, the focus lens will move frequently, making it difficult to achieve a state in which the target subject is in focus, and there is a possibility that shooting will not be able to start. Therefore, when the width of the inference range is equal to or greater than the threshold, the focus adjustment unit 133 determines the average value of the extracted defocus amounts as the target defocus amount. In other words, if the width of the inference range exceeds the threshold, by performing focus adjustment using the average value of the extracted defocus amounts as the target defocus amount, it is expected that fluctuations in the focus state can be reduced and the target subject can be stably focused.
[0092] Note that, since image capture by the image capture device 10 is performed intermittently, subject detection and focus detection are performed each time, and the focus adjustment unit 133 acquires an inference range based on the image capture signal and the subject detection information and focus detection result corresponding to the image capture signal, and performs focus adjustment control. Alternatively, these processes do not need to be performed for all image captures performed by the image capture device 10, i.e., all image signals acquired as live view images, but may be performed only for image signals acquired at predetermined intervals. In addition, the focus adjustment control of the focus adjustment unit 133, which involves inference by the inference unit 132, may be performed only when an operation input related to an instruction to capture an image for recording is made.
[0093] 《Shooting Processing》 The following describes in detail the photographing process performed by the camera body 120 of this embodiment having such a configuration, using the flowchart in Figure 15. The process corresponding to this flowchart can be realized by the camera MPU 125 reading out a corresponding processing program stored in, for example, ROM, expanding it into RAM, and executing it. This photographing process will be described as being started, for example, when an operation input related to a photographing instruction is detected. Note that, in this embodiment, a photographing instruction is described as causing a still image for recording to be captured and recorded, but the present invention is not limited to this, and photographing of a moving image for recording may also be started in response to this instruction.
[0094] In S1501, subject detection unit 130 executes subject detection processing based on a subject detection signal under the control of camera MPU 125. In this embodiment, the target subjects are the head and torso of the person (player) who is the subject, and so subject detection unit 130 detects at least the entire area, head area, and torso area of the subject, for example, by varying the dictionary data used by detection unit 902. Upon completion of subject detection processing, subject detection unit 130 composes and outputs subject detection information.
[0095] In this embodiment, to facilitate understanding of the invention, a single subject (person) is captured within the imaging angle of view as shown in Fig. 7, and focus adjustment control is performed to focus on the head and torso of the subject, but the implementation of the present invention is not limited to this. For example, when multiple subjects are captured within the imaging angle of view, one of the multiple subjects may be selected as a main subject, and focus adjustment processing may be performed on the head and torso of the main subject as a target subject to focus on. The main subject may be determined based on a priority order in descending order of the position of the subject within the area where it is detected, in descending order of the position closest to the center of the imaging angle of view (lowest image height).
[0096] In S1502, the phase difference AF unit 129 executes focus detection processing under the control of the camera MPU 125. When the focus detection processing is completed, the phase difference AF unit 129 generates and outputs a defocus map and a reliability map.
[0097] In S1503, the focus adjustment unit 133 executes a control process for performing focus adjustment control under the control of the camera MPU 125.
[0098] Control Processing The control process executed in this step will now be described in detail with reference to the flowchart of FIG.
[0099] In S1601, under the control of the camera MPU 125, the inference unit 132 derives an inference range for the target subject based on the imaging signal, the subject detection information corresponding to the imaging signal, the defocus map, and the reliability map. As described above, in this step, the input unit 1001 of the inference unit 132 generates input data including the subject map for the target subject based on these data and inputs the input data to the inference unit 1002. The inference unit 1002 infers the range of defocus amounts for the target subject (inference range) from the input data based on the weight parameter information stored in the storage unit 1004. The output unit 1003 then outputs information about the inference range derived by the inference unit 1002.
[0100] In S1602, the focus adjustment unit 133 extracts the defocus amount included in the target area of the defocus map corresponding to the imaging signal based on the inference range derived in S1601. As described above, the processing of this step specifies the extracted defocus amount from among the defocus amounts included in the target area of the defocus map.
[0101] In S1603, the focus adjustment unit 133 determines whether the width of the inference range derived in S1601 is equal to or greater than a first threshold (Dw). If the focus adjustment unit 133 determines that the width of the inference range is equal to or greater than the first threshold, it proceeds to S1604, and if it determines that the width is less than the first threshold, it proceeds to S1605.
[0102] In S1604, the focus adjustment unit 133 determines the average value of the extracted defocus amounts as the target defocus amount.
[0103] On the other hand, if it is determined in S1603 that the width of the inference range is less than the first threshold, the focus adjustment unit 133 determines the maximum value of the extracted defocus amounts as the target defocus amount in S1605. That is, the focus adjustment unit 133 determines the defocus amount corresponding to the nearest end of the extracted defocus amounts as the target defocus amount.
[0104] In S1606, the focus adjustment unit 133 determines the drive amount of the focus lens 104 to be moved in focus adjustment based on the determined target defocus amount.
[0105] In S1607, the focus adjustment unit 133 transmits information about the drive amount of the focus lens 104 together with a drive request to the lens unit 100, and causes the lens MPU 117 to drive the focus lens 104. Based on the drive request, the lens MPU 117 controls the focus drive circuit 116 to move the focus lens 104 by the specified drive amount.
[0106] When the control process is completed in this manner, the photographing process proceeds to S1504.
[0107] In S1504, camera MPU 125 determines whether or not the target subject is in focus, based on the defocus map generated by phase difference AF section 129 for the imaging signal obtained after moving focus lens 104. If camera MPU 125 determines that the target subject is in focus, it proceeds to S1505, and if it determines that the target subject is not in focus, it returns to S1501.
[0108] In S1505, the camera MPU 125 executes control relating to capturing a still image for recording. By this control, a still image for recording is generated and stored in the memory 128.
[0109] As described above, the information processing device of this embodiment enables stable focus adjustment regardless of changes in the state of the subject. That is, the information processing device infers the depth range in which the subject exists based on an imaging signal and the corresponding object map, defocus map, and reliability map, and switches the method for determining the target defocus amount used for focus adjustment based on the inference result. This configuration allows the target defocus amount to be determined while excluding defocus amounts affected by obstructions, etc. For example, even in a scene where a person's arm is stretched out in front of their face, a defocus amount range excluding the arm area can be extracted from the defocus amount based on the focus detection area corresponding to the face. Furthermore, even if a sudden change in the state of the subject causes the inference range to expand, this prevents the subject from becoming unstable in focus.
[0110] In this embodiment, the detection unit 902 detects a subject that is a person, and the inference unit 132 estimates the range of defocus amounts corresponding to the person's head and torso. However, the present invention is not limited to this. That is, the type of subject that is the target of detection and inference is not limited to a person, but can include, for example, vehicles, animals, etc. In such cases, the processing used for detection and inference may be different from when the subject type is a person. Therefore, for example, weight parameters used by the detection unit 902 and the inference unit 1002 may be provided for each type of subject and configured to be changeable.
[0111] In addition, although the present embodiment has been described with reference to a case where the target subject is the head and torso of a person, it goes without saying that the present invention is not limited to this. The part of the subject that is the target subject may be configured to be set by the photographer, for example.
[0112] Although the first threshold value is set in advance in the present embodiment, the present invention is not limited to this. The first threshold value may be adaptively changeable depending on, for example, the distance to the subject, the aperture value of the lens, or the type of subject.
[0113] Furthermore, in this embodiment, the drive amount of focus lens 104 is determined based on the target defocus amount, but the present invention is not limited to this. For example, focus adjustment unit 133 may determine the drive amount for a focus position derived based on a focus position predicted from the history of the most recent focus positions of focus lens 104 and a focus position related to the target defocus amount.
[0114] [Embodiment 2] In the above-described embodiment, a method for determining a target defocus amount from among extracted defocus amounts is switched based on whether the inference range obtained for each imaging signal exceeds a first threshold. However, the present invention is not limited to this. For example, even if the inference range exceeds the first threshold, if the width of the inference range remains stable for a predetermined period of time, using the average of the extracted defocus amounts as the target defocus amount may give the photographer the impression that the target subject is still not properly focused. For this reason, this embodiment describes a method for determining a target defocus amount based on multiple defocus amounts included in the extracted defocus amount when the focus state is unstable due to a sudden change in the state of the target subject. More specifically, the focus adjustment unit 133 of this embodiment determines whether a sudden change in state has occurred in the target subject based on whether the amount of change over time in the width of the inference range is equal to or greater than a predetermined threshold, and then switches the method for determining the target defocus amount.
[0115] The amount of change in the width of the inference range over time can be calculated, for example, for each of the imaging signals sequentially acquired for subject detection, focus detection, and inference, as the difference between the width of the inference range associated with that imaging signal and the width of the inference range associated with the imaging signal acquired at the most recent acquisition timing. In the case where the player's state continuously changes as shown in FIG. 7, the difference in the width of the inference range calculated by the inference unit 132 is, for example, as shown in FIG. 17. The example in FIG. 17 indicates that the amount of change in the width of the inference range over time (the difference from the most recent inference range width) exceeds the threshold Dx during the period from time T(c) to time T(g). In other words, it is assumed that the depthwise distribution of the target subject fluctuates sharply during this period. Therefore, if the defocus amount corresponding to the nearest end is set as the target defocus amount under conditions in which such fluctuations in the amount of change over time occur, the focus lens will frequently move, making it difficult to achieve a state in which the target subject is in focus, and it may be impossible to start shooting. Therefore, when the amount of change over time in the width of the inference range is equal to or greater than the threshold, the focus adjustment unit 133 determines the average value of the extracted defocus amounts as the target defocus amount. In other words, if the amount of change over time in the width of the inference range exceeds the threshold, by performing focus adjustment using the average value of the extracted defocus amounts as the target defocus amount, it is expected that fluctuations in the focus state will be reduced without forcing the focus lens 104 to follow sudden state changes.
[0116] Control Processing The control processing executed in S1503 of the shooting processing of this embodiment will be described in detail below with reference to the flowchart in Fig. 18. Note that in the control processing of this embodiment, steps that perform processing similar to that of the control processing of embodiment 1 are given the same reference numerals and descriptions thereof will be omitted, and the following description will be limited to steps that perform processing characteristic of this embodiment.
[0117] After identifying the extracted defocus amount for the imaging signal acquired at the current acquisition timing in S1602, the focus adjustment unit 133 determines in S1801 whether the amount of change over time in the width of the inference range is equal to or greater than a second threshold (Dx). First, the focus adjustment unit 133 derives the difference between the width of the inference range of the target subject obtained for the imaging signal acquired at the current acquisition timing and the width of the inference range of the target subject obtained for the imaging signal acquired at the most recent acquisition timing as the amount of change over time. The focus adjustment unit 133 then compares the derived amount of change over time with the second threshold to make the determination in this step. If the focus adjustment unit 133 determines that the amount of change over time in the width of the inference range is equal to or greater than the second threshold, the process proceeds to S1604. If the focus adjustment unit 133 determines that the amount of change over time in the width of the inference range is less than the second threshold, the process proceeds to S1605.
[0118] In this way, according to the control process of this embodiment, stable focus adjustment can be performed when it is inferred that a sudden change will occur in the depth direction range in which the target subject exists.
[0119] In this embodiment, the amount of change in the width of the inference range over time is calculated as the difference from the width of the inference range obtained most recently, but the present invention is not limited to this. The amount of change in the width of the inference range over time may be calculated based on information about the inference range for the most recent predetermined period.
[0120] Although the second threshold value is set in advance in the present embodiment, the present invention is not limited to this. The second threshold value may be adaptively changeable in accordance with, for example, the distance to the subject, the aperture value of the lens, or the type of subject, similar to the first threshold value.
[0121] [Variation 1] In the above-described embodiment, the average value of the extracted defocus amounts is determined as the target defocus amount when the target subject is in a state of being expanded in the depth direction, but the present invention is not limited to this. The target defocus amount determined in this state may be the median value of the distribution of the extracted defocus amounts.
[0122] [Variation 2] In the first embodiment described above, when the width of the inference range exceeds the first threshold, the average value of the extracted defocus amounts is determined as the target defocus amount. However, the present invention is not limited to this. For example, when the width of the inference range is below the first threshold, the focus adjustment unit 133 may determine the defocus amount corresponding to the closest extracted defocus amount as the target defocus amount. However, when the width of the inference range exceeds the first threshold, the focus adjustment unit 133 may not need to determine a new target defocus amount. In this case, the focus adjustment unit 133 may determine the target defocus amount by using the target defocus amount determined for the most recently acquired imaging signal for inference. In other words, when the width of the inference range exceeds the first threshold, the focus adjustment unit 133 does not determine a new target defocus amount based on the extracted defocus amount, but instead redetermines the defocus amount most recently determined as the target defocus amount as the target defocus amount.
[0123] In this way, when it is inferred that a state change that affects the depth distribution has occurred in the target subject, the target defocus amount is not changed depending on the defocus amount included in the extracted defocus amount, so stable focus adjustment can be achieved. In other words, when the width of the inferred range exceeds the first threshold, control is performed so that driving of the focus lens 104 is temporarily not performed, so focus adjustment can be stabilized regardless of changes in the posture of the target subject.
[0124] This control related to the determination of the target defocus amount can also be applied to the second embodiment. In other words, in the second embodiment described above, when the amount of change over time of the inference range exceeds the second threshold, the average value of the extracted defocus amounts is determined as the target defocus amount. However, the present invention is not limited to this. For example, when the amount of change over time of the inference range falls below the second threshold, the focus adjustment unit 133 may determine the defocus amount corresponding to the closest one of the extracted defocus amounts as the target defocus amount, but when the amount of change over time exceeds the second threshold, it may not be necessary to determine a new target defocus amount. In this case, the focus adjustment unit 133 may determine the target defocus amount by using the target defocus amount determined for the most recently acquired imaging signal for inference as is.
[0125] [Embodiment 3] In the above-described embodiment, in a situation where a sudden change in the state of the target subject is expected, the average value of the defocus amounts in the target region extracted within the inference range is determined as the target defocus amount, regardless of the value range of the inference range. However, the state change of the target subject does not occur uniformly in the depth direction. That is, the state change of the target subject does not necessarily occur equally on both the side closer to and the side farther away from the imaging device 10 based on the currently focused subject distance. Therefore, if the state change occurs in such a way that the depth range in which the target subject exists temporarily extends to either side, the focus state of the target subject may not be stable if a defocus amount such as the average value is used as the target defocus amount. In this embodiment, a bias in the state change of the target subject is estimated based on the inference range, and the method for determining the target defocus amount is changed based on the estimation result.
[0126] The bias in the state change can be determined, for example, from the positions (focus positions) of the focus lens 104 corresponding to the nearest end and the farthest end of the inference range. The focus position of the focus lens 104 can be derived by converting the defocus amount (maximum or minimum value of the inference range) related to the corresponding end into a drive amount for the focus lens 104 and then adding it to the current focus position of the focus lens 104. In the following description, the focus position corresponding to the nearest end of the inference range will be referred to as the first focus position, and the focus position corresponding to the farthest end will be referred to as the second focus position.
[0127] When determining the target defocus amount, the focus adjustment unit 133 switches the determination method based on the change over time in the first focus position and the second focus position (i.e., the change over time in the position of the focus lens 104) rather than the width of the inference range. In this embodiment, the focus adjustment unit 133 determines that an abrupt change in the state of the target subject has occurred when, for example, the sum of the absolute values of the differences between the average focus position over the most recent predetermined period and the focus position at each time during that period exceeds a predetermined threshold. This state exhibits a so-called "pulsation" behavior in which the focus position instantaneously increases / decreases and returns to its original state in the time direction. In other words, if the fluctuation range of the focus position corresponding to either end within a predetermined period exceeds a predetermined range, the focus adjustment unit 133 can determine that an abrupt change in the state of the subject has occurred in the corresponding direction of the optical axis.
[0128] For example, if neither the first focus position nor the second focus position fluctuates beyond a predetermined range, it can be determined that the state of the subject is stable. In this case, the focus adjustment unit 133 can determine the defocus amount corresponding to the closest one of the extracted defocus amounts as the target defocus amount. Furthermore, if at least one of the first focus position and the second focus position fluctuates beyond a predetermined range, it can be determined that a sudden change in state has occurred in the target subject. In this case, if there is a difference in the fluctuation range of each focus position, the depth direction spread due to the change in state of the target subject will be biased toward the side closer to the imaging device 10 and the side farther away from the focused subject distance.
[0129] When the state change of the target subject is biased in the depth direction, the focus adjustment unit 133 determines the target defocus amount according to the side where the state is considered to be stable (near side / far side). In this embodiment, when a sudden state change of the target subject occurs on the near side, the focus adjustment unit 133 determines the target defocus amount to be a value obtained by offsetting the defocus amount corresponding to the farthest end of the extracted defocus amount by a predetermined amount toward the near side. Furthermore, when a sudden state change of the target subject occurs on the far side, the focus adjustment unit 133 determines the target defocus amount to be a value obtained by offsetting the defocus amount corresponding to the farthest end of the extracted defocus amount by a predetermined amount toward the far side. In this way, focus adjustment control is performed in a manner that reduces the influence of biased, sudden state changes, thereby ensuring stable focus tracking.
[0130] In addition, if the sudden state change that occurs in the target subject is similar on the close side and the far side (the difference between the fluctuation range of the first focus position and the fluctuation range of the second focus position is below a predetermined threshold), the focus adjustment unit 133 may use the average value as the target defocus amount, as in embodiments 1 and 2.
[0131] FIG. 19 illustrates an example of the temporal changes in the first focus position and the second focus position based on the inference range. In the figure, the black circle at each time indicates the focus position (first focus position) corresponding to the nearest end of the inference range at that time, and the white circle indicates the focus position (second focus position) corresponding to the farthest end of the inference range at that time. The value An indicates the average value of the first focus position from time T(i) to time T(o), and the value Af indicates the average value of the second focus position over the same period. In the example shown in the figure, a steep change is observed only on the side of the first focus position corresponding to the nearest end, so a defocus amount that is offset by a predetermined amount toward the near side from the defocus amount corresponding to the farthest end of the extracted defocus amount is determined as the target defocus amount.
[0132] Control Processing The control processing executed in S1503 of the shooting processing of this embodiment will be described in detail below with reference to the flowchart in Fig. 20. Note that in the control processing of this embodiment, steps that perform processing similar to that of the control processing of embodiment 1 are given the same reference numerals and descriptions thereof will be omitted, and the following description will be limited to steps that perform processing characteristic of this embodiment.
[0133] After identifying the extracted defocus amount in S1602, the focus adjustment unit 133 determines in S2001 whether a fluctuation exceeding a predetermined width occurs in at least one of the focus positions corresponding to the nearest end and the farthest end of the inference range related to the most recent predetermined period. If the focus adjustment unit 133 determines that a fluctuation exceeding the predetermined width occurs in at least one of the first focus position and the second focus position related to the inference range, the focus adjustment unit 133 proceeds to S2002. If the focus adjustment unit 133 determines that a fluctuation exceeding the predetermined width does not occur in either the first focus position or the second focus position related to the inference range, the focus adjustment unit 133 proceeds to S1605.
[0134] In S2002, the focus adjustment unit 133 determines which of the fluctuation ranges of the first focus position and the second focus position is larger. If the focus adjustment unit 133 determines that the fluctuation range of the first focus position is larger than that of the second focus position, it proceeds to S2003. If the focus adjustment unit 133 determines that the fluctuation range of the second focus position is larger than that of the first focus position, it proceeds to S2004. If the focus adjustment unit 133 determines that the fluctuation ranges of the first focus position and the second focus position are approximately the same, it proceeds to S1604. Here, the fluctuation ranges of the first focus position and the second focus position are considered to be approximately the same if the difference between them is below a predetermined value. Therefore, a large fluctuation range of either focus position may be determined by a requirement that the difference between the fluctuation ranges be equal to or greater than a predetermined value.
[0135] In S2003, the focus adjustment unit 133 determines a value obtained by subtracting (offsetting) a predetermined amount from the maximum value of the extracted defocus amount as the target defocus amount. That is, the focus adjustment unit 133 determines a value obtained by offsetting a predetermined amount toward the near side from the defocus amount corresponding to the farthest end of the extracted defocus amounts as the target defocus amount.
[0136] On the other hand, if it is determined in S2002 that the fluctuation range of the second focus position is larger, the focus adjustment unit 133 determines in S2004 a value obtained by increasing (offsetting) the minimum value of the extracted defocus amount by a predetermined amount as the target defocus amount. That is, the focus adjustment unit 133 determines a value obtained by offsetting the defocus amount corresponding to the nearest end of the extracted defocus amounts by a predetermined value toward the farther side as the target defocus amount.
[0137] In this way, according to the control processing of this embodiment, when it is inferred that a sudden change will occur in the depth range in which the target subject exists, control can be performed so that stable focus adjustment is performed in accordance with the change in state.
[0138] In this embodiment, whether or not a sudden change in the state of the target subject has occurred is determined using the sum of the absolute values of the differences between the average focus position over the most recent predetermined period and the focus position at each time during that period. However, the present invention is not limited to this. For example, the determination may be made based on the sum of the absolute values of the differences between the focus position at each time and the average focus position over the most recent predetermined period.
[0139] [Variation 3] In the above-described embodiment and modified example, the inferring unit 132 infers the defocus amount range as information indicating the depth range in which the target subject exists. However, the present invention is not limited to this. The defocus amount range is one way of defining the depth range, and it goes without saying that other parameters can be used to identify the depth range of the target subject. Such parameters can include other parameters from which the depth range can be derived, such as the focus position, the subject distance, or the amount of image shift in the focus detection signal.
[0140] [Variation 4] In the above-described embodiment, a defocus map is generated based on the focus detection results of multiple focus detection areas set across the entire imaging angle of view, and subject detection information is generated for a specific subject (person) captured within the imaging angle of view. However, the present invention is not limited to this example. For example, subject detection may be limited to a predetermined region, such as the center of the imaging angle of view. Furthermore, the defocus map may be generated based only on the focus detection results of the focus detection area corresponding to the region where the subject was detected. In this example, a more precise defocus map can be obtained by changing the size and increasing the density of the focus detection areas and arranging them in the region where the subject was detected.
[0141] In order to reduce the calculation load on the inference unit 132, the information input to the inference unit 132 and the input data generated by the input unit 1001 of the inference unit 132 can be configured from, for example, information extracted from an area where the target subject is distributed.
[0142] [Variation 5] In the above-described embodiment and modified example, the learning device 1200 learns image signals, object detection information corresponding to the image signals, a defocus map, and a reliability map to construct an inference model used in the inference unit 132. Furthermore, in these embodiments, the inference unit 132 requires input of a target imaging signal, an object map corresponding to the imaging signal, a defocus map, and a reliability map. However, the present invention is not limited to this. The inference unit 132 can also employ an inference model capable of inferring a depth direction range in which a target object exists without requiring all of these as input. For example, in an embodiment in which a defocus map is limited in advance to information having a predetermined reliability and then learned, or in an embodiment in which the scene to be captured, the object type, and the imaging settings are limited, the imaging signal and the reliability map can be omitted during machine learning and inference. That is, the inference unit 132 can also employ an inference model configured to infer a depth direction range in which the object exists based on information about the distribution of object regions within the imaging field of view and focus detection results of multiple focus detection regions within the imaging field of view.
[0143] Similarly, in the above-described embodiment and modified example, a method for determining a target defocus amount from an extracted defocus amount based on information about an inference range as a target parameter for focus adjustment has been described, but the present invention is not limited to this. Needless to say, the target parameter used for focus adjustment may be other parameters that can be converted into a defocus amount, such as a focus position or a subject distance.
[0144] [Variation 6] In the above-described embodiment and modified example, a defocus amount included in an inference range is extracted from among the defocus amounts included in the target area, and a target defocus amount is determined based on the extracted defocus amount. However, the present invention is not limited to this. For example, the present invention may determine target parameters such as a target defocus amount based on an inference range inferred using information regarding the distribution of subject areas within the imaging field of view and focus detection results as input.
[0145] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0146] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention.
[0147] [Summary of the embodiment and modifications] The disclosure of this specification includes the following information processing device, imaging device, control method, and program. (Item 1) An information processing device that determines target parameters to be used for focus adjustment when photographing a subject, a first acquisition means for acquiring information regarding the distribution of a subject area in an imaging angle of view; a second acquisition means for acquiring focus detection results of a plurality of focus detection areas provided for the imaging angle of view; an inference means for inferring a depth direction range in which the subject exists based on information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; determining means for determining the target parameters based on the inferred depth range; and The determining means switches a method for determining the target parameter depending on the range in the depth direction. 1. An information processing device comprising: (Item 2) The information processing device described in item 1 is characterized in that the determination means extracts focus detection results that fall within the inferred depth direction range from the focus detection results of the multiple focus detection areas, and determines the target parameters based on the extracted focus detection results. (Item 3) The information processing device described in item 2 is characterized in that the method for determining the target parameter includes a first determination method for determining a value corresponding to the nearest end of the extracted focus detection results as the target parameter, and a second determination method for determining a value not corresponding to the nearest end of the extracted focus detection results as the target parameter. (Item 4) Item 3. The information processing device according to item 3, characterized in that the determination means determines the target parameter by the second determination method when the width of the depth-direction range exceeds a first threshold, and determines the target parameter by the first determination method when the width of the depth-direction range is below the first threshold. (Item 5) the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; The determination means determines the target parameter by the second determination method when the amount of change over time in the width of the range in the depth direction exceeds a second threshold, and determines the target parameter by the first determination method when the amount of change over time in the width of the range in the depth direction is below the second threshold. 4. The information processing device according to item 3, (Item 6) The information processing device described in any one of items 3 to 5, characterized in that the second determination method determines a value corresponding to the average value of the extracted focus detection results or a value corresponding to the median value of the depth direction range as the target parameter. (Item 7) the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; the information processing device further includes a determination unit that determines whether or not a change over time exceeding a predetermined width has occurred in each of a first focus position corresponding to a nearest end of the range in the depth direction and a second focus position corresponding to a farthest end of the range in the depth direction, The determination means determines the target parameter by the first determination method when it is determined that no change over time exceeding the predetermined width has occurred in either the first focus position or the second focus position, and determines the target parameter by the second determination method when it is determined that a change over time exceeding the predetermined width has occurred in at least one of the first focus position and the second focus position. 4. The information processing device according to item 3, (Item 8) The second determination method includes: when the width of the change over time of the first focus position is larger than the width of the change over time of the second focus position, a value obtained by offsetting a value corresponding to the farthest end of the extracted focus detection results toward the near side is determined as the target parameter; When the width of the change over time of the second focus position is larger than the width of the change over time of the first focus position, a value obtained by offsetting a value corresponding to the nearest end of the extracted focus detection results toward the farther side is determined as the target parameter. 8. The information processing device according to item 7, (Item 9) the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; The determining means When the width of the range in the depth direction exceeds a first threshold, a value determined for the most recent acquisition timing is determined as the target parameter; When the width of the range in the depth direction is less than the first threshold value, a value corresponding to the nearest end of the extracted focus detection results is determined as the target parameter. 3. The information processing device according to item 2, (Item 10) the target parameter is a defocus amount, the focus detection results of the plurality of focus detection areas are defocus amounts of the respective focus detection areas, The inference means infers a range of defocus amounts as the range in the depth direction. 10. The information processing device according to any one of items 1 to 9, (Item 11) Further, a third acquisition means for acquiring an image signal obtained by capturing the imaging angle of view is provided, the first acquisition means acquires information about the distribution of the subject region based on the image signal acquired by the third acquisition means; The second acquisition means acquires focus detection results of the plurality of focus detection areas based on the image signals acquired by the third acquisition means. 11. The information processing device according to any one of items 1 to 10. (Item 12) Item 12. The information processing device according to item 11, wherein the information about the distribution of the subject area includes information about the position and size of the area of each predetermined part of the subject within the imaging angle of view. (Item 13) 13. The information processing device according to any one of items 1 to 12, wherein the inference means is an inference model that uses, as input, a distribution of areas corresponding to parts of an object of the same type as the subject in an imaging angle of view and a distribution of focus detection results related to each area, and machine-learns a depth range in which each part of the object actually exists. (Item 14) an imaging optical system including a focus lens; An imaging means; a control means for controlling the imaging optical system; An information processing device according to any one of items 1 to 13; and the control means drives the focus lens based on the target parameters determined by the determination means; The imaging means captures an image for recording on the condition that the focus lens has been driven based on the target parameters. An imaging device characterized by: (Item 15) A control method for an information processing device that determines target parameters used for focus adjustment when photographing a subject, comprising: a first acquisition step of acquiring information about the distribution of subject areas in an imaging angle of view; a second acquisition step of acquiring focus detection results of a plurality of focus detection areas provided for the imaging angle of view; an inference step of inferring a depth range in which the subject exists based on information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; determining the target parameters based on the inferred depth range; and In the determining step, a method for determining the target parameter is switched depending on the range in the depth direction. A control method comprising: (Item 16) A program for causing a computer to function as each of the means of the information processing device according to any one of items 1 to 13. [Explanation of symbols]
[0148] 10: Imaging device, 120: Camera body, 125: Camera MPU, 129: Phase difference AF unit, 130: Subject detection unit, 132: Inference unit, 133: Focus adjustment unit
Claims
1. An information processing device that determines target parameters to be used for focus adjustment when photographing a subject, a first acquisition means for acquiring information about a distribution of a subject area in an imaging angle of view; a second acquisition means for acquiring focus detection results of a plurality of focus detection areas provided for the imaging angle of view; an inference means for inferring a depth direction range in which the subject exists based on information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; determining means for determining the target parameters based on the inferred depth range; and The determining means switches a method for determining the target parameter depending on the range in the depth direction.
1. An information processing device comprising:
2. 2. The information processing device according to claim 1, wherein the determining means extracts focus detection results that fall within the inferred depth direction range from the focus detection results of the plurality of focus detection areas, and determines the target parameters based on the extracted focus detection results.
3. 3. The information processing apparatus according to claim 2, wherein the method for determining the target parameter includes a first determination method for determining a value corresponding to the nearest end of the extracted focus detection results as the target parameter, and a second determination method for determining a value not corresponding to the nearest end of the extracted focus detection results as the target parameter.
4. The information processing device according to claim 3, characterized in that the determination means determines the target parameter using the second determination method when the width of the depth-direction range exceeds a first threshold, and determines the target parameter using the first determination method when the width of the depth-direction range is below the first threshold.
5. the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; The determination means determines the target parameter by the second determination method when the amount of change over time in the width of the range in the depth direction exceeds a second threshold, and determines the target parameter by the first determination method when the amount of change over time in the width of the range in the depth direction is below the second threshold.
4. The information processing apparatus according to claim 3,
6. 4. The information processing apparatus according to claim 3, wherein the second determination method determines, as the target parameter, a value corresponding to an average value of the extracted focus detection results or a value corresponding to a median value of the range in the depth direction.
7. the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; the information processing device further includes a determination unit that determines whether or not a change over time exceeding a predetermined width has occurred in each of a first focus position corresponding to a nearest end of the range in the depth direction and a second focus position corresponding to a farthest end of the range in the depth direction; The determination means determines the target parameter by the first determination method when it is determined that no change over time exceeding the predetermined width has occurred in either the first focus position or the second focus position, and determines the target parameter by the second determination method when it is determined that a change over time exceeding the predetermined width has occurred in at least one of the first focus position and the second focus position.
4. The information processing apparatus according to claim 3,
8. The second determination method includes: when the width of the change over time of the first focus position is larger than the width of the change over time of the second focus position, a value obtained by offsetting a value corresponding to the farthest end of the extracted focus detection results toward the near side is determined as the target parameter; When the width of the change over time of the second focus position is larger than the width of the change over time of the first focus position, a value obtained by offsetting a value corresponding to the nearest end of the extracted focus detection results toward the farther side is determined as the target parameter.
8. The information processing apparatus according to claim 7,
9. the acquisition of information regarding the distribution of the subject area by the first acquisition means and the acquisition of focus detection results of the plurality of focus detection areas by the second acquisition means are performed intermittently; the inference means infers the depth direction range for each acquisition timing based on the intermittently acquired information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; The determining means When the width of the range in the depth direction exceeds a first threshold, a value determined for the most recent acquisition timing is determined as the target parameter; When the width of the range in the depth direction is less than the first threshold, a value corresponding to the nearest end of the extracted focus detection results is determined as the target parameter.
3. The information processing apparatus according to claim 2, wherein:
10. the target parameter is a defocus amount, the focus detection results of the plurality of focus detection areas are defocus amounts of the respective focus detection areas, The inference means infers a range of defocus amounts as the range in the depth direction.
2. The information processing apparatus according to claim 1, wherein:
11. a third acquisition means for acquiring an image signal obtained by capturing the imaging angle of view; the first acquisition means acquires information about the distribution of the subject region based on the image signal acquired by the third acquisition means; The second acquisition means acquires focus detection results of the plurality of focus detection areas based on the image signals acquired by the third acquisition means.
2. The information processing apparatus according to claim 1, wherein:
12. 12. The information processing apparatus according to claim 11, wherein the information about the distribution of the subject region includes information about the position and size of the region of each predetermined part of the subject within the imaging angle of view.
13. The information processing device according to claim 1, characterized in that the inference means is an inference model that uses as input the distribution of areas corresponding to parts of an object of the same type as the subject in the imaging angle of view and the distribution of focus detection results related to each area, and machine-learns the depth range in which each part of the object actually exists.
14. an imaging optical system including a focus lens; An imaging means; a control means for controlling the imaging optical system; An information processing device according to any one of claims 1 to 13; and the control means drives the focus lens based on the target parameters determined by the determination means; The imaging means captures an image for recording on the condition that the focus lens has been driven based on the target parameters. An imaging device characterized by:
15. A control method for an information processing device that determines target parameters used for focus adjustment when photographing a subject, comprising: a first acquisition step of acquiring information about a distribution of a subject area in an imaging angle of view; a second acquisition step of acquiring focus detection results of a plurality of focus detection areas provided for the imaging angle of view; an inference step of inferring a depth range in which the subject exists based on information about the distribution of the subject area and focus detection results of the plurality of focus detection areas; determining the target parameters based on the inferred depth range; and In the determining step, a method for determining the target parameter is switched depending on the range in the depth direction. A control method comprising:
16. A program for causing a computer to function as each of the means of the information processing apparatus according to any one of claims 1 to 13.
Citation Information
Patent Citations
Focus adjustment device and focus adjustment method
JP2022137760A