Electronic device, electronic device control method and program
The electronic device employs a pixel-divided image sensor and adaptive focus control to maintain accuracy by detecting phase differences based on occlusion direction, addressing focus accuracy issues when subjects are partially obscured.
Patent Information
- Application Number
- JP2024205743
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-10-27
- Estimated Expiration
- 2041-02-17
AI Technical Summary
Existing technologies face a decrease in focus adjustment accuracy when a subject is occluded, leading to perspective conflicts due to the distribution direction of the occluded area.
An electronic device with an image sensor having pixels divided into different pupil regions to detect phase differences, allowing focus adjustment based on the direction of occlusion, using a control mechanism that switches focus detection methods depending on the occlusion direction.
Prevents a decrease in focus adjustment accuracy by adapting focus detection to the direction of occlusion, thereby maintaining precise focus even when parts of the subject are obscured.
Smart Images

Figure 0007760685000001 
Figure 0007760685000002 
Figure 0007760685000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an electronic device, a control method for an electronic device, and a program. [Background technology]
[0002] A technology is used to detect a subject pattern (for example, a facial area of a person as a subject) from an image captured by an imaging device such as a camera. As a related technology, the technology disclosed in Patent Document 1 is proposed. The technology disclosed in Patent Document 1 performs face area detection, which detects a person's face from an image, and AF / AE / WB evaluation value detection for the same frame, thereby achieving highly accurate focus adjustment and exposure control for the person's face.
[0003] In recent years, deep learning neural networks have been used to detect subjects from images. CNN (Convolutional Neural Network) is used as a neural network suitable for image recognition, etc. For example, Non-Patent Document 1 proposes a technology (Single Shot Multibox Detector) that applies CNN to detect objects in an image. Also, Non-Patent Document 2 proposes a technology (Semantic Image Segmentation) that semantically divides regions in an image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-318554 [Non-patent literature]
[0005] [Non-Patent Document 1] “Liu,SSD:Single Shot Multibox Detector. In: ECCV2016” [Non-patent document 2] “Chen et.al, DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, arXiv, 2016” Summary of the Invention [Problem to be solved by the invention]
[0006] In recent years, a method has been used in which a pupil division function is used to detect phase differences and focus adjustment is performed using the detected phase difference information. When a subject is occluded by some object during shooting, an occluded area may appear in the subject area in the image. Perspective conflict may occur depending on the distribution direction of the occluded area in the subject area. When perspective conflict occurs, there is a problem in that the focus adjustment accuracy decreases. The technology of Patent Document 1 does not solve such a problem.
[0007] Therefore, an object of the present invention is to suppress the occurrence of a decrease in focus adjustment accuracy when an occluded area exists in a subject area. [Means for solving the problem]
[0008] In order to achieve the above object, the electronic device of the present invention comprises an image sensor having a plurality of pixels each having a plurality of photoelectric conversion units that receive light beams passing through different pupil regions of the photographing optical system, and that outputs a signal capable of detecting a first focus state of the photographing optical system from the phase difference between a pair of images in at least a first direction, and a second focus state of the photographing optical system from the phase difference between a pair of images in a second direction different from the first direction, and a control means that performs focus adjustment in accordance with a focus detection result based on the signal output from the image sensor, wherein, when there is an object blocking the subject, the control means performs focus adjustment based on the detection result of the first focus state when the blocking area of the object extends in the first direction, and performs focus adjustment based on the detection result of the second focus state when the blocking area of the object extends in the second direction. [Effects of the Invention]
[0009] According to the present invention, it is possible to suppress the occurrence of a decrease in focus adjustment accuracy when an occluded area exists in a subject area. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a digital camera. [Figure 2] 3A and 3B are diagrams illustrating an example of a pixel arrangement of an image sensor and a direction in which the pixels are pupil-divided; [Figure 3] FIG. 1 is a diagram illustrating an example of a CNN that infers the likelihood of an occluded region. [Figure 4] 5 is a flowchart showing an example of the flow of focus adjustment processing in the first embodiment. [Figure 5] 10A and 10B are diagrams illustrating an example of image data, a distribution of a subject area, and a blocking area. [Figure 6] FIG. 10 is a diagram illustrating an example in which the direction of pupil division of pixels is the Y direction. [Figure 7] 10 is a flowchart showing an example of the flow of focus adjustment processing in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited to the configurations described in the embodiments.
[0012] First Embodiment FIG. 1 is a diagram showing an example of a digital camera C as an electronic device. The digital camera C is a single-lens reflex camera with an interchangeable lens. The digital camera C may also be a camera that does not have an interchangeable lens. Furthermore, the electronic device is not limited to a digital camera and may be any device such as a smartphone or a tablet terminal.
[0013] Digital camera C has a lens unit 100, which is an imaging optical system, and a camera body 120. Lens unit 100 is detachably attached to camera body 120 via mount M (lens mount), which is indicated by a dotted line in the center of FIG. 1. Lens unit 100 has an optical system and a drive control system. The optical system includes a first lens group 101, an aperture 102, a second lens group 103, and a focus lens 104 (focus lens group). Lens unit 100 is a photographic lens that forms an optical image of a subject.
[0014] The first lens group 101 is disposed at the tip of the lens unit 100 and is held so as to be movable in the optical axis direction OA. The diaphragm 102 functions to adjust the amount of light during shooting and also functions as a mechanical shutter that controls exposure time when shooting still images. The diaphragm 102 and the second lens group 103 are movable together in the optical axis direction OA, and move in conjunction with the first lens group 101 to achieve a zoom function. The focus lens 104 is also movable in the optical axis direction OA, and the subject distance (focusing distance) at which the lens unit 100 focuses changes depending on the position. Focus adjustment, which adjusts the focusing distance of the lens unit 100, is performed by controlling the position of the focus lens 104 in the optical axis direction OA.
[0015] The drive control system includes a zoom actuator 111, an aperture actuator 112, and a focus actuator 113. The drive control system also includes a zoom drive circuit 114, an aperture shutter drive circuit 115, a focus drive circuit 116, a lens MPU (microprocessor) 117, and a lens memory 118.
[0016] The zoom drive circuit 114 drives the first lens group 101 and the second lens group 103 in the optical axis direction OA using the zoom actuator 111, thereby controlling the angle of view of the optical system of the lens unit 100. The iris shutter drive circuit 115 drives the iris 102 using the iris actuator 112, thereby controlling the opening diameter and opening / closing operation of the iris 102. The focus drive circuit 116 drives the focus lens 104 in the optical axis direction OA using the focus actuator 113, thereby changing the focal length of the optical system of the lens unit 100. The focus drive circuit 116 also detects the current position of the focus lens 104 using the focus actuator 113.
[0017] The lens MPU 117 controls the zoom drive circuit 114, the aperture shutter drive circuit 115, and the focus drive circuit 116 by performing various calculations and controls related to the lens unit 100. The lens MPU 117 is also connected to the camera MPU 125 via a mount M and communicates commands and data with the camera MPU 125. For example, the lens MPU 117 detects the position of the focus lens 104 and notifies the camera MPU 125 of lens position information in response to a request from the camera MPU 125. The lens position information includes information such as the position of the focus lens 104 in the optical axis direction OA, the position and diameter of the exit pupil in the optical axis direction OA when the optical system is not moving, and the position and diameter of the lens frame that limits the light beam from the exit pupil in the optical axis direction OA. The lens MPU 117 also controls the zoom drive circuit 114, the aperture shutter drive circuit 115, and the focus drive circuit 116 in response to a request from the camera MPU 125. The lens memory 118 pre-stores optical information necessary for autofocus detection. The camera MPU 125 controls the operation of the lens unit 100 by executing programs stored in, for example, an internal nonvolatile memory or the lens memory 118 .
[0018] Like the lens unit 100, the camera body 120 has an optical system and a drive control system. The optical system includes an optical low-pass filter 121 and an image sensor 122. The first lens group 101, aperture 102, second lens group 103, and focus lens 104 of the lens unit 100, and the optical low-pass filter 121 of the camera body 120 form an imaging optical system. The optical low-pass filter 121 is a filter that reduces false colors and moiré in captured images.
[0019] The image sensor 122 includes a CMOS image sensor and peripheral circuits. The image sensor 122 receives incident light from an imaging optical system. The image sensor 122 has m pixels arranged horizontally and n pixels arranged vertically (n and m are integers of 2 or greater). The image sensor 122 has a pupil division function and is capable of performing phase-difference AF (autofocus) using image data. The image processing circuit 124 generates data for phase-difference AF and image data for display and recording from the image data output by the image sensor 122.
[0020] The drive control system includes an image sensor drive circuit 123, an image processing circuit 124, a camera MPU 125, a display 126, a group of operation switches 127, a memory 128, an image sensor plane phase difference detection unit 129, a recognition unit 130, and a communication unit 131. The image sensor drive circuit 123 controls the operation of the image sensor 122 and also A / D converts acquired image signals and transmits them to the camera MPU 125. The image processing circuit 124 performs general image processing performed in digital cameras, such as gamma conversion, white balance adjustment, color interpolation, and compression encoding, on the image data acquired by the image sensor 122. The image processing circuit 124 also generates signals for phase difference AF.
[0021] The camera MPU 125 performs various calculations and controls related to the camera body 120. The camera MPU 125 controls the image sensor drive circuit 123, the image processing circuit 124, the display 126, the operation switch group 127, the memory 128, the image pickup surface phase difference detection unit 129, the recognition unit 130, and the communication unit 131. The camera MPU 125 is connected to the lens MPU 117 via a signal line of the mount M and communicates commands and data with the lens MPU 117. The camera MPU 125 makes various requests to the lens MPU 117. For example, the camera MPU 125 requests lens position information and optical information specific to the lens unit 100 from the lens MPU 117. The camera MPU 125 also requests the lens MPU 117 to drive the aperture, focus lens, zoom, etc. at predetermined drive amounts.
[0022] The camera MPU 125 incorporates a ROM 125a, a RAM 125b, and an EEPROM 125c. The ROM (Read Only Memory) 125a stores a program that controls the imaging operation. The RAM (Random Access Memory) 125b temporarily stores variables. The EEPROM (Electrically Erasable Programmable Read-Only Memory) 125c stores various parameters.
[0023] The display 126 is composed of an LCD (liquid crystal display) or the like, and displays information about the camera's shooting mode, a preview image before shooting, a confirmation image after shooting, an in-focus state display image when focus is detected, etc. The operation switch group 127 includes a power switch, a release (shooting trigger) switch, a zoom operation switch, a shooting mode selection switch, etc. The memory 128 is a removable flash memory that records shot images.
[0024] The image plane phase difference detection unit 129 performs focus detection processing by a phase difference detection method using focus detection data acquired from the image processing circuit 124. The image processing circuit 124 generates pairs of image data formed by light beams passing through two pairs of pupil regions as focus detection data. The image plane phase difference detection unit 129 then detects the amount of focus deviation based on the amount of deviation between each pair of generated image data. The image plane phase difference detection unit 129 performs phase difference AF (image plane phase difference AF) based on the output of the image sensor 122 without using a dedicated AF sensor. The image plane phase difference detection unit 129 may be implemented as part of the camera MPU 125, or may be implemented by a dedicated circuit, CPU, etc.
[0025] The recognition unit 130 performs object recognition based on image data acquired from the image processing circuit 124. For object recognition, the recognition unit 130 performs object detection, which detects the position of the target object in the image data, and region division, which divides the image into an object region and an occluded region where the object is occluded. In this embodiment, the recognition unit 130 determines the region of phase difference information to be used for focus adjustment based on the distribution direction of occlusion information and the detection direction of the amount of focus shift (image shift amount). Hereinafter, the detection direction of the amount of image shift may be referred to as the pupil division direction of the phase difference information.
[0026] The recognition unit 130 of this embodiment uses a CNN (convolutional neural network) to perform object detection and region segmentation. In this embodiment, a CNN that has undergone deep learning for object detection is used for object detection, and a CNN that has undergone deep learning for region segmentation is used for region segmentation. However, the recognition unit 130 may also use a CNN that has undergone deep learning for both object detection and region segmentation.
[0027] The recognition unit 130 acquires an image from the camera MPU 125 and inputs it to a CNN that has undergone deep learning for object detection. The object is detected as an output result of the CNN's inference process. The recognition unit 130 also inputs the image of the recognized object to a CNN that has undergone deep learning for region segmentation. The CNN's inference process detects an occluded area in the image of the object as an output result. The recognition unit 130 may be realized by the camera MPU 125, or may be realized by a dedicated circuit, CPU, or the like. In addition, since the recognition unit 130 performs inference processing using CNN, it is preferable that it has a built-in GPU that is used for the calculation processing of the inference processing.
[0028] This section explains CNN deep learning. CNN deep learning can be performed using any method. This section explains CNN deep learning for object detection. For example, CNN deep learning is achieved by supervised learning, in which an image containing a subject as the correct answer is used as training data, and many learning images are used as input data. In this case, methods such as backpropagation are applied to CNN deep learning.
[0029] The deep learning of CNN may be performed by a predetermined computer such as a server. In this case, the communication unit 131 of the camera body 120 may communicate with the predetermined computer to acquire the deep-learned CNN from the predetermined computer. The camera MPU 125 then sets the CNN acquired by the communication unit 131 in the recognition unit 130. This allows the recognition unit 130 to perform subject detection using the deep-learned CNN. If the digital camera C has a built-in high-performance CPU or GPU suitable for deep learning, or a dedicated processor specialized for deep learning, the digital camera C may perform the deep learning of CNN. However, because deep learning of CNN requires abundant hardware resources, it is preferable that an external device (a predetermined computer) performs the deep learning of CNN and the digital camera C acquires and uses the deep-learned CNN.
[0030] Object detection may be achieved using any method other than CNN. For example, object detection may be achieved using a rule-based method. Furthermore, object detection may use a trained model that has been machine-learned using any method other than deep learning CNN. For example, object detection may be achieved using a trained model that has been machine-learned using any machine learning algorithm such as a support vector machine or logistic regression.
[0031] Next, the operation of the imaging surface phase difference detection unit 129 will be described. FIG. 2 is a diagram showing an example of the pixel arrangement of the image sensor 122 and the direction of pupil division of the pixels. FIG. 2(A) shows a state in which a range of six rows in the vertical direction (Y direction) and eight columns in the horizontal direction (X direction) of a two-dimensional C-MOS area sensor is observed from the lens unit 100 side. The image sensor 122 is provided with color filters in a Bayer array, and green (G) and red (R) color filters are alternately arranged from left to right in the odd-numbered rows of pixels, and blue (B) and green (G) color filters are alternately arranged from left to right in the even-numbered rows of pixels. In the pixel 211, a plurality of photoelectric conversion units are arranged inside the on-chip microlens (microlens 211i) indicated by a circle. In the example of FIG. 2(B), four photoelectric conversion units 211a, 211b, 211c, and 211d are arranged inside the on-chip microlens of the pixel 211.
[0032] In the pixel 211, the photoelectric conversion units 211a, 211b, 211c, and 211d are divided into two in the X and Y directions, and it is possible to read out the photoelectric conversion signals of the individual photoelectric conversion units, and it is also possible to read out the sum of the photoelectric conversion signals of each photoelectric conversion unit independently. The photoelectric conversion signals of the individual photoelectric conversion units are used as data for phase difference AF. The photoelectric conversion signals of the individual photoelectric conversion units may also be used to generate parallax images that form 3D (3-dimensional) images. The sum of the photoelectric conversion signals is used when generating normal captured image data.
[0033] The pixel signals when performing phase-difference AF will be described. In this embodiment, the light beam emitted from the imaging optical system is pupil-divided by photoelectric conversion units 211a, 211b, 211c, and 211d via microlens 211i in FIG. 2A. The areas indicated by two dotted lines in FIG. 2B represent photoelectric conversion units 212A and 212B, respectively. Photoelectric conversion unit 211A is composed of photoelectric conversion units 211a and 211c, and photoelectric conversion unit 211B is composed of photoelectric conversion units 211a and 211d. To perform focus detection based on the amount of image shift (phase difference), a sum signal obtained by adding together the signals output by photoelectric conversion units 211a and 211c and a sum signal obtained by adding together the signals output by photoelectric conversion units 211a and 211d are used as a pair. This enables focus detection based on the amount of image shift (phase difference) in the X direction.
[0034] Here, we will focus on phase-difference AF using focus detection based on the amount of image shift in the X direction. For multiple pixels 211 within a predetermined range arranged in the same pixel row, an image formed by combining the sum signals of photoelectric conversion units 211a and 211c belonging to photoelectric conversion unit 212A is called an AF image A. An image formed by combining the sum signals of photoelectric conversion units 212b and 211d belonging to photoelectric conversion unit 212B is called an AF image B. The outputs of photoelectric conversion units 212A and 212B are pseudo-luminance (Y) signals calculated by adding the outputs of green, red, blue, and green included in the unit array of color filters. However, an AF image A and an AF image B may be formed for each color, red, blue, and green. By detecting the relative image shift amount between the pair of image signals of the AF image A and AF image B generated in this manner through correlation calculation, a prediction [bit], which represents the degree of correlation between the pair of image signals, can be detected. The camera MPU 125 can detect the defocus amount [mm] of a predetermined area by multiplying the prediction by a conversion coefficient. The sum of the output signals of the photoelectric conversion units 212A and 212B generally forms one pixel (output pixel) of the output image.
[0035] The subject detection and occlusion information area division by the recognition unit 130 will be described. In this embodiment, the recognition unit 130 detects (recognizes) a person's face from an image. Any method (for example, the method in Non-Patent Document 1) can be applied as a method for detecting a person's face from an image. However, the recognition unit 130 may detect a subject that can be a target for focus adjustment, such as a person's entire body, an animal, or a vehicle, instead of a person's face.
[0036] The recognition unit 130 performs region segmentation for the detected subject region into occluded regions. Any method (for example, the method of Non-Patent Document 2) can be applied as the region segmentation method. The recognition unit 130 uses a deep-learned CNN to infer the likelihood of an occluded region for each pixel region. However, as described above, the recognition unit 130 may infer the likelihood of an occluded region using a trained model machine-learned by any machine learning algorithm, or may determine the likelihood of an occluded region based on a rule base. When a CNN is used to infer the likelihood of an occluded region, the CNN performs deep learning using occluded regions as positive examples and regions other than occluded regions as negative examples. As a result, the CNN outputs the likelihood of an occluded region for each pixel region as an inference result.
[0037] FIG. 3 is a diagram showing an example of a CNN that infers the likelihood of an occluded region. FIG. 3(A) shows an example of a subject region of an input image input to the CNN. Subject region 301 is detected from the image by the subject detection described above. Subject region 301 includes face region 302, which is the detection target of subject detection. Face region 302 in FIG. 3(A) includes two occluded regions (occluded regions 303 and 304). Occluded region 303 is a region with no depth difference, and occluded region 304 is a region with a depth difference. Occluded regions are also called occlusions, and the distribution of occluded regions is also called occlusion distribution.
[0038] FIG. 3(B) shows an example of the definition of occlusion information. Each of images 1 to 3 in FIG. 3(B) is divided into a white region and a black region, with the white region indicating a positive example and the black region indicating a negative example. The occlusion information obtained by dividing the image of the subject region in FIG. 3(B) is all images assumed to be candidates for training data used when performing deep learning of CNN. Below, we will explain which of the occlusion information in FIG. 3(B) is used as training data in this embodiment.
[0039] Image 1 in Figure 3(B) shows an example of occlusion information when the image is divided into a subject region (face region) and a non-subject region, with the subject region being a positive example and the non-subject region being a negative example. Image 2 in Figure 3(B) shows an example of occlusion information when the image is divided into a foreground occluded region relative to the subject and the other region, with the foreground occluded region being a positive example and the non-occluded region relative to the subject being a negative example. Image 3 in Figure 3(C) shows an example of occlusion information when the image is divided into a occluded region that causes perspective conflict relative to the subject and the other region, with the occluded region that causes perspective conflict being a positive example and the non-occluded region that causes perspective conflict being a negative example.
[0040] In the example of Figure 2(B), photoelectric conversion units 211a, 211b, 211c, and 211d are pupil-divided into two in the X direction. In this case, if occluded areas in the image are distributed in the Y direction, perspective conflicts are likely to occur. Therefore, training data (training images) for occluded areas where perspective conflicts are likely to occur can be generated based on the pupil division direction of multiple photoelectric conversion units in a pixel.
[0041] As shown in image 1 in FIG. 3B, a person's face in an image has a distinctive visibility pattern with small variance, allowing for highly accurate region segmentation. For example, the occlusion information shown in image 1 in FIG. 3B is suitable as training data for the learning process when generating a CNN that detects a person as a subject. In terms of detection accuracy, the occlusion information shown in image 1 in FIG. 3B is more suitable than the occlusion information shown in image 3 in FIG. 3B. However, an image like image 3 in FIG. 3B is suitable as training data for the learning process when generating a CNN that detects an occluded area that causes perspective conflict. Parallax images of images A and B used in focus detection may be used as training data for the learning process when generating a CNN that detects an occluded area that causes perspective conflict. Furthermore, the occlusion information is not limited to the above example and may be generated based on any method for segmenting an occluded area and a non-occluded area.
[0042] FIG. 3(C) shows the flow of deep learning of CNN. In this embodiment, an RGB image is used as the learning input image 310. Furthermore, as the teacher image, a teacher image 314 (teacher image of occlusion information) as shown in FIG. 3(C) is used. The teacher image 314 is an image of occlusion information that causes perspective conflict in FIG. 3(B).
[0043] An input image 310 for training is input to a neural network system 311 (CNN). The neural network system 311 can employ, for example, a layer structure in which convolutional layers and pooling layers are alternately stacked between an input layer and an output layer, and a multi-layer structure in which a fully connected layer is connected downstream of the layer structure. A score map indicating the likelihood of an occluded region in the input image is output from the output layer 312 in FIG. 3(C). The score map is output in the form of an output result 313.
[0044] In CNN deep learning, the error between the output result 313 and the training image 314 is calculated as a loss value 315. The loss value 315 is calculated using, for example, methods such as cross entropy or squared error. Coefficient parameters such as the weights and biases of each node of the neural network system 311 are then adjusted so that the loss value 315 gradually decreases. By performing sufficient CNN deep learning using many training input images 310, the neural network system 311 will output a more accurate output result 313 when an unknown input image is input. In other words, when an unknown input image is input, the neural network system 311 (CNN) will output occlusion information obtained by dividing the input image into occluded and non-occluded regions as the output result 313 with high accuracy. Note that creating training data that identifies occluded regions (areas of overlapping objects) requires a lot of work. For this reason, it is possible to create training data using computer graphics or image synthesis, in which object images are cut out and superimposed.
[0045] As described above, an example was described in which image 3 in FIG. 3B is used as the teacher image 314, in which an area with a depth difference (a foreground area of the subject where the depth difference is equal to or greater than a predetermined value) is designated as the occluded area. Here, an image such as image 2 in FIG. 3B, in which an area with no depth difference (a foreground area of the subject where the depth difference is less than a predetermined value) is designated as the occluded area, may also be used as the teacher image 314. Even if an image such as image 2 in FIG. 3B is used as the teacher image 314, the CNN can infer an area that causes perspective conflict when an unknown input image is input to the CNN. However, the accuracy of inferring an occluded area that causes perspective conflict by the CNN is improved when an image such as image 3 in FIG. 3B, in which an area with a depth difference is designated as the occluded area, is used as the teacher image 314.
[0046] Any method other than CNN can be applied to detect occluded areas. For example, occluded area detection may be realized by a rule-based method. Furthermore, occluded area detection may use a trained model that has been machine-learned by any method other than deep learning CNN. For example, occluded areas may be detected using a trained model that has been machine-learned by any machine learning algorithm such as a support vector machine or logistic regression. This is similar to subject detection.
[0047] Next, the focus adjustment process will be described. Fig. 4 is a flowchart showing an example of the flow of the focus adjustment process in the first embodiment. When performing focus adjustment, the camera MPU 125 determines a reference area for focus adjustment based on the pupil division direction of the phase difference in the imaging surface phase difference detection unit 129 and the direction of the occluded area determined by the recognition unit 130.
[0048] In S401, the recognition unit 130 detects a subject from an image acquired from the image processing circuit 124 via the camera MPU 125, for example, using a CNN that performs subject detection. The recognition unit 130 may detect a subject from an image using a method other than a CNN. In S402, the recognition unit 130 detects an occluded area within the area of the subject detected from the image. At this time, the recognition unit 130 inputs the image acquired from the image processing circuit 124 as an input image to the CNN described in FIG. 3. If the CNN is sufficiently trained, the image output as an output result from the CNN is an image in which occluded areas and areas other than occluded areas can be distinguished. The recognition unit 130 detects the occluded area from the image output from the CNN.
[0049] In S403, the recognition unit 130 determines the distribution direction of the detected occluded areas. For example, the recognition unit 130 may compare the edge integral value in the X direction with the edge integral value in the Y direction and determine the direction with the smaller integral value as the distribution direction of the occluded areas. Figure 5 is a diagram showing an example of the distribution of image data, subject areas, and occluded areas.
[0050] FIG. 5(A) is a diagram showing an example of a case where there is an obstruction (e.g., a rod) extending in the X direction between digital camera C and a subject. An obstruction 502 is reflected in image 500 (image data) acquired from image processing circuit 124 so as to obstruct subject 501. Recognition unit 130 detects an image of subject area 510 from image 500. Subject area 510 includes subject 511 and obstruction 512. Obstruction 512 obstructs subject 511 in the X direction.
[0051] When an image 500 including a subject region 510 is input to the CNN, the recognition unit 130 outputs an output result 520 as an inference result for the subject region 510. The output result 520 includes an occluded region 522 and a region 521 other than the occluded region. The occluded region 522 has a distribution corresponding to the occluding object 512. In the example of FIG. 5(A), the occluded region 522 is distributed in the X direction. The recognition unit 130 compares the edge integral value in the X direction and the edge integral value in the Y direction in the output result 520, and determines that the X direction with the smaller integral value is the distribution direction of the occluded region.
[0052] 5(B) is a diagram showing an example of a case where an obstruction extending in the Y direction is present between digital camera C and a subject. Obstruction 552 is reflected in image 550 acquired from image processing circuit 124 so as to obstruct subject 551. Recognition unit 130 detects subject area 560 from image 550. Subject area 560 includes subject 561 and obstruction 562. Obstruction 562 obstructs subject 511 in the Y direction.
[0053] When the recognition unit 130 inputs an image 550 including a subject region 560 into the CNN, it outputs an output result 570 as an inference result for the subject region 560. The output result 570 includes an occluded region 572 and a region 571 other than the occluded region. The recognition unit 130 determines the distribution direction of the occluded region using a method similar to that shown in FIG. 3(A). In the example of FIG. 3(B), the recognition unit 130 determines that the distribution direction of the occluded region is the Y direction. As described above, the recognition unit 130 can determine the distribution direction of the occluded region.
[0054] Returning to FIG. 4 , the processing from S404 onward will be described. In S404, the camera MPU 125 determines whether the distribution direction of the masked areas is the X direction or the Y direction based on the determination result of S403. If the camera MPU 125 determines in S404 that the distribution direction of the masked areas is the X direction, the flow proceeds to S405. On the other hand, if the camera MPU 125 determines in S404 that the distribution direction of the masked areas is the Y direction, the flow proceeds to S406. In S404, the camera MPU 125 may determine that the distribution direction of the masked areas is the X direction even if the distribution direction of the masked areas does not completely coincide with the X direction, as long as it is within a predetermined angle range with the X direction as the reference. Similarly, in S404, the camera MPU 125 may determine that the distribution direction of the masked areas is the Y direction even if the distribution direction of the masked areas does not completely coincide with the Y direction, as long as it is within a predetermined angle range with the Y direction as the reference.
[0055] In S405, the camera MPU 125 controls focus adjustment based on the image shift amount for the area where the object is detected (object area). If the distribution direction of the occluded area is the X direction, there is a low possibility that perspective conflict will occur in the image shift amount (phase difference) in the X direction. In this case, the image plane phase difference detection unit 129 performs focus detection based on the image shift amount (phase difference) in the X direction, with reference to the entire object area. Then, the camera MPU 125 controls focus adjustment based on the image shift amount detected by the image plane phase difference detection unit 129. Note that the image plane phase difference detection unit 129 may perform focus detection based on the image shift amount (phase difference) in the X direction, excluding the occluded area from the object area. After the processing of S405 is performed, the flowchart in FIG. 4 ends.
[0056] In S406, the camera MPU 125 controls focus adjustment based on the amount of image shift for the area remaining after excluding the occluded areas from the subject area. If the distribution direction of the occluded areas is the Y direction, there is a high possibility that perspective conflict will occur in the amount of image shift (phase difference) in the X direction. In this case, the imaging surface phase difference detection unit 129 performs focus detection based on the amount of image shift (phase difference) in the X direction, with reference to the area remaining after excluding the occluded areas from the subject area.
[0057] If there are multiple correlation calculation blocks in the subject area and a block that does not include an occluded area exists, the imaging surface phase difference detection unit 129 performs focus detection based on the image shift amount of the correlation calculation. On the other hand, if there are multiple correlation calculation blocks in the subject area and no block that does not include an occluded area exists, the imaging surface phase difference detection unit 129 shifts the calculation position of the correlation calculation so that it does not include an occluded area and calculates the image shift amount. Then, the imaging surface phase difference detection unit 129 performs focus detection based on the calculated image shift amount. After the processing of S406 is performed, the flowchart of FIG. 5 ends.
[0058] As described above, according to the first embodiment, the distribution direction of occluded areas in the subject area of a captured image is detected, and the area used for focus detection is controlled according to the detected distribution direction of the occluded areas and the direction of pupil division of each pixel of the image sensor. This prevents a decrease in focus adjustment accuracy due to perspective conflict, even when an occluded area is present in the subject area. In the first embodiment, as shown in FIG. 2(B), an example is shown in which the pupil division direction of each pixel of the image sensor 122 is the X direction, but the first embodiment can also be applied when it is the Y direction.
[0059] Second Embodiment Next, a second embodiment will be described. In the second embodiment, the digital camera C can calculate both the amount of image shift in the horizontal direction (X direction) and the amount of image shift in the vertical direction (Y direction). Therefore, the digital camera C can switch the direction of pupil division between the X direction and the Y direction. The configuration of the second embodiment is the same as the configuration of FIG. 1 described in the first embodiment, so a description thereof will be omitted.
[0060] Calculation of the image shift amount in the X direction is the same as in the first embodiment. Calculation of the image shift amount in the Y direction will now be described. FIG. 6 is a diagram showing an example in which the pupil division direction of the pixel 211 is the Y direction. As in FIG. 2(B), in the pixel 211, the photoelectric conversion units 211a, 211b, 211c, and 211d are divided into two in the X direction and the Y direction. Of the photoelectric conversion units 211a, 211b, 211c, and 211d, the photoelectric conversion unit 212C is made up of the photoelectric conversion units 211a and 211b. Of the photoelectric conversion units 211a, 211b, 211c, and 211d, the photoelectric conversion unit 212D is made up of the photoelectric conversion units 211c and 211d.
[0061] In the second embodiment, a sum signal obtained by adding together the output signals of photoelectric conversion units 211a and 211b belonging to photoelectric conversion unit 212C and a sum signal obtained by adding together the output signals of photoelectric conversion units 211c and 211d belonging to photoelectric conversion unit 212D are used as a pair. This makes it possible to perform focus detection based on the image shift amount (phase difference) in the Y direction. The correlation calculation is the same as in the first embodiment, except that the paired direction is the Y direction instead of the X direction.
[0062] Furthermore, in the second embodiment, the direction of pupil division of pixel 211 can be switched between the X direction and the Y direction. For example, under the control of camera MPU 125, image sensor drive circuit 123 may switch between reading output signals from photoelectric conversion units 212A and 212B and reading output signals from photoelectric conversion units 212C and 212D in the second embodiment. This allows the direction of pupil division of pixel 211 to be switched between the X direction and the Y direction.
[0063] FIG. 7 is a flowchart showing an example of the flow of focus adjustment processing in the second embodiment. Steps S701 to S703 are the same as steps S401 to S403 in FIG. 4, and therefore a description thereof will be omitted. However, in deep learning for CNN, both a teacher image in which occluded regions are distributed in the X direction and a teacher image in which occluded regions are distributed in the Y direction are used. This makes it possible to use an unknown input image as input and to infer the likelihood of the existence of occluded regions distributed in the X and Y directions as an output result using CNN. Deep learning for CNN is the same as in the first embodiment.
[0064] In S704, the camera MPU 125 determines whether the distribution direction of the masked areas is the X direction or the Y direction based on the determination result of S803. If the camera MPU 125 determines in S704 that the distribution direction of the masked areas is the X direction, the flow proceeds to S705. On the other hand, if the camera MPU 125 determines in S704 that the distribution direction of the masked areas is the Y direction, the flow proceeds to S706.
[0065] In S705, the camera MPU 125 switches the direction of pupil division and controls focus adjustment based on the image shift amount in the X direction. When the distribution direction of the occluded area is the X direction, there is a low possibility that perspective conflict will occur in the image shift amount (phase difference) in the X direction. In this case, the image plane phase difference detection unit 129 performs focus detection based on the image shift amount (phase difference) in the X direction. In S705, the image plane phase difference detection unit 129 does not perform focus detection based on the image shift amount (phase difference) in the Y direction. Then, the camera MPU 125 controls focus adjustment based on the image shift amount detected by the image plane phase difference detection unit 129. After the processing of S705 is performed, the flowchart in FIG. 5 ends.
[0066] In S706, the camera MPU 125 switches the direction of pupil division and controls focus adjustment based on the image shift amount in the Y direction. When the distribution direction of the occluded area is the Y direction, there is a low possibility that perspective conflict will occur in the image shift amount (phase difference) in the Y direction. In this case, the image plane phase difference detection unit 129 performs focus detection based on the image shift amount (phase difference) in the Y direction. In S706, the image plane phase difference detection unit 129 does not perform focus detection based on the image shift amount (phase difference) in the X direction. Then, the camera MPU 125 controls focus adjustment based on the image shift amount detected by the image plane phase difference detection unit 129. After the processing of S406 is performed, the flowchart in FIG. 5 ends.
[0067] As described above, according to the second embodiment, in a configuration in which the direction of pupil division of pixel 211 can be switched between the X direction and the Y direction, the direction of pupil division is changed according to the distribution direction of occluded areas, and focus adjustment is controlled based on phase difference information. This prevents a decrease in focus adjustment accuracy due to perspective conflict when an occluded area is present in the subject area, regardless of whether the distribution direction of occluded areas is the X direction or the Y direction.
[0068] The camera MPU 125 may preferentially select a direction based on a direction of contrast (a direction of high contrast) in the subject area excluding the occluded area as the direction of phase difference information when controlling focus adjustment. In this case, the camera MPU 125 performs focus adjustment based on the amount of image shift (phase difference) in the direction of contrast in the subject area excluding the occluded area, rather than the direction of pupil division and the distribution direction of the occluded area. If there is no contrast in the subject area, even if perspective conflict can be avoided, the target subject cannot be detected and focus adjustment becomes impossible. For this reason, the camera MPU 125 performs focus adjustment based on the amount of image shift in the direction of contrast in the subject area excluding the occluded area. This is the same as in the first embodiment.
[0069] Furthermore, if the contrast of the occluded area is low (if the value indicating the contrast is lower than a predetermined threshold), camera MPU 125 does not need to execute the processes of S705 and S706 according to the determination result of S704. This is because even if an occluded area occurs in the subject area due to an obstruction when digital camera C captures an image, the influence of perspective conflict is low at the obstruction and at the boundary between the obstruction and the subject. This allows the processes of S705 and S706 according to the determination result of S704 to be omitted.
[0070] <Modification> In the above-described embodiments, examples have been described in which machine learning is performed by supervised learning using a teacher image in which occluded areas are used as positive examples and areas other than occluded areas are used as negative examples, and the distribution direction of occluded areas is detected from an image. In this regard, in each embodiment, a trained model trained by unsupervised learning may be used. In this case, for example, training images in which occluded areas occluding a subject are distributed in the X direction and training images in which occluded areas occluding a subject are distributed in the Y direction are used for machine learning. The trained model used in each embodiment may be generated by unsupervised machine learning using the training images. The generated trained model is then used to detect occluded areas.
[0071] For example, when an image in which occluding areas that occlude subjects in a subject region are distributed in the X direction is input to the trained model, the occluded areas distributed in the X direction are extracted from the image as features, and the input image is classified as an image in which occluded areas are distributed in the X direction. Similarly, when an image in which occluding areas that occlude subjects in a subject region are distributed in the Y direction is input to the trained model, the occluded areas distributed in the Y direction are extracted from the image as features, and the input image is classified as an image in which occluded areas are distributed in the Y direction. In this way, the distribution direction of occluded areas can be determined.
[0072] By using a trained model that has been machine-learned through unsupervised learning as a trained model for detecting occluded areas, there is no need to prepare training data (trained images). For example, clustering and principal component analysis can be applied as machine learning algorithms for unsupervised learning.
[0073] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the gist of the present invention. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. [Explanation of symbols]
[0074] 120 Camera body 122 Image sensor 125 Camera MPU 129 Imaging surface phase difference detection unit 130 Recognition part 211 pixels 211a~d Photoelectric conversion unit 311 Neural Network System C. Digital camera
Claims
1. an image sensor having a plurality of pixels each having a plurality of photoelectric conversion units that receive light beams passing through different pupil regions of an image pickup optical system, and that outputs signals capable of detecting a first focus state of the image pickup optical system from a phase difference between a pair of images in at least a first direction, and a second focus state of the image pickup optical system from a phase difference between a pair of images in a second direction different from the first direction; a control means for performing focus adjustment in accordance with a focus detection result based on a signal output from the image sensor; The electronic device is characterized in that, when there is an object blocking the subject, the control means performs focus adjustment based on the detection result of the first focus state when the blocking area of the object extends in a first direction, and performs focus adjustment based on the detection result of the second focus state when the blocking area of the object extends in a second direction.
2. The electronic device of claim 1 , wherein the first direction is a horizontal direction and the second direction is a vertical direction.
3. 3. The electronic device according to claim 1, wherein focus adjustment control based on phase difference information in a direction of contrast of the subject area excluding the occluded area is performed with priority over focus adjustment control based on a direction in which the occluded area extends.
4. The electronic device according to claim 1 , wherein when a value indicating a contrast of the shielded area is lower than a predetermined threshold, the focus adjustment control based on the direction in which the shielded area extends is not performed.
5. 5. The electronic device according to claim 1, wherein an occluded area is detected from an image based on the likelihood of the occluded area obtained by inputting an image based on the output from the imaging element into a trained model that has been machine-learned using a plurality of training images in which an occluded area in a subject area is used as a positive example and an area other than the occluded area is used as a negative example.
6. The electronic device according to claim 5 , wherein the occluded region of the teacher image is a foreground region present in a subject region of the teacher image.
7. The electronic device according to claim 5 , wherein the occluded region of the teacher image is a region where the depth difference is equal to or greater than a predetermined value.
8. 5. The electronic device according to claim 1, wherein an image based on the output from the imaging element is input to a trained model that has been machine-learned through unsupervised learning using a plurality of training images in which occluded areas of the subject area are distributed horizontally and a plurality of training images in which occluded areas are distributed vertically, and occluded areas are detected from the image.
9. an image sensor including a plurality of pixels capable of photoelectrically converting light beams that have passed through different pupil regions of an imaging optical system and outputting a pair of signals; a focus detection means capable of detecting a first focus state of the photographing optical system from a phase difference between a pair of images in at least a first direction, and a second focus state of the photographing optical system from a phase difference between a pair of images in a second direction different from the first direction, based on an output of the image sensor; a control means for performing focus adjustment in accordance with a focus detection result based on a signal output from the image sensor; The electronic device is characterized in that, when there is an object blocking the subject, the control means performs focus adjustment based on the detection result of the first focus state when the blocking area of the object extends in a first direction, and performs focus adjustment based on the detection result of the second focus state when the blocking area of the object extends in a second direction.
10. The electronic device of claim 9 , wherein the first direction is a horizontal direction and the second direction is a vertical direction.
11. 11. The electronic device according to claim 9, wherein focus adjustment control based on phase difference information in a direction of contrast of the subject area excluding the shielded area is performed with priority over focus adjustment control based on a direction in which the shielded area extends.
12. The electronic device according to claim 9 , wherein when a value indicating a contrast of the shielded area is lower than a predetermined threshold, the focus adjustment control based on the direction in which the shielded area extends is not performed.
13. 13. The electronic device according to claim 9, wherein an occluded area is detected from the image based on the likelihood of the occluded area obtained by inputting an image based on the output from the imaging element into a trained model that has been machine-learned using a plurality of teacher images in which an occluded area in a subject area is used as a positive example and an area other than the occluded area is used as a negative example.
14. The electronic device according to claim 13 , wherein the occluded region of the teacher image is a foreground region present in a subject region of the teacher image.
15. The electronic device according to claim 13 , wherein the occluded region of the teacher image is a region where the depth difference is equal to or greater than a predetermined value.
16. 13. The electronic device according to claim 9, wherein an image based on the output from the imaging element is input to a trained model that has been machine-learned by unsupervised learning using a plurality of training images in which occluded areas of the subject area are distributed horizontally and a plurality of training images in which occluded areas are distributed vertically, and the occluded areas are detected from the image.
17. A control method for an electronic device capable of detecting a first focus state of an imaging optical system from a phase difference between a pair of images in at least a first direction, and a second focus state of the imaging optical system from a phase difference between a pair of images in a second direction different from the first direction, based on an output of an image sensor having a plurality of pixels, each pixel having a plurality of photoelectric conversion units that receive light beams passing through different pupil regions of the imaging optical system, the method comprising: a control step of performing focus adjustment in accordance with a focus detection result based on a signal output from the image sensor; A control method for an electronic device, characterized in that in the control step, when there is an object blocking the subject, focus adjustment is performed based on the detection result of the first focus state when the blocking area of the object extends in a first direction, and focus adjustment is performed based on the detection result of the second focus state when the blocking area of the object extends in a second direction.
18. A control method for an electronic device having an image sensor including a plurality of pixels capable of photoelectrically converting light beams that have passed through different pupil regions of an imaging optical system and outputting a pair of signals, comprising: a focus detection step capable of detecting a first focus state of the photographing optical system from a phase difference between a pair of images in at least a first direction and a second focus state of the photographing optical system from a phase difference between a pair of images in a second direction different from the first direction, based on an output of the image sensor; a control step of performing focus adjustment in accordance with a focus detection result based on a signal output from the image sensor; A control method for an electronic device, characterized in that in the control step, when there is an object blocking the subject, focus adjustment is performed based on the detection result of the first focus state when the blocking area of the object extends in a first direction, and focus adjustment is performed based on the detection result of the second focus state when the blocking area of the object extends in a second direction.
19. A program for causing a computer to execute each means of the electronic device according to any one of claims 1 to 16.
Citation Information
Patent Citations
Imaging device, control method thereof, program, and storage medium
JP2005318554A
Imaging apparatus
JP2013205675A
Stereoscopic imaging device, stereoscopic imaging system, control method of stereoscopic imaging device, program, and storage medium
JP2015102602A
Imaging device and image display method, program, and program storage medium
JP2016148732A
Image sensor
JP2017184181A