Imaging apparatus and method for controlling the same, program, and storage medium
Patent Information
- Application Number
- JP2023002558
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-11
- Publication Date
- 2026-01-14
AI Technical Summary
Existing imaging devices struggle to accurately determine the main subject when multiple subjects are present, often focusing on the wrong subject and experiencing sudden changes in focal position when the main subject is replaced.
An imaging apparatus that includes subject detection, posture detection, and focus detection means, with a threshold setting mechanism to determine the main subject based on posture and focus state, and a focus adjustment mechanism to maintain focus on the selected subject.
Enables accurate selection and continuous focus on the intended main subject, even when multiple subjects are present, by using posture and focus state analysis to set appropriate threshold values for focus adjustment.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an imaging device. [Background technology]
[0002] In continuous shooting, in which multiple shots are taken in succession, when shooting multiple moving subjects detected by an imaging device, it is necessary to determine a main subject from the multiple subjects and keep the focus on the main subject. Patent Document 1 discloses a method for determining the main subject by detecting two different types of subjects, determining the main subject, and adjusting the focus. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2018-66889 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, as disclosed in Patent Document 1, when determining the main subject from the positional relationship of two different types of subjects, a subject other than the one the photographer intended may be determined to be the main subject and may continue to be focused on. Also, the focal position may suddenly change when the main subject is replaced.
[0005] The present invention has been made in consideration of the above-mentioned problems, and its object is to provide an imaging device that can appropriately select a subject on which to focus when multiple subjects are present. [Means for solving the problem]
[0006] The imaging device according to the present invention is characterized by comprising a subject detection means for detecting a subject, an attitude detection means for detecting the attitude of the subject, a focus detection means for detecting a focus state of the subject, and a setting means for setting a threshold value for determining whether or not to select the subject as a main subject based on the attitude and the focus state of the subject. Effect of the Invention
[0007] According to the present invention, when there are multiple subjects, it is possible to appropriately select the subject on which to focus. [Brief description of the drawings]
[0008] [Figure 1] 1 is a block diagram showing the configuration of a camera according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a pixel array in the camera according to the embodiment. [Diagram 3] 3A and 3B are a plan view and a cross-sectional view of a pixel in the embodiment. [Figure 4] FIG. 2 is an explanatory diagram of a pixel structure according to the embodiment. [Diagram 5] FIG. 4 is an explanatory diagram of pupil division according to the embodiment. [Figure 6] 5A and 5B are diagrams showing the relationship between the defocus amount and the image shift amount in the embodiment. [Figure 7] FIG. 4 is a view showing a focus detection area in the embodiment. [Figure 8] 5 is a flowchart for explaining AF and image capture processing in the embodiment. [Figure 9] 5 is a flowchart for explaining a shooting process in the embodiment. [Figure 10] 5 is a flowchart for explaining subject tracking AF processing in an embodiment. [Figure 11] FIG. 1 is an explanatory diagram of posture information according to an embodiment; [Figure 12A] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 12B] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 12C]5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 13A] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 13B] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 13C] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 14] 5 is a flowchart for explaining subject detection and tracking processing in the embodiment. [Figure 15] 6 is a flowchart for explaining a main subject determination process in the embodiment. [Figure 16] FIG. 2 is a diagram showing an example of a structure of a neural network in the embodiment. [Figure 17A] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 17B] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 18A] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 18B] 5A and 5B are conceptual diagrams illustrating a main subject determination process according to the embodiment. [Figure 19] 10 is a flowchart illustrating a main subject determination process according to a second embodiment. [Figure 20A] 10A and 10B are conceptual diagrams illustrating a main subject determination process according to a second embodiment. [Figure 20B] 10A and 10B are conceptual diagrams illustrating a main subject determination process according to a second embodiment. [Figure 20C] 10A and 10B are conceptual diagrams illustrating a main subject determination process according to a second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0010] (First embodiment) FIG. 1 is a diagram showing the configuration of a digital camera (hereinafter, simply referred to as camera) 100 which is a first embodiment of an imaging device of the present invention.
[0011] In Fig. 1, the first lens group 101 is disposed closest to the subject (front side) in the imaging optical system as an image forming optical system, and is held so as to be movable in the optical axis direction. The aperture 102 adjusts the light amount by adjusting its aperture diameter. The second lens group 103 moves in the optical axis direction together with the aperture 102, and performs magnification change (zoom) together with the first lens group 101 moving in the optical axis direction.
[0012] The third lens group (focus lens) 105 moves in the optical axis direction to adjust the focus. The optical low-pass filter 106 is an optical element for reducing false colors and moire in the captured image. The first lens group 101, the aperture 102, the second lens group 103, the third lens group 105, and the optical low-pass filter 106 configure an imaging optical system.
[0013] The zoom actuator 111 rotates a cam barrel (not shown) around the optical axis, and moves the first lens group 101 and the second lens group 103 in the optical axis direction by a cam provided on the cam barrel, thereby varying the magnification. The aperture actuator 112 drives a plurality of light-shielding blades (not shown) in opening and closing directions to adjust the amount of light of the aperture 102. The focus actuator 114 moves the third lens group 105 in the optical axis direction to adjust the focus.
[0014] A focus driving circuit 126 as a focus adjustment means drives a focus actuator 114 in response to a focus driving command from the camera CPU 121, and moves the third lens group 105 in the optical axis direction. An aperture driving circuit 128 drives an aperture actuator 112 in response to an aperture driving command from the camera CPU 121. A zoom driving circuit 129 drives a zoom actuator 111 in response to a zoom operation by the user.
[0015] In this embodiment, a case will be described in which the imaging optical system, the actuators 111, 112, 114, and the drive circuits 126, 128, 129 are provided integrally with the camera body including the image sensor 107. However, an interchangeable lens having the imaging optical system, the actuators 111, 112, 114, and the drive circuits 126, 128, 129 may be detachably attached to the camera body.
[0016] The electronic flash 115 has a light-emitting element such as a xenon tube or an LED, and emits light to illuminate the subject. The AF (autofocus) auxiliary light emission unit 116 has a light-emitting element such as an LED, and improves focus detection performance for dark or low-contrast subjects by projecting an image of a mask having a predetermined opening pattern onto the subject via a projection lens. The electronic flash control circuit 122 controls the electronic flash 115 to light up in synchronization with the imaging operation. The auxiliary light drive circuit 123 controls the AF auxiliary light emission unit 116 to light up in synchronization with the focus detection operation.
[0017] The camera CPU 121 is responsible for various controls in the camera 100. The camera CPU 121 has a calculation unit, a ROM, a RAM, an A / D converter, a D / A converter, a communication interface circuit, etc. The camera CPU 121 drives various circuits in the camera 100 according to a computer program stored in the ROM, and controls a series of operations such as AF, imaging, image processing, and recording. The camera CPU 121 also functions as an image processing device.
[0018] The image sensor 107 is composed of a two-dimensional CMOS photosensor including a plurality of pixels and its peripheral circuits, and is disposed on the imaging plane of the imaging optical system. The image sensor 107 photoelectrically converts the subject image formed by the imaging optical system. The image sensor drive circuit 124 controls the operation of the image sensor 107, and also A / D converts the analog signal generated by the photoelectric conversion, and transmits the digital signal to the camera CPU 121.
[0019] The shutter 108 has a focal plane shutter configuration, and drives the focal plane shutter in response to a command from a shutter drive circuit built into the shutter 108 based on an instruction from the camera CPU 121. The image sensor 107 is shielded from light while a signal from the image sensor 107 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened and a photographing light beam is guided to the image sensor 107.
[0020] The image processing circuit 125 applies predetermined image processing to image data stored in the RAM in the camera CPU 121. The image processing applied by the image processing circuit 125 includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. Furthermore, the image processing circuit 125 stores the processed image data, the joint positions of each subject, the positions and size information of unique objects, the center of gravity of the subject, and position information of the face and eyes in the RAM in the camera CPU 121. The result of the determination processing may be used for other image processing (for example, white balance adjustment processing).
[0021] The display (display means) 131 includes a display element such as an LCD, and displays information about the imaging mode of the camera 100, a preview image before imaging, a confirmation image after imaging, an index of the focus detection area, and a focused image. The operation switch group 132 includes a main (power) switch, a release (imaging trigger) switch, a zoom operation switch, an imaging mode selection switch, and the like, and is operated by the user. The flash memory 133 records captured images. The flash memory 133 is detachable from the camera 100.
[0022] The subject detection unit 140 as a subject detection means performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 140 uses dictionary data for each subject in order to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 140 performs subject detection while sequentially switching the dictionary data for each subject. In this embodiment, the dictionary data for each subject is stored in the dictionary data storage unit 141. Therefore, multiple dictionary data are stored in the dictionary data storage unit 141. The camera CPU 121 determines which dictionary data from the multiple dictionary data to use for subject detection based on the subject priority set in advance and the settings of the imaging device. For subject detection, detection of a person and detection of organs such as the person's face, eyes, and torso are performed.
[0023] Furthermore, object detection is also performed for non-human objects such as balls, etc. The subject detection unit 140 detects human subjects and objects other than human objects (for example, balls, goal rings, and nets).
[0024] The posture acquisition unit 142, which serves as a posture detection means, performs posture estimation for each of the multiple subjects detected by the subject detection unit 140, and acquires posture information. The contents of the posture information to be acquired are determined according to the type of subject. In this example, since the subject is a person, the posture acquisition unit 142 acquires the positions of multiple joints of the person as the subject. Since the posture acquisition unit 142 can acquire the positions of the person's joints, it is also possible to perform action recognition that determines whether or not the person is performing an action involving a specific movement.
[0025] Any method may be used for estimating the posture, for example, the method described in Reference 1. The details of obtaining posture information will be described later. (Reference 1) Cao, Zhe, et al., "Realtime multi-person 2d pose estimation using part affinity ", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017. Dictionary data storage unit 141 stores dictionary data for each object. Object detection unit 140 estimates the position of the object in the image based on captured image data and dictionary data. Object detection unit 140 may estimate the object's position, size, reliability, etc., and output the estimated information. Object detection unit 140 may output other information.
[0026] Examples of dictionary data for subject detection include dictionary data for detecting a "person" as a subject, dictionary data for detecting an "animal", dictionary data for detecting a "vehicle", dictionary data for detecting a ball, etc. Furthermore, dictionary data for detecting the "whole person" and dictionary data for detecting organs such as the "person's face" may be stored separately in the dictionary data storage unit 141.
[0027] In this embodiment, the subject detection unit 140 is configured by a machine-learned CNN, and estimates the position of a subject included in image data. In this embodiment, the subject detection unit 140 is configured by different CNNs (Convolutional Neural Networks). The subject detection unit 140 may be realized by a GPU (Graphics Processing Unit) or a circuit specialized for estimation processing by CNN.
[0028] The machine learning of CNN may be performed by any method. For example, a specific computer such as a server may perform the machine learning of CNN, and the camera 100 may acquire the trained CNN from the specific computer. For example, the CNN of the subject detection unit 140 may be trained by a specific computer performing supervised learning using image data for training as input and the position of a subject corresponding to the image data for training as teacher data. In this way, a trained CNN is generated. The training of CNN may be performed by the camera 100 or the image processing device described above.
[0029] The sport type acquisition unit 143 as a sport detection means identifies the sport type that the human subject is playing from information from the subject detection unit 140, the dictionary data storage unit 141, and the posture acquisition unit 142. The sport type acquisition unit 143 may have a function of identifying (setting) the sport type in advance according to the photographer's intention.
[0030] Pan / tilt detection unit 144 is configured using a gyro sensor etc., and detects panning and tilting operations of the camera. Body attitude determination unit 145 determines a series of camera operations performed by the photographer from the start to the end of a panning or tilting operation as the body attitude based on the output of pan / tilt detection unit 144. The determination result is notified to photographer intention estimation unit 146.
[0031] Photographer intention estimation unit 146 estimates the photographer's intention to shoot. Photographer intention estimation unit 146 estimates whether the photographer is intentionally trying to switch the main subject based on subject information from subject detection unit 140, information on the type of posture (the above-mentioned action recognition) from posture acquisition unit 142, and history information on the camera operation (panning or tilting operation) by the photographer from body posture determination unit 145. The estimation result is sent to camera CPU 121, and then used for focus drive control on the main subject by each drive command from camera CPU 121.
[0032] Next, the image array of the image sensor 107 will be described with reference to Fig. 2. Fig. 2 shows a pixel array of the image sensor 107 in a range of 4 pixel columns x 4 pixel rows as viewed from the optical axis direction (z direction).
[0033] One pixel unit 200 includes four imaging pixels arranged in two rows and two columns. By arranging a large number of pixel units 200 on the image sensor 107, photoelectric conversion of a two-dimensional subject image can be performed. In one pixel unit 200, an imaging pixel (hereinafter referred to as an R pixel) 200R having a spectral sensitivity of R (red) is arranged at the upper left, and imaging pixels (hereinafter referred to as G pixels) 200G having a spectral sensitivity of G (green) are arranged at the upper right and lower left. Furthermore, an imaging pixel (hereinafter referred to as a B pixel) 200B having a spectral sensitivity of B (blue) is arranged at the lower right. Each imaging pixel includes a first focus detection pixel 201 and a second focus detection pixel 202 divided in the horizontal direction (x direction).
[0034] In the image sensor 107 of this embodiment, the pixel pitch P of the imaging pixels is 4 μm, the number of imaging pixels N is 5575 horizontal columns (x)×3725 vertical rows (y)=approximately 20.75 million pixels, the pixel pitch PAF of the focus detection pixels is 2 μm, and the number of focus detection pixels NAF is 11150 horizontal columns ×3725 vertical rows=approximately 41.5 million pixels.
[0035] In this embodiment, a case where each imaging pixel is divided into two in the horizontal direction will be described, but it may be divided in the vertical direction. Also, the image sensor 107 in this embodiment has a plurality of imaging pixels each including a first and a second focus detection pixel, but the imaging pixel and the first and second focus detection pixels may be provided as separate pixels. For example, the first and second focus detection pixels may be discretely arranged among the plurality of imaging pixels.
[0036] Fig. 3(a) shows one imaging pixel (200R, 200G, 200B) as viewed from the light receiving surface side (+z direction) of the image sensor 107. Fig. 3(b) shows the aa cross section of the imaging pixel of Fig. 3(a) as viewed from the -y direction. As shown in Fig. 3(b), one imaging pixel is provided with one microlens 305 for collecting incident light.
[0037] Furthermore, the imaging pixel is provided with photoelectric conversion units 301, 302 that are divided into N parts (two parts in this embodiment) in the x direction. The photoelectric conversion units 301, 302 respectively correspond to the first focus detection pixel 201 and the second focus detection pixel 202. The centers of gravity of the photoelectric conversion units 301, 302 are decentered on the -x side and +x side, respectively, with respect to the optical axis of the microlens 305.
[0038] An R, G or B color filter 306 is provided between the microlens 305 and the photoelectric conversion units 301 and 302 in each imaging pixel. Note that the spectral transmittance of the color filter may be changed for each photoelectric conversion unit, or the color filter may be omitted.
[0039] Light incident on the imaging pixels from the imaging optical system is collected by the microlens 305, dispersed by the color filter 306, and then received by the photoelectric conversion units 301 and 302 where it is photoelectrically converted.
[0040] Next, the relationship between the pixel structure shown in Fig. 3 and pupil division will be described with reference to Fig. 4. Fig. 4 shows the aa cross section of the imaging pixel shown in Fig. 3(a) viewed from the +y side, and also shows the exit pupil of the imaging optical system. In Fig. 4, the x and y directions of the imaging pixel are inverted with respect to Fig. 3(b) in order to correspond to the coordinate axes of the exit pupil.
[0041] A first pupil region 501, the center of gravity of which is decentered on the +X side of the exit pupil, is an area that is substantially conjugate with the light receiving surface of the photoelectric conversion unit 301 on the -x side of the imaging pixel by the microlens 305. The light flux that passes through the first pupil region 501 is received by the photoelectric conversion unit 301, i.e., the first focus detection pixel 201. A second pupil region 502, the center of gravity of which is decentered on the -X side of the exit pupil, is an area that is substantially conjugate with the light receiving surface of the photoelectric conversion unit 302 on the +x side of the imaging pixel by the microlens 305. The light flux that passes through the second pupil region 502 is received by the photoelectric conversion unit 302, i.e., the second focus detection pixel 202. The pupil region 500 indicates a pupil region that can receive light by the entire imaging pixel including all of the photoelectric conversion units 301, 302 (the first and second focus detection pixels 201, 202).
[0042] FIG. 5 shows pupil division by the image sensor 107. A pair of light beams passing through the first pupil region 501 and the second pupil region 502, respectively, are incident on each pixel of the image sensor 107 at different angles and are received by the first and second focus detection pixels 201 and 202 divided into two. In this embodiment, output signals from the multiple first focus detection pixels 201 of the image sensor 107 are collected to generate a first focus detection signal, and output signals from the multiple second focus detection pixels 202 are collected to generate a second focus detection signal. In addition, the output signals from the first focus detection pixels 201 and the output signals from the second focus detection pixels 202 of the multiple imaging pixels are added to generate an imaging pixel signal. Then, the imaging pixel signals from the multiple imaging pixels are combined to generate an imaging signal for generating an image with a resolution corresponding to the number of effective pixels N. Next, the relationship between the defocus amount of the imaging optical system and the phase difference (image shift amount) between the first focus detection signal and the second focus detection signal acquired from the imaging element 107 will be described with reference to Fig. 6. The imaging element 107 is disposed on the imaging surface 600 in the figure, and as described with reference to Figs. 4 and 5, the exit pupil of the imaging optical system is divided into a first pupil region 501 and a second pupil region 502. The defocus amount d is defined such that |d| is the distance (size) from the imaging position C of the light beam from the subject (801, 802) to the imaging surface 600, and a front-focus state in which the imaging position C is on the subject side of the imaging surface 600 is represented by a negative sign (d<0), and a back-focus state in which the imaging position C is on the opposite side of the subject from the imaging surface 600 is represented by a positive sign (d>0). In a focused state in which the imaging position C is on the imaging surface 600, d=0. The imaging optical system is in focus (d=0) with respect to a subject 801, and in front-focus state (d<0) with respect to a subject 802. The front-focus state (d<0) and back-focus state (d>0) are collectively called a defocus state (|d|>0).
[0043] In the front focus state (d<0), the light beam from the subject 802 that passes through the first pupil region 501 (second pupil region 502) is once collected and then spreads to a width Γ1 (Γ2) centered on the center of gravity position G1 (G2) of the light beam, forming a blurred image on the imaging surface 600. This blurred image is received by each first focus detection pixel 201 (each second focus detection pixel 202) on the image sensor 107, and a first focus detection signal (second focus detection signal) is generated. In other words, the first focus detection signal (second focus detection signal) is a signal that represents an image of the subject 802 at the center of gravity position G1 (G2) of the light beam on the imaging surface 600, blurred by the blur width Γ1 (Γ2).
[0044] The blur width Γ1 (Γ2) of the subject image increases roughly in proportion to an increase in the magnitude |d| of the defocus amount d. Similarly, the magnitude |p| of the image shift amount p (= the difference G1-G2 between the center of gravity positions of the light beams) between the first focus detection signal and the second focus detection signal also increases roughly in proportion to an increase in the magnitude |d| of the defocus amount d. In the back-focus state (d>0), the direction of the image shift between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state, but the same is true.
[0045] In this way, the magnitude of the image shift amount between the first and second focus detection signals increases as the magnitude of the defocus amount increases. In this embodiment, image-surface phase difference detection focus detection is performed to calculate the defocus amount from the image shift amount between the first and second focus detection signals obtained using the image sensor 107.
[0046] Next, the focus detection area of the image sensor 107 from which the first and second focus detection signals are obtained will be described with reference to Fig. 7. In Fig. 7, A(n,m) indicates the n-th focus detection area in the x direction and the m-th focus detection area in the y direction among a plurality of focus detection areas (three in each of the x direction and y direction, totaling nine) set in the effective pixel area 1000 of the image sensor 107. The first and second focus detection signals are generated from output signals from a plurality of first and second focus detection pixels 201, 202 included in the focus detection area A(n,m). I(n,m) indicates an index that indicates the position of the focus detection area A(n,m) on the display 131.
[0047] Note that the nine focus detection areas shown in FIG. 7 are merely examples, and the number, positions, and sizes of the focus detection areas are not limited. For example, one or more areas may be set as focus detection areas within a predetermined range centered on a position specified by a user or a subject position detected by a subject detector. In this embodiment, the focus detection areas are arranged so that focus detection results can be obtained with higher resolution when acquiring a defocus map, which will be described later. For example, a total of 9600 focus detection areas are arranged on the image sensor, divided into 120 horizontally and 80 vertically.
[0048] The flowchart in Fig. 8 shows AF / imaging processing (image processing method) that causes the camera 100 of this embodiment to perform AF (autofocus) operation and imaging operation. Specifically, it shows processing that causes the camera 100 to perform operations from before imaging to display a live view image on the display 131 to capturing a still image. The camera CPU 121, which is a computer, executes this processing in accordance with a computer program. In the following explanation, S means step.
[0049] First, in S1, the camera CPU 121 causes the image sensor drive circuit 124 to drive the image sensor 107, and acquires image data from the image sensor 107. After that, the camera CPU 121 acquires first and second focus detection signals from a plurality of first and second focus detection pixels included in each of the focus detection areas shown in FIG. 7, from the acquired image data. The camera CPU 121 also adds the first and second focus detection signals of all effective pixels of the image sensor 107 to generate an image signal, and has the image processing circuit 125 perform image processing on the image signal (image data) to acquire image data. Note that when the image sensor and the first and second focus detection pixels are provided separately, the camera CPU 121 acquires image data by performing a complementation process on the focus detection pixels.
[0050] Next, in S2, the camera CPU 121 causes the image processing circuit 125 to generate a live view image from the image data obtained in S1, and causes this to be displayed on the display 131. Note that the live view image is a reduced image matched to the resolution of the display 131, and the user can adjust the image capture composition, exposure conditions, and the like while viewing this. Therefore, the camera CPU 121 performs exposure adjustment based on the photometric value obtained from the image data, and displays it on the display 131. The exposure adjustment is realized by appropriately adjusting the exposure time, opening and closing the aperture of the shooting lens, and adjusting the gain for the image sensor output.
[0051] Next, in S3, the camera CPU 121 judges whether or not a switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing a release switch included in the operation switch group 132. If Sw1 is not turned on, the camera CPU 121 repeats the judgment of S3 to monitor the timing at which Sw1 is turned on. On the other hand, if Sw1 is turned on, the camera CPU 121 proceeds to S400 and performs subject tracking autofocus (AF) processing. Here, it performs detection of the subject area from the obtained image capture signal and focus detection signal, setting of the focus detection area, predictive AF processing to suppress the influence of the time lag between the focus detection processing and the image capture processing of the recorded image, and the like. Details will be described later.
[0052] Then, the camera CPU 121 proceeds to S5, where it determines whether or not the switch Sw2, which instructs the start of an imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera CPU 121 returns to S3. On the other hand, if Sw2 is turned on, it proceeds to S300, where it executes an imaging subroutine. Details of the imaging subroutine will be described later. When the imaging subroutine ends, it proceeds to S7.
[0053] In S7, the camera CPU 121 determines whether or not the main switch included in the operation switch group 132 has been turned off. If the main switch has been turned off, the camera CPU 121 ends this process, and if the main switch has not been turned off, the process returns to S3.
[0054] In this embodiment, the subject detection process and AF process are performed after the on-state of Sw1 is detected in S3, but the timing of these processes is not limited to this. By performing the subject tracking AF process in S400 before the on-state of Sw1, it is possible to eliminate the need for the photographer to take preparatory actions before shooting.
[0055] Next, the imaging subroutine executed by the camera CPU 121 in S300 of FIG. 8 will be described with reference to the flowchart shown in FIG.
[0056] In S301, the camera CPU 121 performs exposure control processing and determines the imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.) This exposure control processing can be performed using luminance information acquired from image data of a live view image.
[0057] Then, the camera CPU 121 transmits the determined aperture value to the aperture drive circuit 128 to drive the aperture 102. The camera CPU 121 also transmits the determined shutter speed to the shutter 108 to open the focal plane shutter. Furthermore, the camera CPU 121 causes the image sensor 107 to accumulate electric charge during the exposure period via the image sensor drive circuit 124.
[0058] In S302, the camera CPU 121, which has performed the exposure control process, causes the image sensor drive circuit 124 to read out all pixels of the image sensor 107 image signal for capturing a still image. The camera CPU 121 also causes the image sensor drive circuit 124 to read out one of the first and second focus detection signals from a focus detection area (focus target area) in the image sensor 107. The first or second focus detection signal read out at this time is used to detect the focus state of an image during image playback, which will be described later. By subtracting one of the first and second focus detection signals from the image sensor signal, the other focus detection signal can be obtained.
[0059] Next, in S303, the camera CPU 121 causes the image processing circuit 125 to perform defective pixel correction processing on the imaging data that was read out and A / D converted in S302.
[0060] Furthermore, in S304, the camera CPU 121 causes the image processing circuit 125 to perform image processing and encoding processing such as demosaic (color interpolation), white balance processing, gamma correction (tone correction), color conversion, and edge enhancement on the imaging data after the defective pixel correction processing.
[0061] Then, in S305, the camera CPU 121 records the still image data obtained as a result of the image processing and encoding processing in S304 and one of the focus detection signals read out in S302 in the flash memory 133 as an image data file.
[0062] Next, in S306, the camera CPU 121 associates the camera characteristic information as characteristic information of the camera 100 with the still image data recorded in S305, and records the camera characteristic information in the flash memory 133 and in the memory within the camera CPU 121. The camera characteristic information includes, for example, the following information. Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 125 Information regarding the light receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107 Information regarding vignetting of the imaging light flux within the camera 100 Information on the distance from the mounting surface of the imaging optical system of the camera 100 to the imaging element 107 Information about the manufacturing tolerances of Camera 100 The information on the light receiving sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light receiving sensitivity distribution information) is information on the sensitivity of the image sensor 107 according to the distance (position) on the optical axis from the image sensor 107. Since this light receiving sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, it may be information on these. Furthermore, the light receiving sensitivity distribution information may be information on the change in sensitivity with respect to the incident angle of light.
[0063] Next, in S307, camera CPU 121 records lens characteristic information as characteristic information of the imaging optical system in association with the still image data recorded in S305 in flash memory 133 and a memory in camera CPU 121. The lens characteristic information includes, for example, information on the exit pupil, information on a frame such as a lens barrel that blocks light beams, information on the focal length and F-number at the time of imaging, information on aberration of the imaging optical system, information on manufacturing errors of the imaging optical system, and information on the position of focus lens 105 at the time of imaging (subject distance).
[0064] Next, in S308, the camera CPU 121 records image-related information as information related to the still image data in the flash memory 133 and in a memory within the camera CPU 121. The image-related information includes, for example, information related to a focus detection operation before imaging, information related to the movement of a subject, and information related to focus detection accuracy.
[0065] Next, in S309, the camera CPU 121 displays a preview of the captured image on the display 131. This allows the user to easily check the captured image.
[0066] When the process of S309 ends, the camera CPU 121 ends this imaging subroutine and proceeds to S7 in FIG.
[0067] Next, the subject tracking AF processing subroutine executed by the camera CPU 121 in S400 of FIG. 8 will be described with reference to the flowchart shown in FIG.
[0068] In S401, the camera CPU 121 calculates the amount of image shift between the first and second focus detection signals obtained in each of the multiple focus detection areas acquired in S1 of Fig. 8, and calculates the defocus amount for each focus detection area from the image shift amount. As described above, in this embodiment, a group of focus detection results obtained from focus detection areas arranged on the image sensor at a total of 9,600 points divided into 120 horizontally and 80 vertically is called a defocus map.
[0069] Next, in S402, the camera CPU 121 performs subject detection and tracking processing. The subject detection processing is performed by the above-mentioned subject detection unit 140. Since subject detection may be impossible depending on the state of the obtained image, in such a case, a tracking processing using other means such as template matching is performed to estimate the position of the subject. Details will be described later.
[0070] Next, in S403 , the camera CPU 121 acquires posture information from the positions of each joint of the multiple subjects detected by the subject detection unit 140 .
[0071] Fig. 11 is a conceptual diagram of information acquired by the attitude acquisition unit 142. Fig. 11(a) shows an image to be processed, in which a subject 901 is catching a ball 903. The subject 901 is an important subject in a shooting scene. In this embodiment, by using the subject's attitude information, a subject that is highly likely to be the subject that the photographer intends to focus on is determined. On the other hand, a subject 902 is a non-main subject.
[0072] FIG. 11(b) is a diagram showing an example of posture information of subjects 901 and 902, and the position and size of ball 903. Joint 911 represents each joint of subject 901, and joint 912 represents each joint of subject 902. FIG. 11(b) shows an example in which the positions of the top of the head, neck, shoulders, elbows, wrists, waist, knees, and ankles are acquired as joints, but the joint positions may be some of these, or other positions may be acquired. In addition to the joint positions, information such as axes connecting joints may be used, and any information that represents the posture of the subject may be used as posture information. In the following, a case in which joint positions are acquired as posture information will be described.
[0073] Orientation acquisition unit 142 acquires two-dimensional coordinates (x, y) of joints 911 and 912 in the image. Here, (x, y) are in pixels. Center of gravity position 913 represents the center of gravity position of ball 903, and arrow 914 represents the size of ball 903 in the image. Subject detection unit 140 acquires the two-dimensional coordinates (x, y) of the center of gravity position of ball 903 in the image, and the number of pixels indicating the width of ball 903 in the image.
[0074] Next, in S404, the camera CPU 121 performs a posture type determination process based on the subject detection result in S402 and the posture information of the subject in S403. Here, the posture type determination will be described.
[0075] 12A to 12C are schematic diagrams showing the positions of players in a basketball game. In FIG. 12A to 12C, players 924 and 925 are main subject candidates recognized as subjects taking action poses based on information from pose acquisition unit 142. Here, an action pose means a pose of a subject that a user is predicted to want to capture, such as a player aiming for a shot.
[0076] In FIG. 12A, the player 925 is the main subject, and the AF frame 1900 is set on the player 925. In FIG. 12B, the player 924 holds the ball 903 and is in a position to shoot toward the goal ring 940. The camera CPU 121 detects the player's joints by a known method using the neural network of the posture acquisition unit 142. Then, from the position information of the joints and the information of the dictionary data storage unit 141, it is estimated that the player 924 is in a shooting action posture, and information on the average posture duration time during the shooting action posture is acquired. Note that, for example, as disclosed in JP 2022-135552 A, a method of performing depth detection and joint detection or organ detection on an image to estimate a posture can be used for posture estimation here. FIG. 12C shows a scene in which the player 924 shoots and the ball 903 is heading toward the goal ring 940.
[0077] 13A to 13C are schematic diagrams showing the positions of players in different states in a basketball game. FIG. 13A shows a scene in which players 926 and 927 are present. FIG. 13B shows a scene in which player 926 has ball 903 and player 927 is waiting for a pass from player 926. FIG. 13C shows a scene in which player 926 passes the ball to player 927. In FIGS. 13B and 13C, camera CPU 121 recognizes players 926 and 927 as having a pass action posture based on posture information of players 926 and 927 obtained by posture acquisition unit 142. In addition, information on the average posture duration when in a pass action posture is acquired. The posture type determination and posture duration information acquired at this time are used as a determination element of the reliability of the main subject resemblance for selecting a main subject, which will be described later, and the length of posture duration is proportional to the level of reliability.
[0078] Returning to the explanation of Fig. 10, in S405, the camera CPU 121 performs a sport type determination process. A process of determining the sport type is performed based on the subject detection result in S402, the posture information in S403, the posture type information in S404, and the detected movements of the multiple subjects. The sport type will be described later. Note that instead of determining the sport type, the camera may have a function of selecting the sport type in advance.
[0079] The method of determining the type of sport in S405 will be described. For example, for the image in Fig. 12B, the camera CPU 121 obtains joint connection information based on pose estimation from the pose acquisition unit 142, motion vector information of people and non-people subjects from the sport type acquisition unit 143, and information from the dictionary data storage unit 141. Then, from this information, the type of sport played by the two players 924 and 925 is determined to be basketball. At this time, the type of sport may be determined by taking into account goal rings and court information other than moving objects.
[0080] Next, in S4000, the camera CPU 121 performs a main subject determination process. The main subject is determined using the defocus map obtained in S401, the subject detection result obtained in S402, the posture information obtained in S403, the posture type information obtained in S404, and the sport type information obtained in S405. Details will be described later using the sub-flowchart in FIG.
[0081] Next, in S407, the camera CPU 121 performs predictive AF processing using the focus detection result acquired in S401 and a plurality of defocus amounts that are time-series data on the timing of past focus detection.
[0082] This is a process that is necessary when there is a time lag between the timing of focus detection and the timing of exposure for the captured image. In this process, AF control is performed by predicting the position of the subject in the optical axis direction at the timing of exposure for the captured image, which is a predetermined time after the timing of focus detection. In predicting the subject's image plane position, a multivariate analysis (e.g., the least squares method) is performed using historical data of the subject's past image plane positions and times to find an equation for a prediction curve. The predicted image plane position of the subject can be calculated by substituting the time of exposure for the captured image into the equation for the prediction curve that has been found.
[0083] Furthermore, not only the optical axis direction but also three-dimensional positions may be predicted. Vectors in the XYZ directions are obtained with the screen as the XY plane and the optical axis direction as the Z direction. Specifically, the position of the subject at the timing of exposure of the captured image is predicted from the XY position of the subject obtained by the subject detection and tracking process in S402 and the time series data of the Z direction position based on the defocus amount obtained in S405. Furthermore, prediction may be made from time series data of the joint positions of the subject person. Note that the prediction targets include the main subject, multiple other people, and moving objects other than people.
[0084] Next, in S408, camera CPU 121 calculates the drive amount of the focus lens using the result of the main subject determination process in S4000, the defocus amount obtained in S401, and the result of the predictive AF process in S407. Then, camera CPU 121 drives focus actuator 114 based on this drive amount, and performs focus adjustment process by moving third lens group 105 in the optical axis direction.
[0085] In the focus adjustment process, the focus adjustment is performed on the main subject determined by the main subject determination process in S4000, avoiding sudden acceleration / deceleration focus movement and achieving a smooth focus transition. Also, the focus adjustment may be performed according to the shooting sequence of the imaging device and the control and driving performance of the shooting lens that performs the focus adjustment. For example, in the shooting sequence, when the lens driving time is long, the focus driving time can be secured, so the threshold value of the defocus amount described later is increased to make it easier to drive the lens. Conversely, in the sequence in which the lens driving time is short, the focus driving time cannot be secured, so the threshold value of the defocus amount described later may be decreased to make it difficult to drive the lens. In addition, since the drive amount that can be driven by the focus lens per unit time differs depending on the drive source of the focus lens, the threshold value of the defocus amount described later may be changed depending on the drive source of the focus lens.
[0086] When the process of S408 ends, the camera CPU 121 ends the subroutine of the subject tracking AF process, and proceeds to S5 in FIG.
[0087] Next, the subroutine of the subject detection and tracking process executed by the camera CPU 121 in S402 of FIG. 10 will be described with reference to the flowchart shown in FIG.
[0088] In S2000, the camera CPU 121 sets dictionary data according to the type of subject to be detected based on data detected from the image data acquired in S1 of FIG. 8. Based on the priority of the subject set in advance and the settings of the imaging device, the dictionary data to be used in this process is selected from a plurality of dictionary data stored in the dictionary data storage unit 141. For example, a plurality of types of dictionary data that classify subjects, such as "people", "vehicles", and "animals", are stored as the plurality of dictionary data. In this embodiment, the number of dictionary data to be selected may be one or more. In the case of one, it is possible to repeatedly detect subjects that can be detected by one dictionary data at a high frequency. On the other hand, in the case of selecting a plurality of dictionary data, the dictionary data is set sequentially according to the priority as a detected subject, so that the subjects can be detected sequentially.
[0089] Next, in S2001, subject detection unit 140 detects a subject, either human or non-human, using the image data read in step S1 as an input image and the dictionary data set in step S2000. At this time, subject detection unit 140 outputs information such as the position, size, and reliability of the detected subject. At this time, camera CPU 121 may cause display device 131 to display the above information output by subject detection unit 140.
[0090] In S2001, the subject detection unit 140 hierarchically detects multiple regions of a person, which is a first type of subject, from the image data. For example, if "person" is set as the dictionary data, multiple organs such as the "whole body" region, the "face" region, and the "eyes" region are detected. While local regions such as the person's eyes and face are regions where it is desired to adjust the focus and exposure state as a subject, they may not be detectable due to surrounding obstacles or the direction of the face. Even in such cases, the subject can be robustly detected by detecting the whole body, so the subject is configured to be detected hierarchically.
[0091] Next, in S2002, subject detection unit 140 detects a person or a non-person object that is a second type of subject different from the first type of subject in S2001. For example, dictionary data for detecting a person involved in a sport is selected from a plurality of dictionary data stored in dictionary data storage unit 141. After a person is detected as a subject, the dictionary data is changed to that of a detected object, and the area of the entire object and the center position and size of the object are detected. Note that the detected object may be specified in advance and detected.
[0092] Any method may be used for object detection, for example, the method described in the following document 2 may be used. In this embodiment, the second type of subject is a ball, but it may be another unique object such as a racket. (Reference 2) Redmon, Joseph, et al., "You only look once: Unified, real-time object detection.", Proceedings of the IEEE conference on computer vision and pattern recognition, 2016. Next, in S2003, the camera CPU 121 performs a known template matching process using the subject detection area obtained in S2001 as a template. Using the multiple images obtained in S1 of FIG. 8, a similar area is searched for in the image obtained immediately before using the subject detection area obtained in the past image as a template. As is well known, any information may be used for template matching, such as brightness information, color histogram information, feature point information such as corners and edges, etc. There are various methods for matching and template update, and any method may be used. The tracking process performed in S2003 is performed to realize stable subject detection and tracking process by detecting an area similar to the past subject detection data from the image data obtained immediately before when a subject is not detected in S2001.
[0093] When the process of S2003 ends, the camera CPU 121 ends the subroutine of the subject detection and tracking process, and proceeds to S403 in FIG.
[0094] Next, the subroutine of the main subject determination process performed in S4000 of FIG. 10 will be described with reference to the flowchart shown in FIG.
[0095] In S4001, the camera CPU 121 selects a subject candidate likely to be the main subject from among a plurality of subjects based on the defocus map (focus state) acquired in S401 of Fig. 10 and the posture information in S403. The likelihood of being the main subject is the reliability (probability) of a subject taking the action posture of shooting or passing, as described up to S405 of Fig. 10, calculated based on posture information, being the main subject. In the following, a case will be described in which the probability that the subject is the main subject of the image to be processed is used as the reliability (degree of possibility that the subject is the main subject of the image to be processed) representing the likelihood of being the main subject, but a value other than probability may be used. For example, the reciprocal of the distance between the center of gravity of the subject and the center of gravity of the unique object can be used as the reliability.
[0096] <How to calculate the probability of being the main subject> A method for calculating the probability that an object represents a main subject based on the coordinates and size of each joint will be described below. In the following, a case where a neural network, which is one of the machine learning techniques, is used will be described.
[0097] Fig. 16 is a diagram showing an example of the structure of a neural network. In Fig. 16, 1001 is an input layer, 1002 is a hidden layer, 1003 is an output layer, 1004 is a neuron, and 1005 is a line showing the connection between the neurons 1004. Here, for convenience of illustration, only representative neurons and connection lines are numbered. The number of neurons 1004 in the input layer 1001 is equal to the dimension of the input data, and the number of neurons in the output layer 1003 is two. This corresponds to a two-class classification problem of determining whether or not an object is likely to be the main subject.
[0098] A weight wij is assigned to the line 1005 connecting the i-th neuron 1004 in the input layer 1001 and the j-th neuron 1004 in the hidden layer 1002, and the value zj output by the j-th neuron 1004 in the hidden layer 1002 is given by the following equation.
[0099]
number
[0100] In equation (1), xi represents the value input to the i-th neuron 1004 of the input layer 1001. The sum is taken for all neurons 1004 in the input layer 1001 that are connected to the j-th neuron. bj is called a bias, and is a parameter that controls the ease with which the j-th neuron 1004 fires. Furthermore, the function h defined in equation (2) is an activation function called ReLU (Rectified Linear Unit). It is also possible to use another function, such as a sigmoid function, as the activation function.
[0101] Moreover, the value yk output by the k-th neuron 504 in the output layer 1003 is given by the following equation.
[0102]
number
[0103] In equation (3), zj represents the value output by the j-th neuron 1004 in the hidden layer 1002, where i, k = 0, 1. 0 corresponds to a non-main subject, and 1 corresponds to the likelihood of being the main subject. The sum is taken for all neurons in the hidden layer 1002 that are connected to the k-th neuron. Furthermore, the function f defined in equation (4) is called a softmax function, and outputs a probability value that belongs to the k-th class. In this embodiment, f(y1) is used as the probability that it is the main subject.
[0104] During learning, the coordinates of the person's joints and the coordinates and size of the ball are input. Then, all weights and biases are optimized to minimize the loss function that uses the output probability and the correct label. Here, the correct label takes two values: "1" for the main subject and "0" for a non-main subject. The loss function L can be the binary cross-entropy expressed below.
[0105]
number
[0106] In formula (5), the subscript m represents the index of the subject to be learned. ym is a probability value output from the neuron with k=1 in the output layer 1003, and tm is the correct label. The loss function may be any function that can measure the degree of agreement with the correct label, such as mean square error, in addition to formula (5). By optimizing based on formula (5), it is possible to determine the weight and bias so that the correct label and the output probability value approach each other. The learned weight and bias values are stored in the flash memory 133 in advance, and are stored in the RAM in the camera CPU 121 as necessary. A plurality of types of weight and bias values may be prepared according to the scene. Using the learned weight and bias (the result of machine learning performed in advance), the probability value f(y1) is output based on formulas (1) to (4).
[0107] In addition, when learning, the state before moving on to an important action can be learned as a state of main subject resemblance. For example, in the case of throwing a ball, the state of extending the hand forward when throwing the ball can be learned as one of the states of main subject resemblance. The reason for adopting this configuration is that when an action that is important as a main subject is actually performed, the control of the imaging device needs to be executed accurately. For example, when the reliability (probability value) corresponding to the main subject resemblance exceeds a first predetermined value set in advance, the control of automatically recording the image or video (recording control) is started, so that the photographer can capture the important moment without missing it. At this time, information on the typical time from the state of the learning object to the important action may be used for the control of the imaging device.
[0108] Although the method of calculating the probability using a neural network has been described above, other machine learning methods such as a support vector machine or a decision tree may be used as long as it is possible to classify whether or not an object is likely to be the main subject. In addition, it is also possible to construct a function that outputs a reliability or probability value based on a certain model, not limited to machine learning. It is also possible to use the value of a monotonically decreasing function of the distance between the person and the ball, assuming that the reliability, which is the likelihood of the object being the main subject, increases as the distance between the person and the ball becomes closer.
[0109] Although the ball information was also used to determine the likelihood of a subject being the main subject, it is also possible to determine the likelihood of a subject being the main subject using only the posture information of the subject. Depending on the type of posture information of the subject (e.g., pass, shoot, etc.), it may or may not be better to also use the ball information. For example, in the case of a shoot, the distance between the person and the ball is far, but the photographer may want the subject who took the shot to be the likely main subject, so the likelihood of a subject being the main subject may be determined using only the posture information of the subject, regardless of the ball, or the ball information may also be used to determine the likelihood of a subject being the main subject depending on the type of posture information of the subject.
[0110] Alternatively, data obtained by performing a predetermined transformation such as linear transformation on the coordinates of each joint and the coordinates and size of the ball may be used as input data. Furthermore, frequent switching of the main subject resemblance between two subjects with a defocus difference is often contrary to the photographer's intention. For this reason, frequent switching may be detected from the time series data of the reliability of each subject, and if there are two subjects, the switching may be prevented by increasing the reliability of one of the subjects (for example, the subject closest to the subject). Furthermore, an area including two subjects may be used as an area representing the main subject resemblance.
[0111] As another method, time-series data of the posture information of the person, the positions of the person and the ball, the defocus amount of each object, and the reliability indicating the likelihood of being the main object may be used as input data. Also, the above-mentioned prediction process may be performed, and the reliability may be calculated using the predicted data of the coordinates of the joints of the person and the coordinates and size of the ball at the time of exposure of the captured image as input data. Depending on the image plane movement speed of the object and the time-series change amount of the coordinates of each joint, it may be possible to switch whether or not to use the data on which the prediction process is performed. In this way, when the change in the posture of the object is small, the accuracy of the reliability indicating the likelihood of being the main object can be maintained, and when the change in the posture of the object is large, the result of the prediction process can be used to detect the object indicating the likelihood of being the main object at an earlier time. The above method calculates the reliability of a plurality of first type objects.
[0112] Next, in S4002, the camera CPU 121 determines the reliability of the main subject. If there is a subject that has the highest (higher) reliability indicating the main subject reliability based on posture information among the multiple subjects determined as main subject candidates in S4001 other than the subject aimed at by the photographer with the AF frame 1900, the process proceeds to S4003. If there is no such subject, that is, if the reliability of the main subject corresponding to the current AF frame 1900 is the highest, the process proceeds to S4009.
[0113] Next, in S4003, camera CPU 121 acquires posture type information of the main subject candidate with the highest reliability determined in S4001 (state determination). The posture type information is, as described above, information acquired by posture acquisition unit 142 on the determination of the type of subject action, such as shooting, passing, dribbling, spiking, blocking, receiving, etc., and information on the posture duration of the posture type.
[0114] Next, in S4004, the camera CPU 121 acquires sport type information from the sport type acquisition unit 143 for the main subject candidate determined in S4003 (sport type determination). The sport type is information on what sport is being played, such as basketball, volleyball, soccer, etc., taking into account information from the posture acquisition unit 142 as described above, and a posture duration is set for each sport. At this time, in addition to the posture and movement information of the subject person, information on moving objects such as a ball and information on fixed objects such as a goal ring, net, and goal net may be used as additional information for the sport type determination, and the sport type may be set in advance.
[0115] Next, in S4005, the camera CPU 121 acquires the defocus amount of the main subject candidate determined in S4003.
[0116] Next, in S4006, camera CPU 121 sets a defocus amount threshold for focusing on each main subject candidate without causing a sudden change in focus, based on the posture duration and defocus amount (focus state) estimated from the posture type and sport type information of the main subject candidate determined in S4003, S4004, and S4005. Details of the defocus amount threshold setting are described below.
[0117] 12A to 12C, threshold setting when the sport type is basketball and the position type is a shooting action (shooting position) will be described.
[0118] In basketball, suppose that the game situation changes from the situation shown in Fig. 12A to the situation shown in Fig. 12B. Fig. 12A shows a situation in which two players 924 and 925 exist and the AF frame 1900 is set for player 925. Fig. 12B shows a situation in which player 924 has a ball 903 and is preparing to shoot toward goal ring 940.
[0119] Here, if the control for determining whether to switch from a state in which the AF frame 1900 is located on the player 925 as in FIG. 12A to a state in which the AF frame 1900 is located on the player 924 as in FIG. 12B is applied to the flowchart in FIG. 15, the result is as follows.
[0120] When the situation of the play changes from the situation in FIG. 12A to the situation in FIG. 12B, in S4002 in the flowchart in FIG. 15, the player 924, who is a candidate for the main subject and is closer to the ball, has a higher reliability of being the main subject than the subject where the AF frame 1900 is currently positioned. Therefore, the process proceeds to S4003. In S4003, the posture type of the player 924 in FIG. 12B is determined to be a shooting posture. In S4004, the current sport type is determined to be basketball. In S4005, the defocus amount of the player 924, who is the candidate for the main subject, is obtained. In S4006, a threshold value of the defocus amount for switching the main subject, in other words, a threshold value of the defocus amount of the player 924 for switching the AF frame 1900 from the player 925 to the player 924, is set.
[0121] In the case of basketball, the duration of the shooting posture of player 924 is relatively long, longer than the predetermined time. And since the movement does not change frequently, the main subject is maintained. Therefore, even if AF frame 1900 is switched from player 925 to player 924, there is a high possibility that the focus can be moved to player 924, who is a candidate for the main subject, because time for focus drive can be secured.
[0122] For this reason, in S4006, the threshold value of the defocus amount of the player 924, who is a candidate for the main subject, is set to a large value, for example, 90Fδ (Fδ represents the defocus amount when the best focus position is set to 0 with the aperture value F of the photographing lens and the permissible circle of confusion δ). If the threshold value of the defocus amount is set to a large value, even if the defocus amount of the player 924 is somewhat large, it is determined in S4007 that it is equal to or less than the threshold value, and a decision is made to move the AF frame 1900. This makes it easier to move the AF frame, and in S4008, the AF frame 1900 is moved from the player 925 to the player 924. In this way, the AF frame is quickly moved to the player 924, who is a candidate for the main subject to be focused on.
[0123] Next, suppose that the situation of the game changes from the situation shown in Fig. 12B to the situation shown in Fig. 12C. Fig. 12C shows a situation in which a player 924 shoots and the ball 903 is heading toward a goal ring 940. In this case, in Fig. 12C, there is no other player other than the player 924 who has a high reliability of being the main subject. Therefore, the process proceeds from S4002 to S4009 in Fig. 15, and the AF frame 1900 of the player 924 continues.
[0124] Next, threshold setting when the sport type is basketball and the position type is a pass action (pass position) will be described with reference to Figs. 13A to 13C.
[0125] In basketball, it is assumed that the game situation has changed from the situation shown in Fig. 13A to the situations shown in Fig. 13B and Fig. 13C. Fig. 13A shows a situation where there are two players, a player 926 and a player 927, and the AF frame 1900 is set to the player 927. Fig. 13B shows a situation where the player 926 receives the ball 903. Fig. 13C shows a situation where the player 926 passes the ball 903 to the player 927.
[0126] Here, if the control for determining whether to switch the state in which the AF frame 1900 is positioned on the player 927 as in FIG. 13A to a state in which the AF frame 1900 is positioned on the player 926 is applied to the flowchart in FIG. 15, the result is as follows.
[0127] When the playing situation changes from the situation in FIG. 13A to the situation in FIG. 13B, in S4002 in the flowchart in FIG. 15, the player 926, who is a candidate for the main subject and is closer to the ball, has a higher reliability as a main subject than the subject where the AF frame 1900 is currently positioned. Therefore, the process proceeds to S4003. In S4003, the posture type of the player 926 in FIG. 13B is determined to be a passing posture. In S4004, the current sport type is determined to be basketball. In S4005, the defocus amount of the player 926, who is a candidate for the main subject, is obtained. In S4006, a threshold value of the defocus amount for switching the main subject, in other words, a threshold value of the defocus amount of the player 926 for switching the AF frame 1900 from the player 927 to the player 926, is set.
[0128] In the case of basketball, it is expected that the duration of a posture of a passing player 924 is shorter than the duration of a posture of a shooting action, etc., and shorter than a predetermined time. Since the action may change frequently, it is difficult to maintain the main subject. Therefore, even if the AF frame 1900 is switched from the player 927 to the player 926, the main subject may change after the switch.
[0129] For this reason, in S4006, the threshold value of the defocus amount of the player 926, which is a main subject candidate, is set to a small value, for example, 20Fδ (Fδ represents the defocus amount when the best focus position is set to 0 with the aperture value F of the photographing lens and the permissible circle of confusion δ). If the threshold value of the defocus amount is set to a small value, even if the defocus amount of the player 926 is relatively small, it is determined in S4007 that it is greater than the threshold value, and a determination is made not to move the AF frame 1900. Therefore, it becomes difficult to move the AF frame, and in S4008, the AF frame 1900 is not moved from the player 927 to the player 926. In this way, when the main subject is frequently switched, it is possible to prevent the focus drive from hunting back and forth, and improve the stability of the focus drive.
[0130] Next, suppose that the situation of the game changes from the situation shown in Fig. 13B to the situation shown in Fig. 13C. In this case, since player 927 is holding ball 903 in Fig. 13C, there is no other player other than player 927 who is highly reliable as the main subject. Therefore, the process proceeds from S4002 to S4009 in Fig. 15, and the AF frame 1900 of player 927 continues.
[0131] Next, threshold setting when the sport type is volleyball and the position type is a spike action (spike position) will be described with reference to Figs. 17A and 17B.
[0132] In volleyball, suppose that the game situation changes from that shown in Fig. 17A to that shown in Fig. 17B. Fig. 17A shows a situation in which player 932 jumps in time with the toss of ball 940 by player 934, and player 933 is present across net 905. AF frame 1900 is set for player 932. Fig. 13B shows a situation in which player 933 jumps to block the spike immediately after player 932 hits the spike.
[0133] Here, if the control for determining whether to switch the state in which the AF frame 1900 is positioned on the player 932 as in FIG. 17A to a state in which the AF frame 1900 is positioned on the player 933 is applied to the flowchart in FIG. 15, the result is as follows.
[0134] When the playing situation changes from that of FIG. 17A to that of FIG. 17B, in S4002 in the flowchart of FIG. 15, the player 933, who is a candidate for the main subject and is closer to the ball, may have a higher reliability of being the main subject than the subject where the AF frame 1900 is currently positioned. In that case, the process proceeds to S4003. In S4003, the posture type of the player 932 in FIG. 17B is determined to be a spike posture. In S4004, the current sport type is determined to be volleyball. In S4005, the defocus amount of the player 933, who is a candidate for the main subject, is acquired. In S4006, a threshold value of the defocus amount for switching the main subject, in other words, a threshold value of the defocus amount of the player 933 for switching the AF frame 1900 from the player 932 to the player 933, is set.
[0135] In the case of volleyball, the duration of the spiking player 932's posture is very short. In addition, the movement may change frequently, making it difficult to maintain the main subject. Therefore, even if the AF frame 1900 is switched from the player 932 to the player 933, the main subject may change after the switch.
[0136] For this reason, in S4006, the threshold value of the defocus amount of the player 932, which is a main subject candidate, is set to a small value, for example, 15Fδ (Fδ represents the defocus amount when the best focus position is set to 0 with the aperture value F of the photographing lens and the permissible circle of confusion δ). If the threshold value of the defocus amount is set to a small value, even if the defocus amount of the player 933 is relatively small, it is determined in S4007 that the defocus amount is greater than the threshold value, and a determination is made not to move the AF frame 1900. Therefore, it becomes difficult to move the AF frame, and in S4008, the AF frame 1900 is not moved from the player 932 to the player 933. In this way, when the main subject is frequently switched, it is possible to prevent the focus drive from hunting back and forth, and improve the stability of the focus drive.
[0137] 18A and 18B, threshold setting when the sport type is soccer and the position type is a shooting action (shooting position) will be described.
[0138] In soccer, it is assumed that the situation of the game has changed from the situation shown in Fig. 18A to the situation shown in Fig. 18B. Fig. 18A shows a situation in which player 936 has ball 960, player 935 is waiting for ball 960, and player 937 is waiting as a goalkeeper. The AF frame 1900 is set for player 936. Fig. 18B shows a situation in which player 935, who has received a pass from player 936, shoots toward soccer goal 961.
[0139] Here, if the control for determining whether to switch the state in which the AF frame 1900 is located on the player 936 as in FIG. 18A to the state in which the AF frame 1900 is located on the player 935 as in FIG. 18B is applied to the flowchart in FIG. 15, the result is as follows.
[0140] When the playing situation changes from the situation in FIG. 18A to the situation in FIG. 18B, in S4002 in the flowchart in FIG. 15, the player 935, who is a candidate for the main subject and is closer to the ball, has a higher reliability as a main subject than the subject where the AF frame 1900 is currently positioned. Therefore, the process proceeds to S4003. In S4003, the posture type of the player 935 in FIG. 18B is determined to be a shooting posture. In S4004, the current sport type is determined to be soccer. In S4005, the defocus amount of the player 935, who is a candidate for the main subject, is obtained. In S4006, a threshold value of the defocus amount for switching the main subject, in other words, a threshold value of the defocus amount of the player 935 for switching the AF frame 1900 from the player 936 to the player 935, is set.
[0141] In the case of soccer, the duration of the shooting player 935's posture is relatively long, longer than the predetermined time. And since the movement does not change frequently, the main subject is maintained. Therefore, even if the AF frame 1900 is switched from the player 936 to the player 935, there is a high possibility that the focus can be moved to the player 935, who is a candidate for the main subject, because time for focus drive can be secured.
[0142] For this reason, in S4006, the threshold value of the defocus amount of the player 935, who is a candidate for the main subject, is set to a large value, for example, 80Fδ (Fδ represents the defocus amount when the best focus position is set to 0 with the aperture value F of the photographing lens and the permissible circle of confusion δ). If the threshold value of the defocus amount is set to a large value, even if the defocus amount of the player 935 is somewhat large, it is determined in S4007 that it is equal to or less than the threshold value, and a decision is made to move the AF frame 1900. This makes it easier to move the AF frame, and in S4008, the AF frame 1900 is moved from the player 936 to the player 935. In this way, the AF frame is quickly moved to the player 935, who is a candidate for the main subject to be focused on.
[0143] As described above, in this embodiment, in the case of a sport type and posture type in which the subject's movements do not change frequently and the subject is likely to be maintained as the main subject, the defocus amount threshold is increased to make it easier for the main subject to change (the AF frame is likely to move). On the other hand, in the case of a sport type and posture type in which the subject's movements may change frequently and the subject is likely to be maintained as the main subject, the defocus amount threshold is decreased to make it harder for the main subject to change (the AF frame is unlikely to move). This makes it possible to perform appropriate focus control according to the sport type, posture type, and defocus amount.
[0144] The threshold value of the defocus amount set in S4006 may be set based only on the posture type (action recognition) regardless of the sport. Also, it may be changed depending on the shooting sequence on the imaging device side, the drive processing time of the imaging optical system, etc.
[0145] Second embodiment Next, a second embodiment of the present invention will be described. The configuration of the imaging device in this embodiment is the same as that of the first embodiment shown in Fig. 1, and in the following, the same parts as those in the first embodiment are given the same reference numerals as those in the first embodiment, and the description is omitted, and only the differences from the first embodiment will be described.
[0146] 19, a description will be given of switching of the main subject performed by camera CPU 121 mainly based on information from photographer's intention estimation section 146. Note that the processes shown in steps S4001 to S4005 and S4006 to S4009 are the same as those in Fig. 15, and only steps S5006 to S5008 which differ from Fig. 15 will be described below.
[0147] In S5006, the camera CPU 121 detects whether the photographer is performing a panning / tilting operation using the pan / tilt detection unit 144, and if such an operation is detected, the process proceeds to S5007, or if such an operation is not detected, the process proceeds to S4006.
[0148] Here, an example of the main subject switching operation when a panning / tilting operation is detected will be described with reference to FIGS. 20A to 20C.
[0149] 20A shows a situation in which a player 941 exists within a central image height 1902 of a shooting angle of view 1901, and a player 942 exists at a peripheral image height 1903. An AF frame 1900 is set for the player 941, and the posture of the player 942 has been detected as a candidate for the main subject.
[0150] 20B shows a state in which the photographer intentionally moves the player 942 into the central image height 1902 of the shooting angle of view 1901 by panning / tilting. At this time, the pan / tilt detection unit 144 detects the panning / tilting operation, and the camera CPU 121 switches the player 942 from a main subject candidate to the main subject, and sets the AF frame 1900 to the player 942.
[0151] Fig. 20C shows a shooting scene of player 942, who was switched as the main subject in Fig. 20B. Player 942 continues to be the main subject within central image height 1902, and player 941 within peripheral image height 1903 has his posture detected as a main subject candidate. Note that the range of central image height 1902 may be changed depending on the shooting conditions and subject conditions, and the switching of the main subject may be prioritized according to the posture type of the main subject candidate.
[0152] 20, photographer intention estimation section 146 determines whether the photographer is intentionally trying to switch the main subject, based on camera movement and position information from body orientation determination section 145. If it is determined that the main subject has been switched by the user's intention, the process proceeds to S5008, and if it is determined that the main subject is not being switched by the user's intention, the process proceeds to S4009.
[0153] When switching the main subject in S5008, focusing is performed regardless of the defocus amount of the main subject, but the focus transition to the main subject (focus drive control, etc.) may be changed taking into account responsiveness to panning / tilting operations.
[0154] The disclosure of this specification includes the following imaging device, method, program, and storage medium.
[0155] (Item 1) A subject detection means for detecting a subject; a posture detection means for detecting a posture of the subject; a focus detection means for detecting a focus state of the subject; a setting means for setting a threshold value for determining whether or not to select the subject as a main subject based on the posture and the focus state of the subject; An imaging device comprising:
[0156] (Item 2) 2. The imaging device according to item 1, further comprising a focus adjustment unit for adjusting the focus on the subject selected as the main subject.
[0157] (Item 3) 3. The imaging device according to item 1 or 2, further comprising an acquisition unit that acquires reliability that the subject is a main subject based on a posture of the subject.
[0158] (Item 4) 4. The imaging device according to any one of items 1 to 3, further comprising a state determination means for determining a state of the subject based on a posture of the subject, wherein the setting means sets a threshold for determining whether or not to select the subject as a main subject based on the state of the subject and the focus state.
[0159] (Item 5) 5. The imaging device according to item 4, wherein the state determination means determines the state of what action the subject is performing based on the posture of the subject.
[0160] (Item 6) 6. The imaging device according to any one of items 1 to 5, further comprising a sport determination unit that determines the type of sport the subject is participating in based on the posture of the subject.
[0161] (Item 7) 7. The imaging device according to item 6, wherein the setting means sets the threshold value further based on the type of sport.
[0162] (Item 8) 8. The imaging device according to claim 1, wherein the setting unit sets the threshold value further based on a duration of a posture of the subject.
[0163] (Item 9) 9. The imaging apparatus according to item 8, wherein the threshold value is a threshold value for a defocus amount of the subject.
[0164] (Item 10) 10. The imaging apparatus according to item 9, further comprising a selection unit that selects the subject as the main subject when the defocus amount of the subject is equal to or less than the threshold value.
[0165] (Item 11) 11. The imaging device according to any one of items 8 to 10, wherein the setting unit sets the threshold value to a larger value as the duration of the posture is longer.
[0166] (Item 12) 12. The imaging device according to any one of items 1 to 11, further comprising a detection unit for detecting panning or tilting of the imaging device.
[0167] (Item 13) 13. The imaging device according to item 12, further comprising a second selection means for selecting, as a main subject, a subject that has come closer to the center of the screen by the panning or tilting.
[0168] (Item 14) a subject detection step of detecting a subject; a posture detection step of detecting a posture of the subject; a focus detection step of detecting a focus state of the subject; a setting step of setting a threshold value for determining whether or not to select the subject as a main subject based on the posture and the focus state of the subject; 13. A method for controlling an imaging apparatus comprising:
[0169] (Item 15) Item 15. A program for causing a computer to execute each step of the control method according to item 14.
[0170] (Item 16) A computer-readable storage medium storing a program for causing a computer to execute each step of the control method described in item 14.
[0171] (Other embodiments) The present invention can also be realized by a process in which a program for realizing one or more functions of the above-mentioned embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for realizing one or more functions.
[0172] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0173] 100: camera, 107: imaging element, 121: camera CPU, 126: focus drive circuit, 140: subject detection unit, 142: attitude acquisition unit, 143: sport type acquisition unit, 144: pan / tilt detection unit, 145: body attitude determination unit, 146: photographer intention estimation unit
Claims
1. a subject detection means for detecting a subject; a focus detection means for detecting a focus state of the object detected by the object detection means; a motion detection means for detecting a subject performing a specific motion; a selection means for selecting a main subject based on the detection result of the motion detection means and the focus state; Equipped with The imaging device is characterized in that, when the focus state of the subject detected by the subject detection means is a first focus state, the selection means does not select the subject as a main subject if the subject is performing a first action, and selects the subject as a main subject if the subject is performing a second action.
2. 2. The imaging device according to claim 1, further comprising a focus adjustment unit for adjusting the focus on the object selected as the main object.
3. The imaging device described in Claim 1, characterized in that the movement detection means detects a subject performing an action involving a specific movement based on the posture of the subject.
4. 4. The imaging apparatus according to claim 3, further comprising an acquisition unit that acquires the reliability that the subject is a main subject based on the posture of the subject.
5. An imaging device as described in claim 1, characterized in that the posture of the subject lasts longer during the second action than during the first action.
6. 4. The imaging device according to claim 3, further comprising a game determining unit that determines the type of game the subject is playing based on the posture of the subject.
7. 7. The imaging device according to claim 6, wherein the selection means selects the main subject based further on the type of the sport.
8. 2. The imaging device according to claim 1, further comprising a detection unit for detecting panning or tilting of the imaging device.
9. 9. The imaging apparatus according to claim 8, further comprising a second selection means for selecting, as a main subject, a subject that has come closer to the center of the screen by the panning or tilting.
10. a subject detection step of detecting a subject; a focus detection step of detecting a focus state of the subject detected by the subject detection step; a motion detection step of detecting a subject performing a specific motion; a selection step of selecting a main subject based on the detection result of the motion detection step and the focus state; and In the selection process, when the focus state of the subject detected by the subject detection process is a first focus state, the subject is not selected as a main subject if the subject is performing a first action, and is selected as a main subject if the subject is performing a second action.
11. A program for causing a computer to execute each step of the control method according to claim 10.
12. A computer-readable storage medium storing a program for causing a computer to execute each step of the control method according to claim 10.