Image processing device, image processing method and program
The image processing device uses machine learning-based CNNs to identify and set focus areas for vehicles by considering shooting direction and inclination, addressing the challenge of inaccurate focus detection for non-human subjects, thereby improving image clarity.
Patent Information
- Application Number
- JP2025011233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-02-18
AI Technical Summary
Existing image processing systems struggle to accurately set the focus detection area to the user's desired region for subjects other than human faces, particularly when the subject is a vehicle, due to variations in shooting scenes.
An image processing device equipped with a subject detection unit and a focus area detection unit that utilize machine learning-based CNNs to identify and set the focus area based on the type of vehicle detected, considering shooting direction and inclination, allowing for precise focus area adjustment.
Enables the focus detection area to be accurately set to the user's desired region for various subjects, including vehicles, enhancing image clarity and focus accuracy.
Smart Images

Figure 0007815493000001 
Figure 0007815493000002 
Figure 0007815493000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and a program. [Background technology]
[0002] Within a subject detected by an imaging device, a user (photographer) must set an area where focus detection is desired and track the subject. Related technology is proposed in Patent Document 1. In the technology of Patent Document 1, when the subject is a human face, the eyes within the face are detected, the size of the detected eyes is determined, and a focus detection area is set to the eyes or face. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-123301 Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology of Patent Document 1 mentioned above, when the subject is a person, the focus detection area can be set to the person's eyes or face, but depending on the shooting scene, there is a risk that the focus detection area will be set to an area different from the user's intention for a different subject.
[0005] An object of the present invention is to set the focus detection area to an area that the user desires to set for a detectable subject. [Means for solving the problem]
[0006] In order to achieve the above object, the image processing device of the present invention has a detection means for detecting multiple types of vehicles, an acquisition means for acquiring information regarding at least one of the shooting direction relative to the vehicle or the inclination direction of the vehicle, and a setting means for setting the area to be focused, and is characterized in that the setting means changes the area to be focused depending on the type of vehicle detected by the detection means and the information acquired by the acquisition means. [Effects of the Invention]
[0007] According to the present invention, for a detectable subject, the focus detection area can be set to an area that the user desires to set. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an example of the configuration of an imaging device according to a first embodiment of the present invention. [Figure 2] 2 is a diagram showing a pixel arrangement of an imaging element in the imaging device according to the first embodiment. FIG. [Figure 3] FIG. 3A is a plan view of an imaging pixel in the imaging element of the first embodiment, and FIG. 3B is a cross-sectional view of the imaging pixel in the imaging element of the first embodiment. [Figure 4] 2A and 2B are diagrams for explaining the structure of an imaging pixel in the imaging element of the first embodiment. [Figure 5] FIG. 2 is a diagram for explaining pupil division by the image sensor of the first embodiment. [Figure 6] FIG. 4 is a diagram for explaining the relationship between the defocus amount and the image shift amount in the first embodiment. [Figure 7] FIG. 2 is a diagram for explaining a focus detection area in the first embodiment. [Figure 8] 4 is a flowchart showing the flow of live view shooting in the imaging apparatus according to the first embodiment. [Figure 9] 4 is a flowchart showing the flow of a photographing process in the first embodiment. [Figure 10] 5 is a flowchart showing the flow of subject tracking AF processing in the first embodiment. [Figure 11] 5 is a flowchart showing the flow of subject detection processing and tracking processing in the first embodiment. [Figure 12] 5 is a flowchart showing the flow of focus area detection processing in the first embodiment. [Figure 13] 6 is a flowchart showing the flow of focus detection area setting processing in the first embodiment. [Figure 14] 14A, 14B, 14C, and 14D are diagrams for explaining the focus region detected in the focus region detection process of the first embodiment. [Figure 15] 1A and 1B are diagrams showing examples of scenes in which detection as a focus area in the first embodiment may be effective. [Figure 16] 5 is a flowchart showing the flow of predictive AF processing according to the first embodiment. [Figure 17] 5A and 5B are diagrams for explaining the image plane movement amount of a subject and a predicted curve in the first embodiment. [Figure 18] 18(A), 18(B), 18(C), 18(D), 18(E), and 18(F) are diagrams for explaining the focus movable range in the first embodiment. [Figure 19] FIG. 2 is a diagram for explaining a focus movable range in the first embodiment. [Figure 20] FIG. 4 is a diagram for explaining items to be changed during predictive calculation and focus control in the first embodiment. [Figure 21] 10 is a flowchart showing the flow of focus detection area setting processing according to a second embodiment of the present invention. [Figure 22] 22A and 22B are conceptual diagrams for explaining local regions and focus detection candidate regions in the second embodiment. [Figure 23] 23A, 23B, 23C, and 23D are conceptual diagrams for explaining the setting of focus detection candidate areas in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, each embodiment of the present invention will be described in detail with reference to the drawings. However, the configurations described in each of the following embodiments are merely examples, and the scope of the present invention is not limited to the configurations described in each of the embodiments.
[0010] First Embodiment A first embodiment of the present invention will be described below with reference to the drawings. Fig. 1 is a block diagram showing an example of the configuration of an imaging device (camera 100) according to the first embodiment of the present invention.
[0011] 1, camera 100 has an imaging optical system, a zoom actuator 111, an aperture actuator 112, a focus actuator 114, an electronic flash 115, an AF assist light emitter 116, an image sensor 107, and a shutter 108. Camera 100 also has a CPU 121, an electronic flash control circuit 122, an assist light drive circuit 123, an image sensor drive circuit 124, an image processing circuit 125, a focus drive circuit 126, an aperture drive circuit 128, and a zoom drive circuit 129. Camera 100 also has a display 131, a group of operation switches 132, a flash memory 133, a subject detection unit 140, a dictionary data storage unit 141, and a focus area detection unit 142.
[0012] The imaging optical system is composed of a first lens group 101, an aperture 102, a second lens group 103, a third lens group 105, and an optical low-pass filter 106. The first lens group 101 is arranged closest to the subject (front side) of the imaging optical system as an image-forming optical system, and is held so as to be movable in the direction of the optical axis. The aperture 102 adjusts the amount of light by adjusting its aperture diameter. The second lens group 103 moves in the direction of the optical axis together with the aperture 102, and performs magnification change (zooming) together with the first lens group 101, which moves in the direction of the optical axis. The third lens group (focus lens) 105 moves in the direction of the optical axis to perform focus adjustment. The optical low-pass filter 106 is an optical element that reduces false colors and moire in captured images.
[0013] The zoom actuator 111 rotates a cam barrel (not shown) around the optical axis, and a cam provided on the cam barrel moves the first lens group 101 and the second lens group 103 in the optical axis direction to change the magnification. The diaphragm actuator 112 drives a plurality of light-shielding blades (not shown) in opening and closing directions to adjust the amount of light from the diaphragm 102. The focus actuator 114 moves the third lens group 105 in the optical axis direction to adjust the focus.
[0014] The focus driving circuit 126 drives the focus actuator 114 in response to a focus driving command from the CPU 121, and moves the third lens group 105 in the optical axis direction. The aperture driving circuit 128 drives the aperture actuator 112 in response to an aperture driving command from the CPU 121. The zoom driving circuit 129 drives the zoom actuator 111 in response to a zoom operation by the user.
[0015] In the first embodiment, a case will be described in which the imaging optical system, zoom actuator 111, aperture actuator 112, focus actuator 114, focus drive circuit 126, aperture drive circuit 128, and zoom drive circuit 129 are provided integrally with the camera body. The camera body also includes an image sensor 107. However, an interchangeable lens having the imaging optical system, zoom actuator 111, aperture actuator 112, focus actuator 114, focus drive circuit 126, aperture drive circuit 128, and zoom drive circuit 129 may be detachable from the camera body.
[0016] The electronic flash 115 has a light-emitting element such as a xenon tube or an LED, and emits light to illuminate the subject. The AF assist light emitter 116 has a light-emitting element such as an LED, and projects an image of a mask with a predetermined aperture pattern onto the subject via a projection lens, thereby improving focus detection performance for dark or low-contrast subjects. The electronic flash control circuit 122 controls the electronic flash 115 to turn on in synchronization with the imaging operation. The assist light drive circuit 123 controls the AF assist light emitter 116 to turn on in synchronization with the focus detection operation.
[0017] The CPU 121 is responsible for various controls in the camera 100. The CPU 121 has an arithmetic unit, a ROM, a RAM, an A / D converter, a D / A converter, a communication interface circuit, etc. The CPU 121 executes a computer program stored in the ROM to drive various circuits in the camera 100 and control a series of processes (operations) such as AF processing, imaging processing, image processing, and recording. The CPU 121 functions as an image processing device. In addition to the CPU 121, the image processing device may also include a subject detection unit 140, a dictionary data storage unit 141, a focus area detection unit 142, etc.
[0018] The image sensor 107 is composed of a two-dimensional CMOS photosensor including multiple pixels and its peripheral circuitry, and is disposed on the imaging plane of the imaging optical system. The image sensor 107 photoelectrically converts the subject image formed by the imaging optical system. The image sensor drive circuit 124 controls the operation of the image sensor 107 and transmits to the CPU 121 a digital signal obtained by A / D conversion of the analog signal generated by the photoelectric conversion.
[0019] The shutter 108 has a focal plane shutter configuration, and drives the focal plane shutter in response to commands from a shutter drive circuit built into the shutter 108 based on instructions from the CPU 121. The image sensor 107 is shielded from light while a signal from the image sensor 107 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened and a photographing light beam is guided to the image sensor 107.
[0020] The image processing circuit (image processing unit) 125 applies predetermined image processing to image data stored in the RAM in the CPU 121. The image processing applied by the image processing circuit 125 includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation processing (demosaic processing), and gamma correction processing, as well as signal format conversion processing and scaling processing.
[0021] Furthermore, image processing circuit 125 determines the main subject based on posture information of the subject and position information of an object specific to the scene (hereinafter referred to as a "specific object"). The result of the determination process performed by image processing circuit 125 may be used for other image processing (for example, white balance adjustment process). Image processing circuit 125 saves the processed image data, the joint positions of each subject, position and size information of the specific object, the center of gravity of the subject determined to be the main subject, position information of the face and eyes, etc. in the RAM within CPU 121.
[0022] The display (display means) 131 has a display element such as an LCD (Liquid Crystal Display) and displays information about the imaging mode of the camera 100, a preview image before imaging, a confirmation image after imaging, an index for the focus detection area, and an in-focus image. The operation switch group 132 includes a main switch (power switch), a release switch (photography trigger switch), a zoom operation switch, and an imaging mode selection switch, and is operated by the user. The flash memory 133 records captured images. The flash memory 133 is detachable from the camera 100.
[0023] The subject detection unit 140, which serves as a subject detection means, performs subject detection processing based on subject detection dictionary data generated by machine learning. In the first embodiment, the subject detection unit 140 uses subject detection dictionary data for each subject to detect multiple types of subjects. Each subject detection dictionary data is, for example, data in which the features of the corresponding subject are registered. The subject detection unit 140 performs subject detection by sequentially switching between subject detection dictionary data for each subject. In the first embodiment, the subject detection dictionary data for each subject is stored in the dictionary data storage unit 141. Therefore, multiple subject detection dictionary data are stored in the dictionary data storage unit 141. The CPU 121 determines which of the multiple subject detection dictionary data to use for subject detection based on the subject priorities set in advance and the settings of the camera 100 (imaging device).
[0024] The focus area detection unit 142, which serves as a focus area detection means, detects an area within a subject to be focused (an area to be in focus) based on dictionary data for focus area detection generated by machine learning. In the first embodiment, the focus area detection unit 142 receives as input at least an image signal of an area (hereinafter referred to as a "subject detection area") of a subject (hereinafter referred to as a "detected subject") detected by the subject detection unit 140, and obtains a focus area as an output. The focus area is an area within the detected subject to be focused. In the first embodiment, dictionary data for focus area detection for each subject is stored in the dictionary data storage unit 141. Therefore, a plurality of dictionary data for focus area detection are stored in the dictionary data storage unit 141. Dictionary data for focus area detection used by the focus area detection unit 142, which is associated with the dictionary data for subject detection used by the subject detection unit 140, is selected and used. Details will be described later.
[0025] Dictionary data storage unit 141, which serves as a storage means, stores dictionary data for subject detection and dictionary data for focus area detection for each subject. Subject detection unit 140 estimates the position of the subject in the image based on captured image data and the dictionary data for subject detection. Subject detection unit 140 may also estimate information such as the position, size, and reliability of the subject, and output this estimated information. Subject detection unit 140 may also output other information. Similarly, as described above, focus area detection unit 142 uses image data of the subject detection area as an input image and outputs an area to be focused on within the input image (focus area) based on the dictionary data for focus area detection.
[0026] The subject detection dictionary data used by the subject detection unit 140 includes, for example, person dictionary data for detecting "people" as subjects, animal dictionary data for detecting "animals," vehicle dictionary data for detecting "vehicles," etc. Furthermore, dictionary data for detecting "the whole of a person" and dictionary data for detecting "a person's face" may be stored separately in the dictionary data storage unit 141.
[0027] The focus area detection dictionary data is, for example, dictionary data that, when a "vehicle" is detected as the subject and used as an input image, outputs the area of the head of the vehicle driver or the area of the side of the vehicle's casing, depending on the size of the subject and the shooting settings. The focus area detection unit 142 uses the focus area detection dictionary data. In this way, in the present invention, by outputting the area to be focused (focus area) separately from the subject detection area, it is possible to obtain an image in which the appropriate area is focused depending on the shooting scene. Details will be described later.
[0028] In the first embodiment, the subject detection unit 140 is configured by a CNN (convolutional neural network) that has undergone machine learning (deep learning) and estimates the position of a subject included in captured image data, etc. Furthermore, the focus region detection unit 142 is configured by a CNN (hereinafter referred to as a "trained CNN") that has undergone machine learning (deep learning) and estimates the position at which to focus within a region within a detected subject, etc. In the first embodiment, the subject detection unit 140 and the focus region detection unit 142 are each configured by a CNN that has undergone machine learning using different techniques. The subject detection unit 140 and the focus region detection unit 142 may be realized by a GPU (graphics processing unit) or a circuit specialized for estimation processing using a CNN.
[0029] In the present invention, the machine learning of CNN may be performed by any method. For example, a predetermined computer such as a server may perform machine learning of CNN to generate a trained CNN (i.e., a trained model), and the camera 100 may acquire the trained CNN from the predetermined computer. For example, the predetermined computer may perform machine learning of CNN in the subject detection unit 140 by performing supervised learning using training image data as input and the position of the subject corresponding to the training image data as training data. Alternatively, the predetermined computer may perform machine learning of CNN in the focus region detection unit 142 by performing supervised learning using training image data as input and the position to be focused corresponding to the subject in the training image data as training data. In this manner, trained CNNs (trained models) are generated for the subject detection unit 140 and the focus region detection unit 142.
[0030] As described above, the subject detection unit 140 detects a subject using dictionary data for subject detection. The subject detection unit 140 also detects a subject using dictionary data for subject detection for different types of subjects (people, animals, vehicles, etc.). In the first embodiment, each piece of dictionary data for subject detection used by the subject detection unit 140 is generated by applying a trained CNN that constitutes the subject detection unit 140. The focus region detection unit 142 detects a focus region using dictionary data for focus region detection. The each piece of dictionary data for focus region detection used by the focus region detection unit 142 is also generated by applying a trained CNN that constitutes the focus region detection unit 142.
[0031] Furthermore, the machine learning of the CNN may be performed by the camera 100 (image capture device) or the CPU 121 (image processing device).
[0032] As described above, in the first embodiment, the subject detection unit 140 and the focus region detection unit 142 are each configured by a different machine-learned CNN. However, the present invention is not limited to this, and the subject detection unit 140 and the focus region detection unit 142 may each be configured by a different machine-learned neural network. Furthermore, the subject detection unit 140 and the focus region detection unit 142 may be configured by a trained model other than a trained CNN. For example, the subject detection unit 140 and the focus region detection unit 142 may be configured by a trained model that is machine-learned using any machine learning algorithm such as a support vector machine or logistic regression.
[0033] Next, the pixel array of the image sensor 107 of the imaging device (camera 100) according to the first embodiment will be described with reference to Fig. 2. Fig. 2 shows the pixel array of the image sensor 107, which is in a range of 4 pixel columns x 4 pixel rows, as viewed from the optical axis direction (hereinafter referred to as the "z direction").
[0034] As shown in FIG. 2, one pixel unit 200 includes four imaging pixels arranged in two rows and two columns. Arranging a large number of pixel units 200 on the image sensor 107 enables photoelectric conversion of a two-dimensional subject image. In one pixel unit 200, an imaging pixel 200R having R (red) spectral sensitivity (hereinafter referred to as the "R pixel") is arranged at the upper left, and imaging pixels 200G having G (green) spectral sensitivity (hereinafter referred to as the "G pixel") are arranged at the upper right and lower left, respectively. Furthermore, an imaging pixel 200B having B (blue) spectral sensitivity (hereinafter referred to as the "B pixel") is arranged at the lower right. Each imaging pixel includes a first focus detection pixel 201 and a second focus detection pixel 202, which are divided in the horizontal direction (hereinafter referred to as the "x direction").
[0035] In the image sensor 107 of the camera 100 according to the first embodiment, the pixel pitch P of the imaging pixels is 4 μm, and the number of imaging pixels N is approximately 20.75 million pixels (5575 columns in the x direction × 3725 rows in the vertical direction (hereinafter referred to as the "y direction"). The pixel pitch PAF of the focus detection pixels is 2 μm, and the number of focus detection pixels NAF is approximately 41.5 million pixels (11150 columns in the x direction × 3725 rows in the y direction).
[0036] In the first embodiment, a case is described in which each imaging pixel is divided into two in the horizontal direction, but each imaging pixel may also be divided in the vertical direction. Furthermore, the image sensor 107 in the first embodiment has a plurality of imaging pixels, each of which includes a first focus detection pixel 201 and a second focus detection pixel 202, but the imaging pixels and the first and second focus detection pixels may be provided as separate pixels. For example, the first and second focus detection pixels may be discretely arranged among the plurality of imaging pixels.
[0037] Fig. 3(A) shows one imaging pixel (G pixel 200G in the figure) as viewed from the light receiving surface side (+z direction) of the image sensor 107 of the first embodiment. Fig. 3(B) shows the aa cross section of the imaging pixel in Fig. 3(A) as viewed from the -y direction.
[0038] 3B, one imaging pixel is provided with one microlens 305 for collecting incident light. The imaging pixel is also provided with a photoelectric conversion unit 301 and a photoelectric conversion unit 302 that are divided into N sections in the x direction (divided into two sections in the first embodiment). The photoelectric conversion unit 301 and the photoelectric conversion unit 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively. The centers of gravity of the photoelectric conversion unit 301 and the photoelectric conversion unit 302 are decentered in the -x direction and the +x direction, respectively, with respect to the optical axis of the microlens 305.
[0039] An R, G, or B color filter 306 is provided between the microlens 305 and the photoelectric conversion unit 301 or 302 in each imaging pixel. The spectral transmittance of the color filter may be changed for each photoelectric conversion unit, or the color filter may be omitted.
[0040] Light incident on the imaging pixel from the imaging optical system is collected by the microlens 305, dispersed by the color filter 306, received by the photoelectric conversion unit 301 and the photoelectric conversion unit 302, and then photoelectrically converted.
[0041] Next, the relationship between the structure of the imaging pixel shown in Figures 3(A) and 3(B) and pupil division will be described with reference to Figure 4. Figure 4 shows the aa cross section of the imaging pixel shown in Figure 3(A) as viewed from the +y direction, and also shows the exit pupil of the imaging optical system. In Figure 4, the x and y directions of the imaging pixel are reversed compared to Figure 3(B) to correspond to the coordinate axes of the exit pupil.
[0042] As shown in FIG. 4 , within the exit pupil, a first pupil region 501, whose center of gravity is decentered toward the +X direction, is a region that is made approximately conjugate with the light receiving surface of the photoelectric conversion unit 301 of the imaging pixel on the −x direction side by the microlens 305. The light beam that passes through the first pupil region 501 is received by the photoelectric conversion unit 301, i.e., the first focus detection pixel 201. Furthermore, a second pupil region 502, whose center of gravity is decentered toward the −X direction, is a region that is made approximately conjugate with the light receiving surface of the photoelectric conversion unit 302 of the imaging pixel on the +x direction side by the microlens 305. The light beam that passes through the second pupil region 502 is received by the photoelectric conversion unit 302, i.e., the second focus detection pixel 202. A pupil region 500 indicates the pupil region that can receive light from the entire imaging pixels, which are a combination of the photoelectric conversion units 301 and 302 (first focus detection pixel 201 and second focus detection pixel 202).
[0043] Next, pupil division by the image sensor will be described with reference to FIG. 5. FIG. 5 shows pupil division by the image sensor 107. As shown in FIG. 5, a pair of light beams that pass through a first pupil region 501 and a second pupil region 502, respectively, are incident on each imaging pixel of the image sensor 107 at different angles and are received by the two divided first focus detection pixels 201 and second focus detection pixels 202. In the first embodiment, output signals from the first focus detection pixels 201 of the multiple imaging pixels of the image sensor 107 are collected to generate a first focus detection signal, and output signals from the second focus detection pixels 202 of the multiple imaging pixels of the image sensor 107 are collected to generate a second focus detection signal. In addition, the output signals from the first focus detection pixels 201 and the output signals from the second focus detection pixels 202 of the multiple imaging pixels are added to generate an imaging pixel signal. The imaging pixel signals from the multiple imaging pixels are then combined to generate an imaging signal for generating an image with a resolution corresponding to the number of effective pixels N (number of imaging pixels N).
[0044] Next, the relationship between the defocus amount of the imaging optical system and the phase difference (hereinafter referred to as "image shift amount") between the first focus detection signal and the second focus detection signal acquired from the image sensor 107 will be described with reference to Fig. 6. In Fig. 6, the image sensor 107 is arranged on the imaging plane 600, and the exit pupil of the imaging optical system is divided into a first pupil region 501 and a second pupil region 502, as described with reference to Figs.
[0045] As shown in FIG. 6, the defocus amount d is defined such that |d| is the distance (size) from the imaging position C of the light beam from the subject (801, 802) to the imaging plane 600, and a front-focus state in which the imaging position C is closer to the subject than the imaging plane 600 is represented by a negative sign (d<0). The defocus amount d is also defined such that a back-focus state in which the imaging position C is closer to the subject than the imaging plane 600 is represented by a positive sign (d>0). In the in-focus state in which the imaging position C is on the imaging plane 600, d=0. The imaging optical system is in-focus (d=0) with respect to the subject 801, and in a front-focus state (d<0) with respect to the subject 802. The front-focus state (d<0) and the back-focus state (d>0) are collectively referred to as a defocus state (|d|>0).
[0046] In a front-focus state (d<0), the light beam from the subject 802 that passes through the first pupil region 501 (second pupil region 502) is first focused and then spreads to a width Γ1 (Γ2) centered at the center of gravity G1 (G2) of the light beam, forming a blurred image on the imaging surface 600. This blurred image is received by each first focus detection pixel 201 (each second focus detection pixel 202) on the image sensor 107, and a first focus detection signal (second focus detection signal) is generated. In other words, the first focus detection signal (second focus detection signal) is a signal that represents an image of the subject 802 at the center of gravity G1 (G2) of the light beam on the imaging surface 600, blurred by the blur width Γ1 (Γ2).
[0047] The blur width Γ1 (Γ2) of the subject image increases roughly in proportion to an increase in the magnitude |d| of the defocus amount d between the first focus detection signal and the second focus detection signal. Similarly, the magnitude |p| of the image shift amount p (= the difference G1 - G2 in the center of gravity position of the light beam) of the subject image between the first focus detection signal and the second focus detection signal also increases roughly in proportion to an increase in the magnitude |d| of the defocus amount d. In the back-focus state (d>0), the direction of the image shift of the subject image between the first focus detection signal and the second focus detection signal is opposite to that in the front-focus state, but the same is true.
[0048] In this way, as the magnitude of the defocus amount increases, the magnitude of the image shift amount of the subject image between the first focus detection signal and the second focus detection signal increases. In the first embodiment, "focus detection using an image sensor phase difference detection method" is performed, in which the defocus amount is calculated from the image shift amount of the subject image between the first focus detection signal and the second focus detection signal obtained using the image sensor 107.
[0049] Next, with reference to Fig. 7, the focus detection area of the image sensor 107 that acquires the first focus detection signal and the second focus detection signal will be described. In Fig. 7, A(n, m) indicates the nth focus detection area in the x direction and the mth focus detection area in the y direction out of the multiple focus detection areas (three in the x direction and three in the y direction, for a total of nine) set in the effective pixel area 1000 of the image sensor 107. The first focus detection signal and the second focus detection signal are generated from output signals from the multiple first focus detection pixels 201 and second focus detection pixels 202 included in the focus detection area A(n, m). I(n, m) indicates an index that displays the position of the focus detection area A(n, m) on the display 131.
[0050] Note that the nine focus detection areas shown in FIG. 7 are merely an example, and in the present invention, the number, positions, and sizes of the focus detection areas are not limited to the example in FIG. 7. For example, one or more areas may be set as focus detection areas within a predetermined range centered on a position specified by the user or the position of the subject detected by the subject detection unit 140 (hereinafter also referred to as the "subject position"). In the first embodiment, the focus detection areas are arranged so that higher-resolution focus detection results can be obtained when acquiring a defocus map, which will be described later. For example, a total of 9,600 focus detection areas are arranged on the image sensor 107, divided into 120 horizontal and 80 vertical areas.
[0051] Next, the flow of live view shooting in the imaging device (camera 100) according to the first embodiment will be described. Fig. 8 is a flowchart showing the flow of live view shooting in the camera 100 according to the first embodiment. Specifically, Fig. 8 shows the processing that causes the camera 100 to perform operations from before capturing an image to display a live view image on the display 131 to capturing a still image. The CPU 121 executes the processing of Fig. 8 according to a computer program. In the following description, S means step.
[0052] First, in S1, the CPU 121 causes the image sensor drive circuit 124 to drive the image sensor 107 and acquires image data from the image sensor 107. Thereafter, the CPU 121 acquires, from the acquired image data, first focus detection signals and second focus detection signals from a plurality of first focus detection pixels and second focus detection pixels included in each of the focus detection areas shown in FIG. 7. The CPU 121 also adds the first focus detection signals and second focus detection signals of all effective pixels of the image sensor 107 to generate an image signal, and causes the image processing circuit 125 to perform image processing on the image signal (image data) to acquire image data. Note that if the image sensor pixels and the first focus detection signal and second focus detection pixels are provided separately, the CPU 121 acquires image data by performing interpolation processing on the focus detection pixels.
[0053] Next, in S2, the CPU 121 causes the image processing circuit 125 to generate a live view image from the image data obtained in S1, and causes the generated live view image to be displayed on the display 131. Note that the live view image is a reduced image matched to the resolution of the display 131, and the user can adjust the image capture composition, exposure conditions, etc. while viewing the live view image. Therefore, the CPU 121 performs exposure adjustment based on the photometric value obtained from the image data, and displays it on the display 131. Exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the shooting lens, and adjusting the gain of the image sensor output.
[0054] Next, in S3, the CPU 121 determines whether a switch Sw1 (hereinafter simply referred to as "Sw1"), which instructs the start of an image capture preparation operation, has been turned on by half-pressing a release switch included in the operation switch group 132. If the CPU 121 determines in S3 that Sw1 is not turned on, it repeats the determination made in S3 to monitor the timing at which Sw1 will be turned on. On the other hand, if the CPU 121 determines in S3 that Sw1 is turned on, it advances the process to S400 and performs subject tracking AF processing (subject tracking autofocus processing). The subject tracking AF processing includes detecting a subject area from the obtained imaging signal and focus detection signal, detecting a focus area, setting a focus detection area, and performing predictive AF processing to suppress the influence of the time lag from the focus detection timing to the image exposure timing. Details of the "subject tracking AF processing" that causes the camera 100 to perform subject tracking AF operation will be described later.
[0055] After performing subject tracking AF processing, the CPU 121 proceeds to S5, where it determines whether or not a switch Sw2 (hereinafter simply referred to as "Sw2") that instructs the start of an imaging operation has been turned on by fully pressing the release switch. If the CPU 121 determines in S5 that Sw2 is not turned on, it returns the process to S3. On the other hand, if the CPU 121 determines in S5 that Sw2 is turned on, it proceeds to S300, where it executes imaging processing. Details of the "imaging processing" that causes the camera 100 to perform an imaging operation will be described later. When the imaging processing ends, the CPU 121 proceeds to S7.
[0056] In S7, the CPU 121 determines whether or not a main switch included in the operation switch group 132 has been turned off. If the CPU 121 determines in S7 that the main switch has been turned off, the live view shooting ends. On the other hand, if the CPU 121 determines in S7 that the main switch has not been turned on, the process returns to S3.
[0057] In the first embodiment, the subject tracking AF process is performed after it is detected in S3 that Sw1 is turned on (after it is determined that Sw1 is turned on), but the timing for performing the subject tracking AF process is not limited to this. By performing the subject tracking AF process in S400 before Sw1 is turned on, it is possible to eliminate the need for the photographer to take preparatory actions before shooting.
[0058] Next, a description will be given of the flow of the imaging process executed by the CPU 121 in S300 of Fig. 8. Fig. 9 is a flowchart showing the flow of the imaging process executed by the CPU 121 in S300 of Fig. 8.
[0059] In S301, the CPU 121 performs exposure control processing to determine imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.). This exposure control processing can be performed using brightness information acquired from image data of a live view image. Then, in S301, the CPU 121 transmits the determined aperture value to the aperture drive circuit 128 to drive the aperture 102. Also in S301, the CPU 121 transmits the determined shutter speed to the shutter 108 to open the focal plane shutter. Furthermore, in S301, the CPU 121 causes the image sensor 107 to accumulate charge during the exposure period via the image sensor drive circuit 124.
[0060] In S302, the CPU 121, which has performed the exposure control processing, causes the image sensor drive circuit 124 to read out all pixels of the image sensor 107 image pickup signals for capturing a still image. The CPU 121 also causes the image sensor drive circuit 124 to read out one of the first focus detection signal and the second focus detection signal from the focus detection area (focus target area) within the image sensor 107. The first focus detection signal or the second focus detection signal read out at this time is used to detect the focus state of the image during image playback, which will be described later. By subtracting one of the first focus detection signal and the second focus detection signal from the image pickup signal, the other focus detection signal can be obtained.
[0061] Next, in S303, the CPU 121 causes the image processing circuit 125 to perform defective pixel correction processing on the imaging data that was read out and A / D converted in S302.
[0062] Furthermore, in S304, the CPU 121 causes the image processing circuit 125 to perform image processing and encoding processes such as demosaic processing, white balance adjustment processing, gamma correction processing (gradation correction processing), color conversion processing, and edge enhancement processing on the image data after the defective pixel correction processing.
[0063] Then, in S305, the CPU 121 records the still image data as image data obtained by performing image processing and encoding processing in S304 and one of the focus detection signals read out in S302 in the flash memory 133 as an image data file.
[0064] Next, in S306, CPU 121 associates the camera characteristic information (image capture device characteristic information) as characteristic information of camera 100 (image capture device) with the still image data recorded in S305 and records it in flash memory 133 and memory (RAM) within CPU 121. The camera characteristic information includes, for example, the following information: Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 125 Information about the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107 Information about vignetting of the imaging light beam within the camera 100 Information about the distance from the mounting surface of the imaging optical system of the camera 100 to the imaging element 107. Information about manufacturing errors of the camera 100.
[0065] Information regarding the light sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107 (hereinafter simply referred to as "light sensitivity distribution information") is information regarding the sensitivity of the image sensor 107 according to the distance (position) on the optical axis from the image sensor 107. This light sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, and therefore may be information regarding these. Furthermore, the light sensitivity distribution information may be information regarding changes in sensitivity with respect to the angle of incidence of light.
[0066] Next, in S307, the CPU 121 associates the lens characteristic information (photographing lens characteristic information) as characteristic information of the imaging optical system with the still image data recorded in S305 and records it in the flash memory 133 and the memory (RAM) within the CPU 121. The lens characteristic information includes, for example, the following information: Exit pupil information Information about the frame of the lens barrel, etc. that blocks the light beam -Information about focal length and F-number at the time of shooting - Information about aberrations in the imaging optical system -Information about manufacturing errors in imaging optical systems Information on the position of the focus lens 105 (subject distance) during imaging
[0067] Next, in S308, the CPU 121 records image-related information as information related to the still image data in the flash memory 133 and memory (RAM) within the CPU 121. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.
[0068] Next, in S309, the CPU 121 causes the display 131 to display a preview of the captured image. This allows the user to easily check the captured image. When the processing performed in S309 ends, the CPU 121 ends the imaging processing and proceeds to S7 in FIG. 8.
[0069] Next, a description will be given of the flow of the subject tracking AF process executed by the CPU 121 in S400 of Fig. 8. Fig. 10 is a flowchart showing the flow of the subject tracking AF process executed by the CPU 121 in S400 of Fig. 8.
[0070] In S401, the CPU 121 calculates the amount of image shift of the subject image between the first focus detection signal and the second focus detection signal obtained in each of the multiple focus detection areas obtained in S2, and calculates the defocus amount for each focus detection area from the calculated image shift amount. In this way, the CPU 121 obtains a defocus map by calculating the defocus amount for each focus detection area. As described above, in the first embodiment, the group of focus detection results obtained from the focus detection areas arranged on the image sensor 107, divided into 120 horizontally and 80 vertically, for a total of 9,600 points, is referred to as a defocus map.
[0071] Next, in S402, CPU 121 performs subject detection processing and tracking processing. The above-mentioned subject detection unit 140 performs subject detection processing to detect the subject area. In the subject detection processing, depending on the state of the obtained image, it may be impossible to detect the subject area. In such cases, CPU 121 performs tracking processing using other means such as template matching to estimate the position of the subject. Details of the subject detection processing and tracking processing will be described later.
[0072] Next, in S403, CPU 121 causes focus area detection unit 142 to perform focus area detection processing to detect a focus area. Details of the focus area detection processing will be described later. In the present invention, in S402, subject detection unit 140 (first detection means) performs subject area detection processing (subject detection processing), and in S403, focus area detection unit 142 (second detection means) performs focus area detection processing (focus area detection processing).
[0073] The difference between the subject area detection process and the focus area detection process will be explained below. In the subject area detection process, if the subject is a person, the face area and eye area of the person are detected as the subject area. In addition, if the subject is a vehicle such as a motorcycle, the subject area detection process detects the entire body area of the motorcycle and the helmet area of the driver operating the motorcycle as the subject area. In other words, in the subject area detection process, if the subject is a living thing, the entire body and organs of the living thing are detected, and if the subject is a non-living thing such as a vehicle, parts with a certain function of the non-living thing (for example, the tires of the vehicle, the handlebars of the vehicle, etc.) are detected.
[0074] On the other hand, in the focus area detection process, an area on which the photographer wants to focus (hereinafter referred to as "area to be focused") is detected as a focus area according to the shooting scene (information related to the shooting scene of the subject). For example, if the subject is a person, and the face is photographed relatively large, with a shallow depth of field, and the scene is facing diagonally forward, the focus area detection process detects the area of the eyelashes of the front eye (eyelash area) as the focus area. Also, if the subject is a person, and the face is photographed relatively large, with a shallow depth of field, and the scene is with one eye closed, the focus area detection process detects the area of the open pupil as the focus area. In either shooting scene, the subject area detection process (subject detection process) performed in S402 detects the area of the pupil as the subject area.
[0075] In the first embodiment, the eyelash area is detected as a focus area in an area different from the pupil area, but there are cases where the gaps between the eyelashes are large and focus detection cannot be performed properly. In such cases, the display is performed in the eyelash area, but the focus adjustment may be performed by adding a pre-registered offset amount to the result obtained in the pupil area.
[0076] Similarly, when photographing a motorcycle road race, the subject is often the motorcycle and its driver. When the motorcycle is cornering in a direction toward the photographer, the motorcycle's body leans toward the photographer (toward the viewer). When the motorcycle is cornering in a direction away from the photographer, the motorcycle's body leans away from the photographer (toward the viewer). In such a shooting environment, the area the photographer wants to focus on may be the driver's organs, which are living organisms, or the motorcycle's parts, which are inanimate, depending on the shooting scene, and is not uniquely determined. For example, in a shooting scene where the motorcycle's body is leaning toward the viewer, the area the photographer wants to focus on will be the helmet, and in a shooting scene where the motorcycle's body is leaning toward the viewer, the area the photographer wants to focus on will be the engine or the body near the gas tank. This is because, when capturing an image with a relatively shallow depth of field, if the area of the subject that is in focus is too far back, it will appear unnatural.
[0077] In S402, a specific region of the subject is fixedly detected, and the orientation of the subject is also detected. In S403, a region that the photographer wants to focus on is statistically detected based on the orientation of the subject (e.g., the tilt direction of the motorcycle body), the shooting environment (e.g., the subject size, shallow depth), and the background environment. The region detected in S402 is a first local region (a region corresponding to at least a part of the subject region), and the region detected in S403 is a second local region (a region corresponding to at least a part of the subject region). Also, in S402, the subject detection unit 140 detects, as the subject region (first local region), a region that indicates subject characteristics, such as the entire body or organs of the person when the subject is a person, or parts of the vehicle when the subject is a vehicle. In S403, the focus region detection unit 142 detects, as the focus region (second local region), a region that indicates shooting scene characteristics, such as the subject's pattern, subject size, depth, and tilt direction. The region that indicates shooting scene characteristics is also a region that corresponds to the characteristics of the focus target.
[0078] Next, in S404, the CPU 121 performs a focus detection area setting process to set a focus detection area (area to be focused) using the information on the subject detection area obtained in S402 and the information on the focus area as the area to be focused obtained in S403. In S404, the CPU 121 functions as a local area selection unit. The focus detection area setting process will be described in detail later.
[0079] Next, in S405, the CPU 121 acquires the focus detection result (defocus amount) of the focus detection area set in the focus detection area setting process of S404. The focus detection result acquired in S405 may be selected from the focus detection results calculated in S401 (defocus map acquired in S401) that are closest to the desired area. Also, the focus detection result acquired in S405 may be used to newly calculate the defocus amount using a focus detection signal corresponding to the set focus detection area. Also, the focus detection area for calculating the defocus amount is not limited to one, and multiple focus detection areas may be arranged around the area and used for calculation.
[0080] Next, in S406, the CPU 121 performs predictive AF processing using the defocus amount obtained in S405 and the defocus amount obtained in the past. The predictive AF processing is processing required when there is a time lag between the timing of focus detection and the timing of image exposure, and is processing that predicts the position of the subject a predetermined time after the timing of focus detection and performs AF control. The details of the predictive AF processing will be described later.
[0081] When the predictive AF process performed in S406 ends, the CPU 121 ends the subject tracking AF process and proceeds to S5 in FIG.
[0082] Next, a description will be given of the subject detection process and tracking process executed by the CPU 121 in S402 of Fig. 10. Fig. 11 is a flowchart showing the flow of the subject detection process and tracking process executed by the CPU 121 in S402 of Fig. 10.
[0083] In S2000, the CPU 121 sets dictionary data according to the type of subject to be detected from the data detected from the image data acquired in S2. Specifically, in S2000, dictionary data to be used in the subject detection process and tracking process is selected (set) from multiple dictionary data stored in the dictionary data storage unit 141 based on the subject priority set in advance and the settings of the camera 100 (image capture device). For example, multiple dictionary data are stored by classifying subjects into categories such as "people," "vehicles," and "animals." In the first embodiment, one or more dictionary data may be selected. When one dictionary data is selected, subjects that can be detected using one dictionary data can be detected repeatedly at a high frequency. On the other hand, when multiple dictionary data are selected, the dictionary data can be set sequentially according to the priority of the detected subject, thereby allowing subjects to be detected sequentially.
[0084] Next, in S2001, subject detection unit 140 performs subject detection using the dictionary data set in S2000 and the image data read in S2 as an input image. At this time, subject detection unit 140 outputs information such as the position, size, and reliability of the detected subject as subject detection region information. At this time, CPU 121 may cause display 131 to display the subject detection region information output by subject detection unit 140. Also, in S2001, subject detection unit 140 hierarchically detects multiple regions of the subject as subject detection regions from the image data. For example, if "person" or "animal" is set as the dictionary data in S2000, subject detection unit 140 hierarchically detects multiple regions, such as a "whole body" region, a "face" region, and an "eye" region, as subject detection regions. The detected "whole body" region is the overall region indicating the whole body of the subject, and the detected "face" region and "eye" region are local regions indicating the organs of the subject. While local regions such as the "face" region or "eye" region of a person or animal are regions on which it is desired to focus as a subject, they may not be detectable due to surrounding obstacles or the orientation of the face. In the present invention, even in such cases, subject detection unit 140 is configured to detect the subject hierarchically in order to continue robustly detecting the subject by detecting the entire body. Similarly, when "vehicle" is set as dictionary data in S2000, subject detection unit 140 hierarchically detects the entire region including the vehicle driver and the vehicle body and the region of the driver's helmet (driver's head) as a local region as the subject detection region. In the present invention, when "vehicle" is set as dictionary data, subject detection unit 140 is configured to detect the subject hierarchically by detecting the entire vehicle including the vehicle driver and the vehicle body.
[0085] Next, in S2002, the CPU 121 performs a known template matching process using the subject detection area obtained in S2001 as a template. Using the multiple images obtained in S2, a similar area is searched for in the most recently obtained image using the subject detection area obtained in a past image as a template. As is well known, any information may be used for template matching, such as brightness information, color histogram information, or feature point information such as corners and edges. Various matching methods and template update methods are possible, and any of these methods may be used. The tracking process performed in S2002 is performed to achieve stable subject detection and tracking processes by detecting an area similar to the past subject detection data from the most recently obtained image data if a subject is not detected in S2001.
[0086] When the tracking process performed in S2002 ends, the CPU 121 ends the subject detection process and tracking process, and the process proceeds to S403 in FIG.
[0087] Next, a description will be given of the focus area detection process executed by the CPU 121 in S403 of Fig. 10. Fig. 12 is a flowchart showing the flow of the focus area detection process executed by the CPU 121 in S403 of Fig. 10.
[0088] In S3000, CPU 121 determines whether to perform focus area detection processing. As described above, a focus area is an area within a subject that should be in focus, and focus area detection processing is processing for detecting an area (focus area) different from the subject detection area detected by the subject detection processing described in FIG. 11. Therefore, if it is inappropriate or impossible to detect an area to be in focus within the subject, focus area detection processing is skipped. Focus area detection processing is skipped when the size of the subject area detected in S402 is smaller than a predetermined size or when the depth difference within the subject in the shooting settings or live view settings is smaller than a predetermined value. In these cases, the difference in focus state within the subject area (the difference between the in-focus area and the out-of-focus area) is difficult to visually recognize, so focus area detection processing is skipped.
[0089] If the size of the subject area is smaller than a predetermined size, it becomes difficult to visually recognize the difference in focus state within the subject area. Therefore, in the first embodiment, in S3000, the CPU 121 determines to skip the focus area detection process if the size of the subject area is smaller than a predetermined size.
[0090] As is well known, the depth difference within a subject is determined by the distance to the subject and the aperture diameter of the photographing optical system. The farther the subject distance is and the smaller the aperture diameter, the deeper the depth becomes, and the wider the area within the subject area that is in an acceptable blur (in focus). In other words, the wider the area within the subject area that is in depth. This makes it difficult to visually recognize the difference in focus within the subject area. Therefore, in the first embodiment, at S3000, the CPU 121 determines to skip the focus area detection process if the depth difference within the subject area is smaller than a predetermined value.
[0091] As described above, in S3000, if the CPU 121 determines not to perform the focus area detection process (i.e., if it determines to skip the focus area detection process), it ends the focus area detection process and proceeds to S404 in Figure 10.
[0092] On the other hand, if CPU 121 determines in S3000 to perform focus area detection processing, it advances the process to S3001 and acquires a signal of the subject area. That is, in S3001, CPU 121 acquires image data of all subject detection areas, including the overall area and local areas, hierarchically detected by subject detection unit 140. As described above, the overall area of the subject is the area of the entire body of the subject if the subject is a human or animal, or the area including the vehicle and its driver if the subject is a vehicle such as a motorcycle. Subject detection unit 140, acting as a third detection means, performs the subject detection processing and tracking processing described in S402 based on the image data obtained in S2, and outputs the detection result of the overall area detected as the subject detection area (a signal of the overall area). If there are multiple subject areas (subject detection areas) detected by subject detection unit 140, the focus area detection processing performed in S3002 is performed multiple times.
[0093] Next, in S3002, CPU 121 causes focus area detection unit 142 to perform focus area detection processing to detect a focus area. As described above, focus area detection unit 142, in response to an instruction from CPU 121, detects an area to be focused as a focus area based on the state of the subject in the subject area (subject detection area) detected by subject detection unit 140. In the focus area detection processing, only one area or multiple areas may be detected as the focus area. When multiple areas are detected as focus areas, the imaging device (camera 100) automatically selects the detected multiple areas, or the photographer selects the detected multiple areas, thereby appropriately setting the area to be focused on. At this time, CPU 121 may cause display 131 to display the focus area information output by focus area detection unit 142.
[0094] When the focus area detection process performed in S3002 (if there are multiple subject detection areas, multiple focus area detection processes) is completed, the CPU 121 ends the focus area detection process and proceeds to S404 in FIG.
[0095] Next, a description will be given of the focus detection area setting process executed by the CPU 121 in S404 of Fig. 10. Fig. 13 is a flowchart showing the flow of the focus detection area setting process executed by the CPU 121 in S404 of Fig. 10.
[0096] In S4000, CPU 121 acquires information such as the position, size, and reliability of the subject as information on the subject detection area obtained as the output of the subject detection process and tracking process performed in S402. Next, in S4001, CPU 121 acquires information such as the position, size, and reliability of the focus area as information on the focus area obtained as the output of the focus area detection process performed in S403.
[0097] Next, in S4002, the CPU 121 sets the focus detection area using the subject detection area information obtained in S4000 and the focus area information obtained in S4001. The focus detection area can be set by selecting a focus detection result that is highly reliable and indicates a subject that is relatively close, based on the results of the subject detection area and the focus area within the area set as the focus area. Alternatively, the focus detection area can be set by relocating the focus detection area within the obtained subject detection area and the area set as the focus area, acquiring image data and focus detection signals again, and similarly selecting the focus detection result.
[0098] The following methods can be used to select the area to be used to set the focus detection area from the subject detection area and focus area. When only the subject detection area or the focus area is detected, the detected area is set as the focus detection area. When neither the subject detection area nor the focus area is detected, the focus detection area is set in the same position as the previous focus detection area. When both the subject detection area and the focus area are detected, the focus area takes priority over the subject detection area, and the focus area is set as the focus detection area. Furthermore, when both the subject detection area and the focus area are detected, the subject detection area may be set as the focus detection area, or the focus area may be set as the focus detection area, depending on information about the subject shooting scene. The camera 100 (image capture device) may be configured to display the set focus detection area on the display 131. The camera 100 (image capture device) may be configured to display the subject detection area, focus area, and focus detection area separately or selectively.
[0099] When the setting of the focus detection area performed in S4002 is completed, the CPU 121 ends the focus detection area setting process and proceeds to S405 in FIG.
[0100] Next, the focus area detected by the focus area detection process performed in S403 will be described with reference to Figures 14(A) to 14(D) and 15. Figures 14(A) to 14(D) show examples of scenes that a photographer may want to capture when the subjects are a motorcycle and its driver.
[0101] FIG. 14(A) shows a motorcycle and its driver traveling with the direction of travel being the direction they are approaching the camera 100 (image capture device). FIG. 14(B) shows the motorcycle and its driver approaching the front and about to corner to the left as seen from the driver's perspective. FIG. 14(C) shows a scene in which the motorcycle is about to corner to the left as seen from the driver's perspective, captured from the side. FIG. 14(D) shows a scene in which the motorcycle is about to corner to the right as seen from the driver's perspective, captured from the side. FIGS. 14(A) to 14(D) show an entire region 900 and a local region 901 as the subject region (subject detection region) detected in S402. Similarly, FIGS. 14(A) to 14(D) show local regions 902 and 903 as the focus regions detected in S403.
[0102] 14(A) and 14(B) show how the driver's head is detected as local region 901 for subject detection, and how local region 902 and local region 903 are detected as focus regions. The reason why multiple regions (local region 902 and local region 903) are detected as focus regions is because an image focused on either region may be desired depending on the photographer's preference or intention. In S4002, when setting the focus detection region, CPU 121 determines the priority of the focus region in consideration of the settings of the imaging device, such as driver priority or close-range priority, the position of the detection region within the image in the shooting range, and continuity with the previous focus detection region setting. For example, if CPU 121 determines that close-range priority is to be selected, it sets focus region 903 as the focus detection region.
[0103] FIG. 14(C) shows how the driver's head is detected as local region 901 for subject detection, and local region 902 is detected as the focus region. In the captured scene of FIG. 14(C), the body of the motorcycle is tilted away from camera 100 (image capture device). As a result, there is a depth difference between local region 902, which is the body of the motorcycle, and subject detection region 901. In such a situation, close-range priority is given to images in which the focus is on local region 902, and this is often preferred. Therefore, in the case of the captured scene of FIG. 14(C), the focus region detection process of the present invention does not detect local region 901 of the head as the focus region, but detects local region 902 of the body.
[0104] In the shooting scene in Fig. 14(D), the body of the motorcycle is tilted toward the front, so it is close to the head area, which is an important organ as a subject, and the area to be focused on coincides with it. Therefore, Fig. 14(D) shows how local area 901 detected by subject detection and local area 902 detected by focus area detection overlap.
[0105] Note that the subject detection area and the focus area have been described using Figures 14(A) to 14(D) in which the subject detection area encompasses the focus area, but in the present invention, the size relationship between the areas is not limited to this.
[0106] In this way, in the present invention, not only are important organs such as the head and pupils detected as being important during photography, but by detecting the focus area as an area different from the subject detection area depending on the photography scene, it is possible to perform focus adjustment that is more in line with the photographer's intentions.
[0107] There are various possible shooting scenes in which focus area detection is effective. Figure 15 shows main examples of scenes in which focus area detection is effective.
[0108] As shown in FIG. 15, when taking a portrait of a person, if the right side of the face is closer to the camera 100 (image capture device), the focus is generally set on the right eye of the person. Therefore, the right eye is detected as the important organ during subject detection, and the right eye is also detected during focus area detection. By detecting the focus area of the present invention (the focus area detection process performed in S403 of FIG. 10), the eyelash area of the right eye is detected as the focus area when the person's face is large and the scene is shallow. This allows for an image with high contrast of the eyelashes, making it easier to see the focus state.
[0109] Furthermore, when the subject is a "motorcycle," as explained in Figures 14(A) to 14(D), when photographing from the front in the direction of travel, both subject detection and focus area detection detect the helmet, and focus is adjusted to the helmet. The same is true when the vehicle body is tilted forward. On the other hand, when the vehicle body is tilted backward, subject detection detects the helmet (head) as a vital organ, but detects the body area near the engine as the focus area.
[0110] Furthermore, when the subject is a car (e.g., a race car such as an F1 car), in a scene shot from a slightly elevated position in front of the car's direction of travel, subject detection detects the helmet (head) as a vital organ. However, in the above-described shooting scene, the focus area is detected as a position forward of the driver's seat to ensure the entire car body is within the depth of field. This is because, when shooting a race car such as an F1 car from above and in front of it, the car body has depth, and focusing on the helmet (head) results in the front of the car body being out of the depth of field and becoming blurred. The CNN (hereinafter referred to as the "CNN for focus area detection") constituting the focus area detection unit 142 achieves the above detection by setting and learning a focus area for each shooting direction of the race car image during machine learning. In such a shooting scene, it is conceivable to output necessary depth information as the output of focus area detection. It is sufficient to output aperture value information of the shooting optical system to ensure the entire car body is within the depth of field. By setting a focus detection area based on the focus area detection and setting the aperture value as necessary, an image in which the entire car, extending in the depth direction, is captured within the depth of field can be obtained.
[0111] When the subject is a car (for example, an F1 race car) and the scene is shot from the side in the direction of travel, subject detection detects the helmet (head) as a vital organ, but the side of the car body is detected as the focus area. This is because the helmet (head) is located further back than the car body, for the same reason as when the car body is tilted towards the back of a motorcycle.
[0112] Next, the difference between the machine learning of the focus area detection CNN for realizing focus area detection and the machine learning of the CNN (hereinafter referred to as "subject detection CNN") constituting the subject detection unit 140 for realizing subject detection will be described.
[0113] This section describes the collection of a group of images annotated with training data for focus area detection. First, subject detection is performed on the collected image group for images annotated with training data used in machine learning, and images in which the desired subject is detected are extracted. Images with depth differences within the detected subject area are extracted using the contrast distribution within the subject area and corresponding defocus map information. Training data is then added to the extracted images. Depth differences are determined to exist when the subject area contains areas with high and low contrast, or when the defocus map information indicates areas with low and high defocus amounts. On the other hand, when there is no depth difference within the subject area (when there is no contrast difference or the distribution of defocus amounts is within a predetermined value), the image data is used as a negative sample for learning. By using this machine learning method, a CNN for focus area detection can be realized that performs focus area detection when there is a depth difference within the subject detection area, but does not perform focus area detection when there is no depth difference.
[0114] Although teacher data can be assigned to the extracted images while checking each image, teacher data can be assigned automatically if areas with high contrast or small defocus amounts are known within the subject area. After automatically assigning teacher data, the teacher data can also be manually fine-tuned.
[0115] Data augmentation can be performed to efficiently collect training data. Well-known methods include translation, scaling, rotation, noise addition, and blurring. In the present invention, as a data augmentation method effective for the focus detection area, a method of blurring areas other than the focus area is used, rather than a well-known method of blurring the entire image or the entire subject area. This allows image data corresponding to different depth differences within the subject area to be obtained from an image to which a single piece of training data is added. Furthermore, by varying the degree of blurring for each image, the blur state of areas other than the focus area that serve as training data can be varied, enabling robust learning even when images are captured with different aperture diameters of the imaging optical system. Examples of blurring methods include a method of increasing the blur depending on the distance from the focus detection area, or a method of varying the presence or absence of blur between the focus detection area and other areas, thereby appropriately processing the boundary area. The degree of blurring of areas other than the focus area may also be set depending on the change in aperture diameter. This makes it possible to generate training data that is close to the images that are actually captured.
[0116] Furthermore, in the first embodiment, image data is input to the CNN for focus region detection, but the input data for the CNN for focus region detection is not limited to image data. By inputting information from which depth can be inferred, such as a contrast map or a defocus map, in addition to image data, to the CNN for focus region detection, the focus region can be detected more appropriately. In this case, when performing machine learning on the CNN for focus region detection, a contrast map or a defocus map may be prepared in addition to the image data, and learning may be performed.
[0117] Next, a description will be given of the predictive AF process executed by the CPU 121 in S406 of Fig. 10. Fig. 16 is a flowchart showing the flow of the predictive AF process executed by the CPU 121 in S406 of Fig. 10.
[0118] In S6000, the CPU 121 determines whether the subject is a moving object moving in the optical axis direction. Specifically, the CPU 121 determines whether the subject is moving in the optical axis direction by referring to time-series data of past defocus detection results and determining whether adjacent differences between multiple pieces of time-series data have the same sign. In S6000, if the CPU 121 determines that the subject is a moving object moving in the optical axis direction, it advances the process to S6001. On the other hand, in S6000, if the CPU 121 determines that the subject is not a moving object moving in the optical axis direction, it advances the process to S6012.
[0119] In S6001, the CPU 121 calculates the traveling direction of the subject detected in the latest image data. The orientation of the subject has already been detected by the subject detection process and tracking process performed in S402. Methods for calculating the traveling direction of the subject include a method that uses local detection within the subject (for example, the face or eyes) and a method that uses the posture detection result of the subject.
[0120] First, a method using local detection within a subject (for example, a face or eyes) will be described. Local detection within a subject involves detecting the eyes, head, and torso of a person as a local detection area when the subject is a person. When the subject is a vehicle such as a motorcycle, the local detection area includes the head (helmet) of the driver of the vehicle. A known method will be described for a case where the subject is only a person. When the local detection area is the pupil, and both pupils are detected, the direction of travel of the subject is taken as the optical axis direction. When only the right pupil is detected, the direction of travel of the subject is taken as the rightward direction. When only the left pupil is detected, the direction of travel of the subject is taken as the leftward direction. When the subject is not only a person (for example, when a motorcycle is also included), the position of the pupils may be unknown, or the detected pupils may differ from the direction of travel of the subject. In this invention, in such cases, the direction of travel of the detected subject is estimated from the size of a rectangular frame indicating the entire range of the detected subject and the positional relationship of the local detection area within the detected subject relative to the entire range of the detected subject. Here, a case where the subject is a motorcycle and its driver will be described as an example. If the aspect ratio of the detection range of the entire subject is long and short and the local detection area (in this example, the driver's helmet) is located at the top of the entire range of the detected subject, the subject is considered to be facing forward, and the direction of travel of the detected subject is the optical axis direction. If the aspect ratio of the detection range of the entire subject is short and long and the local detection area is located at the top right of the entire range of the detected subject, the direction of travel of the detected subject is considered to be to the right. As described above, the direction of travel of the detected subject can be calculated from the aspect ratio of the entire range of the detected subject and the positional relationship between the range of the detected subject and the local detection area.
[0121] Even with the above-described method, there is a possibility that the traveling direction of the detected subject cannot be calculated when the traveling direction of the detected subject suddenly changes (for example, when the detected subject suddenly moves upward from a state where it is approaching the optical axis direction, such as when it jumps). In such cases, it is necessary to estimate the traveling direction of the detected subject before the detected subject changes its traveling direction.
[0122] A method of using the results of posture detection of a subject as an estimation method for estimating the direction of travel of a detected subject will be described. There are various methods for detecting the posture of a subject, but in the first embodiment, the subject's joints are first estimated from an image using a deep-learned neural network. The posture information of the subject is detected by connecting the estimated joints. The direction of travel of each subject may be learned in advance, or the direction of travel may be estimated from the amount of movement of each joint between frames. Furthermore, characteristic pre-movements (e.g., movements before a jump) before a change in the direction of travel may be learned in advance. Furthermore, the direction of travel may be estimated in combination with local detection within the subject (e.g., face, eyes, etc.). When the subject is only a person, the direction of travel is estimated from posture information obtained by detecting the joints of the arms and legs before a jump. For example, it is estimated that the direction of travel will change due to the arms being lowered or both legs being bent (e.g., the posture before a jump). Even when the subject is not only a person (e.g., when a motorcycle is also included), the direction of travel of the detected subject is estimated from the relative positions of the joints of the person's arms and legs. When a motorcycle is included, as shown in Figure 14(D), if the right foot is detected, the joint is bent, and the hip and spinal joints are detected, it can be estimated that the detected subject is moving in the direction of the optical axis and to the right. Also, when a motorcycle is included, the moving direction of the detected subject can be estimated by detecting the tilt of the tires and handlebars.
[0123] The estimation of whether the light is traveling in the optical axis direction may be detected by estimating the subject position from the amount of defocus.
[0124] In S6002, CPU 121 predicts the future traveling direction of the subject (predicts the future traveling direction of the subject). Specifically, CPU 121 predicts the traveling direction of the subject from time-series changes in the calculation results of the traveling direction of the subject in past frames. The future traveling direction of the subject may be estimated from the amount of time-series change between frames in the aspect ratio of the detection range of the entire subject (hereinafter simply referred to as the "aspect ratio") and the positional relationship between the detection range of the entire subject and the local detection area (hereinafter simply referred to as the "positional relationship of the subject area"). Here, an example will be described in which the subject is a motorcycle and its driver. If the aspect ratio changes from being vertically long to being horizontally long and the positional relationship of the local detection area changes to the upper right with respect to the detection range of the entire subject, the traveling direction has changed from approaching the optical axis to the right, and it can be estimated that the traveling direction has changed to the right. In this example, the local detection area is the driver's helmet.
[0125] Furthermore, when predicting the direction of travel of a subject by estimating the subject's posture, for example, the movements before a jump can be estimated from posture information based on the joints of the subject's arms and legs and time-series changes in the joints, and it can be predicted that the direction of travel of the detected subject will change upward.
[0126] As described above, in S6002, CPU 121 predicts the future traveling direction of the detected subject based on time-series changes in the calculation results of the traveling direction of the detected subject obtained from multiple past frames. Prediction of the future traveling direction of the detected subject will be described using FIGS. 18(A) to 18(F). FIG. 18(A) is a diagram showing the traveling direction of the detected subject with an arrow. The downward direction is the optical axis direction, and the right side is the rightward direction. When the traveling direction of the detected subject gradually changes as in FIG. 18(A), the traveling direction is calculated from the aspect ratio of the entire range of the detected subject and its positional relationship with local regions, and the next traveling direction of the detected subject is predicted based on the time-series changes in the traveling direction. When the traveling direction of the detected subject suddenly changes as in FIG. 18(C), the subject's posture estimation described above is used to detect the posture of the subject before the sudden change in traveling direction, and the detected subject's next traveling direction is predicted.
[0127] In S6003, the CPU 121 determines whether the image plane velocity of the subject is large. The image plane velocity of the subject is calculated from time-series changes in the image plane position of the subject. In S6003, if the CPU 121 determines that the image plane velocity of the subject is large, the process proceeds to S6004. On the other hand, in S6003, if the CPU 121 determines that the image plane velocity of the subject is not large, the process proceeds to S6012. In S6004, the CPU 121 determines whether there has been a change in the traveling direction of the subject, and if it determines that there has been a change in the traveling direction, the process proceeds to S6005, and if it determines that there has been no change in the traveling direction, the process proceeds to S6008.
[0128] In S6005, the CPU 121 changes the number of historical data items used in the prediction calculation. Specifically, the CPU 121 changes the number of historical data items used for the data on the subject position calculated from the defocus amount and focus position of past frames used when predicting the subject position. FIG. 17 illustrates an example of time-series changes in the subject's image plane position. In FIG. 17, the horizontal axis represents time, the vertical axis represents the amount of movement of the subject on the image plane, the black dots represent historical data on the subject's image plane position based on the results of focus detection, and the dotted line represents a prediction curve obtained by prediction processing. The historical data refers to information on the subject's position on the image plane (subject's image plane position) and the time acquired in the past. This will be explained using the conceptual diagrams of FIGS. 18(A) to 18(F). FIG. 18(A) is a diagram showing the subject's traveling direction with an arrow. FIG. 18(B) corresponds to FIG. 18(A). In Figure 18(B), the horizontal axis represents time, the vertical axis represents the image plane position of the subject, the solid line represents the trajectory of the subject, the black circle represents the image plane position of the subject at the time of focus detection, and the dotted line represents the focus movable range. Figure 18(A) assumes that the camera is shooting from below, with the up-down direction corresponding to the optical axis direction. The figure shows the subject's travel direction estimated by the subject travel direction calculation performed in S6001 and the subject travel direction prediction performed in S6002. Figure 18(A) illustrates an example in which the subject's travel direction approaches the optical axis direction and then changes to the right midway. In Figure 18(B), the time range in which the travel direction is along the optical axis is designated 18-b1, and the time range in which the travel direction includes the rightward direction is designated 18-b2. During the time range 18-b1, the travel direction does not change and the image plane velocity does not change significantly, so the number of historical data points used in the prediction calculation is not changed. On the other hand, in the time range of 18-b2, the direction of travel changes and the image plane velocity also changes, so by reducing the number of historical data items used when the direction of travel is in the optical axis direction, the number of historical data items used in the predictive calculation is reduced, thereby reducing errors in the predictive calculation described below.
[0129] 18(C) and 18(D) show examples in which the subject's traveling direction differs. FIG. 18(C) shows an example in which the subject's traveling direction suddenly changes from the optical axis direction to the right. FIG. 18(D) corresponds to FIG. 18(C). In FIG. 18(D), the horizontal axis represents time, and the vertical axis represents the subject's image plane position. The time range in which the subject's traveling direction is along the optical axis direction is designated 18-d1, and the time range in which the subject's traveling direction is to the right is designated 18-d2. In the examples of FIGS. 18(C) and 18(D), a sudden change in the subject's traveling direction is estimated by subject traveling direction prediction. If the subject's traveling direction changes only to the right from the optical axis direction, the number of historical data used in the prediction calculation is reset immediately before the direction change, so that it is not used. This prevents erroneous predictions that the subject's traveling direction is along the optical axis direction, even if the subject's traveling direction suddenly changes to the right.
[0130] Furthermore, Figures 18(E) and 18(F) show examples in which the subject's traveling direction changes differently. Figure 18(E) is a diagram showing an example in which the subject's traveling direction alternates between the optical axis direction and the right direction, and between the optical axis direction and the left direction. Figure 18(F) corresponds to Figure 18(E). In Figure 18(F), the horizontal axis represents time, and the vertical axis represents the subject's image plane position. In the examples of Figures 18(E) and 18(F), the subject's traveling direction changes, but the subject's image plane velocity does not change, so the number of history data items used in the prediction calculation is not changed. As described above, by changing the number of history data items used in the prediction calculation depending on the subject's traveling direction, the prediction error in the subject's image plane position can be reduced.
[0131] In S6006, the CPU 121 sets the focus movable range. The focus movable range will be explained using FIG. 19. In FIG. 19, the horizontal axis represents time, the vertical axis represents the image plane position of the subject, the solid line represents the focus position, and the dotted line represents the focus movable range. When the subject is moving, if the photographer accidentally moves the AF frame away from the subject and the subject becomes the background, it will take time for the focus to return to the subject if the focus is moved. In the present invention, the range in which the subject will move is estimated from the image plane movement speed of the subject and the subject distance, and the focus movable range is set based on the estimated range in which the subject will move. Therefore, if the subject moves outside the focus movable range, the focus will not move.
[0132] This makes it possible to prevent the subject from suddenly becoming out of focus even if the AF frame accidentally becomes the background due to framing, etc. In the first embodiment, the focus movable range was described, but the focus may not move when the subject is outside the focus movable range, that is, the focus stop time may be changed.
[0133] The setting of the subject's travel direction and the focus movable range will be explained using Figures 18(A) to 18(F). The dotted lines in Figures 18(B), 18(D), and 18(F) indicate the focus movable range. In Figure 18(B), when the subject's travel direction changes from the optical axis direction to the right, the subject is not moving in the optical axis direction, so the focus movable range is set smaller than when moving in the optical axis direction. In Figure 18(D), when the subject's travel direction changes from the optical axis direction to the right, or just before the change, the focus movable range is set smaller than when moving in the optical axis direction.
[0134] In S6007, the CPU 121 changes the focus detection area. Specifically, the CPU 121 changes the focus detection area by predicting the subject's traveling direction in the above-described S6002, so as to widen the focus detection area or move the center of gravity of the focus detection area in the subject's traveling direction other than the optical axis direction. This makes it possible to prevent the subject from moving outside the focus detection area even if the subject's traveling direction changes.
[0135] In S6008, CPU 121 calculates the predicted image plane position of the subject. Specifically, CPU 121 performs multivariate analysis (for example, the least squares method) using historical data of past image plane positions of the subject and time, and predicts the image plane position of the subject by finding an equation for a prediction curve. CPU 121 also calculates the predicted image plane position of the subject by substituting the time of still image capture into the equation for the found prediction curve.
[0136] In S6009, the CPU 121 changes the focus movement speed (image plane movement speed of the focus). Specifically, the CPU 121 changes the image plane movement speed of the focus by estimating the image plane movement speed of the subject from the subject travel direction prediction result obtained in S6002, the predicted image plane position of the subject (predicted image plane position of the subject) obtained in S6008, and history data. In the examples of FIGS. 18(A) and 18(B) described above, the CPU 121 estimates that the image plane movement speed of the subject is decreasing, so changes the image plane movement speed of the focus by decreasing the image plane movement speed of the focus. Also, in the examples of FIGS. 18(C) and 18(D) described above, the CPU 121 changes the image plane movement speed of the focus by setting the image plane movement speed of the focus to 0 because the travel direction of the subject suddenly changes and the subject does not move in the optical axis direction. In this way, in S6009, the CPU 121 changes the image plane movement speed of the focus by setting the image plane movement speed of the focus in accordance with the travel direction of the subject.
[0137] In S6010, the CPU 121 determines whether the subject is within the focus movement range set in S6006, and if it determines that the subject is within the focus movement range, the process proceeds to S6011. On the other hand, if the CPU 121 determines in S6010 that the subject is not within the focus movement range (i.e., outside the focus movement range), the focus is not moved and the predictive AF process ends. In S6011, the CPU 121 moves the focus to an image plane position (predicted image plane position of the subject) corresponding to the predicted subject position (moves the focus lens 105 to the predicted image plane position of the subject), and ends the predictive AF process. In S6012, the CPU 121 moves the focus to an image plane position (image plane position of the subject) corresponding to the subject position calculated from the focus detection result (defocus amount) (moves the focus lens 105 to the image plane position of the subject), and ends the predictive AF process.
[0138] FIG. 20 shows items that are changed during predictive calculation and focus control according to the traveling direction of the subject described in the first embodiment. During predictive calculation and focus control, CPU 121 changes items shown in FIG. 20 such as the "number of history data items to use in predictive calculation," "focus movable range," "focus movement speed," and "focus detection area" according to the traveling direction of the subject. Values for these items shown in FIG. 20 are stored in the ROM of CPU 121 in the imaging device (camera 100) as a table corresponding to changes in the traveling direction of the detected subject. CPU 121 changes the values of these items by referencing the table stored in ROM according to changes in the traveling direction of the detected subject.
[0139] In the present invention, the direction of travel of a detected subject is estimated using the size, positional relationship, and aspect ratio of the detected subject area. Furthermore, in the present invention, future changes in the direction of travel of a detected subject are predicted using information on the size, positional relationship, and aspect ratio of the subject area that change over time. These technical configurations of the present invention reduce the time lag between focus detection timing and image exposure timing, compared to when a similar process is performed using focus detection results, resulting in more accurate photographic results.
[0140] For example, if one were to use only the focus detection results (defocus amount) to estimate that a subject will approach and then move away, it would be difficult to estimate whether the subject will remain stationary or reverse course and move away once it has stopped. However, the present invention uses the subject's image as a shooting scene characteristic to distinguish between multiple images, estimate the direction of travel of the detected subject, and also predict the future direction of travel, thereby achieving less time lag and more accurate shooting results. For example, the multiple images are an image of an "approaching motorcycle," a "sideways motorcycle," and a "moving away motorcycle."
[0141] As described above, in the first embodiment, when a focus area is detected, the focus area is selected as the focus detection area in which focus detection is performed preferentially over the subject detection area, but in the present invention, the method of setting the focus detection area is not limited to this.
[0142] For example, as a method of setting the focus detection area, a mode (second mode) that prioritizes the subject detection area and a mode (first mode) that prioritizes the focus area can be provided in CPU 121 as modes that can be set by the photographer. Specifically, in the first mode, both the subject detection area (first local area) and the focus area (second local area) can be selected as the focus detection area, and the focus area is preferentially set as the focus detection area. In the second mode, the subject detection area (first local area) is set as the focus detection area. In this way, by providing CPU 121 as local area selection means with the second mode that prioritizes the subject detection area and the first mode that prioritizes the focus area, it is possible to easily reflect the photographer's intention regarding the area of the subject that is to be focused on.
[0143] As described above, in the first embodiment, a configuration has been described in which focus area detection is realized by area detection based on machine learning (focus area detection unit 142 configured by machine-learned CNN and detecting focus areas). However, in the present invention, the configuration for realizing focus area detection is not limited to the configuration described in the first embodiment.
[0144] For example, in the present invention, a focus area can be set using information such as the aspect ratio of the subject detection area, the size of the subject detection area, and depth information of the subject using a defocus map (hereinafter referred to as "subject detection information"). When the subject is a person, if the size of the subject detection area is equal to or larger than a predetermined size, the position of the eyelashes can be estimated relative to the pupil area detected as the subject detection area, and the estimated eyelash area can be set as the focus area. When the subject is a motorcycle, the defocus map can be used to detect the tilt direction of the motorcycle body, and the focus area can be switched between setting the focus area on the head or setting the focus area on an area estimated to be the body position of the subject detection area corresponding to the entire motorcycle. Similarly, the distance of the subject can be determined using defocus information of the motorcycle body and defocus information of the helmet (head) as the subject detection area. Furthermore, the aspect ratio of the subject detection area can be used to determine whether the body of a vehicle such as a motorcycle or a car is detected from the front or the side, and then the focus area can be set.
[0145] As described above, by using a configuration that sets the focus area using subject detection information, there is no need to prepare a circuit (CNN for focus area detection) that performs focus area detection using CNN within the imaging device, and focus area detection can be achieved at low cost.
[0146] Second Embodiment Next, a second embodiment of the present invention will be described with reference to Figures 21, 22(A), 22(B), and 23(A) to 23(D). The configuration of the second embodiment is similar to that of the first embodiment, except for the focus detection area setting process performed in the second embodiment, which differs from the focus detection area setting process performed in S4000 to S4002 of the first embodiment. In the focus detection area setting process performed in the second embodiment, a focus detection area is selected from candidate focus detection areas within a detected subject.
[0147] Below, a description of the configuration of the second embodiment, which is the same as that of the first embodiment, will be omitted. Fig. 21 is a flowchart showing the flow of focus detection area setting processing according to the second embodiment of the present invention.
[0148] In S4100, the CPU 121 displays a local region (including a focus region) from the entire subject (entire detected subject) detected in the acquired image. FIGS. 22(A) and 22(B) show display examples of the local region (including a focus region). FIGS. 22(A) and 22(B) show an example in which a motorcycle is detected as a detected subject. In FIGS. 22(A) and 22(B), the region surrounded by dotted lines indicates the region that is a focus detection candidate for the detected subject, and the region surrounded by solid lines indicates the focus detection candidate region. In FIGS. 22(A) and 22(B), the helmet region, headlight region, body logo region, and muffler region surrounded by dotted lines are regions that are focus detection candidate for the detected subject.
[0149] In S4101, the CPU 121, which serves as a designation unit, selects (designates) a focus detection candidate area from a local area (including the focus area) of the detected subject. The user designates a focus detection candidate area from the local area (including the focus area) shown in FIGS. 22(A) and 22(B). In FIG. 22(A), the focus detection candidate area is the helmet area, and in FIG. 22(B), the focus detection candidate area is the body logo area. The focus detection candidate area is designated by changing and selecting the focus detection candidate area by touch operation with the user's finger or button operation. The focus detection candidate area may also be selected by the user's line of sight. The focus detection candidate area may be changed not only when Sw1 is turned on, but also before Sw1 is turned on or during continuous shooting when Sw2 is turned on. Alternatively, instead of the pre-display of the local area as performed in S4100, after the designated position is determined by a touch operation with the user's finger or a button operation, a local area close to the designated position may be selected on the camera 100 side, and the selected local area may be used as a focus detection candidate area.
[0150] As another method, a method for selecting (specifying) a focus detection candidate area from a menu operation within the camera 100 will be described with reference to FIGS. 23(A) to 23(D). FIG. 23(A) shows a setting screen for detecting a detected subject. In FIG. 23(A), a detected subject is selected from detectable subjects (for example, vehicles, animals, and people). (In the example of FIG. 23(A), a vehicle is selected as the detected subject.) In FIG. 23(B), a type of vehicle is further selected from the selected detected subject, a vehicle. (In the example of FIG. 23(B), a motorcycle is selected as the type of vehicle.) In FIG. 23(C), the orientation of the motorcycle, which is the detected subject of the selected type, is specified. (In the example of FIG. 23(C), rightward orientation is selected.) In FIG. 23(D), a focus detection candidate area is selected within the detected subject of the selected type. (In the example of FIG. 23(D), the body logo portion is selected as the focus detection candidate area.) In FIGS. 23(C) and 23(D), a different focus detection candidate area may be selected from a local area (including the focus area) depending on the orientation of the detected subject. As described above, focus detection candidate areas may be selected by menu operation. Alternatively, an image of each detected subject may be selected, and local areas that serve as focus detection candidate areas may be displayed and selected. Furthermore, a three-dimensional model of each detected subject may be displayed as an image, and the three-dimensional model may be made rotatable by user operation, allowing different focus detection candidate areas to be specified depending on the orientation or posture of the detected subject.
[0151] The items that are candidates for local regions (including focus regions) may be changed depending on the orientation of the detected subject. For example, if the front is selected as the orientation of the motorcycle in Figure 23(C), only the helmet and headlight may be displayed as focus detection candidate regions, and if the rightward orientation is selected, the helmet, headlight, body logo, and muffler may be displayed as shown in Figure 23(D).
[0152] Furthermore, the user may set items that are candidates for local regions (including focus regions) depending on the orientation of the detected subject.
[0153] In S4102, the CPU 121 records the focus detection candidate areas designated by the designation means in S4101. The recording may be performed by storing in memory within the CPU 121 the positional relationship on the screen of the focus detection candidate areas designated from the local areas (including the focus area) and the positional relationship in the optical axis direction based on the defocus amount of each local area (including the focus area).
[0154] In S4103, the CPU 121 determines whether the focus detection candidate area specified by the specifying means in S4101 is selectable, and if it determines that the focus detection candidate area is selectable, the process proceeds to S4104. On the other hand, if in S4103 the CPU 121 determines that the focus detection candidate area is not selectable or if the focus detection candidate area is not selected, the process proceeds to S4105. Hereinafter, the determination conditions required by the CPU 121 to determine whether a focus detection candidate area is selectable will be simply referred to as the "determination conditions." The determination conditions include a case where "the specified focus detection candidate area becomes invisible and cannot be detected due to a change in the direction of travel or posture of the detected subject" or a case where "the specified focus detection candidate area is detected, but the posture of the subject has changed at a time different from that specified by the user." Whether the posture of the subject has changed may be determined by comparing the degree of agreement between the posture and direction of travel of the subject at the time of recording in S4102 and the positional relationship with each local area. Specifically, the XY direction on the screen is taken as the XY direction, and the optical axis direction is taken as the Z direction, and the correlation between the magnitude and orientation of the XYZ direction vector at the time of recording is determined. If there is a correlation, it is determined that the posture has not changed, and if there is no correlation, it is determined that the posture has changed. The correlation method calculates the dot product of the vectors to find the angles of the two vectors. If each angle is less than a predetermined value, it is determined that there is a correlation, and if each angle is equal to or greater than a predetermined value, it is determined that there is no correlation. As described in Figures 23(A) to 23(D), if a focus detection candidate area corresponding to the orientation of the detected subject is selected, it is determined that the focus detection candidate area is selectable. On the other hand, if a focus detection candidate area corresponding to the orientation of the detected subject is not selected, if the focus detection candidate area is not visible, or if it is not detected as a focus area according to the shooting scene, it is determined that the focus detection candidate area is not selectable.
[0155] In S4104, the CPU 121 sets the focus detection candidate area specified by the specification means as the focus detection area. When the setting of the focus detection area performed in S4104 is completed, the CPU 121 ends the focus detection area setting process and proceeds to S405 in Fig. 10. In S4105, the CPU 121 sets the local area automatically selected by the camera 100 as the focus detection area, rather than the focus detection candidate area specified by the specification means. When the setting of the focus detection area performed in S4105 is completed, the CPU 121 ends the focus detection area setting process and proceeds to S405 in Fig. 10. The automatic selection method by the camera 100 may be, for example, selecting the local area closest to each local area, or a local area that is prioritized in advance based on the shape, posture, or direction of travel of the subject.
[0156] Although preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the gist of the present invention. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions. [Explanation of symbols]
[0157] 100 cameras 107 Image sensor 121 CPU 126 Focus drive circuit 140 Subject detection unit 141 Dictionary data storage unit 142 Focus area detection unit
Claims
1. detection means for detecting a plurality of types of vehicles; an acquisition means for acquiring information regarding at least one of a photographing direction relative to a vehicle or a tilt direction of the vehicle; a setting means for setting a region to be focused; The image processing device is characterized in that the setting means changes the area to be focused in accordance with the type of vehicle detected by the detection means and the information acquired by the acquisition means.
2. 2. The image processing apparatus according to claim 1, wherein the types of vehicles detected by said detection means include motorcycles and cars.
3. 3. The image processing apparatus according to claim 2, wherein said detecting means detects the area of the head of a driver of a vehicle as the subject area.
4. 3. The image processing apparatus according to claim 2, wherein said detecting means detects a body region of a vehicle as the subject region.
5. The image processing device according to claim 3, characterized in that, when the subject is a motorcycle, the setting means sets the area of the driver's head as the area to be focused in a scene photographed from the front with respect to the direction of travel or in a scene photographed in which the motorcycle body is tilted forward.
6. 5. The image processing device according to claim 4, wherein when the subject is a motorcycle, in a shooting scene in which the motorcycle body is tilted toward the rear, the setting means sets the area of the body as the area to be focused.
7. The image processing device according to claim 2, characterized in that the setting means sets the area to be focused so that, when the subject is a car, the scene is photographed from the front in the direction of travel, and when the scene is photographed from above, the entire car is included within the depth of field.
8. 3. The image processing device according to claim 2, wherein when the subject is a car, the setting means sets the area of the side of the car body as the area to be focused in a scene photographed from the side in the direction of travel.
9. 2. The image processing device according to claim 1, wherein the setting unit sets the area to be focused based on items that are candidates for the area to be focused, according to the orientation of the detected subject that has been set in advance by a user.
10. The image processing device according to claim 1 , wherein the setting unit outputs a region to be focused on for a captured image using dictionary data generated based on machine learning.
11. 2. The image processing apparatus according to claim 1, further comprising a display control means for controlling the display unit to display an indicator indicating the region to be focused.
12. a detection step of detecting a plurality of types of vehicles; an acquisition step of acquiring information regarding at least one of a photographing direction relative to the vehicle or a tilt direction of the vehicle; a setting step of setting a region to be focused, An image processing method, wherein the setting step changes the area to be focused on depending on the type of vehicle and the information.
13. 12. A program for causing a computer to execute each unit of the image processing apparatus according to claim 1.
Citation Information
Patent Citations
Image processing system, image processor, image processing method and image processing program
JP2007172271A
Imaging apparatus
JP2012123301A
Camera and camera system
JP2016080904A
Imaging device
JP2017009942A