Image processing apparatus, imaging apparatus, image processing method and method for controlling imaging apparatus, and program
Patent Information
- Application Number
- JP2022137540
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-09-03
AI Technical Summary
Existing imaging technologies struggle to capture images with both a moving subject and a stationary subject in the same frame without causing motion blur in the stationary subject.
An image processing device that includes an image acquisition unit, region extraction means, recognition means, and information acquisition means to generate photographing parameters that control motion blur in specific regions, allowing for a sense of motion in moving areas while suppressing blur in stationary areas.
Captures images with a sense of motion in moving subjects while effectively suppressing motion blur in stationary subjects, enhancing the dynamic feel of the image.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an imaging device, and more particularly to a technique for capturing an image with a sense of motion in a shooting scene in which areas in which a subject is moving and areas in which a subject is still are mixed. [Background technology]
[0002] Conventionally, there is a method of capturing images with an imaging device such as a digital still camera, in which a long shutter speed is used to capture a moving subject. In a portrait, a person may be captured with a waterfall or fountain in the background, or in a sports scene, the movement of the person's face may be frozen, while the moving limbs may be left hanging to create a sense of motion. In this way, there are cases in which a moving subject and a subject whose movement should be frozen are mixed in one image. Patent Document 1 proposes a method of capturing multiple images continuously at a short shutter speed, aligning the main subject area contained in the multiple images, and synthesizing the multiple images. By synthesizing images captured at a short shutter speed, it is possible to suppress blurring of the subject area that should be frozen. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2006-339903 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, while the method in Patent Document 1 can suppress blurring if the subject is a rigid body whose shape changes little even when it moves, there is an issue with subjects whose shape changes, in that the shapes do not match even after alignment, resulting in blurring.
[0005] Therefore, the present invention aims to enable shooting in a scene in which a single image contains both moving subjects and subjects whose movement is to be frozen, while adding a sense of motion to the moving areas and suppressing motion blur in areas where movement is to be frozen. [Means for solving the problem]
[0006] In order to achieve the above-mentioned object, the image processing device of the present invention comprises an image acquisition means for acquiring a plurality of images captured successively by a first shooting, an area extraction means for extracting a first area and a second area different from the first area from the acquired images, a recognition means for recognizing a shooting scene from the plurality of images, an information acquisition means for acquiring motion information of the first area and the second area from the plurality of images, and a shooting parameter generation means for generating shooting parameters for a second shooting different from the first shooting, wherein the shooting parameter generation means generates the shooting parameters based on the shooting scene and the motion information so that the amount of motion blur in the first area of the image acquired by the second shooting is equal to or greater than a first value. Effect of the Invention
[0007] According to the present invention, it is possible to capture an image with a sense of motion in an area where motion is present, and to suppress motion blur in an area where motion is desired to be stopped. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram for explaining the configuration of a digital camera according to a first embodiment. [Diagram 2] 4 is a diagram for explaining the configuration of an imaging control parameter generating unit in the first embodiment. [Diagram 3] 5 is a flowchart for explaining the operation of a shooting control parameter generating unit in the first embodiment. [Figure 4] FIG. 4 is a diagram for explaining images captured continuously in the first embodiment. [Diagram 5]6A to 6C are diagrams for explaining the effect of superimposing a scene recognition result on a display image in the first embodiment. [Figure 6] 1 is a flowchart for explaining a motion vector calculation method in the prior art. [Figure 7] FIG. 1 is a diagram for explaining a method of calculating a motion vector in the prior art. [Figure 8] 4 is a diagram in which a motion vector is superimposed on an image at time t in the first embodiment. [Figure 9] 5A to 5C are diagrams for explaining a correction gain applied to a motion vector in the first embodiment. [Figure 10] FIG. 4 is a diagram for explaining a defocus map corresponding to an image at time t in the first embodiment. [Figure 11] FIG. 4 is a diagram for explaining an actual captured image in the first embodiment. [Figure 12] FIG. 11 is a diagram for explaining the configuration of an image synthesis processing unit in the second embodiment. [Figure 13] FIG. 11 is a diagram for explaining a synthesis map in the second embodiment. [Figure 14] 10 is a flowchart for explaining a method of determining a composite characteristic in the second embodiment. [Figure 15] 13A to 13C are diagrams for explaining a short-second reference frame and a synthesis order in the second embodiment. [Figure 16] 11 is a flowchart for explaining a method of determining a short second reference frame in the second embodiment. [Figure 17] 11A to 11C are diagrams for explaining the state of movement of a main subject and objects in the second embodiment. [Figure 18] 13A to 13C are diagrams for explaining changes over time in the moving speed of a main subject in the second embodiment. [Figure 19] 13A to 13C are diagrams for explaining a motion amount and a correction ratio of a synthesis mask in the second embodiment. [Figure 20] FIG. 11 is a diagram for explaining the configuration of each processing unit of image capture in the third embodiment. [Figure 21] 11 is a flowchart for explaining an operation during shooting in the third embodiment. [Figure 22] FIG. 11 is a diagram for explaining an exposure period etc. in Example 3. [Figure 23] 11A and 11B are diagrams for explaining an image when a face image is blurred. [Figure 24] FIG. 11 is a diagram for explaining details of an image when a face image is blurred. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] [Example 1] An example of an image processing device according to an embodiment of the present invention will be described in detail below with reference to the drawings. Note that the embodiment described below is an imaging device, and an example in which the present invention is applied to a digital camera as an example of an imaging device will be described. The image processing device referred to in the present invention can also be applied to any electronic device capable of processing captured images. These electronic devices may include, for example, information terminals such as mobile phones, game consoles, tablet terminals, personal computers, watch-type or eyeglass-type devices, and head-mounted displays.
[0010] 1 is a block diagram showing the functional configuration of a digital camera according to an embodiment of the present invention. A control unit 101 is, for example, a CPU, which reads out an operation program for each block of the digital camera 100 from a ROM 102, expands it into a RAM 103, and executes it to control the operation of each block of the digital camera 100. The ROM 102 is a rewritable non-volatile memory, and stores parameters and the like required for the operation of each block in addition to the operation program for each block of the digital camera 100. The RAM 103 is a rewritable volatile memory, and is used as a temporary storage area for data output in the operation of each block of the digital camera 100.
[0011] The optical system 104 forms an image of a subject on the imaging unit 105. The optical system 104 includes, for example, a fixed lens, a variable magnification lens that changes the focal length, and a focus lens that adjusts the focus. The optical system 104 also includes an aperture, which adjusts the aperture diameter of the optical system to adjust the amount of light during shooting. The imaging unit 105 is, for example, an imaging element such as a CCD or CMOS sensor, and performs photoelectric conversion on the optical image formed on the imaging element by the optical system 104, and outputs the obtained analog image signal to the A / D conversion unit 106. The A / D conversion unit 106 applies A / D conversion processing to the input analog image signal, and outputs the obtained digital image data to the RAM 103 for storage.
[0012] The image processing unit 107 applies various image processing such as white balance adjustment, color interpolation, and gamma processing to the image data stored in the RAM 103, and outputs the image data to the RAM 103. The image processing unit 107 also includes a shooting control parameter generating unit 200 (described later), and can generate shooting parameters for the digital camera 100 based on the scene recognition results, motion analysis results, and subject movement direction estimation results for the image data stored in the RAM 103. The shooting parameters generated by the image processing unit 107 are output to the control unit 101, and the control unit 101 controls the operation of each block included in the digital camera 100.
[0013] The recording medium 108 is a removable memory card or the like, and images processed by the image processing unit 107 stored in the RAM 103 and images A / D converted by the A / D conversion unit 106 are recorded as recorded images.
[0014] The display unit 109 is a display device such as an LCD, and performs an electronic viewfinder function by successively updating and displaying image data output in time series via the imaging unit 105. The display unit 109 can also present various types of information in the digital camera 100, such as playing back and displaying images recorded on the recording medium 108. In addition, the display unit 109 can also display, for example, icons based on the scene recognition results of the image data in the image processing unit 107, superimposed on the image.
[0015] Operation input unit 110 includes a user input interface such as a release switch, a setting button, a mode setting dial, etc., and when it detects an operation input made by a user, it outputs a control signal corresponding to the operation input to control unit 101. In addition, in an aspect in which display unit 109 is equipped with a touch panel sensor, operation input unit 110 also functions as an interface that detects a touch operation made to display unit 109.
[0016] The configuration and basic operation of digital camera 100 have been described above.
[0017] Next, the operation of the image processing unit 107, which is a feature of the first embodiment of the present invention, will be described in detail. In the first embodiment of the present invention, an example is described in which a dynamic image is captured of a subject performing a golf swing, with motion blur suppressed in the face area and motion blur in the swinging area. In the following embodiment, golf is used as an example of a sport that involves a swinging motion, but the present invention can also be applied to other sports. Specific examples include, but are not limited to, tennis, badminton, table tennis, baseball, lacrosse, hockey, fencing, kendo, canoeing, and boating. In the first embodiment of the present invention, shooting parameters can be automatically generated such that motion blur occurs in the swinging sports equipment in each sport, and a dynamic image can be captured.
[0018] First, an example of the configuration of the image processing unit 107 will be described with reference to Fig. 2. The image processing unit 107 is composed of a main subject detection unit 201, an area extraction unit 202, a scene recognition unit 203, a motion vector calculation unit 204, a motion vector correction unit 205, a depth change detection unit 206, and a shooting parameter configuration unit 207. Each unit except the shooting parameter generation unit 207 is responsible for at least one of the processes of acquiring feature information (feature information acquisition) and correcting, which will be described later. The image processing unit 107 inputs image data captured by the imaging unit 105 and recorded in the RAM 103, and outputs a scene recognition result 209 and shooting parameters 210.
[0019] Next, the processing of the shooting control parameter generating unit 200 will be described with reference to the flowchart in Fig. 3. The following processing is realized by the control unit 101 controlling each unit of the digital camera 100 in accordance with a program stored in the ROM 102.
[0020] In step S301, the user turns on the digital camera 100. In response to the power being turned on for the digital camera 100, the control unit 101 controls the optical system 104 and the imaging unit 105 to start preparatory shooting. Here, preparatory shooting refers to shooting in which the composition is adjusted and the shooting conditions are set while looking at the electronic viewfinder or rear liquid crystal of the imaging device before the actual shooting. During this preparatory shooting period, the digital camera 100 captures images one after another, and the captured images are displayed on the display device of the display unit 109. The user can prepare for shooting by adjusting the composition and changing shooting parameters such as the exposure time (Tv value), aperture value (Av value), and ISO sensitivity during the actual shooting while looking at the images during the preparatory shooting that are displayed one after another. Here, the actual shooting refers to shooting that is triggered by an action such as the photographer pressing the shutter button, and is executed by the imaging device based on the composition and shooting conditions set in the preparatory shooting. The frame rate in this embodiment is, for example, 120 frames per second. In other words, the imaging unit 105 captures one image every 1 / 120 seconds. Moreover, the shutter speed, which is one of the shooting parameters of the digital camera 100, is set to be as short as possible at this time. An example of images captured continuously is shown in FIG. 4. An image at time t is 401, and an image at time t+1 is 402. FIG. 4 also shows a situation in which a person making a golf swing is being photographed. FIG. 4 shows a person subject 403 and a bird 404 that is not the intended subject but is flying farther away than the person subject 403. The images at each time in FIG. 4 are captured with a short exposure time (sufficiently fast shutter speed), so there is no significant blurring in the captured images. For the sake of explanation, the image size here is 2100×1400 pixels, and the pixel size is 5 μm.
[0021] In step S302, under the control of the control unit 101, the main subject detection unit 201 acquires the image data captured in S301 (image acquisition) and detects the area of the main subject in the image data (main subject area). The main subject detection unit 201 may use a known technique for detecting the area of the main subject, for example, a method such as that disclosed in Patent Document 2 may be used. In addition, by referring to a defocus map, which is distance distribution information described later, it is possible to extract the area of the body (torso, whole body), hands and feet, and even the golf club held, which are present at a depth similar to that of the face. In addition, the main subject detection unit can detect the areas of multiple subjects in the image data, and can determine the main subject area from among them. Here, it is assumed that the main subject detection unit 201 detects the main subject area in the image shown in FIG. 4 as image data, and an area corresponding to the human subject 403 is detected.
[0022] In step S303, under the control of control unit 101, area extraction unit 202 extracts a face area and an area of a swinging object (golf club) from human subject 403 detected by main subject detection unit 201. Known techniques may be used to extract the face area and the golf club area, and in Patent Document 2, each area can be extracted by changing the object to be detected.
[0023] In step S304, the scene recognition unit 203 recognizes the scene captured in the image data under the control of the control unit 101. Here, the type of sport is recognized. A known technique may be used to recognize the type of sport being played in the input scene, and for example, a technique such as that disclosed in Non-Patent Document 1 may be used. It is possible to recognize that not only golf but also various sports such as tennis and baseball are being played. In addition, the scene recognition unit 203 is also capable of outputting whether the recognition of the type of sport was successful or unsuccessful. In addition, a configuration may be adopted in which the detection information of the main subject detected in S302 is used. At this time, even if spectators present in the surroundings are captured, the recognition accuracy can be improved by excluding them from the scene recognition target.
[0024] Furthermore, the scene recognition unit 203 can estimate the direction and range of movement of the golf club from the image data acquired in S301. For the estimation, the recognition result of the type of sport, the positional relationship (facing direction) between the digital camera 100 and the human subject 403, and information on the position and attitude (direction) of the golf club estimated from the image data of the area of the golf club extracted in S303 are used. For example, assume that in the image, a subject playing golf faces the camera, the golf club is located to the left of the screen, and the face of the golf club faces the right of the screen. In this case, the scene recognition unit 203 can estimate the direction and range of movement of the golf club thereafter by setting a prediction result corresponding to the situation in advance, such as the golf club moving to the right along the bottom of the screen. Various scenes are expected during shooting, and the scene recognition unit 203 can respond by preparing settings according to various situations and selecting them. In addition, the method of estimating the position and attitude of the club may be, for example, a method disclosed in Non-Patent Document 2.
[0025] The scene recognition unit 203 compiles information such as whether recognition was successful, the type of recognized sport, and the estimated moving direction and range of a swinging object such as a golf club, and outputs the information as a scene recognition result.
[0026] On the other hand, the image processing unit 107 can superimpose the type of sport recognized according to the scene recognition result and the estimated movement direction and range of a swinging object such as a club on the image data. Under the control of the control unit 101, the image processing unit 107 can display the image with the superimposed information on the display unit 109, as shown in FIG.
[0027] In Fig. 5, 501 is an icon showing the recognized sport, 502 is the estimated movement direction of the club, and 503 is the hatched area showing the estimated movement range of the club. By displaying the scene recognition result on display unit 109 during preparation shooting, the user can confirm that digital camera 100 has recognized the swing sport. Note that the means of notifying the user is not limited to displaying an icon, and may be configured to notify the user by sound, light, or vibration, for example. In this case, digital camera 100 may be provided with a notification means (not shown) corresponding to each notification method.
[0028] In step S305, the control unit 101 checks whether the type of sport output from the scene recognition unit 203 has been recognized. If the recognition is successful, the process proceeds to step S307. On the other hand, if the recognition is unsuccessful, the process proceeds to step S306, where the shooting parameters for the normal sports mode are set, and the parameter determination process is completed. Specifically, the shutter speed is set to a high value at which it is assumed that blurring will hardly occur in sports scenes, and the aperture value and ISO sensitivity are set according to the brightness of the shooting environment.
[0029] In step S307, the motion vector calculation unit 204 calculates the motion vector and the motion vector reliability (motion information) between the images captured continuously under the control of the control unit 101. The motion vector is a vector that represents the horizontal and vertical movement amounts of the subject between the images. The calculation method of the motion vector will be described in detail with reference to Figs. 6, 7, and 8.
[0030] Fig. 6 is a flowchart showing the calculation process of a motion vector and a motion vector reliability by the motion vector calculation unit 204. Fig. 7 is a diagram showing a calculation method of a motion vector by a block matching method. Note that in this embodiment, the block matching method is taken as an example of the calculation method of a motion vector, but the calculation method of a motion vector is not limited to this example, and may be, for example, a gradient method. Each step of this flowchart is executed by the control unit 101 or each unit of the digital camera 100 including the motion vector calculation unit 204 in response to an instruction from the control unit 101.
[0031] In step S601, two captured images that are adjacent in time are input to the motion vector calculation unit 204. In this embodiment, the motion vector calculation unit 204 sets the captured image at time t in Fig. 4 as a base frame, and sets the captured image at time t+1 as a reference frame.
[0032] In step S602, the motion vector calculation unit 204 arranges a reference block 702 of N×N pixels in the reference frame 701 as shown in FIG. 7. In this embodiment, the area in which the reference block 702 is arranged is limited to the face area of the human subject 403 and the area of the golf club extracted by the area extraction unit 202 in step S303. By configuring in this way, it is possible to efficiently analyze only the motion information required to generate the shooting parameters. In particular, the correlation calculation in step S604, which will be described later, is a calculation content that has a large processing load. Therefore, by the motion vector calculation unit 204 performing the calculation only in the necessary area, it is possible to generate the shooting parameters at a higher speed, and it is possible to capture dynamic images in swing sports without missing them.
[0033] In step S603, motion vector calculation unit 204 sets (N+n)×(N+n) pixels around central coordinates 704 of standard block 702 in base frame 701 as search range 705 for reference frame 703, as shown in Fig. 7. As in step S602, search range 705 is set to the periphery of the face area and golf club area of human subject 403. In particular, the search area for the golf club is limited to the movement range of the club estimated by scene recognition unit 203 shown as 503 in Fig. 5. This configuration makes it possible to efficiently analyze only the motion information necessary to generate shooting parameters.
[0034] In step S604, the motion vector calculation unit 204 performs a correlation calculation between the base block 702 of the base frame 701 and the reference block 706 of N×N pixels at different coordinates in the search range 705 of the reference frame 703, and calculates a correlation value. The correlation value is calculated based on the inter-frame absolute difference sum for the pixels of the base block 702 and the reference block 706. In other words, the coordinate with the smallest inter-frame absolute difference sum value is the coordinate with the highest correlation value. Note that the method of calculating the correlation value is not limited to the method of calculating the inter-frame absolute difference sum, and may be, for example, a method of calculating a correlation value based on the inter-frame difference square sum or a normal cross-correlation value. In the example of FIG. 7, it is assumed that the reference block 706 shows the highest correlation.
[0035] In step S605, the motion vector calculation unit 204 calculates a motion vector based on the reference block coordinates showing the highest correlation value obtained in step S604, and sets the correlation value of the motion vector as the motion vector reliability. In the example of FIG. 7, the motion vector is calculated based on the coordinates 704 corresponding to the center coordinates of the base block 702 of the base frame 701 and the center coordinates of the reference block 706 in the search range 705 of the reference frame 703. That is, the coordinate distance and direction from the coordinates 704 to the center coordinates of the reference block 706 are calculated as the motion vector. In addition, the correlation value, which is the result of correlation calculation with the reference block 706 at the time of calculating the motion vector, is calculated as the motion vector reliability. The motion vector reliability increases as the correlation value between the base block and the reference block increases.
[0036] In step S606, the motion vector calculation unit 204 determines whether or not a motion vector has been calculated in the target location in the reference frame 701 where the reference block 702 should be placed, that is, in this embodiment, in the face area and golf club area of the human subject 403 extracted by the area extraction unit 202 in step S303. If the motion vector calculation unit 204 determines in step S606 that the motion vector has been calculated in all target locations, the motion vector calculation process ends. On the other hand, if the motion vector calculation unit 204 determines that the motion vector of the target location has not been calculated, the process returns to step S602 and the subsequent processes are repeated. In addition, the reference block 702 may be configured to be set for each pixel in all pixels included in the face area and golf club area, or may be configured to be set only for the representative pixel in each area. In this embodiment, the description will be given assuming that the reference block 702 is set only for the representative pixel.
[0037] The motion vectors between captured images calculated based on the above-described processing are shown in Fig. 8. In Fig. 8, the arrows indicate the motion vectors, the length of the arrows indicates the magnitude of the motion vectors, and the direction of the arrows indicates the direction of the motion vectors. 801 is the motion vector of the face region of human subject 403, and 802 is the motion vector of the golf club region. It can be seen that the amount of movement between frames is greater than that of the face region because the golf club is swung at high speed. Meanwhile, dashed line 803 is the motion vector of the bird in the background. As explained in step S602, in this embodiment, no reference block is set for bird 404, but in order to explain the characteristics of the control in this embodiment, the motion vector of bird 404 is illustrated as dashed line 803.
[0038] Here, the motion vectors 802 and 803 have the same vector magnitude and direction. In such a situation, when the shooting parameter generating unit 207 described later determines the shutter speed by referring to the calculated motion vector, it is not preferable to refer only to the magnitude and direction of the motion vector within the screen, and there is a possibility that it will be influenced by the motion vector of a moving subject in the background. Therefore, by performing the region extraction and scene recognition in this embodiment in advance, it is possible to eliminate the influence of the motion vector in an unrelated subject region such as 803, and the accuracy of the generated shooting parameters can be improved.
[0039] Through the above processing, the motion vector calculation unit 204 calculates the motion vector and the motion vector reliability between the preparation images that are adjacent in time. At least one of these can be used as motion information in the subsequent processing.
[0040] In step S308, the motion vector correction unit 205 acquires information on the estimated movement direction of the object area of the swinging golf club or the like output by the scene recognition unit 203, and corrects the motion vector output by the motion vector calculation unit 204. Specifically, when the angle between the estimated movement direction generated in step S304 and the motion vector of the golf club area calculated in step S307 is α (0° to 180°), the correction gain represented in FIG. 9(a) may be multiplied by the motion vector of the golf club area. The closer the estimated direction and the calculated motion vector are to the direction (the closer the angle is to 0°), the more the gain amount becomes 1.0, and the calculated motion vector is used as is. On the other hand, if the angle between the estimated direction and the calculated motion vector is significantly different, there is a possibility that the motion vector has not been calculated correctly. The motion vector correction unit 205 suppresses such a motion vector from being used to determine the shooting parameters.
[0041] The following causes may be considered as examples of cases where a motion vector cannot be obtained correctly in a swinging object region. For example, since a swinging object moves at high speed, the image obtained in the preparation shooting may contain motion blur. At this time, the edge of the subject becomes unclear or two-line blur occurs. It is considered that such false edges may cause an incorrect correlation value to be calculated in step S307 during correlation calculation. In addition, when shooting in a dark environment, the accuracy of the motion vector may be reduced by being affected by random noise due to an increase in ISO sensitivity. Therefore, by using not only the motion vector calculated from the difference in pixel values between frames as in this embodiment, but also machine learning and deep learning techniques such as estimation of position and orientation, it is possible to generate highly accurate shooting parameters even for a swinging object. In addition, a configuration may be adopted in which the reliability of the motion vector calculated in step S307 is also utilized as a method of correcting the motion vector. As shown in FIG. 9(b), the motion vector is multiplied by a correction gain such that the gain amount becomes 1.0 as the reliability becomes higher. In addition, in this embodiment, a correction gain that is always 1.0 is applied to the motion vector of the face region.
[0042] In step S309, control unit 101 compares the magnitude of the corrected motion vector output from motion vector correction unit 205 with a predetermined threshold value. If the magnitude of the motion vector of the face region or the motion vector of the golf club region is greater than the threshold value, that is, if the face region where motion blur is not desired is moving or the region of the swinging object where motion blur is desired is moving sufficiently, the process proceeds to step S310.
[0043] On the other hand, if both motion vectors are equal to or smaller than the threshold, the image data is updated and the process from step S302 is repeated. In this embodiment, the description will proceed assuming that the motion vector of the golf club area shown in 802 of FIG. 8 is greater than the threshold. Here, it is assumed that the thresholds to be compared with the motion vectors in the face area and the area of the swinging object can be set to different values. This allows the threshold to be changed depending on whether the user prioritizes preventing motion blur or camera shake of the subject, or whether the user prioritizes large motion blur in the swing area. In other words, the user can take the image he or she desires more flexibly. The threshold set here may also be changed depending on the type of sport.
[0044] Specifically, in golf, the subject rarely runs and changes position before and after the swing, so the threshold value to be compared with the motion vector in the swing area is set small to prioritize motion blur in the swing area. On the other hand, in tennis and badminton, the subject itself may be running during the swing, so the threshold value is set large to prioritize preventing motion blur and camera shake. By configuring in this way, it is possible to generate shooting parameters that correspond to the characteristics of each sport among swing sports, and to shoot dynamic images with greater accuracy.
[0045] The threshold value may also be changed according to the positional relationship (facing direction) between the subject and digital camera 100. Specifically, when a subject making a golf swing is photographed from the front (Z-axis direction), as shown in FIG. 4, the magnitude of the motion vector between frames is likely to change, and motion blur is likely to occur in the area of the golf club. On the other hand, when photographing from behind (X-axis direction), the amount of movement in the horizontal and vertical directions in the image appears smaller. For this reason, it is preferable to make the threshold value looser (larger) than when photographing from the front.
[0046] In addition, in this embodiment, the magnitude of the motion vector between the frames of two images is compared with the threshold value, but the result of accumulating the motion vector between multiple frames may be compared with the threshold value. In this case, by continuously monitoring the motion vector from the start of preparation shooting, it is possible to analyze the average movement even if the swinging object suddenly accelerates or decelerates, and it becomes possible to obtain a more stable threshold judgment result. Therefore, the generated shooting parameters are also stable, the success rate of capturing dynamic images is improved, and usability is improved.
[0047] In step S310, the depth change detection unit 206 detects a change in depth between images captured successively under the control of the control unit 101. The change in depth can be detected by attaching a defocus map as metadata when capturing an image. FIG. 10 shows a defocus map corresponding to 401 in FIG. 4. The defocus map shows the degree of focus on the imaging surface in the form of a grayscale map, with white in the foreground, black in the background, and 50% gray indicating in-focus. It can be seen that the human subject 403 is in focus, and the bird 404 is in the background side of the human subject 403. The defocus map is calculated by a known technique, and may be configured to obtain a defocus map on the imaging surface from a phase difference image obtained from an imaging sensor in which all pixels are phase difference pixels, as disclosed in Patent Document 3, for example.
[0048] The depth change detection unit 206 receives the coordinates of the face area generated by the area extraction unit 202, and calculates the average defocus amount in the face area. It also calculates the average defocus amount in the face area in the defocus map captured at the next time. It then calculates the difference between the average defocus amounts at different times as the depth change amount.
[0049] If the focus lens position is the same between different times, the depth change amount can be treated as the amount of movement of the human subject in the depth direction. If the focus lens position is moved in the depth direction without changing, depth blur will occur in the face area of the human subject. Since not only motion blur but also depth blur is not desired to occur in the face area, the shooting parameter generation unit 207 described later appropriately reduces the aperture value according to the depth change amount to deepen the depth and suppress depth blur. By reducing the aperture value, the depth can be deepened, but in order to limit the amount of light entering, it is necessary to increase the ISO sensitivity in order to shoot with the same brightness. If the ISO sensitivity is made too high, random noise becomes noticeable and the image quality deteriorates, so it is not necessary to simply reduce the aperture value. As in this embodiment, by selecting an appropriate aperture value according to the movement change of the subject for each scene, it is possible to shoot higher quality images.
[0050] In this embodiment, the change in the depth direction is calculated from the defocus map, but the present invention is not limited to this, and any information corresponding to the distance distribution in the depth direction of the subject in the imaging range may be used. For example, the defocus amount distribution after normalization with the focal depth may be used, or a depth map indicating the subject distance of each pixel may be used. In addition, the defocus amount may be two-dimensional information indicating the phase difference (image shift amount) used to derive the defocus amount. In addition, the defocus amount may be a map converted into actual distance information on the subject side via the focus lens position. In other words, any information indicating a change according to the distance distribution in the depth direction is applicable, but is not limited to these.
[0051] In step S311, under the control of the control unit 101, the shooting parameter generating unit 207 generates shooting parameters that allow a dynamic image to be captured of a subject playing a sport involving a swinging motion. Specifically, values for shutter speed, aperture value, and ISO sensitivity are generated. Here, feature information acquired by each unit of the shooting control parameter generating unit 200, such as information on the shooting scene, information on the subject's area, movement information, and information showing changes according to the distance distribution in the depth direction, is utilized.
[0052] First, the shutter speed is determined by referring to the magnitude of the vector output by the motion vector correction unit 205. Here, the motion vector has horizontal and vertical components, but for simplicity, only the horizontal direction will be explained. The vertical direction can be calculated in the same way as the horizontal direction.
[0053] For example, suppose that the horizontal image size is 2100 pixels, and motion blur of 200 pixels, which is about 10%, is desired. If the magnitude of the vector output by the motion vector correction unit 205 is 100 pixels, and the frame rate is 120 frames per second, an image with motion blur of the desired width can be captured by setting the shutter speed to 1 / 60 seconds.
[0054] The value of the shutter speed set here can be adjusted by the shooting parameter generating unit 207 according to the size of the image in the horizontal and vertical directions. The width of the motion blur can be automatically set by the shooting parameter generating unit 207 using information such as the type of sport recognized by the scene recognizing unit 203 and the estimated result of the moving direction and range of a swinging object such as a golf club. Alternatively, the user can set the width of the motion blur via the operation input unit 110. For example, a threshold value (first threshold value) of the desired amount of motion blur is set, and the shooting parameter generating unit 207 can generate shooting parameters so that motion blur equal to or exceeds the set threshold value occurs. On the other hand, a threshold value (second threshold value) of the amount of motion blur that the user can tolerate for an area where motion blur is to be suppressed may also be set. In that case, the shooting parameter generating unit 207 adjusts the shooting parameters so that the motion blur in the area where blur is to be suppressed is equal to or less than the set threshold value. This makes it possible to set shooting parameters that give a sense of motion to an area with movement and suppress motion blur in an area where motion is to be stopped. However, if it is difficult to achieve both, the user may be allowed to set in advance via operation input unit 110 or the like whether to give priority to the expression or suppression of motion blur.
[0055] Next, the aperture value is determined by referring to the change in depth between consecutively captured images output by the depth change detection unit 206. Here, the defocus amount is further normalized by the focal depth (for example, 1Fδ, where F is the aperture value and δ is the allowable circle of confusion diameter, which is twice the pixel size in this embodiment) and is defined as the depth blur amount. If the depth blur amount is greater than 1.0Fδ, the observer can recognize that it is blurred. For example, if the change amount output by the depth change detection unit 206 is 0.05 mm, selecting the aperture value at F5.6 results in 0.05 / (5.6×5×2×10^-3)=0.89Fδ, and the observer cannot recognize that depth blur has occurred.
[0056] Finally, the ISO sensitivity can be selected so that an image can be captured at appropriate brightness when the shutter speed is 1 / 60 seconds and the aperture value is F5.6, given the brightness of the shooting environment. Suppose that the image is captured outdoors on a cloudy day and the EV (Exposure Value) is 11. Since the shutter speed is 1 / 60 seconds, the TV (Time Value) is 6, and since the aperture value is F5.6, the AV (Aperture Value) is 5. Since the EV can be calculated as EV=TV+AV-SV, the SV is 0 and the ISO sensitivity is determined to be 100. Through the above processing, the shooting parameter generation unit 207 generates the shooting parameters.
[0057] In step S312, control unit 101 performs actual shooting according to the shooting parameters determined in step S311, records the actual shot image on recording medium 108, and ends the series of processes. An example of an actual shot image is shown in Fig. 11. By shooting with appropriate shooting parameters, there is almost no motion blur or depth blur in the face of the human subject, while motion blur occurs in the area of the golf club, making it possible to capture a dynamic image.
[0058] As described above, in this embodiment, the motion vector between images is calculated, and the shooting parameters are controlled according to the area where motion blur is desired to occur and the area where motion blur is desired to be suppressed. Therefore, for example, for a subject that is performing a swinging motion, it is possible to shoot a dynamic image by suppressing motion blur of the face and the like, while generating motion blur of an object such as a golf glove.
[0059] In this embodiment, the defocus map is described as being generated based on a group of images that have a parallax relationship, but the method is not limited to this as long as it corresponds to the captured image and the distance distribution of the subject in the imaging range can be obtained. The defocus map generation method may be, for example, a DFD (Depth From Defocus) method that derives the defocus amount from the correlation between two images with different focus or aperture values. Alternatively, the distance distribution of the subject may be derived using information related to the distance distribution obtained from a distance measurement sensor module such as a TOF (Time of Flight) method. Alternatively, it may be a contrast distance measurement method.
[0060] [Example 2] A second embodiment (Example 2) of the present invention will be described below. Note that the embodiment described below is an imaging device similar to Example 1, and an example in which the present invention is applied to a digital camera as an example of an imaging device will be described.
[0061] In the first embodiment, it was possible to capture an image with a sense of motion by controlling the shooting parameters, but in this embodiment, by combining multiple images, an image is generated that combines the sense of motion achieved by long-exposure shooting with the sense of stillness achieved by locally suppressing subject blur.
[0062] The configuration of the digital camera according to the embodiment of the present invention is similar to that shown in the block diagram of FIG. 1 described in the first embodiment, and therefore a description thereof will be omitted.
[0063] In this embodiment, unlike the first embodiment, the image processing unit 107 includes an image synthesis processing unit 1200 as shown in Fig. 12. The process of the image synthesis processing unit 1200 will be described in detail below.
[0064] In the second embodiment, it is assumed that a scene is captured in which there is a moving object (such as a fountain, a waterfall, or a stream of people) in the background and a person as the main subject in the foreground, as shown in Fig. 13(a). In such a scene, the objective is to capture an image in which the background has a sense of motion equivalent to long-exposure photography, and the motion blur in the person area is suppressed. Another objective is to generate an image with a sense of motion by adding a trajectory of the movement in a scene in which the main subject is moving.
[0065] 12, image synthesis processing section 1200 is made up of image storage section 1201 , main subject region detection section 1202 , main subject related feature detection section 1203 , synthesis characteristic control section 1204 , and average synthesis section 1205 .
[0066] The image synthesis processing unit 1200 accumulates a predetermined number of frames of RAW image data captured by the imaging unit 105 or image data after development processing in the image accumulation unit 1201. It is assumed that the input images are continuous frames without non-exposure periods between frames. The image data accumulated in the image accumulation unit 1201 is input to an average synthesis unit 1205, which performs averaging processing of multiple frames in pixel units to generate an image equivalent to a long second.
[0067] The number of frames stored in the image storage unit 1201 is determined by the shutter speed set by the user and the shutter speed at which images are captured.
[0068] In this embodiment, the shutter speed for capturing images is fixed at 1 / 100 seconds. If the shutter speed is set to 1 / 2 seconds by the user, the number of frames captured and accumulated will be 50. Then, by averaging 50 frames of images captured at a shutter speed of 1 / 100 seconds, an image with accumulated blur equivalent to 1 / 2 seconds and brightness equivalent to an image captured at 1 / 100 seconds is generated.
[0069] The image data stored in the image storage unit 1201 is also output to a main subject region extraction unit 1202 and a main subject related feature detection unit 1203 .
[0070] The main subject region extraction unit 1202 extracts the region of the main subject and outputs it as a main subject region map. The main subject is extracted using a known method such as machine learning. In this embodiment, an example will be described in which a person region is extracted as the main subject.
[0071] An example of a main subject region map is shown in Fig. 13. Fig. 13(a) shows an image input to main subject region extraction section 1202. Fig. 13(b) is the extracted main subject (person) map, in which main subject region 1301 has a signal of 1 (white) and the background region has a signal of 0 (black). The extracted main subject region map is output to main subject related feature detection section 1203 and synthesis characteristic control section 1204.
[0072] The main subject region extraction unit 1202 also calculates position information of the main subject. For example, the position of the person's head or the position of the center of gravity of the human body is used as the main subject position. The position information may be calculated by setting a coordinate system centered on the position of the person's head or the position of the center of gravity of the human body, or a coordinate system centered on a certain point in the image. The main subject region extraction unit 1202 generates a main subject region map and calculates the main subject position for multiple frames stored in the image storage unit 1201.
[0073] The main subject related feature detection unit 1203 detects the amount of movement between frames for the main subject region and the main subject background region, and sets the main subject background region based on the input main subject region map. The main subject background region is the background region around the main subject. An example of the main subject background region is shown in FIG. 13(c). The region 1302 shown in white in FIG. 13(c) is the main subject background region. The main subject related feature detection unit 1203 calculates the absolute difference between multiple frames stored in the main subject region image storage unit 1201 in the main subject region 1301 and the main subject background region 1302, and normalizes it by a predetermined area. This is used as an index for detecting the presence or absence of motion blur of the main subject and the presence or absence of a moving subject in the main subject background region.
[0074] The synthesis property control unit 1204 determines synthesis properties in the average synthesis unit 1205 based on the area map information and position information of the main subject output from the main subject area extraction unit 1202, and the movement information (feature information) of the main subject and background area output from the main subject related feature detection unit 1203. The details of the processing of the synthesis property control unit 1204 will be described later, but the synthesis property control unit 1204 outputs a synthesis map used for synthesis and synthesis order information as synthesis information. An example of the synthesis map is shown in Fig. 13(d).
[0075] The average synthesis unit 1205 averages the images output from the image storage unit 1201 based on the synthesis information output from the synthesis characteristic control unit 1204. At this time, images with different shutter speeds are generated for each region by varying the number of synthesis frames for each pixel based on the synthesis map. In this embodiment, the number of frames to be averaged is reduced in the main subject region, outputting pixels with a short shutter speed (corresponding to a short second), and the number of frames to be averaged is increased in the background region, outputting pixels with a long shutter speed (corresponding to a long second).
[0076] The above describes the processing of the image synthesis processing unit 1200. The image data generated as described above is recorded in 108, or in the case of RAW data, is subjected to development processing by a development processing unit included in the image processing unit 107.
[0077] Next, the details of the processing of the composition characteristic control unit 1204 will be described with reference to the flowchart of FIG. 14. Each step of this flowchart is executed by each unit of the digital camera 100 including the image composition processing unit 1200 under the instruction of the control unit 101 or the control unit 101.
[0078] In step S1401, the composition characteristic control unit 1204 sets a short-second reference frame. In the average composition unit 1205, an image including both long-second equivalent pixels and short-second pixels is generated based on the composition information as described above. At this time, the shortest-second pixels are the case where the image of one frame is output as it is without composition. The frame used as the shortest-second pixels is the short-second reference frame.
[0079] A specific example of the short-second reference frame will be described below. FIG. 15(a) shows the images stored in the image storage unit 1201. The horizontal axis represents the time axis on which the images were captured, indicating that 10 images captured in the order of 1 to 10 are stored.
[0080] As the long-second equivalent pixels generated by the average composition unit 1205, an image having an accumulated blur amount captured for 10 frames equivalent to 10 frames by averaging all of the 10 frames of images from 1 to 10 and having the same brightness as before composition can be generated. On the other hand, as the short-second pixels, the pixels of the short-second reference frame are output as they are. Whether to output the short-second pixels or the long-second equivalent pixels in the image is determined by the composition map shown in FIG. 13(d).
[0081] The composition map shown in FIG. 13(d) has a 0-1 signal. When the signal is 0, the long-second equivalent pixels synthesized with the maximum number of synthesis are output, and when the signal is 1, the pixels of the short-second reference frame are output. The signal value M of the composition mask has an intermediate value of 0 < M < 1. In the case of the intermediate value, the number of synthesis N is determined based on the following formula. N=(Max-1)×(1-M)+1
[0082] Here, Max is the maximum number of composite frames. For example, when Max=10 and M=0.8 pixels, N=2, and a signal that averages two images is output.
[0083] The short-second reference frame can be selected from any of the frames used for synthesis, and which frame is to be the short-second reference frame is determined according to a flow described below.
[0084] In the example of Fig. 15(b), the short-second reference frame is the frame 1501 captured last in time. Similarly, in Fig. 15(c), the short-second reference frame is the frame 1502 captured first in time, and in Fig. 15(d), the short-second reference frame is the frame 1503 captured sixth in time.
[0085] A method for setting the short-second reference frame will be described with reference to the flowchart in Fig. 16. Each step of this flowchart is executed by each unit of the digital camera 100, including the image synthesis processing unit 1200, in response to instructions from the control unit 101 or control unit 1101.
[0086] In step S1601, synthesis characteristic control unit 1204 acquires main subject position information for each frame and calculates the amount of positional variation of the main subject between the first frame and the last frame. An example of the amount of positional variation is shown in Fig. 17(a). Reference numeral 1701 denotes the position of the main subject in the first frame, 1702 denotes the position of the main subject in the last frame, and arrow 1703 denotes the amount of positional variation.
[0087] In step S1602, synthesis characteristic control unit 1204 determines whether the main subject has moved between the accumulated frames. Specifically, if the amount of positional variation of the main subject calculated in step S1601 is greater than threshold TH1, it is determined that there has been movement, and if it is less than TH1, it is determined that there has been no movement. If it is determined that there has been no movement of the main subject, proceed to step S1603. If it is determined that there has been movement, proceed to step S1606.
[0088] In step S1603, the synthesis characteristic control unit 1204 determines whether the motion information of the background region output from the synthesis characteristic control unit 1203 is greater than a threshold value TH2. It is assumed that the motion information of the background region is calculated between multiple frames, and if any one of the motions between frames is greater than the threshold value TH2, it is determined that there is motion.
[0089] If it is determined that there is movement in the background, the process proceeds to step S1604, and if it is determined that there is no movement, the process proceeds to step S1605.
[0090] In step S1604, the synthesis characteristic control unit 1204 selects a frame without a moving subject in the background as the short-second reference frame. A scene with a moving background will be described with reference to FIG. 17(b)(c). In FIG. 17(b)(c), 1704 indicates a background region of a person, and 1705 indicates a moving object (bird) other than a person. When there is no subject other than a person in the background region 1704 of a person as in FIG. 17(b), the amount of movement in the background region is small, and when a bird moves and enters the background region 1704 as in FIG. 17(c), the amount of movement is large. Here, the short-second reference frame is selected from frames excluding frames with a large amount of movement in the background region. Specifically, the frame furthest in time from the frame with a large amount of movement is set as the short-second reference frame. An example is shown in FIG. 15(e). In FIG. 15(e), 1504 indicates a plurality of frames with an amount of movement larger than a threshold value TH2. In this example, frame 1505, which is the frame furthest in time from the group of frames with a large amount of movement, is selected as the short-second reference frame. However, when the background area is constantly moving, such as the example of the fountain in FIG. 13(a), the final frame or the frame with the least amount of movement is set as the short-second reference frame.
[0091] By controlling as described above, it is possible to prevent an image that is difficult to use as a still frame, for example an image in which a bird is obscuring a human region in which movement is to be frozen, from being used as a short-second reference frame.
[0092] In the above example, the short-second reference frame is determined based on the amount of motion only in the background region, but the amount of motion in the person region may also be taken into account in the determination.In addition to the determination based on the amount of motion, object detection may be performed, and a frame that does not include any objects other than the main subject in the background region or near the main subject region may be set as the short-second reference frame.
[0093] Step S1605 is for a case where there is no moving subject in the background, in which case there is no significant problem no matter which frame is used as the short-second reference frame, so a predetermined frame (for example, the final frame) is used as the short-second reference frame.
[0094] In step S1606, the synthesis characteristic control unit 1204 judges whether there is a change in the moving speed of the object when the object is moving. FIG. 18 is a graph showing the relationship between the moving speed of the object with time change. The moving speed of the main object is calculated from the positional change amount between the captured frames. In addition, the maximum value Vmax and the minimum value Vmin of the moving speed are calculated. If the difference between the maximum value Vmax and the minimum value Vmin is equal to or greater than the threshold value TH3, it is judged that there is a change in the moving speed of the object, and if it is smaller than the threshold value TH3, it is judged that there is no change in the moving speed. FIG. 18 shows an example in which (a) there is a change in the moving speed, and (b) there is no change in the moving speed. A large change in the moving speed is a scene with a sharp speed in the movement, such as when the main object person jumps or rides on a swing. If it is judged that there is a change in the moving speed, proceed to step S1607, and if it is judged that there is no change in the moving speed, proceed to step S1608.
[0095] In step S1607, the synthesis characteristic control unit 1204 selects the frame in which the subject movement speed is the slowest as the short-second reference frame. In the example of Fig. 18(a), the frame captured at time T1, when the movement speed is the slowest, is selected as the short-second reference frame. By using the time when the movement speed is slow as the short-second reference frame, it is possible to generate an image with a clear distinction between moving and still areas.
[0096] In step S1608, the synthesis characteristic control unit 1204 sets the last captured frame as the short-second reference frame, which makes it possible to generate an image that expresses the trajectory of the main subject's movement.
[0097] The above describes a method for selecting a short-second reference frame. In the above example, a short-second reference frame is selected based on the main subject and the movement information of the main subject's background, and the movement information of the main subject. However, any information related to the main subject may be used to select a short-second reference frame. For example, it is possible to use distance information of the main subject to set the frame with the furthest (or closest) distance as the short-second reference frame. It is also possible to analyze the facial expression of the main subject and set the frame with the highest frequency of smiling as the short-second reference frame.
[0098] 14, in step S1402, the synthesis characteristic control unit 1204 sets the synthesis order based on the short-second reference frame. Examples of the synthesis order are shown in FIGS. 15(b) to 15(d).
[0099] The numbers written on each frame in Figures 15(b) to (d) indicate the synthesis order. A frame with a synthesis order of 1 indicates a frame based on the short second standard. As described above, the average synthesis unit 1205 performs average synthesis of multiple frames, and the synthesis order indicates the order used for average synthesis. For example, when the number of synthesis images is 3, the three images 1 to 3 are synthesized, and when the number of synthesis images is 5, frames 1 to 5 are synthesized.
[0100] In this step, the synthesis order is set as shown in Fig. 15(b) to (d). The method of setting the order is to set the synthesis order in the order of frames that are temporally close to the short-second reference frame, using the short-second reference frame as a reference. Note that if the time difference between the previous and next frames is the same, the later frame is given priority.
[0101] Examples of setting the synthesis order according to the above are shown in Figures 15(b) to (d). Figure 15(b) shows the synthesis order when the short second reference frame is the last frame, Figure 15(c) shows the synthesis order when the short second reference frame is the first frame, and Figure 15(d) shows the synthesis order when the short second reference frame is the sixth frame from the first.
[0102] 13(d) (generate composite map) based on the main subject region map output from main subject region extraction section 1202. Since the main subject region map is binary data of 0 or 1, a low-pass filter is applied to the main subject region map to generate a map having intermediate values between 0 and 1, and this is used as the composite map.
[0103] In step S1404, synthesis characteristic control unit 1204 determines whether the main subject has moved between frames, similarly to step S1602 described above. If the main subject has not moved, the process proceeds to step S1405, and if the main subject has moved, the process ends.
[0104] In step S1405, the synthesis characteristic control unit 1204 corrects the synthesis map generated in step S1403 based on the movement of the main subject and the background. The correction method will be described with reference to FIG. 19. FIG. 19 is a diagram showing the amount of movement of the main subject region or the background region of the main subject and the correction ratio of the mask. The mask correction ratio indicates the ratio of expanding or contracting the synthesis mask, and indicates expansion when the value exceeds 1 and contraction when it is less than 1. FIG. 19(a) shows the relationship between the amount of movement of the main subject region and the correction ratio of the synthesis mask size, and calculates a correction coefficient for correcting the synthesis mask so that the synthesis mask expands as the amount of movement of the main subject region increases. FIG. 13(e) shows an example of expanding the synthesis mask based on the synthesis mask in FIG. 13(d). The main subject does not move in position, but if there is movement, the motion blur part will go beyond the contour of the subject when average synthesis is performed. Therefore, by expanding the synthesis range, it is possible to output a short-second image to prevent blurring of the blurred contour part.
[0105] Meanwhile, Fig. 19(b) shows the relationship between the amount of movement of the background region of the main subject and the mask size, and a correction coefficient is calculated to correct the composite mask so that the greater the movement of the background region, the more it shrinks. Fig. 13(f) shows an example of a composite mask that has been shrunk using the composite mask in Fig. 13(d) as a reference. This corresponds to a case where the main subject does not move, but there is movement in the background such as a fountain, waterfall, or flow of people. When there is movement in the background, if it is synthesized beyond the main subject region, an image is generated in which the background movement is still only around the outline of the main subject. To prevent this, control is performed so that as much as possible of the inside of the main subject is synthesized.
[0106] As described above, after the mask correction coefficient based on the movement of the main subject region and the mask correction coefficient based on the movement of the background region of the main subject are calculated, the two mask coefficients are multiplied together to calculate the final mask correction coefficient.
[0107] In the above example, an example of expanding and contracting the composite mask has been described as an example of a method of correcting the composite mask, but any correction may be performed as long as the composite mask is corrected based on the characteristics of the main subject. For example, a correction may be performed to control the steepness of the gradation of the composite mask. In this case, control may be performed to make the gradation of the composite mask steeper when there is movement in the background, and to make the gradation gentler when there is no movement in the background. This makes it possible to make the difference in motion blur between the main subject area and the background area less noticeable.
[0108] The configuration of this embodiment has been described above. By combining multiple images with the configuration of this embodiment, it is possible to generate an image that combines a dynamic expression achieved by long-exposure shooting and a still expression in which subject blur is suppressed locally.
[0109] In this embodiment, an example has been described in which the number and order of images used for averaging in the average synthesis unit 1205 are controlled, but if multiple images are to be synthesized, other configurations are also possible. For example, it is also possible to generate an image equivalent to a long second by averaging all accumulated images, and to partially synthesize a short second reference image with the image equivalent to a long second.
[0110] In the above embodiment, the captured images are all captured at the same shutter speed, but it is also possible to capture some images at different shutter speeds. FIG. 15(f) shows a plurality of image frames captured continuously and stored in the image storage unit 1201. Frame 1506 shows the final frame, which is captured at half the shutter speed of the other frames 1507. The long-second image is generated by averaging all 11 frames. Frame 1506 is used as the short-second reference frame. At this time, frame 1506 is used after applying a double gain because the exposure is half. It is also possible to select whether to use frame 1503, which has a shorter shutter speed, or another frame with uniform brightness as the short-second reference frame based on the amount of movement of the main subject, in the composition characteristic control unit 1204.
[0111] By capturing a combination of images captured at different shutter speeds in this manner, it becomes easier to generate an image in which blur is suppressed locally when the subject's position is moving quickly or when there is significant motion blur.
[0112] [Example 3] A third embodiment will be described below with reference to Fig. 20 to Fig. 24. In the third embodiment, the photographing method of the first embodiment and the photographing method of the second embodiment are switched depending on the photographing situation. That is, when photographing an image in which a part where subject blur is suppressed and a part where subject blur is to be allowed or emphasized are mixed in one image, the photographing method of the first embodiment and the photographing method of the second embodiment are switched to the photographing method which is most effective.
[0113] Here, the shooting method described in the first embodiment involves shooting a single image with a single exposure, capturing an image that contains a mixture of areas where subject blur is suppressed and areas where subject blur is tolerated or emphasized (hereinafter referred to as "single image shooting").
[0114] In contrast, in the photographing method described in the second embodiment, multiple images are photographed with multiple exposures, and an image portion with less blur from one of the multiple images is used for the portion where subject blur is desired to be suppressed, and an image is generated by combining multiple images for the portion where subject blur is desired to be tolerated or emphasized. Furthermore, the image portion with less blur is combined with a blurred image obtained by combining multiple images, thereby photographing an image in which a portion where subject blur is suppressed and a portion where subject blur is permitted or emphasized are mixed (hereinafter referred to as "multiple image photographing").
[0115] In the case of single-shot photography, if the desired shutter speed can be determined, photography can be performed with a single exposure, and subsequent image processing such as development is also easy, and the image can be displayed promptly after photography. On the other hand, since the area to be blurred (swing area in Example 1) is exposed at a shutter speed that can obtain a sufficient amount of blur, if there is movement in the area to be suppressed from blurring (face area in Example 1) during this exposure period, blurring according to the amount of movement will occur. If the amount of movement is within an allowable range, single-shot photography is sufficient, but when trying to capture a detailed expression in a face image, even a slight amount of movement may be regarded as a blurred image, and the overall image may be unsatisfactory. For example, as shown in Example 1, the golf club area should have a sense of dynamism due to blurring, as shown in FIG. 11, but the face area should have suppressed blurring and a clear image should be obtained. However, if the face area moves even a little during this exposure period, blurring may occur in the face area in the image, as shown in FIG. 23, and the image may become unclear.
[0116] This will be explained with reference to FIG. 24. In FIG. 24, the horizontal axis indicates time, and the double-headed arrow of 2401 indicates the exposure period (period corresponding to the shutter speed) described in the first embodiment. 2402 to 2405 are enlarged views of the face area of the subject during this exposure period, and similarly 2406 to 2409 are enlarged views of the area of the swinging golf club. As shown in FIG. 24, the golf club moves during the exposure period 2401 due to the swing, so the image is captured in a blurred state as shown in FIG. 23, and a dynamic image can be obtained. On the other hand, if the face area is completely still during the exposure period 2401, a dynamic image can be obtained in which the face image is clear and only the golf club is blurred, as shown in FIG. 11. However, there are cases where the face moves. That is, as shown in 2402 to 2405 in the figure, the face image may move slightly. In this case, as shown in FIG. 23, the image shows the club as a dynamic image, but the face image is also blurred. In this case, if the exposure time (shutter speed) is set to a short time, the facial image may be captured clearly, but the golf club will also be stationary, and the intended purpose of capturing an image with a sense of movement will not be obtained.
[0117] On the other hand, if multiple images are taken, as shown in the second embodiment, the facial image is taken with an exposure time shorter than the exposure period, so that an image in which the facial image is still and clearly taken (corresponding to the short-second frame in the second embodiment) can be easily obtained. In addition, the area of the swinging golf club can be obtained by combining multiple captured images to obtain a blurred image with a sense of dynamism. Therefore, by combining with the previous facial image, the desired image with a sense of dynamism can be obtained. However, as explained in the second embodiment, combining multiple images requires a synthesis process, so it takes time to generate a final image, and it is difficult to display the image immediately after the exposure period ends. In addition, if the synthesis process is always assumed, there is a drawback in that the power consumption is always high in devices that use batteries, etc., and the battery life is shortened.
[0118] Therefore, in this embodiment, by switching between single shot and multiple shots depending on the state of movement of the image part where blur is desired to be suppressed, an image is efficiently created that contains a mixture of parts where subject blur is suppressed and parts where subject blur is tolerated or emphasized.
[0119] FIG. 20 shows the configuration of each processing unit for image capture in the third embodiment. Each processing unit can be explained by replacing each unit of the digital camera 100 described in the first embodiment. An image capturing unit 2001 in FIG. 20 corresponds to the image capturing unit 105 in the first embodiment. An image processing unit 2002 performs general image processing in this embodiment. It also includes a shooting control parameter generating unit 2004 and an image synthesis processing unit 2006. A final image 2003 generated in this embodiment is obtained by processing image data acquired by the image capturing unit 2001 in the image processing unit 2002.
[0120] Here, the image capturing unit 2001 will be described as outputting two types of image data. That is, image data (hereinafter, frame image data) 2008 output frame by frame with a short exposure, and image data (hereinafter, captured image data) 2009 resulting from capturing an image with a specified exposure time. These two types of image data can be considered to pass through the same route depending on the configuration and functional operation of the image capturing unit, or can be treated as the same image data by matching the exposure time to the frame interval. For convenience, in this embodiment, the frame image data 2008 and the captured image data 2009 will be described separately.
[0121] A shooting control parameter generation unit 2004 in Fig. 20 receives frame image data 2008 as input, and determines and outputs shooting parameters 2010. The output shooting parameters 2010 have values of shutter speed, aperture value, and ISO sensitivity. In this embodiment, the shutter speed 2011 is extracted from these and input to the image synthesis processing unit 2006. The shooting parameters 2010 are input to a single shooting processing unit 2005 in Fig. 20. The single shooting processing unit 2005 specifies the shutter speed in the image shooting unit 2001 in accordance with the shooting parameters, acquires shot image data 2009, and generates an image as a result of single shooting.
[0122] An image synthesis processing unit 2006 in Fig. 20 receives frame image data 2008, determines image synthesis-related parameters, performs image synthesis processing, and captures multiple images. In this embodiment, a shutter speed 2011 is input to the image synthesis processing unit 2006. This shutter speed 2011 is used to determine synthesis processing parameters as an exposure period for capturing images equivalent to a long second period described in the second embodiment. That is, the image synthesis processing unit 2006 accumulates the frame image data 2008 in an image storage unit equivalent to the image storage unit 1201 in Fig. 12 in the second embodiment for a period equivalent to the shutter speed 2011, and generates an image equivalent to a long second period by synthesizing the accumulated images in an average synthesis unit equivalent to the average synthesis unit 1205 in the second embodiment.
[0123] A switch 2007 in FIG. 20 switches between an image generated by a single shot and an image generated by the image synthesis processing unit 2006 , and outputs one of them as a final image 2003 .
[0124] Next, a flow of photographing in the third embodiment will be described with reference to the flowchart in Fig. 21. Note that the following description will focus on the flow relating to switching of the photographing method performed in the third embodiment, and will omit a description of the flows described in detail in the first and second embodiments. Also, each step of this flowchart is executed by the control unit 101 or each unit of the digital camera 100 including the image processing unit 2002 in response to an instruction from the control unit 101.
[0125] Step S2101 corresponds to the process performed in step S301 in the first embodiment. The user turns on the power of the digital camera 100. In response to the power being turned on for the digital camera 100, the control unit 101 controls the optical system 104 and the image capturing unit 2001 to start preparatory shooting.
[0126] Step S2102 corresponds to the process performed in step S302 in the first embodiment, where image data captured in S2101 is acquired, and a main subject detection unit (not shown) of the image processing unit 2002 detects the main subject. That is, the image processing unit 2002 detects the main subject from the frame image data from the image capturing unit 2001.
[0127] Next, step S2103 is a scene recognition step, which performs scene recognition equivalent to the process performed in step S304 in the first embodiment. That is, the photographing scene is recognized from the main subject detected in step S2102, and conditions required for determining the next photographing parameters are set.
[0128] Next, step S2104 is a shooting parameter determination step, in which shooting parameters are set corresponding to the processing performed in step S306 or step S311 in the first embodiment. That is, the shooting parameters, the shutter speed, the aperture value, and the ISO sensitivity, are determined according to the conditions set in the scene recognition step. The shutter speed determined here is determined as an exposure period for capturing an image in which a portion in which subject blur is suppressed and a portion in which subject blur is tolerable or to be emphasized are mixed as an image at the time of actual shooting.
[0129] Next, in step S2105, the movement of the main subject is determined. That is, whether or not the main subject detected in step S2101 is likely to move during the exposure period corresponding to the shutter speed determined in step S2103 is determined in conjunction with the scene recognition result in step S2102. In general, if the determined shutter speed (exposure period) is long, the main subject is likely to move during that exposure period. Also, even if the exposure time is short, the main subject may move depending on the type of sport to be photographed or a scene involving a specific movement. In particular, when focusing on the face area of the main subject, the determination of whether or not the main subject moves depends on conditions such as whether or not even a slight change in facial expression is captured. In any case, in step S2105, it is determined whether or not the main subject moves during the exposure period, and the subsequent shooting method is switched.
[0130] If it is determined in step S2105 that a movement occurs in the main subject, the process proceeds to step S2106, where digital camera 100 takes multiple shots. On the other hand, if it is determined that no movement occurs in the main subject, the process proceeds to step S2107, where digital camera 100 takes a single shot. If the process proceeds to step S2106 and multiple shots are taken, the movement of the main subject is subsequently determined and an image to be subjected to composition processing is selected, as described in the second embodiment.
[0131] Here, the relationship between the determined shutter speed (exposure period) and the period during which the image capturing section 2001 is actually exposed (actual exposure period) will be described again with reference to FIG.
[0132] FIG. 22(a) shows the relationship between the exposure period and the actual exposure period when shooting one image. 2201 indicates the exposure period, which is determined by the shooting control parameter generating unit 2004 in FIG. 20 or the shooting parameter determination step in FIG. 21, for shooting an image in which a part where subject blur is suppressed and a part where subject blur is to be tolerated or emphasized are mixed as a final image. 2202 indicates the actual exposure period, which shows the exposure period of the shooting image data exposed for the exposure period specified by the image capturing unit 2001 in FIG. 20. In shooting one image, the exposure period and the actual exposure period are equal in length. In FIG. 22(a), the final image data can be obtained by simply performing a predetermined process such as development on the image data after shooting one image.
[0133] On the other hand, FIG. 22(b) shows the relationship between the exposure time and the actual exposure time when multiple images are taken. In the figure, 2203 indicates the exposure period, which is the same as 2201 in FIG. 22(a). FIG. 22(b) shows a case where multiple images are taken every 10 during the exposure period. That is, 2204 shows a case where multiple captured image data from 1 to 10 are taken. In the case of multiple image taking, by taking an image with an exposure time shorter than the exposure period, as described in the second embodiment, it is possible to take an image in which the main subject does not move, as described in the second embodiment. By combining multiple images taken at the same time, it is possible to take an image in which a part where subject blur is suppressed and a part where subject blur is allowed or desired to be emphasized are mixed as the final image, as described in the second embodiment.
[0134] Here, in the case of multiple shots, shooting is repeated with a short exposure time, but image synthesis processing is performed by starting with any image shot with a short exposure and synthesizing as many images as necessary to make the total exposure time equivalent to the exposure period. This is shown in FIG. 22(c). In the figure, 2205 is the exposure period, 2206 is image data with a short exposure time shot including before and after this period, and image data 1 to 13 are shown in the figure. For example, in FIG. 22(c), image data from the third image data 2207 to the twelfth image data 2208 are the objects of synthesis processing. In that case, an image is generated by performing image synthesis processing on 10 images during the shooting period, and an image with the same image effect as that of FIG. 22(b) can be obtained. That is, in the present embodiment, multiple shots are shot to shoot an image in which a part where subject blur is suppressed and a part where subject blur is allowed or emphasized are mixed, and in order to obtain the desired effect in the final image, it is sufficient to have image data of a short exposure obtained during the determined exposure period. Therefore, it is not necessary to synthesize image data other than this exposure period.
[0135] In multiple shots, after all the image data to be combined is obtained, the combined data is processed to obtain the final image data. Although the combined data processing is time-consuming, it is possible to obtain an image in which blurring of the main subject is reliably suppressed.
[0136] As described above, according to this embodiment, by switching between taking a single image and taking multiple images depending on the state of movement of the main subject, it is possible to generate, with an appropriate amount of processing, an image that contains a mixture of areas where subject blur is suppressed and areas where subject blur is tolerated or emphasized.
[0137] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.
[0138] [Other embodiments] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions. [Explanation of symbols]
[0139] 100 Digital Camera 101 Control section 102 ROM 103 RAM 104 Optical system 105 Imaging unit 106 A / D conversion section 107 Image Processing Unit 108 Recording media 109 Display section 110 Operation input section
Claims
1. an image acquisition means for acquiring a plurality of images successively captured by a first photographing; a region extraction means for extracting a first region and a second region different from the first region from the acquired image; recognition means for recognizing a photographed scene from the plurality of images; an information acquisition means for acquiring motion information of the first region and the second region from the plurality of images; a photographing parameter generating means for generating photographing parameters for a second photographing operation different from the first photographing operation, The photographing parameter generating means generates a parameter based on the photographing scene and the motion information, and calculates a value when the amount of motion blur in the first region of the image acquired by the second photographing is equal to or greater than a first value. The image processing device generates the photographing parameters as follows.
2. The photographing parameters include at least one of a shutter speed, an aperture value, and an ISO sensitivity.
2. The image processing device according to claim 1, further comprising:
3. The photographing parameter generating means determines whether the amount of motion blur in the second region is equal to or less than a second value.
2. The image processing method according to claim 1, wherein the imaging parameters are generated so that: Device.
4. 2. The image processing apparatus according to claim 1, wherein the photographed scene is a scene involving a swinging motion.
5. The scene involving a swinging motion is a sports scene, 5. The image processing apparatus according to claim 4, wherein said recognition means determines the type of sport.
6. The photography parameter generating means generates the photography parameters according to the type of sport.
6. The image processing device according to claim 5, wherein the value of the data is changed.
7. The first region corresponds to the swinging object, and the second region corresponds to the face and torso of the person.
10. The image processing device according to claim 6, wherein the image processing device corresponds to at least one of the whole body. 。
8. an imaging means for capturing an object image formed via an optical system; An image processing device according to any one of claims 1 to 7; An imaging device comprising:
9. image acquisition means for acquiring a plurality of images taken successively; a region extraction means for extracting a first region and a second region different from the first region from the plurality of images; recognition means for recognizing a photographed scene from the plurality of images; an information acquisition means for acquiring information about the first region and the second region from the plurality of images; a synthesis means for performing synthesis processing using the plurality of images, the combining means selects an image to be used for the combining process from the plurality of images based on the photographed scene and the information so that the amount of motion blur in the first region is equal to or greater than a first value.
10. 10. The image processing device according to claim 9, wherein the second region corresponds to a region of a main subject.
11. 11. The image processing device according to claim 10, wherein the information acquisition means acquires at least one of the type of the main subject or a subject surrounding the main subject, movement information, depth, facial expression, and position within the image as information corresponding to the area of the main subject.
12. a composite map generating means for generating a composite map based on the information; 12. The image processing apparatus according to claim 11, wherein the combining means performs the combining process based on the combining map.
13. an imaging means for capturing an object image formed via an optical system; An imaging device comprising: the image processing device according to any one of claims 9 to 12.
14. an image acquisition step of acquiring a plurality of images successively captured by a first photographing; a region extraction step of extracting a first region and a second region different from the first region from the acquired image; a recognition step of recognizing a photographed scene from the plurality of images; an information acquisition step of acquiring motion information of the first region and the second region from the plurality of images; a photographing parameter generating step of generating photographing parameters for a second photographing operation different from the first photographing operation, the shooting parameter generation step generates the shooting parameters based on the shooting scene and the motion information so that an amount of motion blur in the first region of the image acquired by the second shooting is equal to or greater than a first value.
15. an image acquisition step of acquiring a plurality of images taken consecutively; a region extraction step of extracting a first region and a second region different from the first region from the plurality of images; a recognition step of recognizing a photographed scene from the plurality of images; an information acquisition step of acquiring information about the first region and the second region from the plurality of images; a synthesis step of performing synthesis processing using the plurality of images, the combining step is characterized in that an image to be used in the combining process is selected from the plurality of images based on the photographed scene and the information so that the amount of motion blur in the first region is equal to or greater than a first value.
16. A program for causing a computer to execute each step of the image processing method according to claim 14 or 15.
17. A computer-readable storage medium storing a program for causing a computer to execute each step of the image processing method according to claim 16.