Image processing apparatus and its control method, program, and storage medium

JP7927429B2Active Publication Date: 2026-10-01CANON KK
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2022026115
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2026-10-01
Estimated Expiration
2042-02-22

AI Technical Summary

Benefits of technology

【0008】 本発明によれば顔とスウィング領域の動き解析結果とスウィング領域の移動方向の推定結果に基づき、シーン毎に異なるスウィング速度·方向を加味して自動でシャッタースピードを含む撮影パラメータ決めることで、躍動感のある画像を撮像することを可能にした撮像装置、撮像方法およびプログラムを提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927429000001
    Figure 0007927429000001
  • Figure 0007927429000002
    Figure 0007927429000002
  • Figure 0007927429000003
    Figure 0007927429000003
Patent Text Reader

Abstract

To provide an imaging apparatus capable of automatically determining a shutter speed based on a motion analysis result of a face and a swing region and an estimation result of a moving direction of the swing region, in consideration of a swing speed and a swing direction different by scene, thereby forming a dynamic image.SOLUTION: An imaging apparatus includes acquisition means configured to acquire an image, subject detection means configured to detect a subject from the image, motion amount detection means configured to detect a motion amount of a first region of the subject and a motion amount of a second region different from the first region, and imaging parameter determination means configured to determine an imaging parameter. The imaging parameter determination means refers to the motion amount of the first region and the motion amount of the second region and determines the imaging parameter so that a blur amount of the first region is less than a first criterion and a blur amount of the second region is greater than a second criterion.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention particularly relates to a technology for capturing an image of a subject playing a sport involving a swing motion such as golf. [[Background Art]]

[0002] In recent years, among imaging devices such as digital still cameras that have been commercialized, some models are equipped with a sports shooting mode. A high shutter speed is automatically set to make it easier to shoot images of a person running, a moving vehicle, or the like, thereby enabling capture of an image with less motion blur of the subject. Technologies have been disclosed in which an imaging device automatically sets a shutter speed for capturing an image with less motion blur of the subject as described above. For example, Patent Document 1 discloses a technology for determining a shutter speed based on a motion analysis result of a face region of a subject. [[Prior Art Documents]] [[Patent Documents]]

[0003] [[Patent Document 1]] Japanese Patent Application Laid-Open No. 2008-301355 [[Patent Document 2]] Japanese Patent Application Laid-Open No. 2002-77711 [[Patent Document 3]] Japanese Patent Application Laid-Open No. 2008-15754 [[Non-Patent Documents]]

[0004] [[Non-Patent Document 1]] Shibuya Aki, Sugiyama Yoshiaki, Ariki Yasuo, "Automatic Classification of Sports Articles and Retrieval of Similar Scenes", Information Processing Society of Japan, Proceedings of the 55th National Convention, pp.65-66 (September 1997) [[Non-Patent Document 2]] Araki Ryosuke, Mano Kosuke, Onishi Takeshi, Hirano Masanori, Hirakawa Tsubasa, Yamashita Takayoshi, Fujiyoshi Hirohito, "Object Grasping Using Object Pose Estimation by Iterative Update Based on Backpropagation of Image Generation Network", Annual Conference of the Robotics Society of Japan, 2020. [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] In sports involving swinging motions with clubs, bats, rackets, etc., such as golf, baseball batting, and tennis (hereinafter referred to as swing sports), it is generally preferable that the subject's face is free from blur. On the other hand, motion blur in the swinging area, such as the club or arm, adds a sense of dynamism and conveys the intensity of the play more impressively to the viewer, resulting in a more favorable image. However, the aforementioned sports shooting mode captures images without motion blur in the swinging area. Furthermore, while the prior art disclosed in Patent Document 1 can capture images without motion blur in the face area, it does not consider images with motion blur in the swinging area.

[0006] This invention was made in view of the above-mentioned problems, and aims to capture dynamic images by automatically determining shooting parameters, including shutter speed, based on the motion analysis results of the face and swing region and the estimation results of the direction of movement of the swing region, taking into account the different swing speeds and directions for each scene. [Means for solving the problem]

[0007] To achieve the above objective, the image processing apparatus according to the present invention includes an acquisition means for acquiring an image, The system comprises: recognition means capable of recognizing the scene in which the image was captured; motion detection means for detecting the amount of motion between multiple frames in a predetermined region of the image acquired by the acquisition means; and shooting parameter determination means for determining shooting parameters. The aforementioned predetermined region is a region corresponding to the object being swung in a sport involving a swinging motion. When the recognition means recognizes that the scene in which the image was taken is a predetermined scene, the shooting parameter determination means determines the shooting parameters such that the amount of blur in the predetermined area is greater than a predetermined standard. [Effects of the Invention]

[0008] According to the present invention, an imaging device, imaging method, and program are provided that enable the capture of dynamic images by automatically determining shooting parameters, including shutter speed, based on the motion analysis results of the face and swing region and the estimation results of the movement direction of the swing region, taking into account different swing speeds and directions for each scene. [Brief explanation of the drawing]

[0009] [Figure 1] A diagram illustrating the configuration of the digital camera in Example 1. [Figure 2] A diagram illustrating the configuration of the imaging control parameter generation unit in Example 1. [Figure 3] A flowchart illustrating the operation of the shooting control parameter generation unit in Example 1. [Figure 4] A diagram illustrating the images taken in succession in Example 1. [Figure 5] A diagram illustrating the effect of superimposing scene recognition results onto the display image in Example 1. [Figure 6] A flowchart illustrating the conventional method for calculating motion vectors. [Figure 7] A diagram illustrating the conventional method for calculating motion vectors. [Figure 8] A figure showing the image at time t in Example 1 with a motion vector superimposed. [Figure 9] A diagram illustrating the correction gain applied to the motion vector in Example 1. [Figure 10] A diagram illustrating the defocus map corresponding to the image at time t in Example 1. [Figure 11] A diagram illustrating the image captured in Example 1. [Figure 12] A diagram illustrating the configuration of the image synthesis processing unit in Example 2. [Figure 13] A diagram illustrating the synthesis map in Example 2. [Figure 14]A flowchart for explaining a method of determining synthesis characteristics in Example 2. [Figure 15] A diagram for explaining a short reference frame and a synthesis order in Example 2. [Figure 16] A flowchart for explaining a method of determining a short reference frame in Example 2. [Figure 17] A diagram for explaining movement states of a main subject and objects in Example 2. [Figure 18] A diagram for explaining temporal changes in the movement speed of the main subject in Example 2. [Figure 19] A diagram for explaining the amount of motion and a correction ratio of a synthesis mask in Example 2. [Figure 20] A diagram for explaining the configuration of each processing unit for image capturing in Example 3. [Figure 21] A flowchart for explaining an operation during capturing in Example 3. [Figure 22] A diagram for explaining an exposure period and the like in Example 3. [Figure 23] A diagram for explaining an image when a face image is blurred. [Figure 24] A diagram for explaining details of an image when a face image is blurred. DESCRIPTION OF EMBODIMENTS

[0010] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The following embodiments do not limit the claimed invention. In addition, although a plurality of features are described in the embodiments, not all of the features are essential to the invention, and a plurality of features may be arbitrarily combined. Furthermore, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant explanations are omitted. EXAMPLES

[0011] A preferred embodiment of the present invention will be described in detail below with reference to the drawings. The embodiment described below is an imaging device, and an example of applying the present invention to a digital camera as an example of an imaging device will be described.

[0012] Figure 1 is a block diagram showing the functional configuration of a digital camera according to an embodiment of the present invention. The control unit 101 is, for example, a CPU, which reads the operation programs of each block in the digital camera 100 from the ROM 102, expands them in the RAM 103, and executes them to control the operation of each block in the digital camera 100. The ROM 102 is a rewritable non-volatile memory that stores the operation programs of each block in the digital camera 100, as well as parameters necessary for the operation of each block. The RAM 103 is a rewritable volatile memory that is used as a temporary storage area for data output during the operation of each block in the digital camera 100.

[0013] The optical system 104 forms an image of the subject on the imaging unit 105. The optical system 104 includes, for example, a fixed lens, a variable magnification lens for changing the focal length, and a focus lens for adjusting the focus. The optical system 104 also includes an aperture, which adjusts the aperture diameter of the optical system to control the amount of light during shooting. The imaging unit 105 is an image sensor such as a CCD or CMOS sensor, which performs photoelectric conversion of the optical image formed on the image sensor by the optical system 104 and outputs the resulting analog image signal to the A / D conversion unit 106. The A / D conversion unit 106 applies A / D conversion processing to the input analog image signal and outputs the resulting digital image data to the RAM 103 for storage.

[0014] The image processing unit 107 applies various image processing, such as white balance adjustment, color interpolation, and gamma processing, to the image data stored in the RAM 103 and outputs the resulting image data back to the RAM 103. The image processing unit 107 also includes a shooting control parameter generation unit 200, which will be described later. This unit recognizes the scene in the image data stored in the RAM 103 and generates shooting parameters for the digital camera 100 based on motion analysis results using the image data and estimation results of the subject's movement direction. The shooting control parameters generated by the image processing unit 107 are output to the control unit 101, which then controls the operation of each block in the digital camera 100.

[0015] The recording medium 108 is a removable memory card or the like, and images processed by the image processing unit 107 stored in the RAM 103, or images converted A / D by the A / D conversion unit 106, are recorded as recorded images. The display unit 109 is a display device such as an LCD, and provides various information in the digital camera 100, such as displaying the subject image captured by the imaging unit 105 as a pass-through to form an electronic viewfinder function, or displaying images recorded on the recording medium 108. It can also display icons based on the scene recognition results of the image data from the image processing unit 107 superimposed on the image.

[0016] The operation input unit 110 includes, for example, a user input interface such as a release switch, a setting button, or a mode setting dial, and when it detects an operation input made by the user, it outputs a control signal corresponding to the operation input to the control unit 101. In a configuration in which the display unit 109 is equipped with a touch panel sensor, the operation input unit 110 also functions as an interface for detecting touch operations made on the display unit 109.

[0017] The above explains the configuration and basic operation of the Digital Camera 100.

[0018] Next, the operation of the image processing unit 107, which is a feature of Embodiment 1 of the present invention, will be described in detail. Embodiment 1 of the present invention describes an example in which a digital camera is generated to capture an image of a subject performing a golf swing, with motion blur suppressed in the face area and motion blur appearing in the swinging area, and the subject is photographed.

[0019] First, an example of the configuration of the shooting control parameter generation unit 200, which is included in the image processing unit 107, will be explained with reference to Figure 2. Figure 2 is a diagram showing an example of the configuration of the shooting control parameter generation unit 200. The shooting control parameter generation unit 200 consists of a main subject detection unit 201, a region extraction unit 202, a scene recognition unit 203, a motion vector calculation unit 204, a motion vector correction unit 205, a depth change detection unit 206, and a shooting parameter configuration unit 207. The shooting control parameter generation unit 200 receives image data 208 captured by the imaging unit 105 and recorded in the RAM 103, and outputs a scene recognition result 209 and shooting parameters 210.

[0020] Next, the processing of the shooting control parameter generation unit 200 will be explained using the flowchart in Figure 3.

[0021] In S301, the user turns on the digital camera 100 and begins preparatory shooting, such as composing the shot. The control unit 101 continuously captures images while maintaining a predetermined frame rate during preparatory shooting. The captured images are displayed on the display unit 109, and the user adjusts the composition while viewing the displayed images. In this embodiment, the frame rate is 120 frames per second. That is, the imaging unit 105 captures one image every 1 / 120th of a second. The shutter speed at this time is set to be as short as possible. An example of continuously captured images is shown in Figure 4. The image at time t is denoted as 401, and the image at time t+1 is denoted as 402. Figure 4 shows an attempt to photograph a person swinging a golf club, and the image includes the person subject 403 and a bird 404 flying further away than the person subject 403, which is not the intended subject. Because the images are taken with a short exposure time, there is no significant blurring in the captured images. The image size is 2100 x 1400 pixels, and the pixel size is 5 μm. In this embodiment, golf is used as an example of a sport involving a swinging motion (swing sport), but this embodiment can also be applied to other sports. Specifically, examples include tennis, badminton, table tennis, baseball, lacrosse, hockey, fencing, kendo, canoeing, and rowing. In each sport, it is possible to automatically generate shooting parameters that create motion blur in the swinging equipment, thereby capturing dynamic images. In this embodiment, sports are used as examples, but scenes involving swinging motions are not limited to these. For example, combat scenes with swords or blades, or scenes of casting with a fishing rod can be considered. In other words, depending on the scene, various variations of objects involved in the swinging motion (swinging objects) can be considered in addition to the sports equipment mentioned above.

[0022] In S302, the main subject detection unit 201, under the control of the control unit 101, detects the main subject in the image data 208 captured in S301. Here, a human subject 403 is detected. The method for detecting the human subject 403 as the main subject can be any known technique, for example, the method disclosed in Patent Document 2. Furthermore, by referring to the defocus map, which is distance distribution information described later, it is possible to extract the area of ​​the body, limbs, and the club being held, which are at a depth similar to that of the face.

[0023] In S303, the region extraction unit 202, under the control of the control unit 101, extracts the face region and the region of the swinging object (golf club) from the human subject 403 detected by the main subject detection unit 201. The method for extracting the face region and the club region can utilize known techniques; as described in Patent Document 2, the respective regions can be extracted by changing the object being detected.

[0024] In S304, the scene recognition unit 203 recognizes the scene in the image data 208 under the control of the control unit 101. Here, it recognizes the type of sport (target scene). For the method of recognizing what sport is being played in the input image data, known techniques can be used, for example, the method disclosed in Non-Patent Document 1 can be used. It can recognize that not only golf, but also tennis and baseball are being played. The scene recognition unit 203 can also output whether it succeeded or failed in recognizing the type of sport. Alternatively, the detection information of the main subject detected in S302 may be used. In that case, even if spectators or other objects in the surroundings are captured in the image, they can be excluded from the scene recognition target, improving the accuracy of the recognition.

[0025] Furthermore, the scene recognition unit 203 estimates the direction and range of movement of the club after the time of shooting, based on the recognition result of the type of sport, the orientation of the digital camera 100 and the person subject 403, and the estimation result of the position and orientation of the club from the image data of the club region extracted in S303. For example, consider the case where in the image, the subject playing golf is facing the camera, the club is located on the left side of the screen, and the club face is pointing to the right side of the screen. In this case, the club will move to the right along the bottom of the screen. Prediction results corresponding to the situation can be set in advance, and the direction and range of movement can be estimated by selecting the setting according to the situation. Also, the method for estimating the position and orientation of the club can be, for example, a method disclosed in Non-Patent Document 2.

[0026] The scene recognition unit 203 outputs a scene recognition result 209 containing information such as whether recognition was successful or not, the type of sport recognized, and the estimated direction and range of movement of swinging objects such as clubs.

[0027] Meanwhile, the image processing unit 107 can superimpose the estimated results of the type of sport recognized and the direction and range of movement of swinging objects such as clubs onto the image data according to the scene recognition result 209. The image data with the estimated results superimposed can be displayed on the display unit 109 under the control of the control unit 101. This is shown in Figure 5. In Figure 5, 501 is an icon indicating the recognized sport, 502 is the estimated direction of movement of the club, and the hatched area of ​​503 is the estimated range of movement of the club. By displaying the scene recognition results on the display unit 109 during preparation shooting, the user can confirm that the digital camera 100 is trying to generate shooting parameters that will allow it to capture dynamic images in swing sports. Note that the means of notifying the user are not limited to displaying icons; for example, the system may also notify the user by emitting sound.

[0028] In S305, the control unit 101 checks whether it has successfully recognized the type of sport output from the scene recognition unit 203. If recognition is successful, the process proceeds to S307 and continues. On the other hand, if recognition fails, the process proceeds to S306, where the shooting parameters for normal sports mode are set and the parameter determination process is completed. Specifically, a high shutter speed is set, and the aperture value and ISO sensitivity are set according to the brightness of the shooting environment.

[0029] In S307, the motion vector calculation unit 204 calculates the motion vector and motion vector reliability between images captured in succession, under the control of the control unit 101. A motion vector is a vector representation of the horizontal and vertical movement of a subject between images. The method for calculating the motion vector will be explained in detail with reference to Figures 6, 7, and 8.

[0030] Figure 6 is a flowchart showing the calculation process of motion vectors and motion vector reliability by the motion vector calculation unit 204. Figure 7 is a diagram showing the method for calculating motion vectors using the block matching method. In this embodiment, the block matching method is used as an example to explain the motion vector calculation method, but the motion vector calculation method is not limited to this example, and for example, the gradient method may also be used.

[0031] In step S601 of Figure 6, the motion vector calculation unit 204 receives two temporally adjacent captured images as input. In this embodiment, the motion vector calculation unit 204 sets the captured image at time t in Figure 4 as the reference frame and the captured image at time t+1 as the reference frame.

[0032] In step S602 of Figure 6, the motion vector calculation unit 204 places an N×N pixel reference block 702 in the reference frame 701, as shown in Figure 7. In this embodiment, the area in which the reference block 702 is placed is limited to the face area of ​​the human subject 403 and the golf club area extracted by the area extraction unit 202 in S303. This configuration allows for efficient analysis of only the motion information necessary to generate shooting parameters. In particular, the correlation calculation in S604, which will be described later, is computationally intensive, so by performing the calculation only in the necessary area, it becomes possible to generate shooting parameters at a faster speed, enabling the capture of dynamic images in swing sports without missing any details.

[0033] In step S603 of Figure 6, the motion vector calculation unit 204 sets the search range 705 for the reference frame 703, as shown in Figure 7, to the center coordinates of the reference block 702 of the reference frame 701 and the surrounding (N+n) pixels of the same coordinates 704. The setting of the search range 705 is also limited to the area around the face of the human subject 403 and the area around the golf club, as in step S602. In particular, the search range for the golf club is limited to the range of club movement estimated by the scene recognition unit 203 shown in step 503 of Figure 5. This configuration allows for efficient analysis of only the motion information necessary to generate the shooting parameters.

[0034] In S604 of Figure 6, the motion vector calculation unit 204 performs a correlation calculation between the reference block 702 of the reference frame 701 and the reference block 706 of N×N pixels at different coordinates within the search range 705 of the reference frame 703, and calculates a correlation value. The correlation value is calculated based on the sum of the absolute values ​​of the inter-frame differences for the pixels of the reference block 702 and the reference block 706. In other words, the coordinate with the smallest value of the sum of the absolute values ​​of the inter-frame differences is the coordinate with the highest correlation value. Note that the method for calculating the correlation value is not limited to calculating the sum of the absolute values ​​of the inter-frame differences; for example, a method of calculating the correlation value based on the sum of squared inter-frame differences or the normal cross-correlation value may also be used. In the example in Figure 7, it is assumed that the reference block 706 has the highest correlation.

[0035] In step S605 of Figure 6, the motion vector calculation unit 204 calculates a motion vector based on the reference block coordinates showing the highest correlation value obtained in S604, and the correlation value of that motion vector is defined as the motion vector reliability. In the example in Figure 7, within the search range 705 of the reference frame 703, the motion vector is determined based on the same coordinate 704 corresponding to the center coordinate of the reference block 702 of the reference frame 701 and the center coordinate of the reference block 706. In other words, the distance and direction between the same coordinate 704 and the center coordinate of the reference block 706 are determined as the motion vector. Furthermore, the correlation value, which is the result of the correlation calculation with the reference block 706 during the calculation of the motion vector, is determined as the motion vector reliability. Note that the motion vector reliability increases as the correlation value between the reference block and the reference block increases.

[0036] In S606 of Figure 6, the motion vector calculation unit 204 determines whether it has calculated motion vectors for the target locations where the reference block 702 should be placed in the reference frame 701, that is, in this embodiment, for the face region of the human subject 403 and the golf club region extracted by the region extraction unit 202 in S303. If the motion vector calculation unit 204 determines that it has calculated motion vectors for all target locations, it terminates the motion vector calculation process. On the other hand, if it determines that it has not calculated motion vectors for the target locations, it returns to S602 and repeats the subsequent processing. Furthermore, the reference block 702 may be configured to be set for each pixel in all pixels included in the face region and the golf club region, or it may be configured to be set only for representative pixels in each region. In this embodiment, the explanation will be based on the configuration where it is set only for representative pixels.

[0037] Figure 8 shows the motion vectors between captured images calculated based on the processing described above. In Figure 8, the arrows indicate motion vectors, the length of the arrows indicates the magnitude of the motion vector, and the direction of the arrows indicates the direction of the motion vector. 801 is the motion vector of the face area of ​​the human subject 403, and 802 is the motion vector of the golf club area. It can be seen that the amount of movement between frames is greater than that of the face area because the golf club is being swung at high speed. On the other hand, the dashed line 803 is the motion vector of the bird in the background. As explained in S602, a reference block is not placed on the bird 404 in this embodiment, but in order to explain the control characteristics of this embodiment, the motion vector of the bird 404 is shown as the dashed line 803. Motion vectors 802 and 803 have the same magnitude and direction. In the shooting parameter generation unit 207, which will be described later, if only the magnitude and direction of the motion vectors within the screen are referred to when determining the shutter speed by referring to the motion vectors calculated here, there is a possibility that the motion vector of the moving subject in the background will have an influence. In this embodiment, by performing region extraction and scene recognition in advance, the influence of motion vectors in unrelated subject areas such as 803 can be excluded, thereby improving the accuracy of the generated shooting parameters.

[0038] Let's return to Figure 3 and continue explaining the process.

[0039] In S308, the motion vector correction unit 205 corrects the motion vector output by the motion vector calculation unit 204 based on the estimated movement direction information of the swinging object region, such as a club, output by the scene recognition unit 203. Specifically, when the angle between the estimated movement direction generated in S304 and the motion vector of the club region calculated in S307 is α (0° to 180°), the correction gain shown in Figure 9(a) should be multiplied by the motion vector of the club region. The closer the direction of the estimated direction and the calculated motion vector are (the closer the angle is to 0°), the greater the gain becomes, and the calculated motion vector will be used as is.

[0040] On the other hand, if the angle between the estimated direction and the calculated direction of the motion vector differs significantly, it is possible that the motion vector has not been calculated correctly, so the system should be discouraged from using such a motion vector to determine shooting parameters. Examples of situations where the motion vector cannot be calculated correctly include the following: The first example is when a swinging object is moving at high speed, and the image obtained in the preparation shot contains motion blur, making the edges of the subject unclear. The second example is when false edges caused by double-line blur result in the calculation of an incorrect correlation value in the correlation calculation. Also, when shooting in a dark environment, the accuracy of the motion vector decreases due to the influence of random noise caused by increased ISO sensitivity.

[0041] As in this embodiment, by using not only motion vectors calculated from differences in pixel values ​​between frames, but also machine learning and deep learning techniques such as position and orientation estimation, it becomes possible to generate highly accurate shooting parameters even for swinging objects. Furthermore, as a method for correcting motion vectors, the reliability of the motion vectors calculated in S307 may also be utilized. As shown in Figure 9(b), a correction gain is multiplied by the motion vector such that the higher the reliability, the greater the gain amount becomes (1.0). In this embodiment, a correction gain of 1.0 is always applied to the motion vectors of the face region.

[0042] In S309, the control unit 101 compares the magnitude of the corrected motion vector output from the motion vector correction unit 205 with a predetermined threshold. If the magnitude of the motion vector in the face region or the motion vector in the club region is greater than the threshold, that is, if the face region where motion blur is not desired is moving, or if the region of the swinging object where motion blur is desired is moving sufficiently, the process proceeds to S310. On the other hand, if the motion vector is less than or equal to the threshold, the image data is updated and the process from S302 is repeated. In this embodiment, we will proceed with the explanation assuming that the motion vector of the club region shown at 802 in Figure 8 is greater than the threshold. Here, the thresholds for comparing the motion vectors in the face region and the region of the swinging object can be set to different values. It is preferable to configure the system so that the threshold is changed according to whether the user prioritizes preventing motion blur or camera shake in the subject, or whether they prioritize greater motion blur in the swing region. By configuring the system to set a threshold, it becomes possible to capture images as desired by the user more flexibly. In addition, the threshold may be changed according to the type of sport. Specifically, in golf and baseball batting, the subject does not run or change position before and after the swing. Therefore, to prioritize minimizing motion blur in the swing region, the threshold for comparing with the motion vector in the swing region is set to be small. On the other hand, in tennis and badminton, the subject itself may move significantly during the swing, so to prioritize preventing motion blur of the subject and camera shake, the threshold is set to be large. By configuring it in this way, it is possible to generate shooting parameters that correspond to the characteristics of each sport, even within swing sports, and to capture more accurate and dynamic images. Alternatively, the threshold can be changed depending on how the subject is facing the digital camera 100. Specifically, when shooting a subject swinging a golf club from the front, as shown in Figure 4, the magnitude of the motion vector between frames tends to change. Therefore, motion blur in the club region is likely to occur, but when shooting from behind, the amount of horizontal and vertical movement in the image appears smaller.Therefore, the threshold can be made looser (larger) than when shooting from the front. In other words, it is effective to change the threshold used to compare the motion vector in the swing region according to the relative positional relationship between the subject and the imaging device.

[0043] Furthermore, in this embodiment, the magnitude of the motion vector between two image frames was compared with a threshold, but it is also possible to configure the system to compare the threshold with the result of accumulating motion vectors between multiple frames (vector accumulation result). By continuously monitoring the motion vector from the start of preparation shooting, it is possible to analyze the average motion even when there is rapid acceleration or deceleration of a swinging object, and a more stable threshold determination result can be obtained. As a result, the generated shooting parameters become more stable, the success rate of capturing dynamic images improves, and usability is improved.

[0044] In S310, the depth change detection unit 206 detects changes in depth between images of continuously captured images under the control of the control unit 101. Changes in depth can be detected by attaching a defocus map as metadata when capturing images. Figure 10 is a defocus map corresponding to 401 in Figure 4. The defocus map shows the degree of focus on the imaging surface in the form of a grayscale map, with white in the foreground, black in the background, and 50% gray indicating focus. It can be seen that the human subject 403 is in focus, and the bird 404 is on the background side of the human subject 403. The defocus map is calculated using known techniques, and for example, as disclosed in Patent Document 3, it is sufficient to configure the system to acquire a defocus map on the imaging surface from a phase difference image obtained from an imaging sensor in which all pixels are phase difference pixels.

[0045] The depth change detection unit 206 calculates distance distribution information (distance information calculation). The depth change detection unit 206 receives the coordinates of the face region generated by the region extraction unit 202 and calculates the average defocus amount in the face region. Furthermore, it also calculates the average defocus amount of the face region in the defocus map captured at the next time point. Then, it calculates the difference in the average defocus amounts at different time points as the depth change amount.

[0046] Assuming the focus lens position remains the same at different times, the change in depth of field can be treated as the amount of movement of the human subject in the depth direction. If the subject moves in the depth direction while the focus lens position remains unchanged, depth blur will occur in the face area of ​​the human subject. Since it is desirable that not only motion blur but also depth blur does not occur in the face area, the shooting parameter generation unit 207, described later, appropriately reduces the aperture value according to the change in depth of field to increase the depth of field and suppress depth blur. Reducing the aperture value can increase the depth of field, but it also limits the amount of incoming light, so the ISO sensitivity needs to be increased to shoot at the same brightness. On the other hand, if the ISO sensitivity is too high, random noise becomes noticeable and the image quality deteriorates, so it is not simply a matter of arbitrarily reducing the aperture value. Therefore, as in this embodiment, by selecting an appropriate aperture value according to the change in the subject's movement for each scene, it is possible to capture higher quality images.

[0047] In this embodiment, the change in the depth direction was calculated from a defocus map, but this is not the only way to do so. Any information corresponding to the distance distribution in the depth direction of the subject within the imaging range is acceptable. For example, it could be the defocus amount distribution after normalization by the depth of field, or a depth map showing the subject distance of each pixel. It could also be two-dimensional information showing the phase difference (amount of image shift occurring between different viewpoints) used to derive the defocus amount. Furthermore, it could be a map converted to actual distance information on the subject side via the focus lens position. In other words, any information showing a change in accordance with the distance distribution in the depth direction is acceptable, and information on the distribution of parallax (parallax distribution information) can also be utilized.

[0048] In S311, the shooting parameter generation unit 207, under the control of the control unit 101, generates shooting parameters that enable the capture of dynamic images of subjects performing sports involving swinging motions. Specifically, it generates values ​​for shutter speed, aperture, and ISO sensitivity. First, the shutter speed is determined by referring to the magnitude of the vector output by the motion vector correction unit 205. Here, the motion vector has horizontal and vertical components, but for simplicity, only the horizontal component will be explained. The vertical component can be calculated in the same way as the horizontal component.

[0049] Let's assume that the horizontal image size is 2100 pixels, and we want motion blur equivalent to approximately 10%, or 200 pixels. If the magnitude of the vector output by the motion vector correction unit 205 is 100 pixels, and the frame rate is 120 frames per second, then by setting the shutter speed to 1 / 60 second, we can capture an image with motion blur of the desired magnitude.

[0050] Next, the aperture value is determined by referring to the change in depth between images in the continuously captured images output by the depth change detection unit 206. Here, the amount of defocus is further normalized by the depth of field (for example, 1Fδ, where F is the aperture value and δ is the allowable circle of confusion diameter, which in this embodiment is twice the pixel size) and defined as the amount of depth blur. It is assumed that when the amount of depth blur is greater than 1.0Fδ, the observer will perceive that the image is blurred. If the amount of change output by the depth change detection unit 206 is 0.05mm, then selecting an aperture value of F5.6 will result in 0.05 / (5.6×5×2×10^-3)=0.89Fδ, and the observer will no longer perceive that depth blur is occurring.

[0051] Finally, the ISO sensitivity should be selected based on the brightness of the shooting environment, so that the image is taken with the appropriate brightness when the shutter speed is 1 / 60 second and the aperture value is F5.6. Let's say you are shooting outdoors on a cloudy day and the EV (Exposure Value) is 11. The shutter speed is 1 / 60 second, so the TV (Time Value) is 6, and the aperture value is F5.6, so the AV (Aperture Value) is 5. From EV = TV + AV - SV, SV is 0, and the ISO sensitivity is determined to be 100.

[0052] In S312, the control unit 101 performs the main shooting according to the shooting parameters determined in S311, records the captured image on the recording medium 108, and ends the series of processes. Figure 11 shows the captured image. Motion blur and depth of field blur do not occur in the face of the human subject. On the other hand, motion blur occurs in the club area, resulting in a dynamic image.

[0053] In this embodiment, the defocus map was described as being generated based on a group of images that have a parallax relationship (a group of images with different viewpoints), but this method is not limited to this method as long as it corresponds to the captured image and allows for the acquisition of the distance distribution of the subject within the imaging range. The method for generating the defocus map may be, for example, the DFD (Depth From DefocuS) method, which derives the amount of defocus from the correlation of two images with different focus and aperture values. Alternatively, the distance distribution of the subject may be derived using information related to the actual distance distribution obtained from measurements of a distance measuring sensor module such as the TOF (Time of Flight) method. Or, it may be based on the contrast distribution information of the captured image obtained by the contrast distance measuring method. [Examples]

[0054] A second embodiment of the present invention will be described below. The embodiment described below is an imaging device, similar to Example 1, and will be explained using an example in which the present invention is applied to a digital camera as an example of an imaging device.

[0055] In the first embodiment, it was possible to capture images with a sense of motion by controlling the shooting parameters, but in this embodiment, by combining multiple images, an image is generated that achieves both a sense of motion through long exposure shooting and a still image with locally reduced subject blur.

[0056] The configuration of the digital camera according to the embodiment of the present invention is the same as the block diagram in Figure 1 described in Example 1, so a detailed explanation is omitted. The user can switch between the first embodiment and a shooting mode specific to this embodiment by operating the operation input unit 110 of the imaging device, etc., according to the shooting scene. Furthermore, the imaging device can be configured to automatically determine the scene and switch the shooting mode.

[0057] Unlike Example 1, in this embodiment, the image processing unit 107 is equipped with an image synthesis processing unit 1200 as shown in Figure 12. The processing of the image synthesis processing unit 1200 will be described in detail below.

[0058] In this embodiment, it is assumed that a scene is being photographed in which there is a moving object in the background (such as a fountain, waterfall, or flow of people) and a person, the main subject, is in the foreground, as shown in Figure 13(a). In such a scene, the objective is to capture an image in which the background has a sense of motion equivalent to long exposure, while suppressing motion blur in the person area. Another objective is to generate an image with a sense of motion by adding a motion trajectory in scenes where the main subject is moving. In Figure 12, the image synthesis processing unit 1200 consists of an image storage unit 1201, a main subject area detection unit 1202, a main subject related feature detection unit 1203, a synthesis characteristic control unit 1204, and an average synthesis unit 1205. Furthermore, each part of the image processing processing unit 1200 executes its respective function under the command of the control unit 101.

[0059] The image synthesis processing unit 1200 stores RAW image data captured by the imaging unit 105 or image data after development processing in the image storage unit 1201 for a predetermined number of frames. The input image is to be a continuous frame with no non-exposure period between frames. The image data stored in the image storage unit 1201 is input to the averaging synthesis unit 1205, which generates an image equivalent to a long exposure by performing an averaging process on a pixel-by-pixel basis across multiple frames.

[0060] The number of frames stored in the image storage unit 1201 is determined by the shutter speed set by the user and the shutter speed used for image capture.

[0061] In this embodiment, the shutter speed for imaging is fixed at 1 / 100 second. If the shutter speed is set to 1 / 2 second by user settings, the number of frames captured and stored will be 50. Then, by averaging 50 images taken at a shutter speed of 1 / 100 second, an image is produced that has accumulated blur equivalent to 1 / 2 second and has brightness equivalent to an image taken at 1 / 100 second.

[0062] Image data stored in the image storage unit 1201 is also output to the main subject area extraction unit 1202 and the main subject-related feature detection unit 1203.

[0063] The main subject region extraction unit 1202 extracts the region of the main subject and outputs it as a main subject region map. The main subject is to be extracted using known methods such as machine learning. In this embodiment, the case of extracting the region of a person as the main subject will be explained as an example. An example of a main subject region map is shown in Figure 13. Figure 13(a) shows the input image to the main subject region extraction unit 1202. Figure 13(b) is the extracted main subject (person) map, in which the main subject region 1301 has a signal of 1 (white) and the background region has a signal of 0 (black). The extracted main subject region map is output to the main subject related feature detection unit 1203 and the composite characteristic control unit 1204.

[0064] Furthermore, the main subject area extraction unit 1202 also calculates the position information of the main subject. For example, the position of a person's head or the center of gravity of the human body can be used as the position of the main subject. The main subject area extraction unit 1202 generates a main subject area map and calculates the position of the main subject for multiple frames stored in the image storage unit 1201.

[0065] The main subject-related feature detection unit 1203 detects the amount of motion between frames in the main subject area and the main subject background area (motion detection). The main subject background area is set based on the input main subject area map. The main subject background area is the background area surrounding the main subject. An example of the main subject background area is shown in Figure 13(c). In Figure 13(c), the area 1302 shown in white is the main subject background area. The main subject-related feature detection unit 1203 calculates the absolute difference between multiple frames accumulated in the main subject area image accumulation unit 1201 in the main subject area 1301 and the main subject background area 1302, and normalizes it to a predetermined area. This is used as an indicator to detect whether or not there is motion blur in the main subject and whether or not there is a moving subject in the main subject background area.

[0066] The composite characteristics control unit 1204 determines the composite characteristics in the averaging composite unit 1205 based on the region map information and position information of the main subject output from the main subject region extraction unit 1202, and the movement information of the main subject and background region output from the main subject related feature detection unit 1203. The details of the processing of the composite characteristics control unit 1204 will be described later, but the composite characteristics control unit 1204 outputs the composite map to be used for composite and the composite order information as composite information. An example of a composite map is shown in Figure 13(d).

[0067] The averaging blending unit 1205 averages the images output from the image storage unit 1201 based on the blending information output from the blending characteristics control unit 1204. At this time, by varying the number of blending frames for each pixel based on the blending map, images with different shutter speeds are generated for each region. In this embodiment, the main subject region outputs pixels with short shutter speeds by reducing the number of frames to be averaged, while the background region outputs pixels equivalent to long exposures with long shutter speeds by increasing the number of average blending frames.

[0068] The processing of the image synthesis processing unit 1200 has been described above. The image data generated in this manner is recorded on 108, or, in the case of RAW data, is processed by the development processing unit provided in the image processing unit 107.

[0069] Next, the details of the processing of the composite characteristic control unit 1204 will be explained using the flowchart in Figure 14.

[0070] In step S1401, the composite characteristics control unit 1204 sets a short-second reference frame. The average composite unit 1205 generates an image that includes both long-second and short-second pixels based on the composite information, as described above. In this case, the shortest-second pixels are output as a single frame image without being composited. The frame used as these shortest-second pixels is the short-second reference frame.

[0071] A specific example of a short-second reference frame is described below. Figure 15(a) shows images stored in the image storage unit 1201. The horizontal axis represents the time axis in which the images were captured, and it shows that 10 images captured in the order of 1 to 10 have been stored.

[0072] As the long-exposure-equivalent pixels generated by the average combining unit 1205, by averaging all of these 10 frames from 1 to 10, an image can be generated that has an accumulated blur amount equivalent to that captured by a long exposure for 10 frames, and has the same brightness as before combining. Meanwhile, as the short-exposure pixels, the pixels of the short-exposure reference frame are output as they are. Whether the average combining unit 1205 outputs short-exposure pixels or long-exposure-equivalent pixels in an image is determined according to the combining map shown in FIG. 13(d).

[0073] The combining map shown in FIG. 13(d) has a value of 0 or 1. When the value is 0, the average combining unit 1205 outputs the long-exposure-equivalent pixels combined with the maximum number of combining sheets; when the value is 1, it outputs the pixels of the short-exposure reference frame. The signal value M of the combining mask has an intermediate value satisfying 0<M<1, and in the case of an intermediate value, the number of combined frames N is determined based on the following formula.

[0074] N=(Max-1)×(1-M)+1 Here, Max is the maximum number of combined frames.

[0075] For example, in the case of a pixel where Max=10 and M=0.8, N=2, and a signal obtained by averaging two images is output.

[0076] The short-exposure reference frame can be selected from any frame used for combining, and which frame is to be used as the short-exposure reference frame is determined according to the flow described later.

[0077] In the example of FIG. 15(b), a case is shown where frame 1501 captured last in terms of time is the short-exposure reference frame. Similarly, in FIG. 15(c), an example is shown where frame 1502 captured first in terms of time is the short-exposure reference frame, and in FIG. 15(d), frame 1503 captured sixth is the short-exposure reference frame.

[0078] A method for setting a short-exposure reference frame will be described based on the flowchart in FIG. 16.

[0079] In step S1601, the main subject area extraction unit 1202 acquires the main subject position information for each frame and calculates the amount of positional change of the main subject between the first and last frames. An example of the amount of positional change is shown in Figure 17(a). 1701 is the position of the main subject in the first frame, 1702 is the position of the main subject in the last frame, and arrow 1703 indicates the amount of positional change.

[0080] In step S1602, the main subject-related feature detection unit 1203 determines whether the main subject has moved between the accumulated frames. Specifically, if the amount of positional change of the main subject calculated by the main subject region extraction unit 1202 in step S1601 is greater than the threshold TH1, it is determined that there has been movement; if it is less than TH1, it is determined that there has been no movement. If it is determined that there has been no movement of the main subject, the process proceeds to step S1603. If it is determined that there has been movement, the process proceeds to step S1606.

[0081] In step S1603, the main subject-related feature detection unit 1203 determines whether the motion information of the background region output is greater than the threshold TH2. The motion information of the background region is assumed to be calculated across multiple frames, and if even one movement between frames is greater than the threshold TH2, it is determined that there is motion. If it is determined that there is motion in the background, the process proceeds to step S1604; if it is determined that there is no motion, the process proceeds to step S1605.

[0082] In step S1604, the composite characteristics control unit 1204 selects a frame with no moving subject in the background as the short-second reference frame. A scene with moving background is explained using Figures 17(b) and 17(c). In Figures 17(b) and 17(c), 1704 represents the background area of ​​a person, and 1705 represents a moving object other than a person (a bird). As in Figure 17(b), if there is no subject other than a person in the background area 1704 of the person, the amount of movement in the background area is small, but when the bird moves and enters the background area 1704 as in Figure 17(c), the amount of movement becomes large. Here, the short-second reference frame is selected from frames excluding frames with a large amount of background movement. Specifically, the composite characteristics control unit 1204 uses the frame furthest in time from the frame with a large amount of movement as the short-second reference frame. An example is shown in Figure 15(e). In Figure 15(e), 1504 shows multiple frames with a motion amount greater than the threshold TH2. In this example, frame 1505, which is the frame furthest in time from the group of frames 1504 with the most motion, is selected as the short-second reference frame. However, in cases where the background area is constantly moving, such as in the fountain example in Figure 13(a), the final frame or the frame with the least amount of motion is used as the short-second reference frame.

[0083] By controlling the process as described above, it is possible to prevent images that are difficult to use as still frames, such as when a bird overlaps the area of ​​a person whose movement you want to freeze, from being used as the short-second reference frame.

[0084] In the example above, the short-second reference frame was determined based only on the amount of movement in the background area. However, a configuration that considers the amount of movement in the person's area in addition to the background area may also be used. Furthermore, in addition to determining based on the amount of movement, object detection may be performed, and a configuration that uses frames in which no objects other than the main subject are included in the background area or near the main subject area as the short-second reference frame may be used.

[0085] Step S1605 is the case where there is no moving subject in the background. In this case, there is no major problem in choosing any frame as the short-second reference frame, so a predetermined frame (for example, the final frame) is used as the short-second reference frame.

[0086] In step S1606, the main subject-related feature detection unit 1203 determines whether there is a change in the subject's movement speed when the subject is moving. Figure 18 is a graph showing the relationship between the subject's movement speed and the change in time. The movement speed of the main subject is calculated from the amount of positional change between multiple captured frames. The maximum value Vmax and minimum value Vmin of the movement speed are also calculated. If the difference between the maximum value Vmax and the minimum value Vmin is greater than or equal to the threshold TH3, it is determined that there is a change in the subject's movement speed; if it is less than the threshold TH3, it is determined that there is no change in movement speed. Figure 18 shows examples where (a) there is a change in movement speed and (b) there is no change in movement speed. A large change in movement speed occurs in scenes with significant speed variations in movement, such as when the main subject is jumping or riding on a swing. If it is determined that there is a change in movement speed, the process proceeds to step S1607; if it is determined that there is no change in movement speed, the process proceeds to step S1608.

[0087] In step S1607, the composite characteristics control unit 1204 selects the frame in which the subject's movement speed is smallest as the short-second reference frame. In the example in Figure 18(a), the frame captured at time T1, when the movement speed is smallest, is selected as the short-second reference frame. By using the time when the movement speed is smallest as the short-second reference frame, it becomes possible to generate an image with a clear distinction between the motion region and the still region.

[0088] In step S1608, the composite characteristics control unit 1204 uses the last captured frame as the short-second reference frame. This makes it possible to generate an image that represents the trajectory of the main subject's movement.

[0089] The above explains the method for selecting a short-second reference frame. In the example above, we described an example in which a short-second reference frame is selected based on the movement information of the main subject and the background of the main subject, as well as the movement information of the main subject. However, any information related to the main subject can be used to select a short-second reference frame. For example, it is possible to use the distance information of the main subject and set the frame with the farthest (or closest) distance as the short-second reference frame.

[0090] Returning to Figure 14, in step S1402, the synthesis order is set based on the short-second reference frame. Examples of synthesis orders are shown in Figures 15(b) to (d).

[0091] The numbers in each frame of Figures 15(b) to (d) indicate the blending order. Frames with a blending order of 1 are short-second reference frames. As mentioned above, the averaging blending unit 1205 performs averaging blending of multiple frames, and the blending order indicates the order in which the frames are used for averaging blending. For example, when there are 3 frames to blend, images 1 to 3 are blended, and when there are 5 frames to blend, frames 1 to 5 are blended.

[0092] In this step, the synthesis characteristics control unit 1204 sets the synthesis order as shown in Figures 15(b) to (d). The synthesis order is set based on the short-second reference frame, with frames being ordered according to their temporal proximity to the short-second reference frame. If the time difference between preceding and succeeding frames is the same, the later frame takes precedence.

[0093] Examples of setting the composition order according to the above are shown in (b) to (d). Figure 15(b) shows the composition order when the short-second reference frame is the last frame, Figure 15(c) shows the composition order when the short-second reference frame is the first frame, and Figure 15(d) shows the composition order when the short-second reference frame is the sixth frame from the beginning.

[0094] In step S1403, a composite map as shown in Figure 13(d) is generated based on the main subject area map output from the main subject area extraction unit 1202. Since the main subject area map is binary data of 0 or 1, a low-pass filter is applied to the main subject area map to generate a map with intermediate values ​​between 0 and 1, and this is used as the composite map.

[0095] In step S1404, similar to step S1602 described above, the main subject-related feature detection unit 1203 determines whether the main subject has moved between frames. If there is no movement of the main subject, the process proceeds to step S1405; if there is movement of the main subject, the process terminates.

[0096] In step S1405, the composite characteristics control unit 1204 corrects the composite map generated in step S1403 based on the movement of the main subject and background. The correction method will be explained using Figure 19. Figure 19 is a diagram showing the amount of movement of the main subject area or the background area of ​​the main subject and the correction ratio of the mask. The correction ratio of the mask indicates the percentage by which the composite mask is expanded or contracted. A value greater than 1 indicates expansion, and a value less than 1 indicates contraction. Figure 19(a) shows the relationship between the amount of movement of the main subject area and the correction ratio of the composite mask size. A correction coefficient is calculated to expand the composite mask as the amount of movement of the main subject area increases. Figure 13(e) shows an example of the composite mask being expanded, using the composite mask in Figure 13(d) as a reference. If the main subject does not move in position but does move, averaging composite will cause the motion blur to extend beyond the contour of the subject. Therefore, expanding the composite range has the effect of outputting short-second images of the blurred contour area as well, thus preventing blurring.

[0097] On the other hand, Figure 19(b) shows the relationship between the amount of movement in the background area of ​​the main subject and the mask size. A correction coefficient is calculated to correct the composite mask to shrink as the movement of the background area increases. Figure 13(f) shows an example of the composite mask being shrunk, using the composite mask in Figure 13(d) as a reference. This corresponds to a case where the main subject does not move in position, but there is movement in the background such as a fountain, waterfall, or flow of people. If there is movement in the background, compositing beyond the main subject area will result in an image where the background movement is only static around the outline of the main subject. To prevent this, the process is controlled to composite only the inside of the main subject as much as possible.

[0098] As described above, after calculating the mask correction coefficient based on the movement of the main subject area and the mask correction coefficient based on the movement of the background area of ​​the main subject, the final mask correction coefficient is calculated by multiplying the two mask coefficients.

[0099] In the above example, we described an example of expanding and contracting the composite mask as an example of a method for correcting the composite mask. However, any correction can be performed as long as the composite mask is corrected based on the characteristics of the main subject. For example, the composite characteristics control unit 1204 may perform a correction that controls the steepness of the gradient of the composite mask. In this case, the control may be such that the gradient of the composite mask is made steeper when there is movement in the background, and smoother when there is no movement in the background. This makes it possible to make the difference in motion blur between the main subject area and the background area less noticeable.

[0100] The configuration of this embodiment has been described above. By combining multiple images using the configuration of this embodiment, it is possible to generate an image that achieves both motion expression through long exposure and stillness expression with locally reduced subject blur.

[0101] In this embodiment, we have described an example of controlling the number and order of images used for averaging in the averaging unit 1205. However, if multiple images are to be combined, other configurations are also possible. For example, it is possible to pre-generate a long-exposure equivalent image by averaging all the accumulated images, and then partially combine the short-exposure reference image with the long-exposure equivalent image.

[0102] Furthermore, although the above embodiment described an example where all captured images were taken with the same shutter speed, it is also possible to configure the system to capture some images with different shutter speeds. Figure 15(f) shows multiple consecutively captured image frames stored in the image storage unit 1201. Frame 1506 is the final frame, and was taken with half the shutter speed of the other frames 1507. The long-exposure equivalent image is generated by averaging and combining all 11 frames. Frame 1506 is used as the short-exposure reference frame. In this case, since the exposure of frame 1506 is halved, it is used after applying a gain of 2. In addition, the composite characteristic control unit 1204 can be configured to select whether to use frame 1503 with an even shorter shutter speed as the short-exposure reference frame, or to use other frames with uniform brightness, based on the amount of motion of the main subject.

[0103] By capturing images with different shutter speeds in this way, it becomes easier to generate images with locally reduced blur when the subject is moving quickly or when motion blur is significant. [Examples]

[0104] The third embodiment will be described below using Figures 20 to 24. The third embodiment involves switching between the shooting method of the first embodiment and the shooting method of the second embodiment depending on the shooting conditions.

[0105] Here, the shooting method described in the first embodiment involves capturing one image in a single exposure (single image acquisition), resulting in an image that contains a mixture of parts where subject blur is suppressed and parts where subject blur is tolerated or even emphasized (hereinafter referred to as "single image capture").

[0106] In contrast, the shooting method described in the second embodiment involves taking multiple images with multiple exposures. For areas where subject blur should be suppressed, one less blurry image is selected from the multiple images. For areas where subject blur should be tolerated or emphasized, an image is generated by combining multiple images (multiple image synthesis). Furthermore, the aforementioned less blurry image is transformed into a blurry image through the synthesis of multiple images, resulting in an image in which areas with suppressed subject blur and areas where subject blur should be tolerated or emphasized are mixed (hereinafter referred to as "multiple image shooting").

[0107] In the case of a single shot, if the desired shutter speed can be determined, the image can be captured in one exposure, and subsequent image processing such as development is simple, allowing the image to be displayed quickly after shooting. On the other hand, because the exposure is set at a shutter speed that provides sufficient blur in the area to be blurred (the swing area in the first embodiment), if there is movement in the area where blur should be suppressed (the face area in the first embodiment) during this exposure period, blur will occur according to the amount of movement. If this amount of movement is within an acceptable range, a single shot is sufficient, but when trying to capture fine expressions in a face image, even a small amount of movement may be considered a blurred image, resulting in an unsatisfactory overall image. For example, as shown in the first embodiment, ideally, an image should be obtained where the club is blurred and the face area is still, as shown in Figure 11. However, if the face area moves even slightly during this exposure period, the face area may become blurred and unclear, as shown in Figure 23.

[0108] This will be explained using Figure 24. In Figure 24, the horizontal axis represents time, and the double arrow at 2401 indicates the exposure period (the period corresponding to the shutter speed) as explained in the first embodiment. 2402 to 2405 are magnified views of the subject's face during this exposure period, and similarly, 2406 to 2409 are magnified views of the swinging club. As shown in Figure 24, during the exposure period 2401, the club moves due to the swing, so the image is captured blurred, as shown in Figure 23, resulting in a dynamic image. On the other hand, if the face remains completely still during the exposure period 2401, a dynamic image can be obtained, as shown in Figure 11, where the face is clear and only the club is blurred. However, the face often moves. That is, as shown in Figures 2402 to 2405, the face image may move slightly. In this case, as shown in Figure 23, the club is a dynamic image, but the face image is also blurred. At this point, if you set a short exposure time (shutter speed), you might be able to capture a clear image of the face, but the club will also become still, and you won't get the dynamic image that was the original goal.

[0109] On the other hand, if multiple images are taken, as shown in the second embodiment, the facial image will be captured with an exposure time shorter than the exposure period, so an image in which the facial image is still and clearly captured (corresponding to a short frame in the second embodiment) can be easily obtained. Also, by combining multiple captured images, a blurred image can be obtained of the swinging club. Then, by combining the blurred image of the club with the facial image, a dynamic image can be obtained. However, the image compositing process takes time to generate the final image, making it difficult to display the image immediately after the exposure period ends. Furthermore, if image compositing is always assumed, devices that use batteries will always consume a lot of power, resulting in the drawback of reduced battery life.

[0110] Therefore, in this embodiment, by switching between taking a single shot and taking multiple shots depending on the motion state of the part of the image where blur must be suppressed, an image is efficiently created in which parts with suppressed subject blur and parts where subject blur is tolerated or even emphasized are mixed.

[0111] Figure 20 shows the configuration of each image capture processing unit in the third embodiment. In the figure, 2001 is the image acquisition unit, which corresponds to 105 in Figure 1 of the first embodiment. 2002 is the image processing unit that performs all image processing in this embodiment, and includes 2004, which corresponds to the image capture control parameter generation unit (Figure 2) in the first embodiment, and 2006, which corresponds to the image synthesis processing unit (Figure 12) in the second embodiment. 2003 shows the final image generated in this embodiment. In this embodiment, the final image 2003 is obtained by processing the image data from the image acquisition unit 2001 in the image processing unit 2002.

[0112] Here, we will explain assuming that the image acquisition unit outputs two types of image data. These are image data (hereinafter referred to as frame image data) 2008, which is output sequentially in frame units with short exposures, and image data (hereinafter referred to as captured image data) 2009, which is the result of shooting with a specified exposure time. Depending on the configuration and operation of the image acquisition unit, these two types of image data can be considered to pass through the same path, or they can be made into the same image data by adjusting the exposure time to match the frame interval. For convenience, the following explanation will refer to the two types of image data separately as frame image data 2008 and captured image data 2009.

[0113] In Figure 20, 2004 is the shooting control parameter generation unit, which corresponds to the shooting control parameter generation unit in the first embodiment (Figure 2). The shooting control parameter generation unit 2004 takes frame image data 2008 as input, determines the shooting parameters 2010, and outputs them. The input frame image data 2008 corresponds to the image data 208 in Figure 2 of the first embodiment. The output shooting parameters 2010 correspond to the shooting parameters 210 in Figure 2 of the first embodiment, and have values ​​for shutter speed, aperture value, and ISO sensitivity. In this embodiment, the shutter speed 2011 is extracted from these and input to the image synthesis processing unit 2006. The shooting parameters 2010 are input to the single-image shooting processing unit 2005 in Figure 20, and the single-image shooting processing unit 2005 specifies the shutter speed in the image shooting unit 2001 according to these shooting parameters, acquires the shooting image data 2009, and generates a single image of the shooting result.

[0114] Similarly, in Figure 20, 2006 is the image synthesis processing unit, which corresponds to the image synthesis processing unit (Figure 12) of the second embodiment. The image synthesis processing unit 2005 takes the frame image data 2008 as input, determines the parameters related to image synthesis, performs image synthesis processing, and takes multiple images. In this embodiment, the shutter speed 2011 is input to the image synthesis processing unit 2006. This shutter speed 2011 determines the parameters of the synthesis processing as the exposure period for shooting equivalent to the long exposure described in the second embodiment. That is, the image synthesis processing unit 2006 stores the frame image data 2008 in the image storage unit corresponding to the image storage unit 1201 in Figure 12 of the second embodiment for a period corresponding to the shutter speed 2011, and the stored images are synthesized in the average synthesis unit corresponding to the average synthesis unit 1205 of the second embodiment to generate an image equivalent to a long exposure.

[0115] In Figure 20, 2007 is a switch that switches between the image generated by single-image capture 2006 and the image generated by the image synthesis processing unit 2006, outputting one of them as the final image 2003.

[0116] Here, we will explain the flow of the third embodiment during shooting using the flowchart in Figure 21. Note that the following explanation will focus on the flow related to switching the shooting method in the third embodiment, and will omit explanations of flows that are explained in detail in the first and second embodiments.

[0117] In the diagram, step S2101 is the main subject detection step, which performs the detection of the main subject, corresponding to the process performed in S302 in Figure 3 of the first embodiment. That is, the main subject is detected from the frame image data from the image acquisition unit 2001.

[0118] Next, step S2102 is a scene recognition step, which performs scene recognition corresponding to the process performed in S304 in Figure 3 of the first embodiment. That is, the shooting scene is recognized from the main subject detected in S2101, and the conditions necessary for determining the next shooting parameters are set.

[0119] Next, step S2103 is a shooting parameter determination step, in which the shooting parameters are set, corresponding to the process performed in S306 or S311 in Figure 3 of the first embodiment. That is, the shooting parameters, namely the shutter speed, aperture value, and ISO sensitivity, are determined according to the conditions set in the scene recognition step. The shutter speed determined here is determined as the exposure period for capturing an image in which the final image contains a mixture of parts where subject blur is suppressed and parts where subject blur is to be tolerated or emphasized.

[0120] Next, in step S2104, the movement of the main subject is determined. Specifically, it is determined, in conjunction with the scene recognition result from S2102, whether the main subject detected in S2101 is likely to move during the exposure period corresponding to the shutter speed determined in S2103. Generally, if the determined shutter speed (exposure period) is long, there is a high probability that the main subject will move during that exposure period. Also, even with a short exposure period, the main subject may move depending on the type of sport being photographed and the scene involving specific movements. In S2104, it is determined whether or not movement occurs in the main subject during the exposure period, and the subsequent shooting method is switched accordingly.

[0121] If it is determined in step S2104 that movement has occurred in the main subject, the process proceeds to step S2105 and multiple images are taken. On the other hand, if it is determined that no movement has occurred in the main subject, the process proceeds to step S2106 and one image is taken. When proceeding to step S2105 and taking multiple images, the process of continuing to determine the movement of the main subject and selecting the image for composite processing is as described in the second embodiment.

[0122] Here, we will again use Figure 22 to explain the relationship between the determined shutter speed (exposure period) and the actual period of exposure in the image capture unit (actual exposure period).

[0123] Figure 22(a) shows the relationship between the exposure period and the actual exposure period when taking a single shot. 2201 represents the exposure period, which is determined by the shooting control parameter generation unit 2004 in Figure 20 and the shooting parameter determination step in Figure 21, and is the exposure period for capturing an image that contains both parts with suppressed subject blur and parts where subject blur is to be tolerated or emphasized as the final image. In the figure, 2202 represents the actual exposure period, which is the exposure period of the captured image data exposed for the specified exposure period by the image acquisition unit 2001 in Figure 20. When taking a single shot, the exposure period and the actual exposure period are of equal length. Figure 22(a) shows that after taking a single shot, the final image data can be obtained simply by performing predetermined processing, such as development processing, on the single image data.

[0124] On the other hand, Figure 22(b) shows the relationship between exposure time and actual exposure time when multiple images are taken. In the figure, 2203 represents the exposure period, which is the same as 2201 in Figure 22(a). Figure 22(b) shows the case where multiple images are taken in increments of 10 during the exposure period. That is, 2204 shows the case where multiple image data from 1 to 10 are taken. When taking multiple images, by taking images with an exposure time shorter than the exposure period (short exposure), it is possible to take images in which the main subject is not moving, as explained in the second embodiment. By combining the images taken simultaneously, as explained in the second embodiment, it is possible to take an image in which the final image contains a mixture of parts where subject blur is suppressed and parts where subject blur is to be tolerated or emphasized.

[0125] In the case of taking multiple shots, the process involves repeatedly taking pictures with short exposure times. However, to perform image compositing, it is sufficient to start with any of the images taken with short exposures and combine enough images so that the total exposure time is equivalent to the exposure period. This is shown in Figure 22(c). In the figure, 2205 is the exposure period, and 2206 is the image data of the short exposure time taken, including before and after this period. In the figure, image data 1 to 13 are shown. For example, in Figure 22(c), the image data from the 3rd image data 2207 to the 12th image data 2208 are to be processed for compositing. In this case, 10 image compositing processes are performed during the shooting period to generate an image, and an image with the same image effect as Figure 22(b) can be obtained. In other words, in this embodiment, when taking multiple shots to capture an image that contains both parts with suppressed subject blur and parts where subject blur is tolerated or emphasized, in order to obtain the desired effect in the final image, it is sufficient to have image data of short exposures obtained during the determined exposure period, and it is not necessary to combine image data outside of this exposure period.

[0126] In multi-shot photography, the final image data is obtained after all the image data for the subject to be combined has been acquired and then processed. Therefore, although the image processing is time-consuming, it is possible to obtain an image with reduced blur in the main subject.

[0127] As explained above, according to this embodiment, by switching between single-shot and multiple-shot photography depending on the movement state of the main subject, it is possible to generate an image with a mixture of parts where subject blur is suppressed and parts where subject blur is tolerated or emphasized, with an appropriate amount of processing power.

[0128] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist.

[0129] [Other embodiments] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions. [Explanation of Symbols]

[0130] 100 Digital Cameras 101 Control Unit 102 ROM 103 RAM 104 Optical system 105 Imaging Unit 106 A / D Conversion Unit 107 Image Processing Unit 108 Recording media 109 Display section 110 Operation Input Section

Claims

1. Means for acquiring images, A recognition means capable of recognizing the scene in which the aforementioned image was taken, Motion detection means for detecting the amount of motion between multiple frames in a predetermined region of the image acquired by the acquisition means, A means for determining shooting parameters, Equipped with, The aforementioned predetermined region is a region corresponding to the object being swung in a sport involving a swinging motion. The image processing apparatus is characterized in that, when the recognition means recognizes that the scene in which the image was taken is a predetermined scene, the shooting parameter determination means determines the shooting parameters such that the amount of blur in the predetermined region is greater than a predetermined standard.

2. The image processing apparatus according to claim 1, characterized in that the recognition means recognizes the type of sport as the scene in which the image was taken.

3. The image processing apparatus according to claim 1, characterized in that the object to be swung in sports involving a swinging motion is a piece of equipment used in each sport.

4. A motion vector calculation means for calculating motion vectors between multiple frames of the aforementioned image, The system includes distance information calculation means for calculating distance distribution information in the aforementioned image, The image processing apparatus according to any one of claims 1 to 3, characterized in that the amount of motion is at least one of the magnitude of the motion vector or the amount of change in the distance distribution information between multiple frames of the image.

5. The image processing apparatus according to claim 4, characterized in that the distance distribution information is one of the following: disparity distribution information obtained from a group of images with different viewpoints, contrast distribution information obtained from a group of images with different focus points, or actual distance distribution measured by the TOF method.

6. The image processing apparatus according to claim 4, characterized in that the distance distribution information includes any of the following: the amount of defocus of the image of the subject present in each pixel of the image, the relative amount of image shift between different viewpoints, or the distance of the subject in the depth direction.

7. The image processing apparatus according to any one of claims 4 to 6, characterized in that the motion vector calculation means limits the range for searching the motion vector to the range in which the predetermined region moves.

8. The system includes motion vector correction means for correcting the aforementioned motion vector, The image processing apparatus according to any one of claims 4 to 7, characterized in that it corrects the motion vector based on the angle between the motion vector and the estimated direction in which the predetermined region moves, or the reliability of the motion vector.

9. The image processing apparatus according to claim 8, characterized in that the shooting parameter determination means calculates the shutter speed when the magnitude of the corrected motion vector is greater than a predetermined threshold.

10. The system includes motion vector accumulation means for accumulating the aforementioned motion vectors across multiple frames, The image processing apparatus according to any one of claims 4 to 9, characterized in that the shooting parameter determination means calculates the shutter speed when the magnitude of the accumulated motion vector is greater than a predetermined threshold.

11. Display means and A display control means for controlling the display of the display means, The image processing apparatus according to any one of claims 1 to 10, characterized in that the display control means causes the display means to display an image on which the recognition result of the recognition means is superimposed.

12. The image processing apparatus according to any one of claims 1 to 11, characterized in that the shooting parameters include at least one of shutter speed, aperture value, and ISO sensitivity.

13. A setting means for setting a predetermined threshold based on the scene in which the aforementioned image was taken, A comparison means for comparing the magnitude of the motion vector in the predetermined region, determined based on the detection result of the motion amount detection means, with the predetermined threshold, Furthermore, The image processing apparatus according to any one of claims 1 to 12, characterized in that, when the comparison means determines that the magnitude of the motion vector in the predetermined region is greater than the predetermined threshold, the shooting parameter determination means determines the shooting parameters such that the amount of blur in the predetermined region is greater than a predetermined standard.

14. Imaging means, An image processing apparatus according to any one of claims 1 to 13, An imaging device having

15. The process of acquiring images, A recognition step that recognizes the scene in which the aforementioned image was taken, A motion detection step for detecting the amount of motion between multiple frames in a predetermined region of the image acquired in the acquisition step, A process for determining shooting parameters, It has, The aforementioned predetermined region is a region corresponding to the object being swung in a sport involving a swinging motion. A control method for an image processing apparatus, characterized in that, when the scene in which the image was taken is recognized by the recognition step as a predetermined scene, the shooting parameter determination step determines the shooting parameters such that the amount of blur in the predetermined area is greater than a predetermined standard.

16. A program for causing a computer to execute each step of the control method described in claim 15.

17. A computer-readable storage medium storing a program for causing a computer to execute each step of the control method described in claim 15.

Citation Information

Patent Citations

  • Image pickup device

    JP2002077711A

  • Image pickup device, image processor and image processing method

    JP2008015754A

  • Imaging apparatus and program therefor

    JP2008301355A

  • Image pickup apparatus, image pickup method, image searching apparatus and image searching method

    JP2009124210A

  • Image capturing apparatus

    JP2010171825A