Image processing device and method, imaging apparatus, and program
Patent Information
- Application Number
- JP2022102865
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-07-11
AI Technical Summary
Existing image capture methods struggle to balance capturing a moving subject with a sense of dynamism while maintaining desirable facial expressions and avoiding blurring, especially when using multiple cameras or predicting complex subject movements.
An image processing apparatus that detects subject areas and feature points in multiple images, aligns and synthesizes them to correct positional deviations, and selectively combines images based on facial expressions and movement to create dynamic, clear images.
The solution enables the capture of images with a sense of dynamism and preferable facial expressions by aligning and synthesizing images to minimize blurring and enhance subject clarity, even in complex movement scenarios.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing device and method, an imaging device, and a program, and more particularly to a technique for aligning and synthesizing a plurality of images obtained by continuous shooting. [Background technology]
[0002] When taking pictures with a camera, it is very difficult to capture a subject that moves around within the field of view with an appropriate composition. When the camera is fixed at a fixed point, it is necessary to take pictures with as short a time as possible, but taking pictures with a short time may result in an image that lacks a sense of dynamism. In addition, if the subject is a person, for example, there is a risk that the subject may be photographed in a generally undesirable state, such as blinking or closing their eyes during the short time shooting period.
[0003] Patent Document 1 discloses a method for using multiple cameras to take photographs periodically and controlling the saving of images in which the subject's eyes are not closed, thereby preventing the retention of images showing undesirable facial expressions such as blinking.
[0004] In addition, when shooting with a handheld camera, the photographer can track the subject, which improves the stability of the composition compared to using a fixed camera. However, shots taken with short exposures still lack a sense of dynamism, and panning shots taken with long exposures require a high level of skill.
[0005] Patent Document 2 discloses a method of predicting the movement of a subject from the motion vector of image data as a panning reference angular velocity, and controlling the drive of an image stabilization means to match the movement of the subject based on the difference with the panning speed of the camera. This method makes it possible to easily capture so-called panning images in which the subject is less blurred and the background is blurred. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] JP 2004-356683 A [Patent Document 2] JP 2019-174608 A Summary of the Invention [Problem to be solved by the invention]
[0007] However, the method of Patent Document 1 requires more installation space for the cameras than when a single camera is used, and this leads to increased running costs.
[0008] In addition, whether shooting with one camera or multiple cameras as in Patent Document 1, shooting with a short exposure time is performed to capture a moment, making it difficult to express a sense of dynamism. In Patent Document 2, the subject's movement is predicted, but this is difficult when the subject's movement is complex, and it is difficult to achieve a good panning shot when the subject moves differently from the prediction during the long exposure time.
[0009] The present invention has been made in consideration of the above problems, and has an object to obtain an image with a sense of dynamism.
[0010] It is a further object of the present invention to obtain an image of a subject having a more pleasing facial expression. [Means for solving the problem]
[0011] In order to achieve the above-mentioned object, the image processing device of the present invention has a first detection means for detecting a subject area of a predetermined subject from each of a plurality of images, a second detection means for detecting a partial area of a predetermined range including the subject area from each of the plurality of images, a feature point detection means for detecting feature points of an image, and a synthesis means for synthesizing the partial areas of the plurality of images so that the feature points of the subject areas match. Effect of the Invention
[0012] According to the present invention, it is possible to acquire dynamic images.
[0013] According to another aspect, an image of a subject having a more pleasing facial expression can be acquired. [Brief description of the drawings]
[0014] [Figure 1] 1 is a block diagram showing the functional configuration of an imaging apparatus according to a first embodiment of the present invention. [Diagram 2] 5 is a flowchart showing a shooting process in the first embodiment. [Diagram 3] FIG. 4 is a conceptual diagram for explaining the effect of the first embodiment. [Figure 4] 10 is a flowchart showing a shooting process in the second embodiment. [Diagram 5] FIG. 11 is a conceptual diagram for explaining an effect of the second embodiment. [Figure 6] FIG. 11 is a block diagram showing the functional configuration of an imaging apparatus according to a third embodiment. [Figure 7] 13 is a flowchart showing a shooting process in the third embodiment. [Figure 8] 13A and 13B are schematic diagrams for explaining a movement amount in the angle of view change processing in the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0016] <First embodiment> A first embodiment of the present invention will be described with reference to FIGS.
[0017] 1 is a block diagram showing the functional configuration of an imaging device 1000 as an example of a device for performing image processing of the present invention. In the first embodiment, the imaging device 1000 is assumed to be a so-called fixed camera, and as an example, a case will be described in which the imaging device 1000 is installed near an attraction in an amusement park, and captures images of users of the attraction enjoying the attraction, and provides the images to the users as content. Note that the present invention is not limited to fixed cameras, and can also be applied to cameras that are primarily used for handheld photography, such as general digital cameras.
[0018] Light from a subject passes through an optical system 100 and forms an image on an image sensor 101. The image sensor 101 is an image sensor including a photoelectric conversion element such as a CMOS sensor or a CCD sensor having a Bayer array color filter, and the formed optical image of the subject is photoelectrically converted at each pixel and read out as an image signal. The read image signal is sent to an image signal processor 10, where various processes are performed thereafter.
[0019] In the image signal processing unit 10, the subject area detection unit 103 detects the subject and its area in the image represented by the read image signal. The cutout unit 104 cuts out a specific part of the read image signal. The feature point detection unit 105 detects edges and the like in the image as feature points. The facial expression determination unit 106 determines the facial expression when a human face is recognized as the subject. The counter 107 counts the number of images captured continuously.
[0020] The first movement amount detection unit 108 detects the amount of movement of feature points between multiple images captured continuously. The second movement amount detection unit 111 detects the amount of movement of feature points in the peripheral area of the subject detected by the subject area detection unit 103. The deformation amount detection unit 109 detects the amount of deformation of the shape of the face caused by a change in the direction of the face when a human face is recognized as the subject. Also, when an object other than a human face is recognized as the subject, it detects the amount of deformation of the subject.
[0021] The combining unit 110 combines a plurality of continuously captured images into one image while adjusting their positions in a predetermined manner. The combining method will be described in detail later.
[0022] The control unit 20 controls each component of the image signal processing unit 10 so as to supply a processed image signal to the development processing unit 112 based on user settings, shooting conditions, etc., and also controls the entire imaging device 1000. When the development processing by the development processing unit 112 is completed, the image signal is stored in the storage unit 113, and a series of shooting operations is completed.
[0023] Fig. 2 is a flowchart showing a series of photographing processes performed by the imaging device 1000 in the first embodiment. Fig. 3 is a conceptual diagram showing the effect obtained by performing the processes shown in Fig. 2. In the compositing process in this embodiment, compositing is performed only on a partial area of a plurality of images captured continuously, and the areas to be composited are selected prior to compositing.
[0024] 2 is started when an instruction to start shooting is issued in the imaging device 1000. Here, for example, the imaging process is started when a sensor other than the imaging device 1000 detects that a cart carrying people has arrived at a predetermined position in an attraction to be photographed. Note that the start of shooting is not limited to this, and for example, shooting may be started in response to a user pressing a shutter button (not shown).
[0025] In S101, the count value N of the counter 107 is set to 1. Then, in S102, the Nth image is captured by the image sensor 101. In S103, the subject area detection unit 103 detects an identifiable subject, such as a person, and its area (subject area) from the captured image. Here, the area of the person's face is detected as the subject area. In S104, the feature point detection unit 105 detects feature points present in the captured image. Feature points are detected not only in the subject area but also in the background area. The feature points detected here are used to detect the amount of movement of the subject and its surroundings, which will be described later.
[0026] In S105, the count value N is compared with a threshold value Th1 indicating a predetermined number of images to be continuously captured, and it is determined whether or not the count value N is less than the threshold value Th1. In this embodiment, continuous shooting is necessary to obtain the effect of combining multiple images, so the threshold value Th1 is set to an integer of 2 or more. If it is determined in S105 that the count value N is less than the threshold value Th1, the process proceeds to S106, where 1 is added to the count value N, and the process returns to S102 to perform the next capture.
[0027] On the other hand, if it is determined in S105 that the count value N is equal to or greater than the threshold value Th1, then the shooting ends and the process proceeds to S107. The method of determining whether or not to continue shooting is not limited to the above. For example, in S105, if shooting is started by pressing the shutter button, it may be determined whether the shutter button has been released, or it may be determined based on the elapsed time since shooting started. In either case, it may be controlled so that at least two images are shot, and if these conditions are not met, the process proceeds to S106, and if these conditions are met, the process proceeds to S107.
[0028] In S107, the second movement amount detection unit 111 detects the movement amount of the feature point existing around the subject area. The movement amount can be detected based on how much the feature point has moved between two consecutive images. Here, since the subject area is the face area of a person, the movement amount of the hands and feet existing around the face of the person is detected in S107. In this embodiment, the case where the movement amount of the feature point is detected between two consecutive images after the continuous shooting is completed will be described, but the present invention is not limited to this, and the movement amount may be detected sequentially when the shooting of the first image is completed and the second and subsequent images are obtained.
[0029] In S108, the synthesis unit 110 sets, as a synthesis area to be used for synthesis, an area in which the movement amount detected by the second movement amount detection unit 111 is greater than a predetermined movement amount for feature points present around the subject area included in a series of images obtained by continuous shooting. Then, in S109, the synthesis unit 110 performs a positioning synthesis process for the set synthesis area. In the positioning synthesis, a known method can be used in which feature points of consecutive images are aligned to the position of the first image to be synthesized and averaged to synthesize. At this time, the synthesis area is positioned so that the feature points of the face area (subject area) match. By doing so, the positional deviation of the face area in the series of images obtained by continuous shooting is corrected, and an area in which the movement amount of the feature points is large around the subject area is synthesized in a state in which the positional deviation of the feature points occurs. In other words, an image in which the blur of the face area (subject area) is small and the feature points around the subject area have a sense of dynamism can be obtained. In addition, for the area not set as the synthesis area, the first image is placed as it is and joined to the area obtained by positioning synthesis.
[0030] Also, when a face is recognized as the subject, the subject area of the composite area may be weighted to the movement of facial feature points to perform alignment and composition processing. This allows an image of the face with less blurring to be obtained. Also, when a face is recognized as the subject, the subject area of one of the multiple images may be selected based on the facial expression determined by the facial expression determination unit 106, or multiple subject areas may be selected to perform alignment and composition processing of the subject areas. At this time, for example, an image with a more favorable facial expression may be obtained by excluding an image with closed eyes or selecting an image determined to be a smile.
[0031] In S110, the development processing unit 112 develops the image combined in S109, stores the developed image in the storage unit 113, and ends the process.
[0032] Next, the effects obtained by the above-mentioned processing will be described with reference to FIG.
[0033] Fig. 3(a) shows an example of two consecutive images, and Fig. 3(b) shows an example of a composite image obtained by combining the two images in Fig. 3(a). A person's face region 310 is detected as a subject region from each image by the subject region detection unit 103. Also, it is shown diagrammatically that the second movement amount detection unit 111 has detected that an arm is moving significantly in a peripheral region 311 between the face region 310 and a frame indicated by a dotted line.
[0034] 3(b), the peripheral area 311 is formed by aligning and combining two images, and the area outside the peripheral area 311 is formed by the first image. For the face area 310, an image selected by weighted alignment and combination processing of two images and / or a selection processing based on the facial expression determined by the facial expression determination unit 106 is used alone or after combination.
[0035] As described above, according to the first embodiment, alignment and synthesis is performed only on the area where the subject is moving. This allows for a clear and dynamic image to be obtained without blurring the background, with a face with a more pleasing expression and less blur, and arms with more dynamic blur.
[0036] In the explanation of this embodiment, the case where there is a single subject has been used as an example, but when multiple subjects are detected, the facial expressions of each subject will be different, so selection of the area to be composited can be performed sequentially for each subject.
[0037] Furthermore, in this embodiment, a fixed camera is used as the imaging device, but the present invention is not limited to this and may be applied to a camera that is intended to be handheld, such as a general digital camera.
[0038] <Second embodiment> Next, a second embodiment of the present invention will be described.
[0039] In the second embodiment, the imaging device 1000 described in the first embodiment with reference to FIG. 1 performs the processing described below.
[0040] Fig. 4 is a flowchart showing a series of photographing processes performed by the imaging device 1000 in the second embodiment. Fig. 5 is a conceptual diagram showing the effect obtained by performing the processes shown in Fig. 4. In the compositing process in this embodiment, when a predetermined subject (e.g., a person's face, etc.) is detected, which images are to be subjected to the compositing process from among a plurality of images captured continuously, and a predetermined area is cut out and composited.
[0041] 2, the photographing process shown in the flowchart of Fig. 4 is started when an instruction to start photographing is issued in the imaging device 1000. In the following description, it is assumed that photographing is started in response to pressing of a shutter button (not shown).
[0042] In S201, the count value N of the counter 107 is set to 1. Then, in S202, the Nth image is captured by the image sensor 101. In S203, the subject area detection unit 103 detects an identifiable subject such as a person and its area (subject area) from the captured image. Here, the area of the person's face is detected as the subject area. In S204, the cropping unit 104 crops out a partial area that includes the subject area detected in S203 and has a size obtained by multiplying the shooting angle of view by a predetermined magnification (1 or less). Here, as an example, the predetermined magnification is a ratio when cropping out a partial area that includes the face area of the detected subject area (face in FIG. 5) and is based on the aspect ratio of the shooting angle of view.
[0043] The predetermined magnification may be a ratio when cutting out a partial region in a similar shape to the detected subject region (in this case, the subject region is set to 1 and the predetermined magnification is greater than 1). At this time, if there is a large change in the size of the partial region to be cut out between images, that is, if the subject is approaching or moving away from the imaging device 1000, a process of resizing the image size so that the size of the subject becomes approximately the same size may be included. In this way, it becomes possible to obtain an image in which the size of the subject is approximately constant and the background appears to flow radially by a synthesis process described later.
[0044] In S205, the feature point detection unit 105 detects feature points present in the partial region cut out in S204.
[0045] In S206, it is determined whether or not the shooting instruction is continuing. Here, it is determined whether or not the shutter button (not shown) is continuing to be pressed. However, in this embodiment, since continuous shooting is required to obtain the effect of combining multiple images, it is also determined whether or not the count value N is less than 2. If it is determined in S206 that the shutter button is continuing to be pressed or the count value N is less than 2, the process proceeds to S207, where 1 is added to the count value N, and the process returns to S102 to perform the next shooting.
[0046] On the other hand, if it is determined in S206 that the shutter button has been released and the count value N is equal to or greater than 2, shooting ends and the process proceeds to S208.
[0047] In S208, the first movement amount detection unit 108 detects the amount of movement between two consecutive images of the region cut out by the cutout unit 104 from the series of continuously captured images. If the amount of movement of the cut out partial region is large, it means that the subject whose image is to be acquired has moved a lot, and it is assumed that the subject has a high sense of dynamism. The detection result of the first movement amount detection unit 108 is used when determining the subsequent synthesis processing conditions.
[0048] S209 is a process that is performed when the subject detected by the subject region detection unit 103 is a person's face, and detects the facial expression and the amount of deformation of the face of the person who is the subject using the facial expression determination unit 106 and the deformation amount detection unit 109. Facial expression detection is performed using a known technique, and it is detected whether the eyes are open, whether the person is smiling, etc., and the detected information is used as information for selecting an image to be used in synthesis, which will be described later. In addition, the amount of deformation of the face is detected by detecting whether the orientation of the face is within a predetermined range, and the detected information is used as information for selecting an image to be used in synthesis, which will be described later.
[0049] In S210, the synthesis unit 110 selects an image to be used for synthesis from images N=1 to N. Specifically, the images to be selected are those having partial regions whose movement amount between the images obtained in S208 is equal to or greater than a predetermined threshold value and whose facial expression and facial deformation amount obtained in S209 are judged to be usable for synthesis processing.
[0050] It is preferable that the selected images are consecutive images so that the movement of the background has continuity. Therefore, if consecutive images are not selected under the above judgment conditions, for example, a threshold value for judging that the image next to the image selected as the image containing the partial region to be synthesized can be changed to make it easier to select. Alternatively, an image sandwiched in time between two non-consecutive images selected as the image containing the partial region to be synthesized may be selected even if it does not satisfy the above judgment conditions. In S211, the synthesis unit 110 aligns and synthesizes the partial regions to be synthesized in the selected images. Note that the alignment and synthesis process performed here is performed by synthesizing the images using the same process as S109 in FIG. 2.
[0051] Finally, in S212, the development processing unit 112 develops the image combined in S211, stores the developed image in the storage unit 113, and ends the process.
[0052] Next, the effects obtained by the above-mentioned processing will be described with reference to FIG.
[0053] Fig. 5(a) shows three images captured in succession. Since the subject is a person, a face region 501 is detected as the subject region by the subject region detection unit 103 in S203. Furthermore, partial regions 502 are cut out in S204, and Fig. 5(b) shows images of the partial regions 502 cut out from each of the three images shown in Fig. 5(a). These three partial regions 502 meet the selection conditions in S210 and are therefore to be combined.
[0054] Fig. 5(c) shows an image in which the three partial regions 502 shown in Fig. 5(b) are synthesized in S211 by the synthesis unit 110. In this way, this process makes it possible to obtain an image that gives a sense of dynamism to the subject.
[0055] 5 shows a case where the subject moves within the imaging plane, but the present invention is not limited to this. As described in S204, when the subject moves in the optical axis direction of the imaging device 1000 (when the change in the size of the subject area is greater than a threshold), it is also possible to obtain an image of the subject of a constant size with a dynamic feel, in which the background flows radially, by resizing.
[0056] In this way, by performing alignment synthesis on partial areas selected based on the amount of movement, facial expression, and facial deformation of the subject, a dynamic image can be obtained in which the background other than the subject flows while the subject's face is less blurred.
[0057] As described above, according to the second embodiment, it is possible to capture an image of a subject with a more pleasing expression while expressing a sense of dynamism that was difficult to achieve with conventional techniques.
[0058] This embodiment may be applied to a fixed camera, a camera attached to a moving object, or a camera that is assumed to be handheld, such as a general digital camera. In the case of a camera attached to a moving object, a partial region to be used for alignment synthesis may be selected according to the moving speed of the moving object.
[0059] <Third embodiment> Next, a third embodiment of the present invention will be described.
[0060] 6 is a block diagram showing the functional configuration of the imaging device 2000 in the third embodiment. The imaging device 2000 is a so-called digital camera, and in this embodiment, the imaging device 2000 mainly performs handheld shooting, and in particular performs panning shooting in which the photographer shoots while chasing a moving object. However, the panning shooting performed by the imaging device 2000 in this embodiment is different from the conventional panning shooting in which shooting is performed with one exposure, and an image in which the background is blurred and the main subject is clearly displayed is obtained by aligning and combining multiple images captured continuously.
[0061] The functional configuration of the imaging device 2000 in the third embodiment is obtained by adding an angle-of-view changing operation detection unit 214 to the imaging device 1000 described in the first embodiment with reference to Fig. 1. The angle-of-view changing operation detection unit 214 detects whether the photographer has performed an operation to change the angle of view, such as panning, in order to fit the subject within the screen during panning shooting. The angle-of-view changing operation detection unit 214 detects whether an angular velocity equal to or greater than a predetermined value is applied during shooting, using an angular velocity sensor (not shown) provided in the imaging device 2000 for image blur correction.
[0062] In the third embodiment, the feature point detection unit 105 also performs the feature point detection process on the live view image read from the image sensor 101 in the shooting preparation state, in addition to the feature point detection process on the captured image. Furthermore, the movement amount detection unit 208 detects the amount of subject movement between successive live view images. From these detection results, the cropping unit 104 can grasp whether the angle of view has been changed by panning, tilting, etc., and how much the direction in which the image capture device 2000 faces and the direction of the subject match, and can change the crop size of the partial region as described later. Other functional configurations are similar to those of the image capture device 1000 in FIG. 1, so the same reference numbers are used and the description is omitted.
[0063] 7 is a flowchart for explaining the angle of view change process in this embodiment performed by the imaging device 2000. This process starts when the imaging device 2000 is in a shooting preparation state and before the photographer determines the shooting angle of view, i.e., when a so-called aiming operation is started.
[0064] In S301, the angle-of-view changing operation detection unit 214 judges whether or not the angle of view is changed by panning, tilting, etc., based on the angular velocity detected by the angular velocity sensor. When performing panning, panning or tilting is usually not performed after issuing a shooting instruction, but is started immediately before issuing a shooting instruction to try to fit a subject to be shot within the angle of view, so that the result of this judgment can be used to judge whether or not panning is performed. In this embodiment, the angle-of-view changing operation detection unit 214 uses an angular velocity sensor, but the present invention is not limited to this, and even when a mode for performing panning is selected as the camera setting, it may transition to S302.
[0065] If it is determined in S301 that the angle of view has not been changed, the process proceeds to S305, and the size of the partial region to be cut out in S309 (to be described later) is set to "medium."
[0066] On the other hand, if it is determined in S301 that the angle of view has not been changed, the process proceeds to S302.
[0067] In S302, the suitability of the angle of view change operation is determined. More specifically, the subject is detected on the above-mentioned live view image, and it is determined whether the magnitude of the movement amount P of the detected subject is greater than a threshold value P0. If it is determined that the magnitude of the movement amount P is greater than the threshold value P0, the process proceeds to S303, and if it is determined that the movement amount P is equal to or less than the threshold value P0, the process proceeds to S304.
[0068] In S303, the size of the partial region to be cut out in S309 (described later) is set to "small." On the other hand, in S304, the size of the partial region to be cut out is set to "large."
[0069] Here, the amount of movement P of the subject between live view images determined in S302 of Fig. 7 will be described with reference to Fig. 8. Fig. 8(a) and Fig. 8(b) respectively show two consecutive live view images superimposed on each other in a state where the amount of movement P of the subject on the live view images is large and a state where the amount of movement P is small.
[0070] 8(a) shows a state in which the amount of movement P of the subject between live view images is greater than the threshold value P0, indicating that the subject cannot be tracked well by panning. Therefore, if a wide partial region is cut out and then aligned and combined, there is a high possibility that part of the partial region will be removed from the captured image. Therefore, by setting the size of the partial region to be cut out small in S303, the target subject can be easily captured in the captured image.
[0071] 8(b) shows a state in which the amount of movement P of the subject between live view images is equal to or less than the threshold value P0, indicating that the subject can be successfully tracked by panning. Therefore, it is expected that the deviation of the subject between images captured in succession thereafter will be small, and images with similar compositions can be acquired, and it is considered that even if a large partial region is cropped, the partial region is unlikely to deviate from the captured image. Therefore, by setting the size of the partial region to be cropped large in S304, it is possible to acquire a panning image with a wider angle of view.
[0072] After setting the cropping amount, in S306, it is determined whether or not shooting has been instructed by, for example, pressing a shutter button (not shown). If shooting has not been instructed, the process returns to S301 and repeats the above-mentioned processing. If it is determined in S306 that shooting has been instructed, the process proceeds to S307, where the count value N of the counter 107 is set to 1. Then, in S308, the Nth image is captured by the image sensor 101. In S309, the subject area detection unit 103 detects a predetermined subject and its area (subject area) from the captured image. Then, in S310, the cropping unit 104 crops out a partial area that includes the subject area detected in S309 and has any of the sizes set in S303 to S305.
[0073] Here, the size is the magnification multiplied by the ratio when cutting out a partial region in a similar shape to the detected subject region, and is set to 1 or more relative to the subject region. As an example, a partial region is cut out vertically and horizontally with the detected subject region as the center, which is 1.2 times larger when the set size is "medium," 1.1 times larger when the set size is "small," and 1.3 times larger when the set size is "large." Note that the magnification is not limited to these values and can be changed as appropriate.
[0074] In S311, the feature point detection unit 105 detects feature points that exist within the partial region cut out in S310.
[0075] In S312, it is determined whether or not the shooting instruction is still in progress. Here, the same determination as in S206 in the second embodiment is performed. If the shooting instruction is still in progress in S312, the process proceeds to S313, where 1 is added to the count value N, and the process returns to S308 to perform the next shooting.
[0076] On the other hand, if it is determined in S312 that the shooting instruction is not continuing, shooting ends and the process proceeds to S314.
[0077] In S314, the deformation amount detection unit 109 uses the feature points detected in S311 to determine the amount of deformation of the subject detected by the subject region detection unit 103. Here, for example, the state of the subject in the image obtained immediately after the shutter button was pressed in S306 is assumed to be the state of the subject intended by the photographer, and is set as a reference image, and the amount of deformation between the subject detected from the reference image and subjects detected from images other than the reference image is determined. Note that the reference image is not limited to this, and multiple obtained images may be displayed so that the photographer can select one.
[0078] In S315, the synthesis unit 110 selects an image to be used for synthesis from images N=1 to N. Specifically, the image is selected if the deformation amount of the subject detected in S314 is smaller than a predetermined threshold value. Here, similar to step S210, it is preferable that the selected images are continuous images so that the movement of the background has continuity, so it is possible to make it easier to select continuous images using a concept similar to that of S210. In S316, the synthesis unit 110 aligns and synthesizes partial regions to be synthesized in the selected images. Note that the alignment and synthesis process performed here is performed by synthesizing using the same process as S109 in FIG. 2.
[0079] Finally, in S317, the development processing unit 112 develops the image combined in S314, stores the developed image in the storage unit 113, and ends the process.
[0080] As described above, according to the third embodiment, when performing panning, a good panning image can be obtained by utilizing an appropriate size of cropping amount and alignment synthesis.
[0081] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.
[0082] The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0083] <Summary> The disclosure of this embodiment includes the following configuration.
[0084] (Configuration 1) A first detection means for detecting a subject area of a predetermined subject from each of a plurality of images; a second detection means for detecting a partial region of a predetermined range including the subject region from each of the plurality of images; A feature point detection means for detecting feature points of an image; a synthesis means for synthesizing the partial regions of the plurality of images so that feature points of the subject region match; 13. An image processing device comprising:
[0085] (Configuration 2) The image processing device according to configuration 1, wherein the second detection means detects, as the partial region, a region in which a movement amount of a feature point detected around the subject region between successive images is greater than a predetermined movement amount.
[0086] (Configuration 3) a determination means for determining a facial expression when the subject area is a face area of a person, The image processing device according to configuration 1 or 2, characterized in that the synthesis means selects, for the subject regions of the partial regions, subject regions of the plurality of images whose facial expressions satisfy predetermined conditions, and synthesizes the selected subject regions.
[0087] (Configuration 4) The method further comprises: determining an amount of deformation of a face when the subject area is a face area of a person; The image processing device according to any one of configurations 1 to 3, characterized in that the synthesis means selects, for the subject regions of the partial regions, subject regions among the plurality of images in which the deformation amount is smaller than a predetermined threshold value and synthesizes the selected subject regions.
[0088] (Configuration 5) 5. The image processing device according to any one of configurations 1 to 4, wherein the combining means further combines the combined partial region with a region in any one of the plurality of images excluding the partial region.
[0089] (Configuration 6) a third detection means for detecting an amount of movement of the partial region between successive images among the plurality of images, 5. The image processing device according to any one of configurations 1 to 4, wherein the combining means selects, from among the partial regions of the plurality of images, partial regions in which the amount of movement is greater than a predetermined amount of movement, and combines the partial regions.
[0090] (Configuration 7) a third detection means for detecting an amount of movement of the subject area between successive images among the plurality of images; and a setting unit for setting a size of the partial region based on the amount of movement, The image processing device according to configuration 1, characterized in that when the movement amount is a second movement amount larger than the first movement amount, the setting means sets the size of the partial region so that the partial region is smaller than when the movement amount is the first movement amount.
[0091] (Configuration 8) The method further includes a determination unit that determines an amount of deformation between the subject area of a reference image serving as a reference among the plurality of images and the subject area of an image other than the reference image, 8. The image processing device according to claim 7, wherein the synthesis means selects and synthesizes a partial region including the subject region in which the amount of deformation is smaller than a predetermined threshold value.
[0092] (Configuration 9) An imaging means; An image processing device according to any one of configurations 1 to 8; An imaging device comprising:
[0093] (Configuration 10) a first detection step of detecting a subject region of a predetermined subject from each of a plurality of images; a second detection step of detecting a partial region of a predetermined range including the subject region from each of the plurality of images; a feature point detection step of detecting feature points of an image; a synthesis step of synthesizing the partial regions of the plurality of images so that feature points of the subject regions match; 13. An image processing method comprising:
[0094] (Configuration 11) A program for causing a computer to function as each of the means of the image processing device according to any one of configurations 1 to 8.
[0095] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0096] 100: imaging device, 10: image signal processing unit, 20: control unit, 101: imaging element, 103: subject region detection unit, 104: cropping unit, 105: feature point detection unit, 106: facial expression determination unit, 107: counter, 108: first movement amount detection unit, 109: deformation amount detection unit, 110: synthesis unit, 111: second movement amount detection unit, 214: angle of view change operation detection unit
Claims
1. First detection means for detecting a subject area of a predetermined subject from each of a plurality of images; Second detection means for detecting a partial area including the subject area from each of the plurality of images; Feature point detection means for detecting feature points of an image; Combining means for combining the partial areas of the plurality of images so that the feature points of the subject areas coincide; An image processing apparatus characterized by comprising the above.
2. The image processing apparatus according to claim 1, wherein the second detection means detects, as the partial area, an area in which the amount of movement of feature points detected around the subject area between consecutive images is larger than a predetermined amount of movement.
3. When the subject area is an area of a human face, further comprising determination means for determining the facial expression; The image processing apparatus according to claim 1, wherein the combining means selects and combines, for the subject area of the partial area, a subject area among the subject areas of the plurality of images that satisfies a predetermined condition for the expression.
4. When the subject area is an area of a human face, further comprising determination means for determining the amount of deformation of the face; The image processing apparatus according to claim 1, wherein the combining means selects and combines, for the subject area of the partial area, a subject area among the subject areas of the plurality of images that has a deformation amount smaller than a predetermined threshold value.
5. The image processing apparatus according to claim 1, wherein the combining means further combines the combined partial area with an area excluding the partial area in any of the plurality of images.
6. Further comprising third detection means for detecting the amount of movement of the partial area between consecutive images among the plurality of images; The image processing apparatus according to claim 1, wherein the combining means selects and combines the partial areas among the partial areas of the plurality of images that have an amount of movement larger than a predetermined amount of movement.
7. Third detection means for detecting the amount of movement of the subject area between consecutive images among the plurality of images; and Setting means for setting the size of the partial area based on the amount of movement. The setting means sets the size of the partial region such that the partial region becomes smaller when the amount of movement is a second amount of movement greater than the first amount of movement than when the amount of movement is the first amount of movement. The image processing apparatus according to claim 1, characterized in that.
8. The apparatus further includes determination means for determining a deformation amount between the subject region of a reference image as a reference among the plurality of images and the subject region of an image other than the reference image. The combining means selects a partial region including the subject region in which the amount of deformation is smaller than a predetermined threshold value and combines it with the partial region of the reference image. The image processing apparatus according to claim 7, characterized in that.
9. An imaging means; The image processing apparatus according to any one of claims 1 to 8 An imaging apparatus characterized by having.
10. A first detection step of detecting a subject region of a predetermined subject from each of a plurality of images; A second detection step of detecting a partial region including the subject region from each of the plurality of images; A feature point detection step of detecting feature points of an image; A combining step of combining the partial regions of the plurality of images so that the feature points of the subject regions coincide with each other An image processing method characterized by having.
11. A program for causing a computer to function as each means of the image processing apparatus according to any one of claims 1 to 8.