Image synthesis device, image synthesis method, and program
The image synthesis device addresses composition deviation by detecting moving subjects and correcting composite distortion, ensuring high-quality HDR synthesis without artifacts.
Patent Information
- Application Number
- JP2024025994
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-09-03
AI Technical Summary
Existing HDR composition technologies struggle with composition deviation due to subject movement, leading to artifacts such as gray artifacts and local noise, especially when using composition masks under certain imaging conditions.
An image synthesis device that acquires multiple images under different exposure conditions, detects moving subject areas, performs blur removal, and corrects composite distortion by generating extended differences and background color information to prevent artifacts.
Prevents artifacts like gray artifacts and local noise, enabling high-quality HDR synthesis even with subject movement.
Smart Images

Figure 2025128945000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image synthesis technique and the like. [Background technology]
[0002] High Dynamic Range (HDR) synthesis technology is known as an image synthesis technology for capturing images with a wide dynamic range. For example, Patent Document 1 discloses an image synthesis device that can generate an appropriate synthesized image even if the subject moves during capture. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6823942 Summary of the Invention [Problem to be solved by the invention]
[0004] The image composition device in Patent Document 1 introduces a composition mask to avoid composition deviation due to movement of the subject. However, when HDR composition is performed using the composition mask, there are cases where appropriate composition cannot be performed within the application range of the composition mask depending on the imaging conditions. [Means for solving the problem]
[0005] According to one aspect of the present invention, an image synthesis device for generating a synthesized image by synthesizing at least two images from a series of multiple images in which the same subject is captured includes an acquisition unit that acquires multiple images including a first image captured under a first exposure condition, a second image captured at a second timing under a second exposure condition that is higher than the first exposure condition, and a third image captured at a third timing under the first exposure condition, an area detection unit that compares pixel values of the first image and the third image to detect a moving subject area where a moving subject is depicted, and a pixel detection unit that compares pixel values corresponding to the moving subject area with the pixel values corresponding to the moving subject area of the second image. The image processing device includes a combining unit that performs a moving subject blur removal process to bring pixel values closer to those corresponding to the subject region, and combines the target image after the moving subject blur removal process with a second image, and a correction unit that corrects composite distortion occurring in the target image, wherein the correction unit calculates an extended difference that spatially diffuses the moving subject region, generates first correction information for the region for which the composite distortion is to be corrected based on the second image and the extended difference, calculates second correction information that corrects the composite distortion based on at least the first image or the third image, and corrects the composite distortion based on the target image, the first correction information, the second correction information, and the extended difference. [Effects of the Invention]
[0006] According to the video editing device of the present invention, even when a synthesis mask is used to avoid synthesis shifts due to subject movement, it is possible to prevent the occurrence of artifacts due to quantization distortion and perform high-quality HDR synthesis. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of an image processing apparatus. [Figure 2] FIG. 10 is a schematic diagram showing an example of an artifact. [Figure 3] 10 is a flowchart showing an example of a processing flow of the image processing apparatus. [Figure 4] 10 is a flowchart showing an example of the flow of gray artifact correction pre-processing. [Figure 5] 10A and 10B are diagrams showing a specific example of binary extended composite mask generation processing. [Figure 6] 10A to 10C are diagrams showing a specific example of gray artifact region drawing processing. [Figure 7] 10A to 10C are diagrams showing a specific example of background color extraction processing. [Figure 8] 10A and 10B are diagrams showing a specific example of background color complementation and expansion processing. [Figure 9] 10A and 10B are diagrams showing a specific example of gray artifact filling processing. [Figure 10] FIG. 1 is a block diagram showing an example of the functional configuration of a smartphone. [Figure 11] FIG. 1 is a diagram showing an example of a recording medium. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, an example of an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same elements are denoted by the same reference numerals, and duplicated descriptions may be omitted. Furthermore, the components described in this embodiment are merely examples and are not intended to limit the scope of the present invention.
[0009] [Embodiment] An example of an embodiment for realizing the image processing technology and image synthesis technology of the present invention will be described below.
[0010] In this specification, the term "image" may refer to, for example, information having a data structure in which the values of each RGB channel constituting the image are spatially distributed. It may also refer to, for example, information expressed in a form that can be visually grasped. In this image processing device, it is not necessary to distinguish between images and information.
[0011] 1 is a block diagram showing an example of the functional configuration of an image processing device 1 according to one aspect of the present embodiment. The image processing device 1 may also be called an information processing device or an image synthesis device. The image processing device 1 includes, for example, an acquisition unit 10, a region detection unit 11, a synthesis unit 12, and an artifact correction unit 15. These are, for example, functional units (functional blocks) included in a processing unit (processing device) and control unit (control device) (not shown) of the image processing device 1, and are configured with a processor such as a CPU or DSP, or an integrated circuit such as an ASIC.
[0012] The imaging unit 20 is a device having a function of capturing, for example, a moving image or the like made up of a plurality of images (for example, a first image, a second image, and a third image in chronological order). For example, the first image and the third image may have the same exposure conditions (for example, exposure conditions such as EV value and shutter speed), and the second image may have an exposure condition that results in a higher brightness than the first image or the second image. Hereinafter, the first image and the third image may be referred to as "low brightness images," and the second image may be referred to as "high brightness image." The imaging unit 20 may be, for example, an imaging device using a pixel sensor such as a CMOS, etc. The imaging unit 20 may store the captured image and imaging conditions in a storage unit (not shown).
[0013] The acquisition unit 10 has a function of acquiring, for example, a plurality of frame images captured by the imaging unit 20. Note that the acquisition unit 10 may also acquire a plurality of images stored in a storage unit (not shown).
[0014] The area detection unit 11 has a function of, for example, aligning the entire image to correct movement between input screens in each image (for example, uniform movement across the entire image caused by camera shake, etc.) based on a series of input images acquired by the acquisition unit 10. The area detection unit 11 also has a function of detecting a moving subject area, which is an area where a moving subject is depicted, by comparing pixel values between low-luminance images (for example, the first image and the third image) or between high-luminance images.
[0015] The synthesis unit 12 generates a synthesis mask based on, for example, pixel values in the moving subject region of the high-brightness image detected by the region detection unit 11 and pixel values corresponding to the moving subject region of the low-brightness image.The synthesis unit 12 then has a function of executing a reference image complementation process to align the position of the moving subject by overwriting the pixel values of one of the high-brightness image and the low-brightness image in the region specified by the generated synthesis mask with the pixel values of the other (referred to as the "reference image"). The synthesis mask is a mask for avoiding a phenomenon in which a subject appears double or triple overlapped (ghost phenomenon) when an image is synthesized in which the subject, for example, a person or a car, exhibits local movement within the image. Hereinafter, an image from which moving subject blur has been removed based on the synthesis mask is referred to as a "reference image complement image."
[0016] The synthesis unit 12 also has a function of executing an HDR synthesis process that performs HDR synthesis based on a low-brightness reference image complement image (an image based on a low-brightness image partially replaced with a high-brightness image by a synthesis mask) in which the positioning of the moving subject has been performed, and a high-brightness reference image complement image (an image based on a high-brightness image partially replaced with a low-brightness image by a synthesis mask), for example, based on a mixing ratio determined according to the difference in brightness values of both images.
[0017] The display unit 23 is a device having a function of displaying a composite image, and may be, for example, a display device.
[0018] For details of these functional units, see Patent Document 1, for example, and they may be configured in a similar manner.
[0019] It should be noted that "displaying" an image is a type of "outputting" an image. The "output" of an image can include not only the display of the image on the device itself (display output), but also, for example, the output of the image to another functional unit on the device itself (internal output), the output of the image to a device other than the device itself (external device) (external output) or transmission (external transmission), etc.
[0020] The artifact correction unit 15 has a function of, for example, performing gray artifact correction preprocessing to estimate an area where artifacts may occur in an HDR synthesis image (a reference image complemented image in which a moving subject has been aligned) before HDR synthesis and to calculate various information for correcting the estimated area to an appropriate pixel value. The artifact correction unit 15 also has a function of, for example, performing gray artifact fill processing to fill in the reference image complemented image based on the results of the gray artifact correction preprocessing and correct the pixel values to an appropriate value where artifacts do not occur. The reference image complemented image that has been filled in based on the results of the gray artifact correction preprocessing may be referred to as a "gray artifact corrected image."
[0021] FIG. 2 is a conceptual diagram showing an example of an artifact that may occur when the image synthesis technique described in Patent Document 1 is used, for example.
[0022] The top part of Figure 2 shows an example of pattern (A) in which HDR synthesis is performed on a moving image of a person shaking their head. The pattern on the left (A-1) is an example of an image with ideal HDR compositing. In this example, the outline of the person, which would be blurred if the compositing mask was not applied, is clearly drawn. The trees and sky in the background, as well as the person's hair, are also clearly drawn with appropriate colors. The central pattern (A-2) is an example of an image that results when a high-brightness image is composited as a reference image in an area where a composite mask has been applied. In this example, a gray area (artifact) has occurred in the upper left outline of the person, encroaching on the sky area. This occurs because the sky area in the high-brightness image is blown out (for example, the pixel values are close to (R, G, B) = (255, 255, 255)), so the sky color is not properly reproduced and is composited as a gray color. Pattern (A-3) on the right is an example of an image that is synthesized using a low-brightness image as the reference image in an area where a synthesis mask has been applied. In this example, an artifact has occurred in which part of the texture of the hair is crushed to black. This is because the dark parts of the low-brightness image are crushed to black (for example, the pixel values are close to (R, G, B) = (0, 0, 0)), so the fine texture is not properly reproduced and is crushed to black when synthesized.
[0023] In this specification, for example, gray artifacts that occur when a high-brightness image corresponding to pattern (A-2) is used as a reference image and a reference image complemented image is created and synthesized are referred to as "gray artifacts." Also, for example, artifacts that occur when textures become dark and crushed, corresponding to pattern (A-3), are referred to as "local noise."
[0024] The bottom part of Figure 2 shows an example of pattern (B) in which HDR synthesis is performed on a moving image of a landscape in which trees are swaying in the wind. The pattern on the left (B-1) is an example of an image with ideal HDR compositing. In this example, the outlines of the trees, which would otherwise be blurred if the compositing mask was not applied, are clearly drawn. The sky and buildings in the background are also clearly drawn with appropriate colors. The central pattern (B-2) is an example of an image obtained by combining a high-brightness image with a reference image in the area where the combining mask is applied. In this example, gray artifacts appear at the boundaries between the trees and the sky, and between the sky and the building, encroaching on the sky area. The pattern on the right (B-3) is an example of an image that is synthesized using a low-brightness image as the reference image in the region where the synthesis mask is applied. In this example, local noise occurs in the leaf area.
[0025] In this specification, a method for reducing gray artifacts out of the two types of artifacts described above will be described.
[0026] FIG. 3 is a flowchart showing an example of the image processing procedure in this embodiment. The processing in the flowchart of FIG. 3 is realized, for example, by the processing unit of the image processing device 1 reading out the code of an HDR merging program stored in a storage unit (not shown) into a RAM (not shown) and executing the code.
[0027] Each symbol S in the flowchart of FIG. 3 represents a step. Furthermore, the flowchart described below merely shows one example of the procedure for image processing in this embodiment, and it goes without saying that other steps may be added or some steps may be deleted.
[0028] First, the acquisition unit 10 of the image processing device 1 executes an image acquisition process (S10). In the image acquisition process, the acquisition unit 10 acquires, for example, a captured image from the imaging unit 20 as an input. Then, for example, the area detection unit 11 performs a registration process to perform overall registration between the series of acquired images (S12). Then, for example, the region detection unit 11 executes a process for detecting a moving subject region, and generates a composite mask based on the detected moving subject region (S14). The image acquisition process, the positioning process, and the moving subject region detection process may be performed according to the process shown in FIG. 7 of Patent Document 1, for example.
[0029] Then, for example, the synthesis unit 12 executes a reference image complementation process (S16a). In the reference image complementation process, for example, the synthesis unit 12 removes moving subject blur using a synthesis mask based on the synthesis process of FIG. 7 described in Patent Document 1.
[0030] For example, in an area where a high-exposure image is prioritized as a reference image complement image (a high-brightness image is used as the reference image), lowering the brightness during HDR compositing can result in gray artifacts. Hereinafter, among the images to which a compositing mask is applied to generate a reference image complement image, a high-exposure image may be referred to as a "high-brightness target image," and a low-exposure image may be referred to as a "low-brightness target image."
[0031] Thereafter, for example, the artifact correction unit 15 refers to the high-brightness target image or the reference image complemented image and determines whether or not to remove the gray artifact (S20). For example, if a setting flag for removing gray artifacts is set, the artifact correction unit 15 may determine to remove gray artifacts.
[0032] If it is determined that gray artifacts are to be removed (S20: YES), for example, the artifact corrector 15 executes gray artifact correction preprocessing (S30).
[0033] FIG. 4 is a flowchart showing an example of the procedure for gray artifact correction pre-processing. First, the artifact correction unit 15 executes a binary extended composite mask generation process (S310). FIG. 5 illustrates a specific example of binary extended composite mask generation processing, which will be described below.
[0034] In the binary extended composite mask generation process, the artifact correction unit 15 first refers to a composite mask determined based on a high-brightness target image and a low-brightness target image, for example, in the moving subject region detection process. In the composite mask, for example, pixels with a value of "0" are represented by black, which has the lowest brightness value, and pixels with higher values are represented by a color closer to white, which has a higher brightness value. Areas where subject blur occurs (areas to which the composite mask is applied) are represented by pixels with high brightness values in the difference image. For example, if the input image includes buildings, roadside trees, and the sky, the area where subject blur occurs corresponds to the roadside trees.
[0035] For example, the artifact correction unit 15 refers to the composite mask and then performs a composite mask expansion process. In the composite mask expansion process, the artifact correction unit 15, for example, binarizes the composite mask, applies a blur filter, and generates an expanded composite mask. Then, the expanded composite mask is binarized again to generate a binary expanded composite mask.
[0036] In the composite mask expansion process, the artifact correction unit 15 may generate an expanded composite mask by, for example, binarizing the composite mask and applying a Gaussian filter.
[0037] In addition, in the composite mask expansion process, the artifact correction unit 15 may not distinguish between an expanded composite mask and a binary expanded composite mask. For example, the expanded composite mask may be a mask image obtained by spatially diffusing and binarizing the composite mask through processing such as blurring. For example, the binary expanded composite mask may be a mask image obtained by spatially diffusing the composite mask through processing such as blurring.
[0038] In the dilated composite mask or the binary dilated composite mask, for example, areas where pixel values are represented by white correspond to areas where gray artifacts may occur.
[0039] Returning to FIG. 4, for example, after executing the binary extended composite mask generation process, the artifact correction unit 15 executes the gray artifact region drawing process (S320). FIG. 6 illustrates a specific example of the gray artifact region drawing process, which will be described below. In the following, for example, of the pixel values (R, G, B) in a color image, an image that references the R value will be referred to as an "R channel image," an image that references the G value will be referred to as a "G channel image," and an image that references the B value will be referred to as a "B channel image."
[0040] In the gray artifact region drawing process, the artifact correction unit 15 determines whether or not the pixel value (R value) in the R channel image of the high-brightness target image exceeds a blown-out highlight threshold "α." Then, for example, if the pixel value exceeds the blown-out highlight threshold "α," the "R channel threshold image" is generated, which takes "255" and if the pixel value is equal to or less than the blown-out highlight threshold "α," the "R channel threshold image" is generated, which takes "0." Similarly, the artifact correction unit 15 generates a "G channel threshold image" based on the G channel image of the high-brightness target image, and also generates a "B channel threshold image" based on the B channel image of the high-brightness target image.
[0041] The blown-out highlight threshold "α" is a threshold indicating the conditions under which blown-out highlights occur, and may be set to a value close to "255" (for example, "230"), for example.
[0042] Also, for example, if the pixel value of the R channel image exceeds the whiteout threshold "α", it may take a value close to "255" (for example, "250"), and if it is equal to or less than the whiteout threshold α, it may take a value of "0" to generate an "R channel threshold image".
[0043] The same applies to the other channels.
[0044] Then, the artifact correction unit 15 multiplies the binary extended composite mask by the threshold image of each channel, and then generates a gray artifact region image by combining the images of each channel obtained by the multiplication. The artifact correction unit 15 may, for example, take the logical product of the extended synthesis mask and the threshold image of each channel.
[0045] In other words, areas where the pixel values of the gray artifact area image are high indicate areas where subject blur occurs and where blown-out highlights are likely to occur when synthesized based on a high-brightness target image (using the high-brightness target image as the reference image). The gray artifact region image can be said to be an image that extracts and expresses a region that may become a gray artifact in each of the (R, G, B) channels.
[0046] In the gray artifact region drawing process, the artifact correction unit 15 may generate a gray artifact region image by multiplying the binary extended synthesis mask by each channel value of the high-brightness target image, for example.
[0047] Returning to FIG. 4, for example, after executing the gray artifact region drawing process, the artifact correction unit 15 executes the brightness conversion parameter calculation process (S330). In the brightness conversion parameter calculation process, the artifact correction unit 15 calculates values (vR, vG, vB) corresponding to whiteout in the low-brightness target image in order to accurately calculate the area of the low-brightness target image that corresponds to the area where whiteout occurs in the high-brightness target image.
[0048] For example, if it is calculated that blown-out highlights will occur in the "x" percentile region in the R channel image of a high-brightness target image based on the blown-out highlight threshold "α," the average or median value of the top pixel values occupying the "x" percentile region in the R channel image of a low-brightness target image may be set as the blown-out highlight equivalent value "vR" for the R channel image. The same applies to the other channels.
[0049] The artifact correction unit 15 may not execute the luminance conversion parameter calculation process, but may use manually set values corresponding to blown-out highlights (vR, vG, vB), for example.
[0050] Furthermore, in the brightness conversion parameter calculation process, for example, if it is calculated that blown-out highlights will occur in the "x" percentile region in the R channel image of a high-brightness target image based on the blown-out highlight threshold "α," the average value or median value of the top pixel values occupying the "x×0.9" percentile region in the R channel image of a low-brightness target image may be set as the blown-out highlight equivalent value "vR" for the R channel image. The same applies to the other channels.
[0051] For example, after executing the luminance conversion parameter calculation process, the artifact correction unit 15 executes the background color extraction process (S340). FIG. 7 shows a specific example of the background color extraction process, which will be described below.
[0052] In the background color extraction process, the artifact correction unit 15 determines, for example, whether the pixel value of the R channel image of the low-brightness target image exceeds the blown-out highlight equivalent value "vR." The same applies to the other channels. Then, a background color image (1ch) is calculated by extracting an area in one channel (for example, the "B channel image") from the RGB channel images that has pixel values exceeding the value corresponding to blown-out highlights (for example, "vB"). Similarly, a background color image (2ch) is calculated by extracting an area in two channels from the RGB channel images that has pixel values exceeding the value corresponding to blown-out highlights, and a background color image (3ch) is calculated by extracting an area in all channels from the RGB channel images that has pixel values exceeding the value corresponding to blown-out highlights.
[0053] These background color images are the source information for appropriately complementing pixel values to suppress gray artifacts. The background color can be said to be a color that fills in areas where blown-out highlights occur in a high-brightness target image based on values equivalent to blown-out highlights in a low-brightness target image, for example.
[0054] In order to prevent further artifacts from occurring, for example, edge extraction processing may be performed on the low-brightness target image, and background color extraction processing may not be performed around the extracted edges.
[0055] Additionally, the background color images may be overlapped with each other.
[0056] In the background color extraction process, the artifact correction unit 15 may generate a background color image by multiplying the extended synthesis mask by each channel value of the low-brightness target image, for example.
[0057] Returning to FIG. 4, for example, after the artifact correction unit 15 has executed the background color extraction process, it then executes the background color complementation and enlargement process (S350). FIG. 8 illustrates a specific example of the background color complementation and expansion process.
[0058] In the background color complementation and expansion process, the artifact correction unit 15 complements and expands the background color image (1ch) extracted from the low-brightness target image, for example, independently for each (R, G, B) value to generate a background color complemented image (1ch). The same process is performed for other background color images.
[0059] An example of the complementary and expanded algorithm will now be described. The complementary and expanded algorithm may be, for example, as follows. (1) Based on the background color image P(0), a binarized image Q(0) is generated by binarizing P(0). (2) The background color image P(0) and the binarized image Q(0) are smoothed (for example, by applying box blur) to calculate bP(0) and bQ(0). (3) For each pixel value in the image, generate an image where P(1) = bP(0) ÷ bQ(0). (4). Reduce the size of P(1) and generate Q(1) from the reduced P(1). (5) Repeat steps (2) to (4) to expand the complementary region. (6) Based on P(n) and Q(n), the enlargement and synthesis are repeated to generate a background color complement image of the same size as P(0).
[0060] The background color complementation and expansion process can prevent, for example, the color information of the background color image from locally having too much influence on the corrected image.
[0061] Returning to FIG. 3, for example, after gray artifact correction pre-processing is performed (S30), the artifact correction unit 15 performs gray artifact filling processing (S40).
[0062] In the gray artifact filling process, the artifact correction unit 15 performs gray artifact removal processing based on, for example, the reference image complemented image, the gray artifact region image, the extended synthesis mask, and the background color complemented image, and generates a gray artifact corrected image.
[0063] Let CMPb(x,y) be the pixel value at position (x,y) of the R channel reference image complemented image before gray artifact removal, CMPa(x,y) be the pixel value at position (x,y) of the R channel gray artifact corrected image after gray artifact removal, and CPT(n,x,y) be the pixel value at position (x,y) of the R channel background color complemented image (n.ch). Also, let M(x,y) be the pixel value at position (x,y) of the extended synthesis mask, and GA(x,y) be the pixel value at position (x,y) of the R channel gray artifact region image. In this case, CMPa(x,y) can be calculated, for example, using the following formula: CMPa(x,y)=SUM_(n=1,2,3){w(x,y)×CPT(n,x,y)}+(1-w(x,y))×CMPb(x,y) Here, the weight w(x,y) = M(x,y) ∧ GA(x,y)
[0064] The same applies to the other G and B channels.
[0065] In practice, calculation may be skipped for areas where the weight w(x, y) is "0" to speed up the processing.
[0066] Furthermore, for example, the weight w(x, y) may be calculated using the pixel value at the position (x, y) of the binary extended composite mask as M(x, y).
[0067] Alternatively, instead of using all background color complement images, only a specific channel (e.g., "3ch") may be used, or only two predetermined channels (e.g., "1ch" and "2ch") may be used.
[0068] If it is determined that the gray artifacts are not to be removed (S20: NO), for example, the artifact corrector 15 skips the gray artifact correction pre-processing (S30) and the gray artifact filling-out process (S40).
[0069] For example, the artifact correction unit 15 may perform gray artifact correction preprocessing without making the determination in S20. For example, if the area of a region in the gray artifact region image where each pixel value exceeds a threshold exceeds a predetermined ratio of the gray artifact region image, the artifact correction unit 15 may determine to remove the gray artifact. Furthermore, for example, if the average value of each pixel value in the gray artifact region image exceeds a threshold, the artifact correction unit 15 may determine to remove the gray artifact.
[0070] Then, the composition unit 12 executes, for example, HDR composition processing (S16b). In the HDR composition processing, for example, the composition unit 12 composes and outputs an HDR image based on the reference image complement image to which an exposure conversion function (which may be referred to as a mixture ratio determined according to the difference in luminance values of the images) has been applied. In addition, in the processing of step S16b, for example, if gray artifact removal is selected (S20: YES), the synthesis unit 12 may perform HDR synthesis processing using the gray artifact corrected image instead of the reference image complemented image.
[0071] Then, the synthesis unit outputs the HDR synthesized image (S18).
[0072] FIG. 9 illustrates and explains a specific example of the effect of the gray artifact filling process. 9, for example, in the HDR composite image using the reference image complemented image, gray artifacts occur around the roadside trees that overlap with the sky. In contrast, in the HDR composite image using the gray artifact-corrected image to which a background color complemented image weighted based on the extended composite mask and the gray artifact region image is added, the gray artifacts around the roadside trees are corrected to an appropriate sky blue, and the artifacts disappear.
[0073] The above-described embodiment shows an example of an information processing device according to the present invention. The information processing device according to the present invention is not limited to the image processing device 1 according to the embodiment. For example, in each of the above-described embodiments, an example in which the imaging unit 20 acquires frame images has been described, but images may also be acquired from another device via a network. Furthermore, the image processing device 1 according to each of the above-described embodiments may be configured to operate together with an optical image stabilization device.
[0074] [Example] Next, examples of a terminal, an electronic device (electronic equipment), and an information processing device to which the image processing device 1 described above is applied or which includes the image processing device 1 described above will be described. Here, as an example, an embodiment of a smartphone, which is a type of mobile phone with a camera function (with an image capturing function), will be described, however, it goes without saying that the embodiments to which the present invention can be applied are not limited to this embodiment.
[0075] FIG. 10 is a diagram illustrating an example of the functional configuration of the smartphone 1000. As shown in FIG. The smartphone 1000 includes, for example, a processing unit 100, a memory unit 200, an imaging unit 310, an inertial measurement unit 320, an operation unit 330, a display unit 340, a sound input unit 350, a sound output unit 360, and a communication unit 370.
[0076] The processing unit 100 is a processing device that comprehensively controls each unit of the smartphone 1000 in accordance with various programs such as system programs stored in the memory unit 200, and performs various processes related to imaging and synthesis processing, and is configured with processors such as a CPU, GPU, and DSP, and integrated circuits such as ASIC.
[0077] The processing unit 100 includes an acquisition unit 10, a region detection unit 11, a synthesis unit 12, an artifact correction unit 15, and a display control unit 18 as main functional units. These functional units correspond to the functional units included in the image processing device 1 in FIG. 1, for example. The display control unit 18 also has a function of causing the display unit 340 to display and output the processing results of the processing unit 100, for example.
[0078] The storage unit 200 is a storage device configured to include a volatile or non-volatile memory such as a ROM, an EEPROM, a flash memory, or a RAM, a hard disk drive, or the like.
[0079] The storage unit 200 stores, for example, an HDR composite image program 210 and a camera image temporary storage unit 220.
[0080] The HDR composite image program 210 is a program that is read by the processing unit 100 and executed as HDR composite image program processing. The HDR composite image program 210 may also be called a camera application program.
[0081] The camera image temporary storage unit 220 is, for example, a buffer (frame buffer) in which captured images (image sensor images) captured by the imaging unit 310 and the like are stored.
[0082] The imaging unit 310 is an imaging device configured to be able to capture an image of any scene, and is configured with an imaging element (semiconductor element) such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary MOS) image sensor. The imaging unit 310 forms an image of light emitted from an object to be imaged on the light-receiving plane of the imaging element using a lens (not shown), and converts the brightness of the light in the image into an electrical signal by photoelectric conversion. The converted electrical signal is converted into a digital signal by an A / D (Analog-Digital) converter (not shown) and output to the processing unit 100.
[0083] The inertial measurement unit 320 includes, for example, a gyro sensor that detects angular velocities around three axes (pitch, roll, and yaw) and an acceleration sensor that detects inertial forces in the axial directions of the three axes (pitch, roll, and yaw). The detection results of the inertial measurement unit 320 are output to the processing unit 100 as needed.
[0084] The operation unit 330 is configured to have input devices such as operation buttons and operation switches that allow the user to input various operations to the smartphone 1000. The operation unit 330 also has a touch panel (not shown) that is configured integrally with the display unit 340, and this touch panel functions as an input interface between the user and the smartphone 1000. The operation unit 330 outputs an operation signal to the processing unit 100 in accordance with a user operation.
[0085] The display unit 340 is a display device configured to include an LCD (Liquid Crystal Display), an OLED (Organic Electro-luminescence Display), or the like, and performs various displays based on display signals output from the display control unit 18.
[0086] The sound input unit 350 is a sound input device including a microphone, an A / D converter, etc., and performs various sound inputs based on sound input signals input to the processing unit 100.
[0087] The sound output unit 360 is a sound output device including a D / A converter, a speaker, etc., and outputs various sounds based on the sound output signal output from the processing unit 100.
[0088] The communication unit 370 is a communication device for transmitting and receiving information used within the device to and from an external information processing device. As a communication method for the communication unit 370, various methods can be applied, such as a wired connection via a cable conforming to a predetermined communication standard such as Ethernet or USB (Universal Serial Bus), a wireless connection using a wireless communication technology conforming to a predetermined communication standard such as Wi-Fi (registered trademark) or 5G (fifth generation mobile communication system), and a connection using short-range wireless communication such as Bluetooth (registered trademark).
[0089] The processing unit 100 of the smartphone 1000 performs image capturing processing in accordance with the HDR composite image program 210 stored in the storage unit 200 .
[0090] The image capturing process may be performed, for example, according to the flowcharts of Figures 3 and 4. In this case, the acquisition unit 10 may acquire the image, for example, by the imaging unit 310. Furthermore, the composition unit 12 may cause the display unit 340 to display the composition process result and / or the gray artifact filling process result via the display control unit 18, for example.
[0091] In the above embodiments, the present invention is applied to an information processing device, a terminal, an electronic device (electronic equipment), and a smartphone, which is an example of an image synthesis device, but is not limited thereto. The present invention can be applied to various devices such as a video camera, a workstation, a tablet terminal, and a server.
[0092] [Actions and Effects of the Embodiments] The image synthesis device (e.g., image processing device 1) in this embodiment acquires a plurality of images including a first image (e.g., a low-brightness image) captured under a first exposure condition (e.g., low exposure), a second image (e.g., a high-brightness image) captured at a second timing under a second exposure condition (e.g., high exposure) that is higher than the first exposure condition, and a third image (e.g., a high-brightness image) captured at a third timing under the first exposure condition, compares pixel values of the first and third images to detect a moving subject region (e.g., a synthesis mask) in which a moving subject is depicted, and performs a moving subject blur removal process (e.g., a reference image interpolation process) that brings pixel values corresponding to the moving subject region closer to pixel values corresponding to the moving subject region of the second image, and This shows an example of a configuration including a correction unit that corrects composite distortion that occurs in a target image when combining a target image (e.g., a reference image complement image) after blur removal processing with a second image, and the correction unit calculates an extended difference (e.g., a binary extended composite mask) that spatially diffuses the moving subject area, generates first correction information (e.g., a gray artifact area image) regarding the area in which the composite distortion of the composite image is to be corrected based on the second image (e.g., a high-brightness target image) and the extended difference, calculates second correction information (e.g., a background color image) that corrects the composite distortion based on at least the first image or a third image (e.g., a low-brightness target image), and corrects the composite distortion based on the target image, the first correction information, the second correction information, and the extended difference. This makes it possible to identify areas where composite distortion may occur in the target image before correction using the extended difference, and to effectively correct composite distortion from the target image by using first correction information based on an image captured under the first exposure conditions and second correction information based on an image captured under the second exposure conditions.
[0093] [Recording Media] In the above embodiments, various programs and data related to image processing are stored in the storage unit 200, and the image processing in each of the above embodiments is realized by the processing unit reading and executing these programs. In this case, the storage unit of each device may have, in addition to internal storage devices such as ROM, EEPROM, flash memory, hard disk, and RAM, non-transitory tangible recording media (recording media, external storage devices, storage media) such as a memory card (SD card), CompactFlash (registered trademark) card, memory stick, USB memory, CD-RW (optical disc), and MO (magneto-optical disc), and the above various programs and data may be stored in these recording media. These storage media are examples of computer-readable non-transitory recording media (storage media).
[0094] FIG. 11 is a diagram showing an example of a recording medium in this case. In this example, the image processing device 1 is provided with a card slot 410 for inserting a memory card 430, and a card reader / writer (R / W) 420 for reading information stored on the memory card 430 inserted into the card slot 410 or writing information to the memory card 430.
[0095] Under the control of the processing unit, the card reader / writer 420 writes the programs and data recorded in a storage unit (not shown) to the memory card 430. The programs and data recorded in the memory card 430 can be read by an external device other than the image processing device 1, so that the image processing in the above-described embodiment can be realized in the external device.
[0096] The recording medium can also be applied to various devices such as a terminal (smartphone) equipped with the image processing device 1 described in the above embodiment, an information processing device, an electronic device (electronic equipment), and an image synthesis device. [Explanation of symbols]
[0097] 1. Image processing device 10 Acquisition Department 11 Area detection unit 12 Synthesis section 15 Artifact Correction Unit 1000 smartphones
Claims
1. An image synthesis device that generates a synthetic image by synthesizing at least two images from a series of multiple images captured of the same subject, an acquisition unit that acquires the plurality of images including a first image captured under a first exposure condition, a second image captured at a second timing under a second exposure condition that has higher exposure than the first exposure condition, and a third image captured at a third timing under the first exposure condition; a region detection unit that compares pixel values of the first image and the third image to detect a moving subject region where a moving subject is depicted; a combining unit that performs a moving subject blur removal process to bring pixel values corresponding to the moving subject region closer to pixel values corresponding to the moving subject region of the second image, and combines the target image after the moving subject blur removal process with the second image; a correction unit that corrects a composite distortion occurring in the target image; Equipped with The correction unit calculating a spatially diffused dilation difference of the moving subject region; generating first correction information relating to an area in which the composite distortion is corrected based on the second image and the dilation difference; calculating second correction information for correcting the composite distortion based on at least the first image or the third image; correcting the composite distortion based on the target image, the first correction information, the second correction information, and the dilation difference; Image synthesis device.
2. 2. The image synthesis device according to claim 1, The correction unit calculating a first dilated difference obtained by spatially diffusing the moving subject region and a second dilated difference obtained by binarizing the first dilated difference; generating the first correction information based on the second image and the second dilation difference; correcting the composite distortion based on the target image, the first correction information, the second correction information, and the first dilation difference; Image synthesis device.
3. 3. The image synthesis device according to claim 2, The correction unit performing a first threshold determination process for each color channel of the second image using a threshold that may cause overexposure in the target image; generating the first correction information based on a result of the first threshold determination process and the second extension difference; Image synthesis device.
4. 4. The image synthesis device according to claim 3, The correction unit calculating a corresponding value corresponding to the threshold value in the first image or the third image based on a result of the first threshold determination process; calculating the second correction information by a second threshold determination process based on the corresponding value; Image synthesis device.
5. 5. The image synthesis device according to claim 4, the second correction information is composed of a plurality of channels based on the results of the second threshold determination process for each color channel of the first image or the third image; Image synthesis device.
6. 5. The image synthesis device according to claim 4, the correction unit complements and enlarges the second correction information for each color channel to a size that covers the composite image. Image synthesis device.
7. 7. The image synthesis device according to claim 6, The correction unit a product of the first extension difference and the first correction information is used as a weight; adding the second correction information to the target image in accordance with the weight; Image synthesis device.
8. An image synthesis method for generating a synthetic image by synthesizing at least two images from a series of multiple images captured of the same subject, comprising: acquiring the plurality of images including a first image captured under a first exposure condition, a second image captured at a second timing under a second exposure condition with higher exposure than the first exposure condition, and a third image captured at a third timing under the first exposure condition; comparing pixel values of the first image and the third image to detect a moving subject region in which a moving subject is depicted; performing a moving subject blur removal process to bring pixel values corresponding to the moving subject region closer to pixel values corresponding to the moving subject region of the second image, and combining the target image after the moving subject blur removal process with the second image; calculating a dilated difference by spatially diffusing the moving subject region; generating first correction information relating to an area in which a composite distortion of the target image is corrected based on the second image and the dilation difference; calculating second correction information for correcting the composite distortion based on at least the first image or the third image; correcting the composite distortion based on the target image, the first correction information, the second correction information, and the dilation difference; An image synthesis method comprising:
9. An image synthesis device that synthesizes at least two images from a series of multiple images captured of the same subject to generate a synthesized image, acquiring the plurality of images including a first image captured under a first exposure condition, a second image captured at a second timing under a second exposure condition with higher exposure than the first exposure condition, and a third image captured at a third timing under the first exposure condition; comparing pixel values of the first image and the third image to detect a moving subject region in which a moving subject is depicted; performing a moving subject blur removal process to bring pixel values corresponding to the moving subject region closer to pixel values corresponding to the moving subject region of the second image, and combining the target image after the moving subject blur removal process with the second image; calculating a dilated difference by spatially diffusing the moving subject region; generating first correction information relating to an area in which a composite distortion of the target image is corrected based on the second image and the dilation difference; calculating second correction information for correcting the composite distortion based on at least the first image or the third image; correcting the composite distortion based on the target image, the first correction information, the second correction information, and the dilation difference; A program that executes the following.
Citation Information
Patent Citations
Image synthesis device, image synthesis method, image synthesis program, and storage medium
JP6823942B2