High dynamic range image acquisition with multiple exposure

The method allows low-cost digital cameras to generate and display HDR images by employing image registration, fusion, and ghost removal processes, addressing the computational challenges of existing technologies and enhancing image detail.

DE102012007838B4Active Publication Date: 2025-10-09QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102012007838
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2011-04-20
Filing Date
2012-04-19
Publication Date
2025-10-09
Estimated Expiration
2032-04-19

AI Technical Summary

Technical Problem

Existing digital cameras with limited processing power struggle to efficiently generate and display high dynamic range (HDR) images from multiple exposures due to the complexity and computational demands of image registration, ghost removal, and tone mapping, making it impractical for low-cost cameras to produce HDR images in-camera.

Method used

A method and apparatus for generating HDR images using image registration, fusion, and ghost removal processes that can be performed on low-power CPUs, allowing in-camera processing and display of HDR images by clustering pixels into patches, selecting replacement images, and converting bit formats, while utilizing look-up tables for weight calculation and tone mapping.

Benefits of technology

Enables the generation and display of HDR images on low-cost digital cameras using limited processing power, reducing computational time and cost, and improving image detail in shadow and highlight regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for blending a plurality of digital images of a scene, the plurality comprising more than two images, comprising: Taking pictures at different exposure levels; Registering counterpart pixels of each image to each other; Deriving a normalized image exposure level for each image; Using the normalized image exposure level in an image merging process; Using the image blending process to blend a first selected image and a second selected image to create an intermediate image, wherein the image blending begins before capturing all of the plurality of images, and Repeating the image merging process using the previously generated intermediate image instead of the first selected image and another selected image instead of the second selected image until all images are merged, and outputting the last generated intermediate image as the merged output image.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority under 35 USC 119(e) to U.S. Provisional Application Serial No. 61 / 326,769, entitled "Multiple Exposure High Dynamic Range Imaging," filed April 22, 2010, and is a continuation-in-part of U.S. Patent Application Serial No. 12 / 763,693, entitled "Multiple Exposure High Dynamic Range Imaging," filed April 20, 2010, which claims priority to U.S. Provisional Patent Application Serial No. 61 / 171,936, entitled "Multiple Exposure HDR," filed April 23, 2009. BACKGROUND OF THE INVENTION 1. Field of the Invention

[0002] This application relates to the acquisition and processing of images displaying the full range of grayscale colors appearing in a physical scene, often referred to as a "high dynamic range" or "HDR" image. More specifically, it relates to a system and method for image acquisition and processing of an HDR image in a digital image capture device such as a consumer digital camera. 2. Discussion of the state of the art

[0003] Images captured by digital cameras are mostly low-dynamic-range (LDR) images, where each image pixel has a limited number of digital bits per color. The number of digital bits per pixel is called the digital pixel bitwidth value, which is typically 8 bits. Such 8-bit pixels can be used to form an image with 256 different shades of gray for each color at each pixel location. In an LDR image of a scene, shadowy areas of the scene are shown as completely black (black saturation), bright, sunlit areas of the scene as completely white (white saturation), and scene areas in between are shown in a range of gray tones. A high-dynamic-range (HDR) image is one that has digital pixel bitwidth values ​​of more than 8 bits; 16 bits per pixel is one possible value. In such an image, the full range of gray tones that appear in a physical scene can be displayed.These gray tones provide image details present in the shadows, highlights, and midtones of the scene that are missing in the LDR image. Thus, in an HDR image, scene details are present in dark image fields that are in shadow due to their proximity to tall buildings and trees, in light fields directly illuminated by bright sunlight, and in mid-light fields exposed between these two extremes.

[0004] An HDR image can be created by capturing numerous LDR images of a scene taken at different exposure levels. These multiple LDR images are called a bracketed image series. A low exposure level correctly captures the gray tones in scene areas fully illuminated by bright sunlight, and a high exposure level correctly captures the gray tones in scene areas completely shielded from the sun and sky by buildings and trees. At the low exposure level, however, the shadow areas of the scene are completely black, with black saturation, and show no detail, and the midtone areas lose detail. Furthermore, at the high exposure level, the highlights of the scene are completely white, with white saturation, and show no detail, and the midtone areas again lose detail.A third mid-level image is often also taken, which correctly captures mid-level gray tones. By blending these three LDR images, an HDR image can be created that represents the entire grayscale range of the scene.

[0005] Deriving an HDR image from a bracketed image series currently requires a complex implementation using an expensive computing engine. This is due to the need to perform three separate processing operations to correctly blend the bracketed image series into a single HDR image, and a fourth to convert the resulting image, now consisting of pixels with digital pixel bitwidth values ​​of more than 8 bits per color, into an image that can be displayed on common 8-bit-per-pixel-per-color displays. These four processing operations are: “Image registration” for accurate alignment of the many images to each other; “Image blending” to merge the many images with the right weighting; “Ghost removal” to remove shifted renditions of scene objects or “ghost images” that would appear in the blended HDR image due to the movement of these objects during the time in which the multiple images were captured; and Tone mapping to prepare the final HDR image for presentation on common display devices limited to displaying 8-bit-per-pixel-per-color image pixels.

[0006] Executing these four processing operations requires a large number of floating-point operations in a short period of time, as can be seen from a review of the article "High Dynamic Range Imaging Acquisition, Display, and Image-Based Lighting" by Erik Reinhard, Sumanta Pattanaik, Greg Ward, and Paul Debevec, published by Morgan Kaufmann Publishers, copyright 2005 by Elsevier, Inc. This is particularly the case with the image blending and ghost removal processing operations. Thus, powerful and expensive computing engines (central processing units, or CPUs) must be used. Their cost may be tolerable when used with professional digital cameras, but they represent an impractical solution for inexpensive "point-and-shoot" digital cameras, which include CPUs with limited processing power.

[0007] An HDR image can be created from a bracketed series of images captured by a low-cost digital camera by uploading the image series from the camera to a general-purpose computer, such as a personal computer (PC). An image processing application, such as Adobe Photoshop, can be used to combine the required complex HDR image combination process on a desktop. This approach is not efficient or suitable and does not meet the requirements of restoring an HDR image to the camera's built-in display shortly after it is captured.

[0008] US 2010 / 0 328 315 A1 describes a method and apparatus for generating a radiation map in the generation of high dynamic range (HDR) images by calculating mean difference curves for a sequence of images taken with different exposures and a transformation curve from the mean difference curve by an algorithm approximating the Debevec function, by means of which a radiation map can be calculated from the taken images.

[0009] WO 2010 / 123 923 A1 describes techniques for creating a high dynamic range (HDR) image in a consumer digital camera from a series of images of a scene captured at different exposure levels, and displaying the HDR image on the camera's integrated display. The approach uses blending of images of the series to incorporate both shadow and highlight detail of the scene and removes "ghosts" that arise in the blended HDR image due to movement in the scene during the capture of the series. The low computational complexity of the image blending and deghosting processes of the present invention, as well as the ability to initiate image blending and deghosting before all images in the series are captured, can significantly reduce the time required to generate and display a tone-mapped HDR image.

[0010] There is therefore a need for an in-camera method and apparatus that can quickly generate an HDR image from a bracketed series of images and display it on the in-camera display device shortly after capture, using a CPU with limited processing power. SUMMARY OF THE INVENTION

[0011] Systems, methods, and computer program products for generating a composite high dynamic range image from multiple exposures are disclosed.

[0012] Embodiments of the invention can detect local motion between a reference image and a plurality of comparison images, cluster relevant pixels into patches, select a replacement image from a plurality of candidate replacement images as a source of replacement patch image data, replace patch image data in a composite image, and

[0013] Convert the original format to a format with fewer bits per pixel for display.

[0014] The invention can be implemented as a method for blending a plurality of digital images of a scene, including capturing the images at different exposure levels, registering counterpart pixels to each image, deriving a normalized image exposure level for each image, and using the normalized image exposure levels in an image blending process. The image blending process includes using the image blending process to blend a first selected image and a second selected image to create an intermediate image, and if the plurality consists of two images, outputting the intermediate image as a blended output image.When the plurality consists of more than two images, the image merging process includes repeating the image merging process using the previously generated intermediate image in place of the first selected image and a further selected image in place of the second selected image until all images are merged, and outputting the last generated intermediate image as the merged output image.

[0015] The method may also include selectively converting the images from a lower bits-per-pixel format to a higher bits-per-pixel format prior to using the image merging process, and then selectively converting the merged output image to a predetermined lower bits-per-pixel format.

[0016] The image fusion process fuses the counterpart pixels of two images and includes deriving a luma value for a pixel in the second selected image, using the luma value of a second selected image pixel as an index into a lookup table to obtain a weight value between the numbers zero and one, using the weight value, the normalized exposure level of the second selected image, and the second selected image pixel to generate a processed second selected image pixel, selecting a first selected image pixel corresponding to the second selected image pixel, using the first selected image pixel and the result of subtracting the weight value from one to generate a processed first selected image pixel, adding the processed first selected image pixel to the processed second selected counterpart image pixel to generate a fused image pixel,and repeating the above processing sequence until every second selected image pixel is merged with its first selected counterpart image pixel.,

[0017] The method may include the feature that the obtained weighting value decreases as the luma value used as an index into the lookup table increases. The method may also include using a different lookup table for each image to obtain the weighting value. The method may also begin image blending before all images of the plurality are captured. The method may also begin image blending immediately after the second image of the plurality is captured.

[0018] The embodiments further include a method for removing displaced representations of scene objects that appear in a blended image produced by a digital image blending process applied to a plurality of images taken at different exposure levels and at different times, wherein the counterpart pixels of the images are registered to each other.The method includes normalizing luma values ​​of the images to a specific standard derivative and mean, detecting local motion between at least one reference image and at least one comparison image, clustering comparison image pixels with local motion into patches, selecting corresponding patches from the reference image, generating a merged binary image by logically ORing the patches generated from certain reference images together, and fusing the merged image with the reference images, wherein each reference image is weighted by a weight value calculated from the merged binary image to generate an output image.

[0019] The method may detect local motion by determining an absolute luma variance between each pixel of the reference image and the comparison image to generate a difference image, and detecting difference image zones whose absolute luma variances exceed a threshold. Clustering may involve finding sets of connected detected image blobs using morphological operations and bounding each set by a polygon. The selected images used as reference images may include the image with the lowest exposure level, the image with the highest exposure level, or any images with intermediate exposure levels. The luma value of the images to be processed may be downgraded before processing.

[0020] An embodiment selects reference images for a candidate reference image having an intermediate exposure value, calculating a sum of patches of saturated zones and a ratio of the sum of saturated patches to total patches, selecting the candidate reference image as the reference image if the ratio is less than or equal to a parameter value, for example, 0.03 to 1, and selecting an image of less than the intermediate exposure value as the reference image if the ratio is greater than the parameter value.

[0021] The invention may also be embodied as an image capture device that captures a plurality of digital images of a scene at different exposure levels, comprising an image registration processor that registers counterpart pixels of each image of the plurality to each other, and an image blender that combines multiple images of the plurality to produce a single image. The image blender may comprise an image normalizer that normalizes the image exposure level for each image, and an image merger that uses the normalized exposure level to blend a first selected image and a second selected image to produce an intermediate image, and, if the plurality consists of two images, outputs the intermediate image as a blended output image.If the plurality consists of more than two images, the image merger may repeatedly merge the previously generated intermediate image in place of the first selected image and a further selected image in place of the second selected image until all images are merged, and outputs the last generated intermediate image as the merged output image.

[0022] The image capture device may further comprise a gamma correction converter that selectively converts the images from a lower bits-per-pixel format to a higher bits-per-pixel format prior to image blending, and selectively converts the blended output image to a predetermined lower bits-per-pixel format.

[0023] The image capturing device may further include a luma conversion circuit that outputs the luma value of an input pixel, a lookup table that outputs a weighting value between the numbers zero and one for an input luma value, a processing circuit that receives a derived luma value for a pixel in the second selected image from the luma conversion circuit, obtains a weighting value from the lookup table for the derived luma value, and generates a processed second selected image pixel from the second selected image pixel, the normalized exposure level of the second selected image, and the weighting value, and a second processing circuit that receives as inputs the result of subtracting the weighting value from one and the first selected image pixel corresponding to the second selected image pixel and generates a processed first selected image pixel, and an addition circuit,which adds the processed first selected image pixel to the processed second selected image pixel to produce a merged image pixel.

[0024] The image capture device may output a lower weighting value as the luma value input to the lookup table increases. It may also use a first lookup table for the first selected image and a second lookup table for the second selected image. The image capture device may begin blending the plurality of images before capturing all of the plurality of images. It may begin blending the plurality of images immediately after capturing the second of the plurality.

[0025] The invention can be further embodied as an image capture device that captures a plurality of digital images of a scene at different exposure levels and at different times, comprising an image registration processor that registers the counterpart pixels of each image to each other, an image mixer that combines multiple images to create a mixed image, and a ghost remover. The ghost remover removes spatially shifted representations of scene objects that appear in the mixed image and may include an image normalizer that normalizes the image exposure level for each image to a specific standard deviation and mean, a reference image selection circuit that selects at least one reference image from the images, a local motion detector circuit that detects local motion between the reference image and at least one comparison image, a cluster circuit,clusters the comparison image pixels with local motion into patches and selects corresponding patches from the reference image, and comprise a connected binary image generator circuit that logically ORs the patches generated from certain reference images together, wherein the image mixer mixes the mixed image with the reference images in a final stage of processing, with each reference image weighted by a weighting value calculated from the connected binary image to produce an output image.

[0026] The local motion detector circuit may detect local motion by determining an absolute luma variance between each pixel of the reference image and the comparison image to generate a difference image, and by identifying difference image zones that have absolute luma variances that exceed a threshold. The clustering circuit may cluster comparison image pixels by finding groups of connected image blobs using morphological operations and bounding each group by a polygon. The selected scene images used as reference images may include the scene image with the lowest exposure level, the scene image with the highest exposure level, or all scene images with intermediate exposure levels. The image capture device may also include circuitry for downscaling the luma values ​​for the images.The reference image selecting circuit may select a reference image from the scene images for a candidate reference image having an intermediate exposure value by calculating a sum of patches of saturated zones and a ratio of the sum of saturated patches to total patches, selecting the candidate reference image as the reference image when the ratio is less than or equal to a parameter value, for example, from 0.03 to 1, and selecting an image of less than an intermediate exposure value as the reference image when the ratio is greater than the parameter value.

[0027] Another embodiment is a digital camera that captures a plurality of digital images of a scene at different exposure levels at different times and generates a tone-mapped high-dynamic-range image therefrom. The camera may include an image registration processor that registers counterpart pixels of each image of the plurality to each other, an image blender that combines multiple images of the plurality to generate a single image, a ghost remover that removes displaced representations of scene objects appearing in the blended image, and a tone-mapping processor that maps the processed blended image pixels to display pixels with the number of digital bits that can be presented on a built-in digital camera image display.

[0028] The image mixer may include an image normalizer that normalizes the image exposure level for each image, and an image merger that uses the normalized exposure level to merge a first selected image and a second selected image to generate an intermediate image, and if the plurality consists of two images, outputs the intermediate image as the merged image. If the plurality consists of more than two images, the image mixer may repeatedly merge the previously generated intermediate image in place of the first selected image and another selected image in place of the second selected image until all images have been merged, and outputs the last generated intermediate image as the merged image.

[0029] The ghost remover may include an image normalizer that normalizes the image exposure level for each image to a specific standard deviation and mean, a reference image selection circuit that selects at least one reference image from the images, a local motion detector circuit that detects local motion between the reference image and at least one comparison image, a clustering circuit that clusters comparison image pixels with local motion into patches and selects corresponding patches from the reference image, and a linked binary image generator circuit that logically ORs the patches generated from certain reference images together, wherein the image mixer combines the mixed image with the reference images in a final processing step, wherein each reference image is weighted by a weight value calculated from the linked binary image to generate a processed mixed image.

[0030] The digital camera may implement the image registration processor, the image mixer, the ghost remover, and / or the tone mapping processor by using at least one programmable general-purpose processor, or at least one programmable application-specific processor, or at least one application-specific fixed-function processor.

[0031] The invention can be implemented as a method for detecting local motion between a reference image and a plurality of comparison images. The method can comprise defining a number of luma thresholds that determine a group of reference image luma value ranges from which a particular comparison image is selected, the number corresponding to the number of comparison images, and defining a difference threshold function that sets thresholds according to luma values ​​of the reference image. Then, for each particular comparison image, the method can comprise generating an intermediate detection map characterizing the local motion between the reference image and the particular comparison image by applying the thresholds to a difference image formed by comparing the reference image and the particular comparison image.Finally, the method may output a final detection map that is a union of the intermediate detection maps. The method may select a comparison image with relatively lower luma within relatively higher reference image luma value ranges, and vice versa. The luma threshold set may vary with difference image luma values. The intermediate detection map may be generated by considering pixels with luma values ​​in a range corresponding to the determined comparison image, calculating the standard deviation and mean of the considered pixels, normalizing the respective pixels to have the same standard deviation and mean for the reference image and the determined comparison image, generating an absolute difference map of the considered pixels, and applying the difference threshold function to the difference map to detect local motion detections.

[0032] Another embodiment of the invention is a method for clustering pixels into patches, comprising applying a morphological dilation operation to a binary image of relevant pixels, with a 5x5 square structuring element, for example; applying a morphological closure operation to the binary image, with a 5x5 square structuring element, for example; applying a binary labeling algorithm to distinguish different patches; describing each patch by a boundary polygon; and outputting the patch description. In this embodiment, the relevant pixels may share detected characteristics, which may include non-uniform luma values ​​from at least one image comparison or detected local motion. The boundary polygon may be an octagon.

[0033] In another embodiment, a method is provided for selecting a surrogate image from a plurality of candidate surrogate images as a source of surrogate patch image data. The method may include calculating a weighted histogram of luma values ​​of edge region pixels of a particular patch of a reference image, dividing the histogram into a plurality of zones according to thresholds determined from relative exposure values ​​of the candidate surrogate images, calculating a scoring function for each histogram zone, selecting the zone with the maximum score, and outputting the corresponding candidate surrogate image. The histogram weighting may increase the influence of oversaturated and undersaturated luma values. In this embodiment, candidate surrogate images with relatively low exposure values ​​may be selected to replace patches of reference images with relatively high exposure values, and vice versa.The reference image can be a medium-exposure image with boosted luma values. The weighting function for a given histogram zone can be defined as the ratio of the number of pixels in the given histogram zone, taking into account the size of the given histogram zone, to the mean difference of the histogram inputs from the luma value mode for the given histogram zone.

[0034] Additionally, the invention may be implemented as a method for replacing patch image data in a composite image, comprising smoothing a patch image boundary polygon, upscaling the smoothed patch images to full image resolution, and merging the composite image and the upscaling smoothed patch image to generate an output image. The patch image boundary polygon may be an octagon. The patch image boundary octagon may be represented in 8-bit format, and a low-pass filter, for example, 7x7, is applied.

[0035] Another embodiment is a method for converting images from an original format to a lower bits-per-pixel format, including performing tone mapping that adapts to image luma values ​​and outputting the tone-mapped image. The method may further include performing non-linear local tone mapping that maps a pixel according to the average luma values ​​of its neighbors into individual color components, each represented by a predetermined number of bits. The method may further allocate additional grayscale levels for more frequent luma values. Images may be converted to a display format having pixels with the number of bits that can be presented on a built-in digital camera image display.The method may also define a family of local mapping lookup tables and a single global mapping lookup table, and use a selected lookup table for each image transformation operation, where the selection of the lookup table is based on the image scene characteristics. The method may further construct a brightness histogram of a composite image and define a mapping according to the distribution of brightness values. The brightness histogram may be estimated using a combination of histograms of the downgraded images used to form the composite image. Histogram equalization may define the mapping.

[0036] The invention can be embodied as a system for detecting local motion between a reference image and a plurality of comparison images. The system can include a processor that defines a number of luma thresholds that determine a group of reference image luma value ranges from which a particular comparison image is selected, the number corresponding to the number of comparison images, and a difference threshold function that sets thresholds according to luma values ​​of the reference image. Then, for each particular comparison image, the processor can generate an intermediate detection map indicative of local motion between the reference image and the particular comparison image by applying the thresholds to a difference image formed by comparing the reference image and the particular comparison image. Finally, the processor can output a final detection map that is a union of the intermediate detection maps.The system can select a comparison image with relatively lower luma within relatively higher reference image luma value ranges, and vice versa. The luma threshold set can vary with difference image luma values. The intermediate detection map can be generated by considering pixels with luma values ​​in a range corresponding to the specific comparison image, calculating the standard deviation and mean of the considered pixels, normalizing the considered pixels so that they have the same standard deviation and mean for the reference image and the specific comparison image, generating an absolute difference map of the considered pixels, and applying the difference threshold function to the difference map to detect local motion detections.

[0037] Another embodiment of the invention is a system for clustering pixels into patches, comprising a processor that can apply a morphological closure operation to a binary image of relevant pixels, with a 5x5 square structuring element, for example, apply a morphological closure operation to the binary image, with a 5x5 square structuring element, for example, apply a binary labeling algorithm to distinguish different patches, describe each patch by a boundary polygon, and output the patch description. The relevant pixels can share detected characteristics, which can include inconsistent luma values ​​from at least one image comparison or detected local motion. The boundary polygon can be an octagon.

[0038] In another embodiment, a system for selecting a replacement image from a plurality of candidate replacement images as a source of replacement patch image data is provided. The system may include a processor that calculates a weighted histogram of luma values ​​of fringe field pixels of a particular patch of a reference image, divides the histogram into a plurality of zones according to thresholds determined by relative exposure values ​​of the candidate replacement images, calculates a scoring function for each histogram zone, selects the zone with the highest score, and outputs the corresponding candidate replacement image. The histogram weighting may increase the influence of oversaturated and undersaturated luma values. In this embodiment, candidate replacement images of relatively low exposure values ​​may be selected to replace patches of reference images of relatively high exposure values, and vice versa.The reference image may be a medium-exposure image with downgraded luma values. The weighting function for a particular histogram zone may be defined as the ratio of the number of pixels in the particular histogram zone, taking into account the size of the particular histogram zone, to the mean difference of the histogram inputs from the luma value mode for the particular histogram zone.

[0039] Additionally, the invention may be implemented as a system for replacing patch image data in a composite image, comprising a processor that smooths a patch image boundary polygon, upscales the smoothed patch images to full image resolution, and merges the composite image and the upscaled smoothed patch image to generate an output image. The patch image boundary polygon may be an octagon. The patch image boundary octagon may be represented in 8-bit format, and a low-pass filter, for example, 7x7, is applied.

[0040] Another embodiment is a system for converting images from an original format to a format with a lower number of bits per pixel, comprising a processor that performs tone mapping processing that adapts to image luma values ​​and outputs the tone-mapped image. In the system, the processor may also perform non-linear local tone mapping that maps a pixel according to the average luma values ​​of its neighbors into individual color components, each represented by a predetermined number of bits. In the system, the processor may further allocate additional grayscale levels for more common luma values. The system may convert images into a display format that has pixels with the number of bits that can be presented on a built-in digital camera image display.The system may also include the processor defining a family of local mapping lookup tables and a single global mapping lookup table, and using a selected lookup table for each image transformation operation, wherein the lookup table selection is based on the characteristics of the image scene. The system may further include the processor constructing a brightness histogram of a composite image and defining a mapping according to the distribution of brightness values. The brightness histogram may be estimated using a combination of histograms of the downgraded images used to form the composite image. Histogram equalization may define the mapping. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are not intended to be to scale. In the drawings, each identical or nearly identical component shown in different figures is represented by the same reference numeral. For the sake of clarity, not every component may be identified in each drawing. In the drawings: Fig. 1 a block diagram of a digital camera or other image capture device that captures a plurality of digital images of a scene at different exposure levels and at different times and displays those images on the image display device built into the camera; Fig. 2 a high-level block diagram of processing modules as implemented in a digital camera; Fig. 3 is a block diagram of the 2-image merging engine according to embodiments of the invention; Fig. 4 is a flowchart illustrating the complete image mixing process sequence of an image mixer processing method according to embodiments of the invention; Fig. 4A details the image generated by the 2-image merging engine of the Fig. 3 used process according to embodiments of the invention; Fig. 5 is a block diagram of the ghost removal processing module according to embodiments of the invention; Fig. 6 is a flowchart showing the processing sequence of a ghost removal processing method; Fig. 7 is a photograph showing ghost images created by using three photographs taken consecutively and stitched together in accordance with embodiments of the invention; Fig. 8 the photo of the Fig. 7 after removing the ghost images according to embodiments of the invention; Fig. 9 shows a photograph requiring HDR processing according to embodiments of the invention; Fig. 10 a photo with markings corresponding to the photo of the Fig. 9 according to embodiments of the invention; Fig. 11 a photo with a patch selection corresponding to the photo of the Fig. 9 according to embodiments of the invention; Fig. 12 is an illustration of a bounding box for pixel patches according to embodiments of the invention; Fig. 13 is an illustration of a boundary diamond for pixel patches according to embodiments of the invention; Fig. 14 is a representation of a bounding octagon for pixel patches according to embodiments of the invention; Fig. 15 is an illustration of the HDR process according to embodiments of the invention; Fig. 16 is an illustration of the algorithmic HDR flow according to embodiments of the invention; Fig. 17 is an illustration of the algorithmic image fusion process according to embodiments of the invention; Fig. 18 illustrates an example of a piecewise linear function for a weight lookup table according to embodiments of the invention; Fig. 19 is an illustration of the ghost removal process according to embodiments of the invention; and Fig. 20 is a diagram illustrating an exemplary division of the mean image histogram into three different zones according to embodiments of the invention. DESCRIPTION OF THE EMBODIMENTS

[0042] The invention will now be described in more detail hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific embodiments in which the invention may be practiced. The invention may, however, be embodied in many different forms and is not to be construed as limited to the embodiments set forth herein; on the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Among other things, the invention may be embodied as methods or devices.Accordingly, the invention may take the form of an all-hardware embodiment in the form of modules or circuits, and an all-software embodiment in the form of software executing on a general-purpose microprocessor, an application-specific microprocessor, a general-purpose digital signal processor, or an application-specific digital signal processor, or an embodiment combining software and hardware aspects. Thus, in the following description, the terms "circuit" and "module" are used interchangeably to refer to a processing element that performs an operation on an input signal and provides an output signal therefrom, regardless of the hardware or software form of its implementation.Likewise, the terms "register," "registration," "align," and "alignment" are used interchangeably to refer to the process of causing similar objects to correspond to one another and to align them properly, regardless of whether the mechanism used to achieve such correspondence is implemented in hardware or software. The following detailed description is therefore not intended to be limiting.

[0043] Throughout the specification and claims, the following terms have the meanings explicitly assigned to them unless the context clearly indicates otherwise. The phrase "in one embodiment," as used herein, does not necessarily refer to the same embodiment, although it may. As used herein, the term "or" is an inclusive "or" operator and is equivalent to the term "and / or" unless the context clearly indicates otherwise. The term "based on" is non-exclusive and permits basing on additional, undescribed factors unless the context clearly indicates otherwise. Additionally, the meanings of "a," "a," "and," and "the" include plural referents throughout the specification. The meaning of "in" includes "in" and "on."Also, the use of “comprise,” “have,” “have,” “include,” and variations thereof is intended to include the items listed below and equivalents thereof, as well as additional items.

[0044] Fig. 1 shows a digital camera or other image capture device including an optical imaging system 105, an electronically controlled shutter 180, an electronically controlled lens aperture 185, an optical image sensor 110, an analog amplifier 115, an analog-to-digital converter 120, an image data signal processor 125, an image data storage unit 130, an image display device 106, and a camera controller 165. The image data storage unit could be a memory card or an internal non-volatile memory. Data from images captured by a camera can be stored on the image data storage unit 130. In this embodiment, it may also include an internal volatile memory for temporary image data storage and intermediate image processing results.This volatile memory may be distributed among the individual image data processing circuits and need not be architecturally arranged in a single image data storage unit such as image data storage unit 130. The optical system 105 may be a single lens, as shown, but is typically a set of lenses. An image 190 of a scene 100 is formed in visible optical radiation on a two-dimensional surface of an image sensor 110. An electrical output 195 of the sensor carries an analog signal resulting from the scanning of individual photodetectors of the surface of the sensor 110 onto which the image 190 is projected. Signals proportional to the light intensity incident on the individual photodetectors are obtained at output 195. The analog signal 195 is applied via an amplifier 115 to an analog-to-digital converter 120 by means of the amplifier output 102.Analog-to-digital converter 120 generates an image data signal from the analog signal at its input and applies it to image data signal processor 125 via output 155. The photodetectors of sensor 110 typically detect the intensity of light incident on each photodetector element in one of two or more individual color components. Earlier detectors detected only two separate colors of the image. Detection of components of three primary colors, such as red, green, and blue (RGB), is now common. Image sensors are available today that detect more than three color components.

[0045] Several processing operations are performed on the image data signal from the analog-to-digital converter 120 by the image data signal processor 125. The processing of the image data signal in this embodiment is shown in Fig. 1 as being performed by multiple image data signal processing circuits in image data signal processor 125. However, these circuits may be implemented by a single integrated circuit image data signal processor chip, which may include a general-purpose processor executing algorithmic operations defined by stored firmware, multiple general-purpose processors executing algorithmic operations defined by stored firmware, or dedicated processing logic circuits as shown. Additionally, these operations may be implemented by multiple interconnected integrated circuit chips, but a single chip is preferred. Fig. Figure 1 illustrates the use of series-connected image data signal processing circuits 135, 140, 145, and 150 to perform many algorithmic processing operations on the image data signal from the analog-to-digital converter 120. The result of these operations is stored non-volatile digital image data that can be displayed either on the internal image display device 106 of the digital camera in Fig. 1 or an external display device. This viewing can be accomplished either by physically removing a memory card from the digital camera and reinserting it into an external display device, or by electronically communicating the digital camera with an external display device using a Universal Serial Bus (USB) connection or a wireless local Wi-Fi or Bluetooth network.

[0046] Additional processing circuitry, such as that shown by points 175 between circuits 145 and 150, may be included in the digital camera's image data signal processor. The serial structure of the image data signal processor of the embodiment is known as a "pipeline" architecture. This architectural configuration is used as the embodiment of the invention, but other architectures may be used. For example, an image data signal processor having a "parallel architecture" may be used, in which one or more image data signal processing circuits are arranged to receive processed image data signals from a plurality of image data signal processing circuits, rather than after they have been processed serially by all previous image data signal processing circuits. A combination of a partly parallel and partly pipeline architecture is also possible.

[0047] The series of image data signal processing circuits of the image data processor 125 is called an “image processing pipeline.” Embodiments add image data signal processing circuits that are Fig. 2, to those routinely included in the image processing pipeline of a digital camera. Image data signal processing circuits routinely included in the image processing pipeline of a digital camera include circuits for white balance correction (WBC), lens shading correction (LSC), gamma correction (GC), color transformation (CTM), dynamic range compression (DRC), demosaicing, noise reduction (NR), edge enhancement (EE), scaling, and lens distortion correction (LDC). As shown in Fig. 2, embodiments add an image registration processor (IRP) circuit 210, an image mixer (IM) circuit 220, a ghost remover (GR) circuit 230, and a tone mapping processor (TMP) circuit 235 to the complement of image data signal processing circuits discussed above. The image memory 200 of the Fig. 2 stores the digital data of a series of two or more images of a scene, each series image consisting of pixels comprising digital bits, which digital bits have been processed by the image data signal processing circuits in the image processing pipeline. The image memory 200 could share the memory with the image memory 130 of the Fig. 1, memory resources used for temporary image data storage and intermediate image processing results, or entirely separate volatile or non-volatile memory resources could provide the memory resources used by image memory 200.

[0048] With reference to Fig. 1, the camera controller 165, via line 145 and control / status lines 160, causes the electronic shutter 180, the electronic aperture 185, the image sensor 110, the analog amplifier 115, and the analog-to-digital converter 120 to capture a series of images of a scene and convert them into digital image data. These images are captured at various exposure levels, processed by image data signal processing circuits, and stored in the image memory 200 of the Fig. 2. The image registration processor 210 reads the digital image data of the image series stored in the image memory 200 and registers counterpart pixels of each image of the image series with each other, with one image (generally the image with the middle exposure time) serving as a reference image. Image registration is performed before image blending in order to align all the series images pixel by pixel. Due to camera movement that occurs during the capture of an image series, such alignment is of the utmost importance to the image blender 220. Fig. 2 is necessary so that it can correctly combine burst pixels and form an image with the full range of gray tones of the captured scene. Such an image is often referred to as a "high dynamic range" or "HDR" image. In the image blender 220, each burst pixel of each captured burst image is combined with its counterpart pixel in each captured burst image. Thus, an image pixel representing a particular position at the edge of or within the body of an object appearing in a first burst image is blended with its counterpart located at the same edge or within the body of the same object appearing in a second burst image. In this context, the location of a pixel in an image refers to the object of which it is a part, not to the fixed coordinate system defined by the vertical and horizontal outer edges of the image.

[0049] The image registration processor 210 generally uses a first image, captured at a nominal camera exposure setting, as a reference image to which all other images in the series are aligned. A number of techniques are currently used for image alignment and registration. A good example is described in "High Dynamic Range Video" by S.B. Kang, M. Uyttendaele, S. Winder, and R. Szeliski, Interactive Visual Media Group, Microsoft Research, Redmond, WA, 2003.

[0050] The described procedure handles both camera motion and object motion in a scene. For each pixel, a motion vector is calculated between consecutive burst images. This motion vector is then refined using additional techniques, such as hierarchical homography, to handle degeneracy cases. Once the motion of each pixel is determined, frames can be warped and registered to the selected reference image. The images can then be blended into an HDR image by the image blender 220.

[0051] It should be noted that, for practical purposes, in the present application, motion vectors are derived from NxM blocks instead of motion vectors for each pixel. Additionally, given the motion vectors, a general transformation [x' y'] = f(x,y) is derived, where x' and y' are the new locations of a given point {x,y}. The image mixing process

[0052] The image mixer 220 has the ability to mix an unlimited number of images, but uses an image blending engine that blends the pixels of two images at once. The 2-image blending engine of the mixer 220 is Fig. 3. In one embodiment of the invention, the merging engine merges the pixels of a first 8-bit image, whose digital image data appears on input line 300, and the pixels of a second 8-bit image, whose digital image data appears on input line 305. Images with bit widths greater than 8 bits, for example, 10 bits, or narrower, for example, 7 bits, can be used. The flowchart of the Fig. 4 illustrates the complete image mixing process of the image mixer 220, and Fig. Figure 4A shows in detail block 420 of the Fig. 4, which is generated by the 2-image fusion machine of the Fig. 3 process used.

[0053] With reference to Fig. 3, the image blending process blends two images during each blending operation. The first two images to be blended are both taken from the acquired image series, with each image of the acquired image series previously registered to a reference image taken at a nominal camera exposure setting. For the initial blending operation, the series image with a lower exposure level provides the first-image digital image data input, which is available on line 300 of the Fig. 3 appears, and the series image with a higher exposure level provides the digital second-image image data input, which appears on line 305. For a second and all subsequent image blending operations, a subsequent image of the series is blended with the result obtained from a previous image blending operation. For these subsequent blending operations, the digital image data of a subsequent image of the series serves as the digital second-image image data input, which appears on 305, and the digital blended image data result serves as the digital first-image image data input, which appears on line 300 of the Fig. 3 appears. In all cases, the following image in the series was exposed at a higher exposure level than its direct predecessor in the series.

[0054] The digital second-image data on line 305 is initially processed in two ways. (1) The luma conversion circuit 320 extracts the luma, the black and white component of the combined red, green, and blue (RGB) component data comprising the digital second-image data 305, and outputs the luma component of each image data pixel on line 325. (2) The image normalizer 310 normalizes the exposure level of each RGB component of the second-image data on line 305 to the exposure level of a reference image and outputs, on line 302, each image data pixel, for each color component, normalized to the reference image exposure level. It should be noted that the reference image used is not necessarily the same reference image used for the previously described registration process. For one embodiment of the invention, the exposure level of the darkest image in the series, i.e.of the least exposed image as the reference exposure level, and all other images in the series are normalized to it. For example, if the acquired image series consists of three images: a dark image exposed for 1 / 64 s, a middle image exposed for 1 / 16 s, and a light image exposed for 1 / 2 s, the normalized value of each pixel of the middle image appearing on line 302 would be: AveragePixelValueNormalized=AveragePixelValueInput / ((1 / 16) / (1 / 64))=AveragePixelValueInput / 4; and the normalized value of each pixel of the bright image appearing on line 302 would be: BrightPixelValueNormalized=BrightPixelValueInput / ((1 / 2) / (1 / 64))=BrightPixelValueInput / 32

[0055] Therefore, for this embodiment of the invention: Exposure levelNormalized=Exposure levelContinuous shot / Exposure levelLeast exposed continuous shot and the normalized value of each pixel of the second image image data input on line 305 and output on line 302 is: SecondImagePixelValueNormalized=SecondImagePixelValueInput / SecondImageExposureLevelNormalized

[0056] The luma component of each digital second image data pixel present on line 325 of the Fig. 3 is input to the lookup table (LUT) 315 to obtain a per-pixel weighting parameter, Wi, on lines 330 and 335. The luma component of each digital second-image pixel serves as an index into the LUT 315 and causes a weighting parameter value Wi between the numbers zero and one to be output on lines 330 and 335 for each input luma value. This value is output in the form of a two-dimensional matrix in which: W(m,n)=255-Luma(m,n);

[0057] Luma(m, n) is the luma component of each second-image digital data pixel at image coordinates (m, n), which for this embodiment of the invention can reach a maximum value of 255 because the embodiment merges pixels from 8-bit burst images, and 255 = one, which represents the 100% output value of the table 315. Second-image pixel values, which serve as indices into the LUT 315, are 8-bit digital values ​​with a number range of 0 to 255. Therefore, defining 255 as one enables a direct mapping from the input index value to the output weighting parameter value and reduces the weighting parameter application computation workload. Other values ​​of one can be chosen. For example, if the second-image pixel values ​​of the luma component, which serve as indices into the LUT 315, are 10-bit digital values, with a number range of 0 to 1023, it would be appropriate and convenient to assign the value of 1023 to one.

[0058] For this embodiment, the weighting parameter value output from LUT 315 decreases linearly as the second-image pixel value of the luma component, which serves as an index into LUT 315, increases. Other LUT functions, such as trapezoidal functions in which the weighting parameter value obtained from LUT 315 remains at a predetermined value and begins to decrease linearly when the second-image pixel value index of the luma component decreases below a threshold, may also be used.The choice of LUT 315 features is based on the observation that when blending two images, one highly saturated due to exposure at a high exposure level, perhaps with a long exposure time, and the other dark due to exposure at a lower exposure level, perhaps with a short exposure time, it is desirable to apply a low weight to the highly saturated pixels of the high-exposure image while applying a high weight to the counterpart pixels of the low-exposure image. This results in a blended image with fewer highly saturated pixels, as many may have been replaced by correctly exposed counterpart pixels. The result is a blended image with greater detail in its highlight areas, while pixels belonging to the shadow zone are taken primarily from the higher-exposure image.

[0059] The embodiments are not limited to the use of a single LUT 315. A plurality of LUTs may be used. In this case, a different LUT may be assigned to each burst image to obtain the weighting value, or two or more burst images may be assigned the same LUT from a plurality of supplied LUTs. These LUTs may, for example, be populated with weight parameter values ​​that respond to burst exposure levels.

[0060] The present application provides an improvement in certain embodiments not available in the parent application. Image sensors are typically of at least 12-bit resolution, while current display devices only have 8-bit resolution, so a non-linear gamma operation is applied to original pixels during capture to limit the range to 8 bits for the display device. Therefore, in certain embodiments provided in the present application, merging is preferably performed on a "gamma-corrected version" of two input images Im1 and Im2, where the gamma correction operation is a LUT that moves the RGB image pixels back from 8 bits (i.e., the non-linear range) to 16 bits (i.e., the linear range). Therefore, Merged image = Gamma correction (ImI) × (I−W) + Gamma correction (Im2) × W

[0061] Since the exposure takes place in the linear range, the normalization to the exposure time should also take place in this range.

[0062] Preferred embodiments of the present invention may be implemented by minor modifications to the embodiments described in the parent application, as will now be described. With reference to the implementation of the parent application, which may be described, for example, as Fig. 3 of this application, gamma correction function blocks 301 and 302 are now added beforehand to the image normalizer, as shown in Fig. 3 of the present application. These blocks can be selectively activated by flags, e.g., DG_FLAG and DG_FLAG2. Also, during the first execution of the iterative process described above, the first image data input (300) is converted into the gamma correction domain (i.e., linear). In the next iterations using a 2-image fusion engine, it is not necessary to apply a gamma correction to the first image if the fused output from a previous stage is used and is already in the gamma correction domain. Likewise, the implementation of the Fig. 4A of the parent application, according to preferred embodiments of the present invention, by adding an additional pixel gamma correction function to the normalized pixel (block 429). This block can also be selectively enabled / disabled by an external flag. Likewise, when embodiments of the present invention apply images to the inputs 500 and 505 of the Fig. 5, these images can also be processed to be in the gamma correction range if necessary.

[0063] During the merging operation, the weighting parameter is applied, on a pixel-by-pixel basis, to each color component of the normalized digital second image data appearing on line 302, and 1 minus the weighting parameter, e.g., (1 - Wi), is applied to each color component of the digital first image data appearing on line 300. The pixel-by-pixel merging operation is defined by the following equation: Merged image data pixel = (I−Wi)×(1st image data pixel)+Wi×(Normalized 2nd image data pixel)

[0064] The processing blocks of the Fig. 3, equation (7) is expressed as follows: The luma of the second-frame data on line 305 is derived by the luma conversion circuit 320 and used by the LUT 315 to generate the weighting parameter Wi on lines 330 and 335. The multiplier 307 multiplies the normalized second-frame digital image data, normalized by the image normalizer 310, by Wi on line 330 and outputs the result on line 355. Wi is also applied to the data subtractor 340 through line 335, which outputs (1 - Wi) on line 345. The first-frame digital image data on line 300 is multiplied by (1 - Wi) on line 345 by the multiplier 350, and the result is output on line 365. The normalized and weighted second-picture digital image data on line 355 is added to the weighted first-picture digital image data on line 365 by adder 360.Adder 360 outputs merged image pixels on line 370. These pixels are stored in frame buffer 375 and output as merged 2-frame image data on line 380.

[0065] The image generated by the 2-image merging machine of the Fig. 3 used pixel merging process is in processing block 420 of the Fig. 4A. The data of the first image to be merged enters the process at 423, and the data of the second image to be merged enters the process at 413. The pixel merging process begins at 445. At 427, a pixel from the second image is selected. The selected second image pixel is normalized at 429, and its luma is derived at 450. The luma of the second image pixel is used to obtain the weighting parameter Wi, 485, from a LUT at 465. The normalized second image pixel is multiplied by the weighting parameter Wi, 485, at 455. The normalized and weighted second-image pixel enters an addition process in 475. The first-image counterpart pixel of the selected second-image pixel is selected in 460 and multiplied by (1 - Wi) in 470. The weighted first-image pixel enters the addition process in 475 and is merged with the normalized and weighted second-image pixel.If more image pixels remain to be merged, as determined at decision point 480 and indicated by "No" at 431, the next pixel merging cycle is initiated at 445, resulting in a next first image pixel and a next second image pixel being selected at 460 and 427, respectively, and merged as described.

[0066] If there are no more image pixels remaining to be merged, as determined at decision point 480 and indicated by "Yes" at 433, but there are more images to be blended in the acquired image series, as determined at decision point 495 and indicated by "No" at 487, another second image is selected from the remaining, unblended series images, the data from which serves as the next second image data to be blended. In the implementations of the parent application, the selected second image has a higher exposure level than the previous second image selection, but in the embodiments of the present application, such constraints do not apply; the images may be processed in any order, regardless of the relative exposure level. Such a selection is made by the image selection process 425 in response to the "No" at 487.Additionally, the blended 2-image output image data at 493 is selected by the first image data selection process 435 as the next first image data to be blended because decision point 430 indicates "yes" to process 435 at 441 in response to blended 2-image image data being available at 493. If blended 2-image image data is not available at 493, as would be the case at the beginning of an image blending process, decision point 430 would signal to processing block 440, by placing a "no" at 437, that it should select a first image from the captured image series to be blended with a lower exposure level than the second image to be blended (in the embodiments of the parent application) selected from the captured image series by processing block 425.In this case, information about the selected second image exposure level is transmitted to the selection block 440 in 407, information about a lowest exposed continuous shooting exposure level is transmitted to the selection block 440 in 443, and selected continuous shooting data is transmitted to the processing block 440 in 439. As shown in the flowchart of FIG. Fig. 4 and used in the embodiments of the invention of the present application, the continuous shot with the lowest exposure level may not necessarily be selected as the first image. Furthermore, block 440 of the parent application is no longer necessary.

[0067] When there are no more images to be blended in the captured image series, the process ends with the blended HDR image output appearing in 497.

[0068] As described above regarding the implementation of the root application, which is Fig. As shown in Figure 4A, a flag can be added at image output point 497 to enable / disable a gamma operation for preferred embodiments of the present invention. For all burst images, this gamma operation is OFF except for the last one, when the final output is provided.

[0069] The flow chart of the Fig. 4 of the present invention illustrates the complete image blending process of the image blender 220. Block 420, which represents the image blending process performed by the 2-image blending engine of the Fig. 3 represents the 2-image fusion process used in Fig. 4 and in detail in Fig. 4A illustrates. Fig. 4 also includes the processing that precedes the 2-image merging process block 420. This processing includes capturing a series of images, each image exposed to a different exposure level, in processing block 400; registering duplicate burst image pixels to one another in processing block 405; determining the least exposed burst image (in the parent application) in processing block 410; and calculating a normalized exposure level according to equation (3) described above, which may be related to the least exposed image of the series or to another image in this improved embodiment, for each image in the series, in processing block 415.This calculated normalized exposure level is used by 429 of processing block 420 to normalize each second-image pixel as previously described by equation (4) before it is multiplied by a weighting parameter and merged with a weighted first-image pixel according to equation (7) previously described.

[0070] The image blending process of image blender 220 uses only summation and multiplication, thus avoiding computationally intensive division operations and allowing it to be implemented by a fixed-point arithmetic engine. Additionally, because the blending approach used is based on consecutive two-image blending operations, there is no need to wait until all burst images have been captured before starting a blending operation. The blending process can begin immediately after just two images in the burst have been captured. These properties allow the blending process to quickly generate an HDR image from a burst of images using a arithmetic engine with limited processing power. The ghost removal process

[0071] In the parent application, the ghost remover removes 230 displaced renditions of scene objects, or ghost images, that are present in the mixed HDR image output data in 497 of the Fig. 4 due to the movement of these objects during the time the images of the series are captured. Essentially, ghost images appear due to the blending of series images in which an object shown in a first-series image has moved relative to its first-series image location coordinates, as shown in a second-series image. Consequently, in the blended image, the object may appear in many locations, where the locations depend on the object's speed and direction of movement. Embodiments of the invention described in the parent application use a single, two-phase approach to mitigating ghost images.

[0072] Given a blended image, HDR(i, j), generated from a weighted sum of images from an acquired aligned image series of K images, the embodiment, in a first processing stage, first calculates the variance of the luma value of each pixel of HDR(i, j) as follows: V(i,j)=EfcW(i,j,k)×P2(i,j,k)−HDR(i,j)2 where: V(i, j) = The variance of the luma value of the blended HDR image pixel, HDR(i, j), located at image coordinates (i, j) with respect to the value of the Kth burst image pixel located at image coordinates (i, j), over the K aligned images of the acquired image series; HDR(i, j) = The luma value of the blended HDR image pixel located at image coordinates (i, j); W(i, j, k) = A normalizing weight applied to the level of the kth burst pixel located at image coordinates (i, j) to normalize the burst pixel level range to the pixel level range of the blended HDR image; and P(i, j, k) = The value of the Kth serial image pixel located at the image coordinates (i, j); and then replaces a blended HDR image pixel whose luma variance exceeds a first threshold by its duplicate pixel from an aligned reference image, HDR ref , where the reference image is selected from the acquired image series to obtain first processed mixed image data, HDR 1.verarbeitete , to generate.

[0073] The first phase of the ghost removal processing described above is based on the observation that, if there is no local motion in the blended series images, the variance of a pixel in the blended HDR image output across the K aligned images of the acquired image series is low, as defined by equation (8) above. The only significant error associated with this assumption is the alignment error. Since alignment is inherently a global process, it cannot compensate for local motion of a local image object, and thus, the motion of a local object manifests itself as high-amplitude pixel variance zones.By analyzing the high-amplitude pixel variance zones in the two-dimensional variance data generated by equation (8), zones of local motion in the blended HDR image output data can be defined, and blended HDR image output data pixels with variances above a predefined threshold can be replaced with duplicate pixels from the reference image. The selected reference image is often the least exposed burst image, but can be a burst image exposed at a higher exposure level.

[0074] The first phase of the ghost removal processing of the embodiment generates first processed mixed image data, HDR 1.verarbeitete , with less ghosting. However, some residual ghosting remains. A second processing phase improves these results by comparing the content of HDR 1.verarbeitete with HDR refIn this second phase, residual ghosting is detected by analyzing the pixel-to-pixel result obtained by subtracting the luma from HDR ref from the luma of HDR 1.verarbeitete Another threshold is applied based on the maximum value of the differences between the luma of HDR 1.verarbeitete and the luma of HDR ref Each HDR 1.verarbeitete -Data pixel that exceeds the second threshold is replaced by its counterpart HDR ref -Data pixels replaced, resulting in second processed mixed image data, HDR 2.verarbeitete , with fewer ghost images. The procedure used by this embodiment of the second processing phase of the invention can be summarized as follows: (a) Create D0 = ABS(Luma(HDR 1.verarbeitete ) - Luma(HDR ref )); (b) Determine threshold2. = Max(D0) = DM0. (c) Replace each HDR 1.verarbeitete-Data pixel that exceeds the threshold2. by its counterpart-HDR ref -Data pixels, resulting in HDR 2.verarbeitete mixed image data (d) Compare HDR 2.verarbeitete with HDR ref and create Max(DI) = Dm1, where Dm1 = Max((ABS(Luma(HDR 2.verarbeitete ) - Luma(HDR ref ))) (e) If Dm1 > 60% of DM0, Threshold2 is too large and HDR 2.verarbeitete can be too much like HDR ref look like this. Then: (f) Segment DM0 into 2 levels, where DM00 = a value < 0.5DM0, and DM01 = a value > 0.5DM0 (g) Determine, in percent, the amount of HDR 2.verarbeitete Image field with respect to the full image field, which exceeds DM01, SIZE_1 (h) Determine, in percent, the amount of HDR 1.verarbeitete Image field with respect to the full image field that exceeds DM00, SIZE_0, where SIZE_1 should be >= SIZE_0. (i) Calculated SIZE RATIO = SIZE_1 / SIZE_=0 (j) If (SIZE_0 > 40% || (SIZE_RATIO > 2 && SIZE_1 > 8 %)) (k) Take burst shots again (l) Otherwise (m) Replace pixels from HDR 2.verarbeitete Image fields that exceed DM01 by their counterpart HDR ref -Pixels, resulting in HDR 2.verarbeitete mixed image data (n) end

[0075] The ghost removal process can be applied to all aligned, captured, bracketed series of two or more images. In addition, two or more HDR ref-Images are used by the process. For example, if the process is applied to a series of three images, where the exposure level of a first series image is lower than the exposure level of a second series image and the exposure level of the second series image is lower than the exposure level of a third series image, the fields of the zones of the second series image corresponding to mixed image data with variances that exceed the first threshold can be used to perform HDR between two reference images. ref1 and HDR ref2 In this example, second burst zone fields that are saturated are summed, and a ratio of the sum of saturated fields to the total field of the second burst image is used to achieve HDR ref for the rest of the ghost removal processing. If the ratio is less than or equal to a parameter value, for example 0.03 to 1, then HDR ref2= second burst image selected. If the ratio is greater than the parameter value, HDR ref1 = first series image selected. Furthermore, the above selection procedure, or others of a similar nature that respond to other image features, such as the size of image fields with object movement above a predetermined threshold, or spatial frequency details above a predetermined threshold, can be used to create an HDR ref1 for the first phase of ghost removal processing and additionally another HDR ref2 for the second phase of ghost removal processing.

[0076] Fig. 5 is a block diagram of one embodiment of the ghost removal processing module 230 of Fig. 2. HDR composite image pixel data is input to the luma conversion circuit 515 and the first pixel replacement circuit 540 on line 505 of the Fig. 5. The luma conversion circuit 515 converts HDR composite image pixel data into composite image luma pixel data, and inputs composite image luma pixel data to the variance calculation circuit 550 via line 525. Although not in Fig. 2, aligned images of the captured bracketed series of two or more images are input to the ghost removal module 230 on line 500 of the Fig. 5, which is connected to the reference image selection circuit 510. In this embodiment, the reference image selection circuit 510 selects the least exposed continuous image as the reference image HDR ref off, but a series image exposed with a higher exposure level could be selected. Line 520 transmits HDR ref pixel data is also applied to the variance calculation circuit 550. In addition, line 520 applies HDR refpixel data to the 1st pixel replacement circuit 540, 2nd pixel replacement circuit 585 and luma conversion circuit 560. From the HDR ref -pixel data on line 520 and the blended image luma pixel data on line 525, the variance calculation circuit 550 generates output blended image luma pixel variance data on line 530. This luma pixel variance data is applied to the first pixel replacement circuit 540 via line 530. On line 535, a first threshold is also applied to the first pixel replacement circuit 540. From these inputs, the first pixel replacement circuit 540 replaces pixels of the blended image pixel data on line 505 whose luma variance exceeds the first threshold on line 535 with counterpart pixels from the HDR ref -Data on line 520 to provide first processed mixed image pixel data, HDR 1.verarbeitete , on line 545, which are the output of a first processing phase.

[0077] The output of the first processing phase, HDR 1.verarbeitete , on line 545, is converted to HDR by the luma conversion circuit 565 1.verarbeitete Luma pixel data is converted. The output of circuit 565 appears on line 595 and is connected to comparison circuit 575. HDR 1.verarbeitete on line 545 is also applied to the 2nd pixel replacement circuit 585. Line 520 applies HDR ref pixel data to the luma conversion circuit 560 and the 2nd pixel replacement circuit 585. The pixel data to luma conversion circuit 560 converts HDR ref -Pixel data in HDR ref -Luma pixel data and delivers the HDR ref -Luma pixel data to the comparison circuit 575 via line 570. The comparison circuit 575 calculates the difference between each HDR 1.verarbeitete -Luma data pixels and its counterpart-HDR ref-Luma data pixels and generates a second threshold based on the maximum value of the differences. This second threshold is applied to the second pixel replacement circuit 585 via line 580. The second pixel replacement circuit 585 replaces each HDR 1.verarbeitete -Data pixel on line 545 that exceeds the 2nd threshold by its counterpart HDR ref -Data pixels on line 520, with the resulting 2nd processed mixed image data, HDR 2.verarbeitete , on line 590, are the ghost-reduced output of a second processing phase.

[0078] The ghost removal processing module of the Fig. 5 used two-phase ghost removal process is shown in the flowchart of Fig. 6. In block 600, a bracketed series of two or more images is acquired, each image exposed at a different exposure level and at a different time. In block 605, these images are registered to each other such that counterpart burst pixels correspond. In block 610, a reference image is selected from the acquired scene images, and its pixel data is passed to processing blocks 625, 640, 660, and 695 via processing path 615. The selected reference image is often the least exposed burst image, but may be a burst image exposed at a higher exposure level. It need not be the same reference image used by image registration processor 210, which generally uses a first image acquired at a nominal camera exposure setting as a reference image to which all other images in the series are aligned.The image mixer described above, whose image mixing process is shown in the flow chart of the . Fig. 4, performs the processing in block 620, wherein the mixed HDR image output image data pixels of the Fig. 4 enter processing blocks 650 and 640. In block 650, the luma of the blended data pixels is generated and passed to block 625, where the variance of each blended image data pixel luma is calculated by block 650, compared to its counterpart reference image data pixel from block 610. This variance is passed to processing block 640 and used by block 640, along with a first threshold entering processing block 640 along path 635, blended image data pixels from processing block 620 entering processing block 640 along path 650, and reference image data pixels entering processing block 640 along path 615, to replace all blended image data pixels with a variance exceeding the first threshold with their counterpart reference image data pixels, generating first-processed blended image data pixels. First-processed blended image data pixels are the result of a first phase of ghost removal processing.

[0079] 1st processed blended image data pixels are passed to processing blocks 665 and 695 along processing path 655. Processing block 665 generates the luma of the 1st processed blended image data pixels, while processing block 660 generates the luma of each reference image data pixel from reference image data pixels entering processing block 660 via path 615. Processing block 670 calculates the difference, on a pixel-by-pixel basis, between the luma value of each 1st processed blended image data pixel and the luma value of its duplicate reference image data pixel and provides these differences to processing block 675. Processing block 675 determines the maximum value of these differences, and processing block 680 generates a 2nd threshold based on this maximum value. A processing block 695 receives this 2nd threshold via path 685, along with the reference image data pixels via path 615 and 1.processed blended image data pixels via path 655 and replaces each 1st processed blended image data pixel that exceeds this 2nd threshold with its corresponding reference image data pixel counterpart, thus generating enhanced deghosted 2nd processed blended image data on processing path 695. This 2nd processed blended image data, the result of a 2nd phase of deghosting processing, is provided as input to a tone mapping processor, such as 235 of the . Fig. 2, used.

[0080] The embodiments of the present invention differ from the embodiments of the parent application described above. The blending scheme presented in the parent application was intended to work best in a static scene, but when this is not the case and there is some local motion, such as a car, people, or tree branches moving due to a blowing wind, ghosting may be noticed, as mentioned above. A separate ghosting removal process is then required. The embodiments of the present invention include a ghosting removal process that is improved over the previous ghosting removal process, as described below.

[0081] Fig. Figure 7 shows a non-limiting example of the ghost image generated by blending three images taken at different times and with different exposures, according to an embodiment of the present invention. The ghost images, indicated by arrows 710, 720, and 730, are clearly visible and should be eliminated as much as possible. Complete elimination of the ghost images in all cases may be impossible, but the embodiments to be described eliminate the ghost images in many cases, and the remaining residue is better than that presented by other prior art solutions.

[0082] The parent application's ghost removal is based on the assumption that, if there is no local motion, the difference image generated from two aligned images will have low amplitudes. Unfortunately, local motion is not compensated for by the registration process (which is global in nature) and therefore actually appears as a high-amplitude zone in the difference image.

[0083] Therefore, it is noted in the present application that by walking through the difference image, one can locate these high-amplitude zones and replace the regular output with an original patch taken from a reference image, preferably either the dark image or the center image. This replacement is performed using the image mixer again, with the inputs being the HDR (with ghosting) and the reference image, from which patches with a correct weighting W are taken, calculated as follows: HDRout=(I−W)×HDR+W×ref_image

[0084] The resulting HDRout image produced by this process should now be ghost-free.

[0085] Now the main steps provided for generating the weight W are given: (a) Downscaling the lumas of the (typically, but not necessarily, three) input images to, for example, a 180x240 resolution. (b) Normalizing the lumas to a specific standard deviation and mean. (c) Create d = abs(dark_luma - medium_luma) > threshold difference image (d) Create d1 = abs(bright_Luma - medium_Luma) > threshold difference image (e) Generate a binary image G = (d OR d1) (f) Dilating, filtering, and upscaling G to produce a smoothed version of W for use by the image mixer.

[0086] In Fig. 8, a non-limiting example of the resulting ghost-free mixing according to one embodiment is provided. Arrows 810, 820, and 830 point to the respective fields previously pointed to by arrows 710, 720, and 730, respectively, in Fig. 7. The ghost images of a moving person and moving vehicles were essentially completely removed. It should be noted that this procedure also eliminates registration noise. For example, if a door edge is not correctly aligned, the difference between aligned images would be high. Consequently, the edge is replaced by a patch taken from the reference image. Therefore, it makes no difference in the ghost removal process of the present application where the noise comes from, e.g., either from local motion or from misregistration. In both cases, it is replaced by a patch. Selection of the reference image for patch implantations

[0087] In general, it is advantageous to take the replacement patches from a higher-exposure image (typically the center image in a series of three images), but these patches may consist of saturated zones. If this is the case, a patch can be taken from the darkest image. The criterion for selecting patches from the center image is based on calculating the overall percentage of saturated zones given the total field of patches. If this is less than, for example, 3%, the center image serves as the reference image from which patches are taken; otherwise, the dark image is taken. The disadvantage of using patches from the dark image is that the dark image has the lowest signal-to-noise ratio and is noisy. This noise is increased by the final tone mapping.

[0088] The decision to select the patch (either from the dark image, the light image, or the middle image in the typical case where three input images are available) is made per image segment and not on the entire image. This requires that the connected binary image G is first segmented. For this purpose, an exemplary and non-limiting image is selected according to a method as described in Fig. 9, which is the middle image. Fig. 10 shows the same image after segment marking, where the markings 1010 and 1020 correspond to the markings 910 and 920 of the Fig. 9. In Fig. 11 shows the middle patch 1110 and the dark patch 1120, corresponding to the markings 910 and 920, respectively.

[0089] According to embodiments of the present invention, the patching process therefore comprises: (a) Labeling - Finding sets of connected binary blobs and assigning a color to each ( Fig. 10). (b) Polygon boundary - Bounding each segment by a polygon, such as an octagon. (c) Patch selection - Checking the zone of each polygon to determine if more than (typically) 3% of its field is saturated in the center image. Dark image patches were taken for saturated segments, while medium image patches were taken for non-saturated segments.

[0090] Polygon bounding is well known in the art but is described here. A linear bin is one whose interior is specified by a finite number of linear inequalities. For example, in 2D, a bin Z could be specified by k inequalities: f i (x,y) = a i x + b i y + c i< 0 (i=1,k), all of which must be true for a point (x,y) to be in the zone. If any of the inequalities were to fail, the test point would be outside Z.

[0091] Now with reference to Fig. 12, a box is a rectangular zone whose edges are parallel to the coordinate axes, and is therefore defined by its maximum and minimum extension for all axes. Thus, a 2D box is given by all (x,y) coordinates x min ≤ x ≤ x max and y min ≤ y ≤ y max The inclusion of a point P = (x, y) in a box is tested by verifying that all inequalities are true; if any of them fail, the point is not within the box. In 2D, there are four inequalities, and on average, a point is rejected after two tests.

[0092] The bounding box of a 2D geometric object is the box of minimal field that encloses the object. For any collection of linear objects (points, segments, polygons, and polyhedra), its bounding box is given by the minimum and maximum coordinate values ​​for the point set S of the object's vertices: (x min , x max , y min , y max These values ​​are easily computed in O(n) time with a single scan of all vertices, sometimes while reading or calculating the object's vertices. The bounding box is the computationally simplest of all linear bounding boxes, and the most widely used in many applications. At runtime, the inequalities involve no arithmetic and only compare raw coordinates with the precomputed min and max constants.

[0093] Now with reference to Fig. 13, after some arithmetic, the simplest non-trivial expressions are those that simply add and subtract raw coordinates. In 2D, one has the expressions p=(x+y) and q=(xy), which correspond to lines with slopes of (-1) and 1, respectively. For a 2D set S of points, the minimum and maximum over S of each of these two expressions can be calculated to give (p min , p max , q min , q max ). Then the “boundary diamond” for the set S is the zone given by the coordinates (x,y) containing the inequalities: p min ≤ (x+y) ≤ p max and q min ≤ (xy) ≤ q max Geometrically, it is a rectangle rotated 45 degrees and resembles a diamond.

[0094] The "bounding diamond" involves slightly more computation than the bounding box. However, after calculating both, it can be determined that one is better than the other. All included minima and maxima (8 of them) can be computed in O(n) time with a single scan of the set S. Then, the fields of the bounding box B and the bounding diamond D can be compared, and the smaller bin can be used if desired. Since everything is rectangular, one has area(B) = (x max -x min )*(y max -y min ), and field(D) = (p max -p min )*(q max -q min ).

[0095] Furthermore, both containers, the box and the diamond, could be used to obtain an even smaller combined “bounding octagon” bounded by all 8 inequalities as in Fig. 14. Typically, inclusion or intersection with the bounding box is tested first, followed by the bounding octagon. If one wants to test a point P = (x, y) for inclusion in a polygon Z, the only overhead is the computation of the expressions (x + y) and (xy) just before testing for diamond inequalities, after P has been found to be within the bounding box.

[0096] Therefore, given the HDR image with ghosting, the ghost-free image (HDRout) for the three-image bracketing case (which is the most common case) is generated by two blends as follows: HDRout=HDR(1−Wd)+Wd×Dark_Image+HDR(1−Wm)+Wm×Medium_Image

[0097] The weights Wd and Wm are determined by the bounding octagons of the respective segments after a certain smoothing.

[0098] It should be noted that the embodiments of the present invention employing this ghost removal process differ from the embodiments described in the parent application in that pairwise whole image merges are not required to accumulate a relatively ghost-free image montage. Instead, interfering image zones are detected, and replacement patches are merged from similar zones in available alternative images. The tone mapping process

[0099] The enhanced ghost-removed image has a bit width of 16 bits and comprises three color components: a red component, a green component, and a blue component. This RGB 16-bit data is intended to be displayed on a built-in display device of a 245-bit digital camera. Fig. 2, which is an 8-bit display device, so the enhanced, ghost-removed image must be converted from 16-bit RGB data to 8-bit RGB data. The process of converting image data from one bit width to a narrower bit width, such as from 16 bits to 8 bits, while retaining the relative grayscale levels represented in the wider bit width data in the resulting 8-bit data, is called "tone mapping." There are many such tone mapping processes that can be used. This embodiment of the invention uses a single tone mapping approach originally intended to map 12-bit wide image data into 8-bit wide image data. Therefore, this approach first removes the 4 least significant bits from the second processed composite image data, leaving 12-bit RGB image data.Three lookup tables (LUTs) are used to map the remaining 12-bit RGB data to the required 8-bit RGB data:. A normal gain LUT, A high-gain LUT; and A maximum gain LUT

[0100] The correct LUT for use in the 12-bit to 8-bit tone mapping process must be selected to accurately represent the image grayscale presented in the 8-bit data format. The selection criteria depend on the size of the image field populated with pixels whose average value is below a predefined pixel value, or "dark," compared to the rest of the image field. The lower the average pixel value in the dark field of the image, the higher the selected LUT gain.

[0101] The process of LUT selection is as follows: 1. Shift the 12-bit RGB image 4 bits to the right. This results in an 8-bit image; 2. Generation of the luma component of the resulting 8-bit image; 3. Calculate the average value, Mn, of all pixels in the 8-bit image whose luma is less than a dark field threshold, Td. A digital value of 20 out of a maximum digital value of 255 (the maximum 8-bit value) can be used for Td; 4. If the sum of all pixels with a luma < Td is less than a field threshold, P%, use the normal gain LUT, otherwise: 5. If Mn is given and predefined thresholds pixel value thresholds T1 are less than T2: 6. If Mn < T1, use the maximum gain LUT; 7. If Mn is between T1 and T2, use the high gain LUT 8. If Mn > T2, use the normal gain LUT

[0102] Td = 20 of 255, T1 = 5 of 255 and T2 = 10 of 255 are examples of the thresholds that can be used in the above tone mapping LUT selection process.

[0103] The tone mapping procedure used by this embodiment is designed to enhance low-illumination areas, while the three LUTs behave the same for high-illumination areas. This has been found to produce good results when applied to the resulting ghost-free HDR.

[0104] To summarize, the enhanced HDR resolution of the parent application increases the dynamic range of a given scene by fusing and blending the information from three different images of a given scene. Each image is captured at a different exposure setting, e.g., nominal exposure, overexposure, and underexposure. The underexposed image can be referred to as the dark image, the nominally exposed image can be referred to as the medium image, and the overexposed image can be referred to as the bright image.

[0105] The HDR algorithm improves the dynamic range of the scene by performing the following:

[0106] Dark areas in the middle image are replaced by pixels from the bright image to brighten and enhance the detail of the scene.

[0107] Saturated fields in the center image are replaced with pixels from the dark image to recover burned-out details.

[0108] The overall methodology of the embodiments of the contrasting present invention will now be summarized. In Fig. Figure 15 shows an exemplary and non-limiting implementation of HDR according to one embodiment. While shown in black and white format to comply with USPTO filing constraints, this should not be considered limiting the embodiments of the present invention; all images may be full color without departing from the scope of the invention. Furthermore, any number of images may be input, although three is typical. Three input images with different exposures, e.g., image 1510, which is the dark image, image 1520, which is the middle image, and image 1530, which is the light image, are merged to produce a high dynamic range result shown in image 1540.

[0109] The HDR algorithm now has three main phases, which are related to Fig. 16. The three phases are: an image registration phase 1610 for aligning the three input images; an image fusion stage 1620 for blending the three aligned images together to produce a high dynamic range image, as well as integrated ghost removal, using the principles explained in more detail above; a tone mapping phase 1630 for mapping the high dynamic range result into an 8-bit range, which is typically required to display the result on common display devices, but should not be considered as limiting the scope of the present invention.

[0110] The goal of image registration is to align the three images to the same set of coordinates. To align the images, two registration procedures are integrated: the first aligns the dark image with the center image, and the second aligns the light image with the center image.

[0111] The registration phase 1610 detects and compensates for the global motion of the scene between two different captured frames. This phase is required for the HDR algorithm because the camera is assumed to be handheld and thus susceptible to the effects of shaky holding. The alignment scheme consists of four phases:

[0112] Motion vector extraction - a set of motion vectors is extracted between the two images;

[0113] Global motion estimation – a global transformation model, typically but not necessarily affine, is assumed between the images. A sampling consensus (RANSAC) algorithm is applied to the motion vectors to estimate the most probable transformation parameters.

[0114] Image distortion - according to the estimated global transformation, a hardware-based distortion mechanism typically transforms the dark or bright image to the mean image coordinates; and,

[0115] Unified field of view - Due to camera movement, there may be some differences between the fields of view of the images. In this phase, the maximum field of view existing in all three images is calculated while maintaining the original aspect ratio of the input images.

[0116] The image fusion phase 1620 blends the three images together (phase 1620 is shown in more detail in Figure 17). The blending is performed as follows: using, for example, the center image as a reference, the dark image contributes information in overexposed (1710) fields, and the bright image contributes information in underexposed fields (1720). This blending rule is used when the scene is static, as mentioned above. However, if there is local motion in the scene, as shown in certain examples above, the blending can lead to visible artifacts in the HDR result, known as ghosting artifacts. To overcome these motion-related artifacts, a ghosting treatment mechanism (1730) is applied as part of the image fusion phase.

[0117] The basic operation of image blending takes two images with different exposures and blends them according to a pixel-by-pixel blending factor. To describe the steps of the image blending procedure, the less exposed image is labeled I1 and the image with greater exposure is labeled I2. The exposure value of each image is labeled ExpVal1 and ExpVal2, respectively. The exposure value in computational photography is calculated according to the following formula: ExpVal=ISO⋅ExpTimeF#2

[0159] where ISO represents the ISO level, ExpTime represents the exposure time and F# represents the F-number of the optical system.

[0118] The following phases are applied within the image blending scheme. First, a preprocessing phase, which includes:

[0119] If I1 or I2 is given in the gamma range (not in the linear range), a gamma correction operation is applied to display the input images in the linear range; and

[0120] The brighter image, I2, is normalized to the exposure value of the darker image, I1. The manipulations to the input image can be summarized as: {I1upd=Gamma correction(I1)I2upd=Gamma correction(I2).ExpVal1ExpVal2

[0121] Second, the calculation of blending weights takes place. To determine the weights, the brightness values ​​(luma, denoted as Y) of the brighter image, I2, are used as an input to a weighting LUT. This can be formulated as W = LUT(Y2). The weighting LUT can be described as a general mapping, but is usually implemented as a piecewise linear function, as in Fig. 18, where the piecewise linear graph 1810 is shown.

[0122] Finally, the mixing is performed, with the actual mixing operation being carried out according to the following formula: Iout=(1−W)⋅I1upd+W⋅I2upd

[0123] If the example in Fig. 18 is used as the weight LUT, the blending operation takes dark pixels from I2upd, bright pixels of I1upd and performs pixel-by-pixel merging between the two images for average luma values. Note that, unlike the embodiments of the parent application, all images in the series can be processed in any order.

[0124] According to an embodiment of the present invention, a ghost removal solution is provided below as described in Fig. 19. The ghosting treatment mechanism attempts to detect fields of local motion between the three HDR inputs. In these fields, the HDR fusion results must not include a merging of the images, as this may lead to a ghosting artifact, e.g., a walking person may be seen twice or more, as in Fig. 7. Instead, only a single image is selected to represent the HDR fusion result in the specified field (i.e., a patch). Accordingly, the ghosting treatment mechanism has the following phases, beginning in step 1900: - Motion detection - Identify whether there is local motion between the input HDR images. This phase is performed per pixel, starting in step 1910: - Definition of ghost patches - in this phase, the pixels experiencing motion are clustered into patches (image blobs) using morphological operations, starting in step 1920; - Patch selection - each of the identified patches must be represented by a single input image. In this phase, a scoring function is used to decide whether the information develops over the exemplary bright, medium, or dark image, beginning in step 1930; and - Patch Correction - in this phase, hardware-based patch correction is typically used to replace the ghost patch with the selected input image, starting in step 1940.

[0125] Each of the four phases is discussed in more detail below, beginning with the motion detection phase. After the image registration phase, which fuses the dark and light images to the mean image coordinates, the motion detection scheme attempts to detect differences between the three aligned images. The underlying assumption is that these changes can arise from motion in the scene.

[0126] To efficiently detect local motion, a downscaled version of the aligned images can be used. For example, a 1:16 downscale factor is typical for 12-megapixel images. During this phase, any type of difference detection algorithm can be used to detect which pixels are different between the images. A simple detection scheme distinguishes between images in pairs, i.e., medium and dark and medium and light, and may include the following steps: - Calculate STD (standard deviation) and mean for each downgraded image: - Normalize the images to have the same STD and mean for each downgraded image; - Generating an absolute difference map of the two downgraded images: and, - Using a threshold on the difference map to distinguish between actual detections and registration noise.

[0127] This method, although simple and straightforward, can lead to unreliable and noisy difference mapping under certain conditions. For example, in saturated fields, the bright image histogram is highly saturated (most pixel values ​​are close to 255), while the center image histogram is only partially saturated (only a portion of the pixels are close to 255). Due to the truncation of image intensities (in the bright image), a systematic error can be introduced into the calculations of the STD and mean of the bright image. Furthermore, the difference mapping can identify the difference between the truncated pixel in the bright image and the partially saturated pixel in the center image as a detection, regardless of the presence of moving objects.

[0128] Therefore, a more sophisticated motion detection scheme that produces higher motion detection performance than the one described above is described for additional embodiments of the present invention. This method is based on the assumption that within the HDR frame, the bright image can only affect darker fields, while the dark image can affect brighter fields. Therefore, it is not necessary to use the information from the unnecessary image in the motion detection scheme. This motion detection algorithm follows these phases: (a) Two brightness thresholds are defined: BrightTH - Information from the bright image is used when the brightness values ​​are in the range [0, BrightTH], and DarkTH - Information from the dark image is used when the brightness values ​​are in the range [DarkTH, 255]. (b) Defining a difference threshold function: Instead of using a single threshold for the entire difference map, a threshold function is used, where in this case the threshold depends on the brightness of the pixel, so that a higher threshold is used for higher lumas and vice versa. (c) Processing the difference image of the bright and medium image: - Pixel classification: only consider pixels with brightness values ​​in the range [0, BrightTH], the pixels are called BM_Pixel (Bright Medium Pixel); - Calculate the STD and the mean on the set of BM pixels; - Normalize the BM_Pixels in both images to have the same STD and mean for each image; - Generating an absolute difference map of the BM_Pixels; and - Using the threshold function on the difference map to identify detections. (d) Processing the difference image of the dark and middle image: - Pixel classification: only consider pixels with brightness values ​​in the range [DarkTH, 255], the pixels are called DM_Pixel (Dark Medium Pixel); - Calculate STD and mean on the set of DM_pixels; - Normalize the DM_Pixels in both images to have the same STD and mean for each image; - Generating an absolute difference map of the DM_Pixels; and - Use the threshold function on the difference map to detect detections. (e) The final detection image is obtained as a union of the dark-medium detections and the light-medium detections.

[0129] While the above description discusses two images relative to a reference image, it should be understood that more than two may be used without departing from the teachings of the invention. For example, if five images are used, e.g., lightest, light, middle, dark, and darkest, various blends relative to the middle image can be created to achieve desired results.

[0130] After the motion detection phase, several pixels are typically identified as having inconsistent brightness values. The goal is to cluster these detections into patches (or image blobs). Transforming the detected pixels into a whole image patch enables consistent ghosting processing for neighboring pixels at the cost of identifying motionless pixels as detections. A non-restrictive clustering technique, simple but effective, for determining ghosting patches is: (a) Applying a morphological dilation operation to the binary image of the detections, with a structuring element, such as a 5x5 square structuring element; (b) applying a morphological closure operation to the binary image of the detections with a structuring element, such as a 5x5 square structuring element; (c) applying a binary labeling algorithm to distinguish between different ghost patches; and (d) for efficient software implementation, describing each ghost patch by its bounding octagon.

[0131] According to embodiments of the present invention, the patch selection phase decides on the most suitable input image to replace the detected patch. An incorrect patch selection can lead to visible artifacts at the edge of the patch field. There are several heuristics that assist in selecting the most suitable input image: - The patching artifact is more visible in flat fields than in fields with detail or texture. The human eye is more sensitive to image changes in DC fields (e.g., undetailed or "flat") compared to fields with additional high-pass content (texture, edges, etc.). - Patch selection should be as consistent as possible with the blending decisions at the patch's edges. Following this rule will better tailor the patch to its surroundings. - Patching artifacts are especially visible when edge pixels are overexposed or underexposed. - If some pixels are overexposed and some are underexposed, the majority of pixels should influence the input image selection. For example, if there are significantly more overexposed than underexposed pixels at the edge of the patch, the dark image should be selected, and vice versa, if applicable.

[0132] Based on these heuristics, the patch selection algorithm consistently selects the most suitable input image for replacing the ghost patch. The patch selection algorithm is based on histogram analysis or the mean downgraded image brightness values, as in Fig. 20. The basic idea is to divide the histogram into three different fields: underexposed field 2010, in which the bright image is selected, correctly exposed field 2020, in which the medium image is selected, and overexposed field 2030, in which the dark image is selected. As with the motion detection algorithm, predefined thresholds BrightTH 2015 and DarkTH 2025 define the three fields 2010, 2020, and 2030. If more images are used, additional thresholds can be used without departing from the scope of the invention.

[0133] According to this embodiment, several parameters are defined before the actual algorithmic description:

[0134] AreaB, AreaM, AreaD = the brightness range of each of the three histogram zones and described as: {AreaB=[0,LightTH]AreaM=[LightTH+1,DarkTH−1] AreaD=[DarkTH,255]

[0135] AreaSizeB, AreaSizeM, AreaSizeD = the size of the brightness range of each zone.

[0136] NumPixelsB, NumPixelsM, NumPixelsD = the number of pixels in each zone.

[0137] EffAnzPixelB, EffAnzPixelM, EffAnzPixelD = the number of pixels in each zone, taking into account the size of each zone, which can be formulated as: {EffNumberPixelB=NumberPixelB⋅AreaSizeBAreaSizeB+AreaSizeM+AreaSizeDEffNumberPixelM=NumberPixelM⋅AreaSizeMMareaSizeB+AreaSizeM+AreaSizeDEffNumberPixelD=NumberPixelD⋅AreaSizeDDareaSizeB+AreaSizeM+AreaSizeD

[0138] ModeB, ModeM, ModeD = the brightness value in each histogram zone where the histogram is maximum.

[0139] DiffB, DiffM, DiffD = the mean difference of the histogram inputs from the mode brightness value for each histogram zone. ie: DiffB=1NumberPixelsB⋅∑i∈AreaBHist(i)⋅|i−ModeB|p where Hist(i) represents the histogram frequency for the i-th brightness value and p defines the difference metric (best results were obtained with p = 1 or 2). It is understood that DiffM and DiffD can be represented in the same way.

[0140] According to this embodiment, the patch selection algorithm comprises the following phases: (a) Calculating a weighted histogram of the brightness values ​​of pixels in the border areas of a specific patch within the downgraded center image. The histogram is weighted such that pixels with oversaturated or undersaturated brightness values ​​have a greater impact on the histogram frequency counts. This allows overexposed and underexposed zones of the histogram to have a greater impact on the patch selection algorithm, thereby overcoming the associated artifacts. (b) Dividing the histogram into three zones according to BrightTH, DarkTH parameters. (c) Performing the following calculations: Calculating the range size for each histogram zone (RangeSizeB, RangeSizeM, RangeSizeD); Calculating the number of pixels for each histogram zone (NumPixelsB, NumPixelsM, NumPixelsD); Calculating the effective number of pixels for each histogram zone (EffNumPixelsB, EffNumPixelsM, EffNumPixelsD); Calculating the mode of the histogram in each histogram zone (ModeB, ModeM, ModeD); and Calculating the average difference to the mode in each histogram zone (DiffB, DiffM, DiffD), (d) Calculate a weighting function for each histogram zone. The weighting function for selecting the bright image is defined as: RatingsB=EffAnzPixelBDiffB

[0141] Weighting functions for the mid-tone and dark-tone selections are defined similarly. Furthermore, similar techniques can be used when more than three images are used and adapted accordingly without departing from the scope of the invention.

[0142] (e) After calculating all three scoring functions, the zone with the maximum score defines the patch selection result. The scoring function comprises a partition of two calculated measurements described below, although it would be immediately clear to a person of ordinary skill in the art that other scoring functions are possible without departing from the scope of the invention: EfficientPixelB - if the edge of the patch consists primarily of pixels with low brightness values ​​(i.e., in the bright region of the histogram), the bright image is selected. Instead of using the number of pixels itself, its weighted version is used, which also takes into account the size of each histogram region (effective number of pixels); and DiffB - As discussed above regarding patch selection heuristics, when blending is a possible artifact, it is preferred that such an artifact appear in a field with visible texture or detail rather than in flat fields. Measurements of the amount of "flatness" across the patch edge histogram are therefore necessary. In each histogram zone, a Diff value is calculated to measure how spread out the histogram actually is. Smaller Diff values ​​suggest that the edge pixels (which are part of the same histogram zone) have similar brightness values, and can therefore be described as "flat." A spread out histogram, on the other hand, is more likely to represent texture or detail. The inverse relationship between the score and the Diff value supports favoring the input image that best matches the flat zone of the histogram.

[0143] According to this embodiment, after ghost detection and patch generation based on the brightness values ​​of the downgraded images, the patch correction is applied to the full-size blended result of the input images. The correction phase consists of the following phases: (a) Smoothing bounding octagons - Instead of representing the bounding octagon as a binary image, an 8-bit or other sub-maximum-resolution representation is used, and a low-pass filter (LPF), e.g., a 7x7 LPF, is applied to the bounding octagon to smooth the edges of the selected patch. This ultimately results in a smoother blend between the introduced patch and its HDR-driven surroundings. The smooth patch image is called the patch mask. The applied low-pass filter can be determined by parameter values. (b) Upscaling the smooth patch image back to full image resolution. The resulting image is called the patch mask, W patch , designated. (c) Merging between the current HDR blend result (referred to as HDR cur ) and the selected patch input image (as I patch referred to) according to the patch mask as follows: HDRout=(1−Wpatch)⋅HDRcur+Wpatch⋅Ipatch

[0144] According to this embodiment, at the end of the image fusion phase, the resulting high-dynamic-range image is represented as a linear RGB image with 12 bits per color component. The tone mapping task is to convert the 12-bit representation to an 8-bit representation, starting in step 1950 of the Fig.19. This phase is required to enable the image to be presented on common display devices. The main challenge in this phase is to perform intelligent tone mapping that retains the visual value of the image fusion process, even in the 8-bit representation.

[0145] In PC-based HDR algorithms, the tone mapping phase is typically tailored and optimized for each HDR scene and regularly requires human intervention. The described method delivers an in-camera, real-time, hardware-based tone mapping solution that can adaptively change its behavior according to the characteristics of the captured scene.

[0146] While there are many possible techniques for performing tone mapping, the disclosed tone mapping algorithm is based on two different transformations controlled by predefined LUTs: (a) Global Mapping - Performing a gamma-like mapping on the HDR fusion result (while still preserving the 12-bit-per-color component). This mapping is typically the opposite of the gamma correction operation used at the beginning of the fusion phase of the HDR algorithm. This choice is made because it is advantageous to maintain a similarity in color and ambiance between the input images. (b) Local Mapping - Performing a non-linear local mapping that maps a pixel according to the average brightness values ​​of its neighbors in an 8-bit representation per color component. Such a tone mapping operator is excellent at dynamic range compression while maintaining local contrast, whereas tone mapping operations that use only pixel information tend to corrupt local contrast.

[0147] Since the LUTs of these mappings are predefined, a simple way to adaptively modify the tone mapping behavior is to define a family of local mapping LUTs (with a single global mapping LUT) and use a selected LUT representative of each HDR operation. However, this requires some additional heuristics and preprocessing analysis to determine in which scenario a specific LUT should be used. For example, the brightness values ​​of the input images can be used to determine in which LUT they should be used.

[0148] A more sophisticated solution to the challenge of adaptive tone mapping, provided by contrasting embodiments of the present invention, is to perform an online change of the mappings to capture the full dynamic range of the scene. Therefore, according to these embodiments, an additional global mapping is introduced on the brightness component of the HDR fusion result immediately after the gamma transformation. The additional global transformation is derived from the unity transformation (Y=X) to assign additional grayscale levels for more frequent brightness values. For example, if a dominant part of the image is bright, more grayscale levels are assigned for bright brightness levels at the expense of dark brightness levels.By using this adaptive mapping with a predefined gamma LUT and local mapping LUT, you can enjoy colors similar to those in the source image, adjust the brightness levels (without affecting the scene colors) to the adaptive mapping, and add variations to the HDR tone mapping result by defining the local mapping LUT.

[0149] The proposed adaptive tone mapping algorithm includes the following phases: (a) Construct a brightness histogram of the HDR image after gamma LUT conversion. Since this can be computationally expensive, an estimated histogram can be generated using a combination of histograms from the downgraded input images. The estimated histogram is obtained using the following steps: (a1) Converting the three images to the gamma range with exposure compensation, whereby the exposure compensation ensures that the brightness levels in all images are aligned; { I˜dark=IdarkI˜medium=Gamma(Gamma−Correction(Imedium)⋅ExpValuedarkExpValuemedium)I˜bright=Gamma(Gamma−Correction(Ibright)⋅ExpValuedarkExpValuebright) (a2) Calculate a brightness level histogram for the three images Ĩ dunkel , Ĩ mittel , Ĩ hell , and (a3) Combine the three histograms into a single HDR histogram by using two predefined thresholds BrightTH and DarkTH- HistHDR(i)={Histhell(i)i∈[0,LightTH]Histhell(i)i∈[LightTh+1,DarkTH−1]Histdark(i)i∈(DarkTH,255) where Hist HDR represents the combined histogram, and Hist hell , History mittel , History dunkel display the histogram of the input images after exposure compensation and gamma LUT. (b) Define a mapping according to the distribution of brightness values. A wider range of output levels should be assigned to the most densely populated zones of the brightness histogram. One of the most popular techniques for defining such a mapping is known as histogram equalization. A similar concept is used here: (b1) Normalize the histogram, p(i)=HistHDR(i)∑jHistHDR(j) (b2) Calculate the cumulative distribution function, Pr(i)=∑j≤ip(i) (b3) Define the mapping as T(i) = α Pr(i) + (1-α) i where α∈[0,1] is a strength factor that merges between the histogram equalization transform and the unity transform. The strength factor is useful in cases where histogram equalization is too aggressive and may lead to image quality degradation.

[0150] Having thus described several aspects of embodiments of the invention, it will be understood that various changes, modifications, and improvements will occur to those skilled in the art. Such changes, modifications, and improvements are intended to be part of this disclosure and are intended to be within the scope of the invention. Accordingly, the foregoing descriptions and drawings are given by way of example only.

[0151] Consistent with the common practice of those skilled in the art of computer programming, embodiments are described below with reference to operations performed by a computer system or similar electronic system. Such operations are sometimes referred to as computer-executed. It is understood that symbolically represented operations include the manipulation by a processor, such as a central processing unit, of electrical signals representing data bits, and the maintenance of data bits at storage locations, such as in system memory, as well as other signal processing. The storage locations where data bits are maintained are physical locations that have particular electrical, magnetic, optical, or organic properties corresponding to the data bits.

[0152] When implemented in software, the elements of the embodiments are primarily the code segments to perform the necessary tasks. The non-transitory code segments may be stored in a processor-readable medium or a computer-readable medium, which may include any medium capable of storing or transmitting information. Examples of such mediums include an electronic circuit, a semiconductor storage device, read-only memory (ROM), flash memory or other non-volatile memory, a floppy disk, a CD-ROM, an optical disk, a hard disk, a fiber-optic storage device, a radio frequency (RF) link, etc. User input may include any combination of a keyboard, mouse, touch screen, voice control input, etc.User input may similarly be used to direct a browser application running on a user's computing device to one or more network resources, such as web pages, from which computing resources may be accessed.

[0153] While the invention has been described in connection with specific examples and various embodiments, it will be understood by those skilled in the art that many modifications and adaptations of the invention described herein are possible without departing from the scope of the invention as hereinafter claimed. Therefore, it is to be understood that this application is only an example and not a limitation on the scope of the invention as hereinafter claimed. This description is intended to cover any variations, uses, or adaptations of the invention that generally follow the principles of the invention and including such departures from the present disclosure as are within the range of known and customary practice in the art to which the invention belongs.

Claims

[1] A method for blending a plurality of digital images of a scene, the plurality comprising more than two images, comprising: Taking pictures at different exposure levels; Registering counterpart pixels of each image to each other; Deriving a normalized image exposure level for each image; Using the normalized image exposure level in an image merging process; Using the image blending process to blend a first selected image and a second selected image to create an intermediate image, wherein the image blending begins before capturing all of the plurality of images, and Repeating the image merging process using the previously generated intermediate image instead of the first selected image and another selected image instead of the second selected image until all images are merged, and outputting the last generated intermediate image as the merged output image. [2] The method of claim 1, further comprising: selectively converting the images from a lower bits-per-pixel format to a higher bits-per-pixel format prior to using the image merging process, and selectively converting the merged output image to a predetermined lower bits-per-pixel format. [3] The method of claim 1, wherein the image merging process merges the counterpart pixels of two images and comprises: Deriving a luma value for a pixel in the second selected image; Using the luma value of a second selected image pixel as an index into a lookup table to obtain a weight value between the numbers zero and one; Using the weighting value, the normalized exposure level of the second selected image, and the second selected image pixel to generate a processed second selected image pixel; selecting a first selected image pixel corresponding to the second selected image pixel; Using the first selected image pixel and the result of subtracting the weight value from one to generate a processed first selected image pixel; adding the processed first selected image pixel to the processed second selected counterpart image pixel to produce a merged image pixel; and Repeat the above processing sequence until every second selected image pixel is merged with its first selected counterpart image pixel. [4] The method of claim 3, wherein the obtained weight value decreases as the luma value used as the index into the lookup table increases. [5] The method of claim 3, further comprising using a different lookup table for each image to obtain the weight value. [6] The method of claim 1, wherein the image blending begins immediately after the second image of the plurality is captured. [7] A processor-readable medium comprising instructions for removing displaced representations of scene objects appearing in a blended image produced by a digital image blending process applied to a plurality of images acquired at different exposure levels and at different times, the images having their counterpart pixels registered to one another, wherein execution of the instructions by a processor enables actions comprising: Normalizing the luma values ​​of the images to a specific standard deviation and mean; Detecting a local movement between at least one reference image and at least one comparison image; Clustering comparison image pixels with local motion into patches; selecting corresponding patches from the reference image; Creating a connected binary image by logically ORing the patches generated from certain reference images together; and Merging the blended image with the reference images, wherein each reference image is weighted by a weight value calculated from the blended binary image to produce an output image. [8] The data carrier of claim 7, wherein detecting local movement further comprises: Determining an absolute luma variance between each pixel of the reference image and the comparison image to generate a difference image; and Identify difference image zones with absolute luma variances that exceed a threshold. [9] The data carrier of claim 7, wherein clustering comprises: Finding sets of recognized image blobs using morphological operations; and Limit each set by a polygon. [10] A data carrier according to claim 7, wherein the selected images used as reference images comprise at least one of: the image with the lowest exposure level, the image with the highest exposure level, and an image with an intermediate exposure level. [11] A data carrier according to claim 7, wherein the luma value of the images to be processed is downgraded before processing. [12] A data carrier according to claim 7, further comprising selecting reference images by: for a candidate reference image of an intermediate exposure value, calculating a sum of patches of saturated zones and a ratio of the sum of saturated patches to total patches; and Selecting the candidate reference image as the reference image when the ratio is less than or equal to a parameter value, and selecting an image of less than the intermediate exposure value as the reference image when the ratio is greater than the parameter value. [13] An image capture device that captures a plurality of digital images of a scene at different exposure levels, the plurality comprising more than two images, comprising: an image registration processor that registers counterpart pixels of each image of the plurality to each other; an image mixer that combines many images of the plurality to produce a single image, wherein the mixing of the plurality of images begins before all images of the plurality are captured, and wherein the image mixer comprises: an image normalizer that normalizes the image exposure level for each image; and an image merger that uses the normalized exposure level to merge a first selected image and a second selected image to produce an intermediate image, and repeatedly merges the previously generated intermediate image in place of the first selected image and another selected image in place of the second selected image until all images are merged, and outputs the last generated intermediate image as the merged output image. [14] The image pickup device of claim 13, further comprising a gamma correction converter that selectively converts the images from a lower bits-per-pixel format to a higher bits-per-pixel format prior to image blending, and selectively converts the blended output image to a predetermined lower bits-per-pixel format. [15] An image pickup device according to claim 13, further comprising: a luma conversion circuit that outputs the luma value of an input pixel; a lookup table that outputs a weight value between zero and one for an input luma value; a processing circuit that detects a derived luma value for a pixel in the second selected image from the luma conversion circuit, obtains a weighting value from the lookup table for the derived luma value, and generates a processed second selected image pixel from the second selected image pixel, the normalized exposure level of the second selected image, and the weighting value; a second processing circuit receiving as inputs the result of subtracting the weight value from one and the first selected image pixel corresponding to the second selected image pixel and generating a processed first selected image pixel; and an addition circuit that adds the processed first selected image pixel to the processed second selected image pixel to produce a fused image pixel. [16] An image pickup device according to claim 15, which outputs a lower weighting value as the input luma value to the lookup table increases. [17] An image pickup device according to claim 15, using a first look-up table for the first selected image and a second look-up table for the second selected image. [18] An image pickup device according to claim 13, which starts mixing the plurality of images immediately after capturing the second image of the plurality.

Citation Information

Patent Citations

  • Method and unit for generating a radiance map

    US20100328315A1

  • Multiple exposure high dynamic range image capture

    WO2010123923A1