Generation of HDR images from corresponding raw camera images and SDR images
A multi-stage framework using guided filtering and static 3D-LUT mapping enhances SDR images to produce high-quality HDR images, addressing the challenges of spatial resolution and compression artifacts in existing conversion methods.
Patent Information
- Application Number
- JP2024561876
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-21
- Filing Date
- 2023-03-08
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-03-08
AI Technical Summary
Existing image processing technologies struggle to efficiently convert standard dynamic range (SDR) images into high dynamic range (HDR) images while maintaining visual quality and detail, particularly due to issues with spatial resolution, bit depth, and compression artifacts.
A multi-stage framework is employed to generate HDR images from SDR and camera raw images, involving guided filtering with varying kernel sizes to enhance SDR images, followed by static 3D-LUT-based dynamic range mapping, and local reshaping functions to improve contrast and saturation.
The framework effectively generates high-quality HDR images with improved spatial details and visual appeal, comparable to or exceeding device-generated HDR images, by leveraging both SDR and camera raw image data.
Smart Images

Figure 0007819363000082 
Figure 0007819363000083 
Figure 0007819363000084
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 333,374, both filed April 21, 2022, and European Patent Application No. 22169192.6, each of which is incorporated by reference in its entirety.
[0002] [Technical field] The present disclosure relates generally to image processing operations. More particularly, one embodiment of the present disclosure relates to video codecs. [Background technology]
[0003] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, for example, from the darkest black (dark) to the brightest white (highlight). In this sense, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referred" intensity. Unless a particular meaning is explicitly designated as having special significance at any point in the description herein, it should be inferred that the term may be used in either sense, for example, interchangeably.
[0004] As used herein, the term high dynamic range (HDR) refers to a DR width that spans approximately 14 to 15 orders of magnitude or more of the human visual system (HVS). In practice, the DR that allows humans to simultaneously perceive a wide range of intensities may be somewhat shortened compared to HDR. As used herein, the terms extended dynamic range (EDR) or visual dynamic range (VDR), individually or interchangeably, may refer to a DR that is perceptible within a scene or image by the human visual system (HVS), including eye movements, and that allows for some light-adaptive changes across the scene or image. As used herein, EDR may refer to a DR that spans 5 to 6 orders of magnitude. Although EDR may be somewhat narrower compared to true scene-referential HDR, it still represents a wide DR width and is sometimes referred to as HDR.
[0005] In practice, an image comprises one or more color components of a color space (e.g., luma Y and chroma Cb and Cr), each color component being represented with n bits of precision per pixel (e.g., n=8). Using non-linear luminance coding (e.g., gamma coding), an image with n≦8 (e.g., a color 24-bit JPEG image) can be considered a standard dynamic range image, and an image with n>8 can be considered an extended dynamic range image.
[0006] The reference electro-optical transfer function (EOTF) of a given display characterizes the relationship between the color values (e.g., luminance) of an input video signal and the output screen color values (e.g., screen luminance) produced by the display. For example, ITU Rec. ITU-R BT.1886, "Reference electro-optical transfer function for flat panel displays used in HDTV studio production" (March 2011), incorporated herein by reference in its entirety, defines the reference EOTF for flat panel displays. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term "metadata" herein refers to any auxiliary information transmitted as part of a coded bitstream that assists a decoder in rendering the decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters, as described herein.
[0007] The term "PQ" as used herein refers to perceptual luminance amplitude quantization. The human visual system responds highly nonlinearly to increases in light levels. A person's ability to see a stimulus is affected by the luminance of that stimulus, its size, the spatial frequencies that make up the stimulus, and the luminance level to which the eye is adapted at the particular moment the stimulus is viewed. In some embodiments, a perceptual quantizer function maps linear input gray levels to output gray levels that more closely match the contrast sensitivity thresholds of the human visual system. An exemplary PQ mapping function is described in SMPTE ST 2084:2014 High Dynamic Range EOTF of Mastering Reference Displays (hereinafter "SMPTE") (incorporated herein by reference in its entirety), which states that given a fixed stimulus size, for all luminance levels (e.g., stimulus levels), the minimum visible contrast step at that luminance level is selected according to the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).
[0008] Displays supporting a luminance of 200 to 1,000 cd / m² or nits are representative of low dynamic range (LDR) (also known as standard dynamic range (SDR)) compared to EDR (or HDR). EDR content can be displayed on EDR displays that support a higher dynamic range (e.g., 1,000 nits to 5,000 nits or more). Such displays can be defined using an alternative EOTF that supports high luminance capabilities (e.g., 0 to 10,000,000 nits or more). Examples of such EOTFs are defined in SMPTE 2084 and Rec. ITU-R BT.2100, “Image parameter values for high dynamic range television for use in production and international program exchange” (06 / 2017). As recognized by the inventors, improved techniques for generating high-quality video content data with high dynamic range, high local contrast, and vivid colors are desirable.
[0009] WO 2022 / 072884 discloses an adaptive local reshaping method for SDR-HDR upconversion. A luma codeword in an input image is used to generate a global index value for selecting a global reshaping function for the relatively low dynamic range input image. Image filtering is applied to the input image to generate a filtered image. The filtered values of the filtered image provide a measure of local luminance levels in the input image. The global index value and the filtered values of the filtered image are used to generate local index values for selecting a specific local reshaping function for the input image. A relatively high dynamic range reshaped image is generated by reshaping the input image with the specific local reshaping function selected using the local index value.
[0010] Chinese Patent Application Publication No. 112200719 discloses an image processing method including the steps of: obtaining a standard dynamic range (SDR) image of a first resolution; obtaining a guide map of an SDR image of a second resolution and filter coefficients of the SDR image of the second resolution according to the SDR image of the first resolution; and performing guided filtering according to the guide map of the SDR image of the second resolution and the filter coefficients of the SDR image of the second resolution to obtain an HDR image of a second resolution, where the second resolution is higher than the first resolution. This disclosure can reduce the calculation complexity of the filter coefficients, realize fast mapping from an SDR image to an HDR image, and solve the problem of large calculation volume in the mapping method from an SDR image to an HDR image.
[0011] WO 2020 / 210472 discloses a method for generating a high dynamic range (HDR) image, including: (a) denoising a short exposure time image, the denoising including applying a first guided filter to the short exposure time image, the guided filter utilizing the long exposure time image as its guide; (b) scaling at least one of the short exposure time image and the long exposure time image after the denoising to place the short exposure time image and the long exposure time image on a common radiance scale; and (c) merging the short exposure time image with the long exposure time image after the scaling to generate an HDR image.
[0012] U.S. Patent Application Publication No. 2015 / 078661 discloses an algorithm for improving the performance of a conventional tone mapping operator (TMO) by calculating both a contrast discard score and a contrast loss score for a first tone-mapped image produced by the TMO. The two contrast scores can be used to optimize the performance of the TMO by reducing noise and improving contrast. An algorithm for generating an HDR image by converting a nonlinear color space image to a linear color space format, aligning the image to a reference, optionally deghosting the aligned image, and merging the aligned (and potentially deghosted) image to create the HDR image. Merging can be performed using exposure fusion, HDR reconstruction, or other suitable techniques.
[0013] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, it should not be assumed that problems identified with one or more approaches have been recognized in any prior art based on this section, unless otherwise specified. Summary of the Invention
[0014] The invention is defined by the independent claims. The dependent claims relate to optional features of some embodiments of the invention. [Brief explanation of the drawings]
[0015] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference symbols refer to similar elements and in which: [Figure 1] 1 illustrates an exemplary transformation from a camera raw image and an SDR image to a corresponding HDR image. [Figure 2A]1 illustrates an exemplary generation of an enhanced SDR image as an intermediate image. [Figure 2B] 1 illustrates an exemplary configuration of a static three-dimensional look-up table (3D-LUT). [Figure 2C] 1 illustrates an exemplary image enhancement. [Figure 3A] 10 shows an example compressed and filtered SDR image in guided filtering with a relatively small kernel. [Figure 3B] 10 shows exemplary codeword plots for JPEG SDR, filtered SDR, and / or camera raw images. [Figure 3C] 10 shows exemplary codeword plots for JPEG SDR, filtered SDR, and / or camera raw images. [Figure 3D] 10 shows an example compressed and filtered SDR image in guided filtering with a relatively large kernel. [Figure 3E] 1 shows an exemplary two-dimensional map with smoothed binarized local standard deviations. [Figure 3F] 1 shows an exemplary compressed and fused SDR image. [Figure 3G] 10 shows an example mapped HDR image generated using a static 3D-LUT with and without darkness adjustment post-processing. [Figure 3H] 1 shows exemplary curves for generating a hybrid shifted sigmoidal lookup table (HSSL LUT). [Figure 3I] 1 illustrates an exemplary local reshaping function. [Figure 3J] 1 illustrates an exemplary local reshaping function. [Figure 4A] 1 illustrates an exemplary process flow. [Figure 4B] 1 illustrates an exemplary process flow. [Figure 5] FIG. 1 illustrates a simplified block diagram of an exemplary hardware platform on which a computer or computing device described herein may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0016] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, obscuring, or obscuring the present disclosure. overview
[0017] An image capture device may be used to capture a camera raw image. The camera raw image stores original (physical or optical) sensor information (or sensor readout) that has not yet been processed by a camera image signal processor (ISP). Image signal processing operations may be performed on the camera raw image using the camera ISP of the image capture device to generate a (JPEG) SDR image (e.g., a thumbnail image having a relatively small spatial resolution and a relatively low dynamic range) corresponding to the camera raw image. The camera raw image and the SDR image may be contained in a single (raw) image container or file or digitally packaged. Examples of such image containers may include, but are not necessarily limited to, any of a DNG (Digital Negative) file, a TIFF file, a BMP file, a JPEG file, etc.
[0018] Under the techniques described herein, camera-raw and SDR images can be converted into relatively high-quality corresponding HDR images with a visually pleasing appearance comparable to or better than device-generated HDR images. These techniques can be implemented by the image acquisition device itself or by a separate computing device as part of an off-ISP image processing tool. A non-limiting example of such an off-ISP image processing tool can include, but is not necessarily limited to, Adobe Photoshop, which can be operated by a user to utilize camera-raw images as a starting point for creating the user's own preferred image(s).
[0019] In many operating scenarios, JPEG SDR images have a relatively pleasing picture appearance, but are produced at a relatively lower spatial resolution and / or a relatively lower bit depth compared to the camera raw image by using the image capture device's (e.g., dedicated) camera ISP to perform a number of local mapping / processing operations on different spatial regions or image blocks on the camera raw image.
[0020] In contrast, camera raw images, while not necessarily visually pleasing, contain relatively rich image data, including the image detail of the highest resolution image supported by the image sensor used to generate the camera raw image.
[0021] Under the techniques described herein, a relatively high-quality (or visually pleasing) HDR image can be generated from an SDR image using many details from the camera raw image by utilizing both the SDR image and the camera raw image. Compared to the SDR image, the camera raw image can be used to help create an HDR image with better image details. Meanwhile, given that the SDR image has a relatively pleasing picture appearance (or SDR picture mode) supported by the image acquisition device or camera, the SDR image can be used together with the camera raw image to help provide or present the best visual appearance in the resulting HDR image.
[0022] In some operational scenarios, a framework for implementing these techniques may include multiple (e.g., three) stages or major components. The first stage generates a filtered SDR image from a camera raw image and an SDR image (e.g., JPEG) using local linear mapping and / or guided filtering with different (spatial filtering) kernel sizes, and generates an enhanced SDR image by fusing the filtered SDR images. The second stage converts the enhanced SDR image into an HDR image using static SDR-HDR mapping. The static mapping may be defined or specified using operational parameters learned or optimized in a training process using training RAW / SDR / HDR images. The training images used to train the optimized operational parameters for static SDR-HDR mapping may include, but are not limited to, captured SDR and HDR images or spatially aligned SDR and HDR images generated from pictures. The third stage applies local reshaping or hybrid shifted sigmoid function(s) for image enhancement to further enrich the contrast, image detail, color saturation, etc., of the HDR image generated in the second stage. Note that the local linear mapping tools used in the first stage and the image enhancement tools used in the third stage may be applied to different image processing applications or operational scenarios and are not necessarily limited to converting RAW / SDR images into corresponding HDR images as described herein. For example, some or all of these local linear mapping and / or image enhancement tools may be used individually or in combination to provide a relatively rich set of operational parameters for a user to select from and customize the converted HDR image to different preferred visual looks.
[0023] In some operational scenarios, training camera raw / SDR / HDR images may be acquired or captured using a mobile device. One or more such mobile devices may be operated to film or capture (training) camera raw images and both SDR and (training) HDR images, for example, in a captured Profile 8.4 video. The collected / captured training camera raw / SDR images and the collected / captured training HDR images may be used to establish or specify training mappings or pairs that include the training camera raw images, training SDR images corresponding to the training camera raw images, and training HDR images corresponding to the training camera raw images and training SDR images, respectively. These training images may then be used to optimize operational parameters in an image processing or enhancement tool as described herein to generate optimized operational parameters.
[0024] The multi-stage framework as described herein can be deployed or adapted to different image capture devices or cameras with a similar training process. Once these image processing tools, including prediction / mapping models or optimized operating parameters for the stages of the framework, are generated or trained by the training process(es) as described herein, these image processing tools or modules can be deployed, installed, downloaded, and / or executed on the mobile device side or on a cloud-based server side.
[0025] Exemplary embodiments described herein relate to generating an image of a first dynamic range from an input image of a second dynamic range different from the first dynamic range. Using a raw camera image as a guide image, guided filtering is applied to the first image to generate an intermediate image. The first image of the first dynamic range is generated from the raw camera image using an image signal processor of an image capture device. Dynamic range mapping is performed on the intermediate image of the first dynamic range to generate a second image of a second dynamic range different from the first dynamic range. The second image of the second dynamic range is used to generate a specific local reshaping function index value for selecting a specific local reshaping function for the second image of the second dynamic range. The specific local reshaping function is applied to the second image of the second dynamic range to generate a locally reshaped image of the second dynamic range. Exemplary Video Delivery Processing Pipeline
[0026] 1 illustrates an exemplary transformation framework for transforming camera raw images (as well as SDR images) into corresponding HDR images. By way of example and not limitation, the transformation framework includes three main components or processing blocks. Some or all of the transformation framework or the components / processing blocks therein may be implemented in software, hardware, a combination of software and hardware, etc.
[0027] The first major component or processing block of the transformation framework is for receiving a raw camera image and an SDR image and processing the received images to generate an intermediate image. The received SDR image may be generated by a camera ISP of the image acquisition device, while a camera of the image acquisition device simultaneously captures or acquires the raw camera image. The received SDR image can serve as a thumbnail image representative of the relatively high-resolution, relatively high-dynamic-range raw camera image or a relatively low-resolution, low-dynamic-range image. An exemplary intermediate image may be, but is not necessarily limited to, an enhanced SDR image.
[0028] In some operational scenarios, the camera raw image and the SDR image may be digitally packed or packaged together in an input image container or file received as input by a first major component or processing block. Such an image container or file may contain a relatively rich set of image data, including, but not limited to, the camera raw image in the form of an original raw sensor image, the SDR image in the form of a JPEG SDR thumbnail image, etc. The camera raw image represents the light rays captured by the optical sensor, and the SDR image represents a processed (e.g., SDR, thumbnail, preview, etc.) image look produced by the camera ISP performing a relatively large amount of global / local image processing on the camera raw image. The techniques described herein can utilize both the camera raw image and the SDR image to produce intermediate images, such as enhanced SDR images, for subsequent major components or processing blocks (or stages) to use or perform further image processing operations.
[0029] The second main component or processing block receives as input the enhanced image generated by the first main component or processing block and converts it into an HDR image using one of various types of dynamic and / or static (image transformation / mapping) models. By way of illustration and not by way of example, the static model may be implemented as a static three-dimensional lookup table (3D-LUT) and constructed from a tensor product B-spline (TPB)-based mapping / function / model in the RGB domain with relatively low computational complexity or cost. Additionally, optionally, or alternatively, the 3D-LUT described herein may be constructed from a three-dimensional mapping table (3DMT). A 3D-LUT constructed using a 3DMT may introduce more irregularities and discontinuities and may therefore be more likely to cause banding or color artifacts. In comparison, a 3D-LUT constructed using a TPB-based model may have less irregularity and smoother continuity in the local neighborhood of the codeword space or in nearby codeword nodes or entries of the 3D-LUT, thus providing a more perceptually pleasing appearance. An exemplary TPB-based mapping / function / model is described in U.S. Provisional Patent Application No. 62 / 908,770, filed October 1, 2019, by Guan-Ming Su, Harshad Kadu, Qing Song, and Neeraj J. Gadgil, entitled "Tensor-product B-spline predictor," the contents of which are incorporated herein by reference in their entirety as if fully set forth herein. Exemplary operations for constructing a lookup table, such as a 3D-LUT, from a 3DMT for SDR-HDR mapping or conversion are described in U.S. Patent Application No. 17 / 054,495, filed May 9, 2019, the contents of which are incorporated by reference in their entirety as if fully set forth herein.
[0030] A third main component or processing block within the framework receives as input the (mapped or transformed) HDR image generated by the second main component or processing block and enhances or improves local contrast, local image detail, and local saturation in the mapped HDR image to generate an output HDR image. Similar to an SDR image output from the camera ISP, the HDR image output from the camera ISP is also enhanced for a visually pleasing appearance by local image processing tools used by or in conjunction with the camera ISP. The framework as described herein, or a third main component or processing block therein, can be used to support or provide the same or similar level of visually pleasing appearance compared to that of a native HDR image output by the camera ISP. In some operating scenarios, the framework can be used to provide an HDR image enhancement component that enhances the (mapped or transformed) HDR image to make it appear perceptually or visually appealing, for example, outside of or instead of the camera ISP or a local image processing tool operating in conjunction with the camera ISP. Raw Image Processing
[0031] FIG. 2A illustrates an exemplary first major component or processing block ("raw image processing") of FIG. 1 for generating an enhanced SDR image as an intermediate image. As shown in FIG. 2A, resolution and bit-depth adjustments may be performed on a (JPEG) SDR image and a camera raw image corresponding to or resulting in the (JPEG) SDR image, respectively. Using enhancement information extracted from the adjusted camera raw image, a first guided filter having a relatively small spatial kernel may be performed on the adjusted SDR image to generate a first filtered SDR image. Also, using enhancement information extracted from the adjusted camera raw image, a second guided filter having a relatively large spatial kernel may be performed on the adjusted SDR image to generate a second filtered SDR image. A local standard deviation may be determined in the adjusted (JPEG) SDR image and used to derive fusion weighting factors. Based on the (fusion) weighting factors, image fusion may be performed on the first and second filtered SDR images to generate an enhanced SDR image.
[0032] As mentioned above, an image container / file storing raw images or camera raw images may also include camera ISP-generated images, such as SDR images. In many operating scenarios, SDR images (e.g., JPEG, etc.) generated by the camera ISP from the camera raw images already provide relatively high visual quality (appearance) because the camera ISP has performed numerous image processing operations, such as local tone mapping, noise reduction, detail enhancement, and color enrichment. To generate a visually pleasing HDR image corresponding to or from the raw image or image stored in the image container / file, the (JPEG) SDR image can be used as a (e.g., good) starting point. The (JPEG) SDR may discard some details or suffer from compression artifacts during (camera) ISP processing. Under the image fusion framework implemented by the techniques described herein, some or all missing detailed information in an SDR image as a result of ISP processing or compression can be obtained or reacquired from sensor raw data stored in a camera raw image, which has higher spatial resolution and may contain all image data, information, and / or details prior to all image processing operations used to derive the (JPEG or compressed) SDR image. Under this image fusion framework, the image data, information, and / or details extracted from the camera raw image can be used as additional information to be merged or fused together with the image data, information, and / or details represented or included in the (JPEG) SDR image to generate an intermediate image. The intermediate image can be an enhanced SDR image compared to the (JPEG) SDR image and can be used instead of the (JPEG) SDR image to generate a relatively high-quality HDR image by applying an SDR-HDR transformation to the enhanced SDR image.
[0033] The camera raw (or sensor) image and the (JPEG) SDR image may have different spatial resolutions, and the camera raw image often has a higher spatial resolution than the (JPEG) SDR image. The camera raw (or sensor) image and the (JPEG) SDR image may have different bit depths. To combine or fuse image data or information from both the camera raw image and the SDR image, the SDR image may be upsampled to the same resolution as the camera raw image. If applicable, one or both of the camera raw image and the SDR image may be converted to a common bit depth, such as 16 bits, i.e., each codeword or codeword component in the image may be represented by a 16-bit value. Using these adjustment operations, the spatial resolutions and bit depths of the (adjusted) camera raw image and the SDR image are matched.
[0034] Let the i-th pixel of the (adjusted) raw camera (or sensor) image be r i and the i-th pixel of the corresponding (adjusted JPEG) SDR image is denoted by s i It is expressed as:
[0035] Several image enhancement tasks can be performed to generate an enhanced SDR image as an intermediate image. First, some or all missing or lost information in the (JPEG) SDR image resulting from or resulting from relatively small spatial resolution, quantization distortion associated with relatively low bit depth (e.g., 8-bit JPEG), and image compression artifacts from lossy compression in generating JPEG SDR can be restored. Second, some or all clipped information (or clipped codeword values) in the (JPEG) SDR image caused by the camera ISP performing tone mapping to fit the (specified) SDR dynamic range can also be restored.
[0036] Artifacts can arise from spatial resolution resampling operations, particularly near edges separating different image details (e.g., contours of human figures, contours of eyes, etc.). As a previously downsampled JPEG-SDR image (e.g., as part of camera ISP processing, etc.) is upsampled back to a higher spatial resolution, up to the spatial resolution of the camera raw image, these edges separating image details can become blurred and can exhibit visual distortions, such as artifacts like zigzag "pixelation." As part of the image enhancement task, these (edge) artifacts can be treated as noise in the SDR image. Original or relatively accurate edge information in the camera raw image can be extracted and used in conjunction with edge-preserving filtering to enhance SDR images containing distorted edges.
[0037] In some operating scenarios, a guided filter may be used as an edge-preserving filter to perform edge-preserving filtering relatively efficiently. Because edges are visually indicated or depicted with relatively small image details or relatively small spatial scales / regions, the guided filter may be set or configured with a relatively small kernel size.
[0038] Artifacts can also result from tone mapping, which the camera ISP uses to map the raw camera image to an SDR image, especially in highlight and dark regions caused by clipping (of codeword value ranges). This can occur when the camera ISP adjusts the input (e.g., relatively high) dynamic range of all codewords or pixel values across the captured scene to an SDR-like output (e.g., relatively low) dynamic range. The camera ISP may make a tradeoff to preserve or maintain more codewords in the mid-tone regions of the captured scene at the expense of using or clipping fewer codewords in the highlight or dark regions of the captured scene.
[0039] A local tone mapping that maps a local area (e.g., a particular image detail, a particular image coding unit, a particular image block, etc.) of the camera raw image to a corresponding local area of the (JPEG) SDR image is assumed to be linear. Local image data or information in the corresponding local areas of the camera raw image and the (JPEG) SDR image can be collected to construct or reconstruct a linear model for the local tone mapping. This "local" mapping can be performed using a sliding window that moves through each local neighborhood within multiple local neighborhoods that make up some or all of the (JPEG) SDR image. A second guided filter can be used to aid in the construction of the local mapping. Because the highlight or dark regions from which this local tone mapping attempts to restore missing clipped codewords or pixel values are relatively large compared to edges that delineate or separate different image details, a relatively large kernel size can be used for the second guided filter.
[0040] In the case of a guided filter, the input image (as mentioned above, the i-th pixel is s i ) and the guide image (as mentioned above, the i-th pixel is i It is assumed that there is a local (e.g., local neighborhood) linear correlation between the rectified raw camera image (sometimes denoted as Φ ) and the rectified raw camera image (sometimes denoted as Φ ). The local neighborhood (Φ i ), the local linear correlation in the neighborhood is a linear function (α i ,β i ) can be modeled or represented using:
number
[0041] Input image (s i) and predicted image (i-th pixel is s ^ i The difference between the noise or noise image (the i-th pixel is denoted as n i can be modeled or represented as:
number
[0042] The linear function used to model the local linear correlation, or the slope α therein i and offset β i The operating parameters are as follows: i and s ^ i This can be determined or found by minimizing the difference between:
number
[0043] Optimal solution
number
number
[0044] To avoid sudden changes or discontinuities, the linear function used to model the local correlation in the local neighborhood between the input image and the guide image is the sum of the optimized motion parameters calculated for some or all pixels in the local neighborhood of the ith pixel, as follows:
number
number
[0045] The final (guided filtered or predicted) value of the i-th pixel may be given as:
number
[0046] The aforementioned guided filter operation as shown in equations (1)-(6) above can be expressed as follows:
number
[0047] Without limitation, to reconstruct image or spatial details, including their associated edges, lost due to resampling (or downsampling and / or upsampling), the guided filter in equation (7) above may be deployed or applied in a relatively small neighborhood with a relatively small value of B. The relatively small value of B used to restore image or spatial details lost due to resampling may be denoted as BS. Example values of BS may include, but are not necessarily limited to, 12, 15, 18, etc.
[0048] The deployment or application of this guided filter with a relatively small value of BS can be expressed as follows:
number
[0049] 3A shows an exemplary (JPEG) SDR image on the left and an exemplary filtered SDR image on the right using a guided filter with a relatively small neighborhood (or spatial kernel) that incorporates enhancement information from a raw camera image that corresponds to or gives rise to the (JPEG) SDR image. As shown, a large amount of detail has been restored by the enhancement information from the raw camera image via a guided filter with a relatively small neighborhood (or spatial kernel).
[0050] 3B shows three exemplary plots of codeword values in corresponding rows of a (JPEG) SDR image, a filtered SDR image, and a camera raw image. More specifically, a first plot shows compressed codeword values for rows of the (JPEG) SDR image, a second plot shows filtered codeword values for rows of the filtered SDR image (corresponding to rows of the (JPEG) SDR image), and a third plot shows camera raw codeword values for rows of the camera raw image (corresponding to rows of the (JPEG) SDR image). The first and second plots are overlaid in FIG. 3C for ease of comparison. As shown, compared to the compressed codeword values for the (JPEG) SDR image rows, the guided filter increases the total number of codeword values in the filtered codeword values for the filtered SDR image rows, thereby causing the filtered SDR image to contain more of the image detail represented in the camera raw image. The guided filter therefore effectively increases the bit depth of the filtered SDR image, significantly reducing banding artifacts that would otherwise be caused by the relatively small number of compression codewords used in (JPEG) SDR images. Enhance spatial detail with clipping
[0051] As previously mentioned, assuming that the mapping implemented using the camera ISP to convert raw camera images to a (JPEG) SDR mapping is locally linear, a (second) guided filter can be deployed or applied to restore spatial detail lost due to (codeword) clipping, for example, in highlight or dark regions. Spatial regions affected by clipping in a (JPEG) SDR image typically occur within or across a large spatial area. Because pixels within the spatial region have constant values (e.g., upper or lower bounds of a codeword space comprising available codewords for coding), a small kernel size may not be able to restore the detail lost due to clipping; instead, a linear function trained to map different non-constant (codeword or pixel) values to a constant (codeword or pixel) value may be used. Such a linear function is ineffective at restoring the detail lost due to clipping. Therefore, a relatively large neighborhood or kernel size may be used in the (second) guided filter to cover a relatively large area that includes both unclipped and clipped pixels. Correspondingly, a relatively large value, denoted BL, may be selected for the parameter B in equation (7) above. Examples of values for BL may include, but are not necessarily limited to, any of 512, 1024, 2048, and the like.
[0052] The deployment or application of this guided filter with a relatively large value of BL can be expressed as follows:
number
[0053] 3D shows an exemplary (JPEG) SDR image on the left and an exemplary filtered SDR image on the right using a guided filter with a relatively large neighborhood (or spatial kernel) that incorporates enhancement information from the raw camera image that corresponds to or gives rise to the (JPEG) SDR image. As shown, visually noticeable spatial details, such as clouds depicted in the image, that are lost due to clipping in the (JPEG) SDR image are restored in the filtered SDR image by the enhancement information from the raw camera image via the (second) guided filter, while the filtered SDR image changes some colors and details depicted in the (JPEG) SDR image in unclipped spatial regions or areas, such as buildings and trees.
[0054] In some operating scenarios, the relatively large kernel size used in guided filtering can be set equal to or even larger than the entire picture / frame. This is equivalent to having a global mapping from the raw camera image to a (JPEG) SDR image. This model can be implemented with relatively high computational efficiency and relatively low computational complexity, but the resulting filtered SDR image may contain less clipped spatial and / or color detail within the previously clipped spatial region / area compared to a filtered SDR image generated with a relatively large kernel size but smaller than the entire picture / frame. In this case, a simple and very fast global linear regression model can be implemented. Image fusion from multiple enhanced or filtered images
[0055] Although guided filters can be used to reconstruct lost details, filtered images produced from guided filters or filtering with different kernel sizes have different advantages and disadvantages. For example, guided filters with relatively small kernel sizes may not restore spatial or color details due to clipping, while guided filters with relatively large kernel sizes may affect color or luminance and even blur image details in unclipped spatial regions / areas.
[0056] To take advantage of both small and large kernel sizes, image fusion can be used to combine different filtered images together into an overall filtered image. Because guided filters with relatively large kernel sizes are intended to increase or restore missing spatial and / or color detail in clipped areas of a JPEG SDR image, the guided filters can be limited to apply to clipped spatial regions / areas. In some operating scenarios, clipped spatial regions / areas are detected within the (JPEG) SDR image. The weighting coefficient map used to achieve image fusion can be generated to limit or prevent adverse effects from a second guided filter in unclipped spatial regions / areas.
[0057] 4A shows an exemplary process flow for generating a weighting coefficient map for a combination of a first guided filter having a relatively small kernel size and a second guided filter having a relatively large kernel size, which process flow can be performed iteratively to generate a specific weighting coefficient map for each color channel of the color space in which the JPEG SDR image and / or filtered image is represented.
[0058] Block 402 involves calculating the local standard deviation: d for each (e.g., ith) SDR pixel in the (JPEG) SDR image. iThe local standard deviation, denoted as B d ×B d Pixel neighborhood Ψ i can be calculated using a standard deviation filter with the filter kernel set to:
number
[0059] B d A non-limiting example of a value of may include, but is not necessarily limited to, 9. The local standard deviation may be represented as a two-dimensional local standard deviation map, with each pixel in the two-dimensional local standard deviation map having a pixel value as the local standard deviation calculated for the pixel.
[0060] Block 404 involves binarizing the local standard deviations output from the standard deviation filter. A binarization threshold, denoted as Td, may be selected or used to binarize the local standard deviation map or the local standard deviations therein into a binary mask as follows:
number
[0061] Threshold T d A non-limiting example of a value for may include, but is not necessarily limited to, 0.01.
[0062] Block 406 applies a Gaussian spatial filter to create smooth spatial transitions in the spatial distribution of the binary values 0 and 1 in the binary mask;
number
number
[0063] The spatially smoothly transitioned binarized local standard deviation can be represented as a two-dimensional map, as shown in FIG. 3E, in which each pixel has a pixel value as the spatially smoothly transitioned binarized local standard deviation for the pixel.
[0064] Block 408 involves generating and normalizing or renormalizing weighting coefficients used to combine filtered images generated from guided filters having relatively small and relatively large spatial kernel sizes.
[0065] Spatially smoothly transitioned binarized local standard deviation
number
number
number
number
[0001] .
[0066] a first guided filter having a relatively small spatial kernel to be applied to the filtered pixel or codeword values of the first filtered SDR image generated from the first guided filter;
number
number
[0067] to be applied to the filtered pixel or codeword values of the second filtered SDR image generated from the second guided filter having a relatively large spatial kernel.
number
number
[0068] Block 410 is as follows:
number
number
number
[0069] 3F shows an exemplary (JPEG) SDR image with clipped and resampled codewords or pixel values on the left, and an exemplary enhanced or fused SDR image produced by fusing guided-filtered SDR images produced with guided filtering using a relatively small spatial kernel and a relatively large spatial kernel. This enhanced or fused (guided-filtered) SDR image can be used as an intermediate image and provided to the second major component or processing block of FIG. 1 ("Enhanced SDR to HDR Conversion") as part of the input for constructing an HDR image corresponding to the (JPEG) SDR image and / or the raw camera image. Enhanced SDR to HDR conversion
[0070] The second major component or processing block of Figure 1 ("Enhanced SDR to HDR Conversion") receives the enhanced SDR image from the first major component or processing block of Figure 1 ("Raw Image Processing") and converts or maps (each SDR pixel or codeword value of) the enhanced SDR image to (each HDR pixel or codeword value of) an HDR image corresponding to the (JPEG) SDR image and the camera raw image used to generate the enhanced SDR image. In some operating scenarios, the SDR-to-HDR conversion or mapping may be performed via a static 3D-LUT, which may, for example, map R.709 RGB SDR pixel or codeword values as input to R.2020 RGB HDR pixel or codeword values as output.
[0071] 2B shows an example configuration of the static 3D-LUT described herein. For example, the configuration of the 3D-LUT may include three stages: building a TPB SDR-HDR conversion / mapping model, building a static 3D-LUT using the TPB model, and performing 3D-LUT post-processing operations such as node correction operations.
[0072] The TPB model may be constructed using TPB basis functions that ensure smooth transitions and continuity, thereby eliminating or avoiding discontinuities within local neighborhood(s) that may cause abrupt color change artifacts in the transformed / mapped HDR image. Applying the TPB model and directly using the TPB coefficients and TPB basis functions in image transformation operations may result in relatively high computational complexity or cost. In some operational scenarios, a computationally efficient (e.g., static) 3D-LUT may be constructed or pre-constructed from the TPB model. Instead of directly using the TPB coefficients and TPB basis functions of the TPB model at runtime, the 3D-LUT may be used at runtime to map or transform the image. Additionally, optionally, or alternatively, the 3D-LUT may be refined into a final static 3D-LUT that can be used to enhance dark regions in the mapped or transformed HDR image, for example, by modifying the 3D-LUT in a post-processing stage.
[0073] By way of example and not limitation, two image acquisition devices of the same type and / or model and / or manufacturer may be used to capture the same (physical) scene to generate two image data sets.
[0074] For example, one of the two image acquisition devices may operate to generate a plurality of raw image files / containers (e.g., DNG still image or image files, etc.) that include a plurality of (device-generated) raw camera images and a plurality of (device-generated JPEG) SDR images, where each raw image file / container in the plurality of raw image files / containers includes a camera raw image and a (JPEG) SDR image generated from the camera raw image through camera ISP processing by the image acquisition device.
[0075] Simultaneously or nearly simultaneously, the other of the two image acquisition devices may operate to generate an HDR video (e.g., Profile 8.4 video, etc.) including a plurality or series of (device-generated) HDR images that correspond to or depict the same scene as the plurality of camera raw images and the plurality of (JPEG) SDR images, each HDR image in the plurality of HDR images corresponding to or depicting the same scene as a respective camera raw image in the plurality of camera raw images and a respective (JPEG) SDR image in the plurality of (JPEG) SDR images.
[0076] The plurality of HDR images, the plurality of camera images, and the plurality of (JPEG) SDR images may be used as or collected into an image dataset including a plurality of image combinations, each image combination in the plurality of image combinations including an HDR image in the plurality of HDR images, a camera raw image to which the HDR image corresponds, and a (JPEG) SDR image generated from the camera raw image through camera ISP processing.
[0077] In some operating scenarios, the images in some or all of the image combinations may be acquired or captured freehand by two image acquisition devices without space or movement constraints. In some operating scenarios, the images in some or all of the image combinations may be acquired by two image acquisition devices with space or movement constraints, such as mounted or configured via a tripod.
[0078] Because the multiple device-generated (JPEG) SDR images and the multiple device-generated HDR images in the image dataset are acquired using two image acquisition devices, a device-generated (JPEG) SDR image in the multiple device-generated (JPEG) SDR images may not be spatially aligned with a corresponding device-generated HDR image in the multiple device-generated HDR images. Feature points are extracted from the multiple device-generated (JPEG) SDR images and the multiple device-generated HDR images and used to perform a spatial alignment operation on the multiple device-generated (JPEG) SDR images and the multiple device-generated HDR images to generate a plurality of spatially aligned device-generated (JPEG) SDR images and a plurality of spatially aligned device-generated HDR images. Some or all of the SDR pixels in each spatially aligned device-generated (JPEG) SDR image in the multiple spatially aligned device-generated (JPEG) SDR images are spatially aligned with some or all of the corresponding HDR pixels in each spatially aligned device-generated HDR image in the multiple spatially aligned device-generated HDR images.
[0079] A plurality of spatially aligned device-generated (JPEG) SDR images and a plurality of spatially aligned device-generated HDR images may be used as training SDR and HDR images to train or optimize operational parameters that specify or define a TPB transformation / mapping model that predicts or outputs predicted, mapped, or transformed HDR pixel or codeword values from SDR pixel or codeword values as input.
[0080] For example, SDR pixel or codeword values in each spatially aligned device-generated (JPEG) SDR image in a plurality of spatially aligned device-generated (JPEG) SDR images may be used by a TPB model to predict a mapped / transformed HDR pixel or codeword value as an output. A prediction error may be measured or generated as a difference between the (predicted / mapped / transformed) HDR pixel or codeword value and a corresponding HDR pixel or codeword value in each spatially aligned device-generated HDR image in the plurality of spatially aligned device-generated HDR images in the image dataset. The operating parameters of the TPB model may be optimized by minimizing the prediction error in a model training phase. The TPB model with the optimized parameters may be used to predict a transformed / mapped HDR pixel or codeword value from the SDR pixel or codeword value.
[0081] Exemplary feature point extraction, image space alignment, and TPB model training are described in U.S. Provisional Patent Application No. 63 / 321,390, "Image Optimization in Mobile Capture and Editing Applications," by Guan-Ming Su et al., filed March 18, 2022, the contents of which are incorporated by reference in their entirety as if fully set forth herein.
[0082] The optimized operating parameters for the TPB model (or inverse reshaping) are calculated in the RGB domain or color space in which SDR and HDR images are represented:
number
[0083] As described above, the TPB model with optimized TPB coefficients and basis functions can be directly used to predict HDR values from SDR values with relatively high computational complexity and cost. In some operating scenarios, to reduce runtime computational load, the TPB model may not be directly deployed at runtime to predict HDR values in a mapped / transformed HDR image from (untrained) SDR values in an SDR image. A static 3D-LUT may be constructed before runtime from the TPB model with optimized TPB coefficients and basis functions, with each dimension of the 3D-LUT being for predicting component HDR codeword values in a respective dimension or color channel of the output HDR color space in which the mapped / transformed HDR image is represented.
[0084] By way of example and not limitation, each dimension or color channel of an input SDR color space in which an SDR image is represented, such as an RGB color space, may be partitioned into N (e.g., equal, unequal, etc.) partitions, intervals, or nodes, where N is a positive integer greater than 1. Example values of N may include, but are not necessarily limited to, any of 17, 33, 65, etc. These N partitions, intervals, or nodes in each such dimension or color channel may be represented by N respective component SDR codeword values in the dimension or color channel. Each partition, interval, or node in the N partitions, intervals, or nodes of a dimension or color channel in the N partitions, intervals, or nodes may be represented by a corresponding component SDR codeword value in the N respective component SDR codeword values.
[0085] The static 3D-LUT contains NxNxN entries. Each entry in the 3D-LUT contains a different combination of three component SDR codeword values as a lookup key. Each of the three component SDR codeword values is included in the N component SDR codeword values in one of the dimensions or color channels of the input SDR color space. Each entry has or outputs three mapped / transformed component HDR codeword values, as predicted by the TPB model, in each of the three dimensions or color channels of the output HDR color space as lookup values.
[0086] Given an (input) SDR codeword value including three component SDR codeword values (e.g., for each pixel of an input SDR image, for the i-th pixel of an input SDR image, etc.), the nearest neighboring partition, interval, or node represented by the respective component SDR codeword value (e.g., generated by quantizing or dividing the entire codeword range, such as a normalized value range of [0, 1], by the total number of partitions, intervals, or nodes in each color channel) that is closest to the component SDR codeword value of the input SDR codeword value may be identified in the partition, interval, or node represented in the static 3D-LUT. Each nearest neighboring partition, interval, or node may be identified or indexed using node indices (a, b, c) (where 0≦a, b, c≦N) and corresponds to a particular entry in the 3D-LUT. The predicted component HDR codeword value in channel ch may be output as a part / component / dimension of the lookup value in a particular entry of the 3D-LUT using node indices identified by or transformed from the component SDR codeword values of the input SDR codeword value. Predicted component HDR codeword values in a particular channel, denoted ch, of the output HDR color space from some or all entries (of the 3D-LUT) corresponding to the nearest neighboring partitions, intervals, or nodes (e.g., two, four, six, eight, etc.) may be used, e.g., through bilinear interpolation, trilinear interpolation, extrapolation, etc., for channel ch.
number
[0087] In some operating scenarios, the 3D-LUT generated from the TPB model may have a dark lift problem, which increases the luminance level in the predicted HDR codeword values. To address these issues, a 3D-LUT post-processing operation can be performed to make manual and / or automatic adjustments to the 3D-LUT to make the dark nodes (or predicted HDR codeword values in dark regions) darker than they would otherwise be.
[0088] First, the predicted component RGB HDR codeword values for the same node index (a, b, c) predicted by the static 3D-LUT for all dimensions or channels of the output HDR color space are given as follows:
number
number
number
[0089] Therefore, a 3D-LUT for the YCbCr HDR color space can be constructed from the 3D-LUT generated for the output RGB HDR color space.
[0090] Second, the Y codeword value range [0,Y D ), the dark range along the Y axis / dimension / channel of the YCbCr HDR color space can be partitioned into L intervals, where Y D An example value of y may include, but is not necessarily limited to, 0.3 in the overall normalized Y codeword value range (or domain) 0001. The i-th interval (where i=0, 1, ..., L-1) is the Y codeword value subrange
number
number
[0091] Third, each set Ω i About the set Ω i The lookup values of the entries or nodes in the static RGB 3D-LUT corresponding to all the YCbCr entries of the node in are reduced as follows (the darker the corresponding overall RGB node or codeword value, the more the component RGB codeword values in each channel or component ch of the overall RGB node or codeword are reduced):
number
[0092] An example value for D may include, but is not necessarily limited to, 0.065. The final 3D-LUT containing these dark codeword adjustments
number
[0093] FIG. 3G shows an example HDR image transformed by a static 3D-LUT with and without darkness adjustment post-processing. As shown, the darkness-enhanced 3D-LUT
number
number
[0094] A camera ISP in an image capture device as described herein may be used to perform many different enhancement processing operations in order to provide a visually pleasing appearance in the output picture / image. Enhancement processing operations performed with such a camera ISP may include, but are not necessarily limited to, local contrast enhancement and image detail enhancement.
[0095] To implement or simulate the same or similar visually pleasing appearance in the mapped or transformed HDR image as described herein, an enhancement module can be implemented in the third component or processing block of FIG. 1 to create or output a final HDR image with a relatively visually pleasing appearance. The enhancement module can operate in conjunction with or without other available solutions for enhancing an image. These other available solutions may include smear masking, local Laplacian filtering, proprietary or standard image enhancement operations, etc. In some operating scenarios, compared to these other available solutions, the techniques described herein can be used to implement a relatively simple and effective solution for enhancing an image within the same dynamic range, e.g., for enhancing a mapped or transformed HDR image generated from the second component or processing block of FIG. 1 into a final enhanced mapped or transformed HDR image within the same dynamic range or HDR.
[0096] 2C illustrates an exemplary third major component or processing block of FIG. 1 for image enhancement. As shown in FIG. 2C, this framework includes, among other things, Y (or luminance contrast) local reshaping and RGB (or R / G / B saturation) local reshaping. Prior to performing these local reshaping operations with respect to an input HDR image (e.g., a mapped or transformed HDR image generated from the second major component or processing block of FIG. 1), a family of Hybrid Shifted Sigmoid functions with linear mapping (HSSL) LUTs may be built or constructed.
[0097] In some operating scenarios, each dimension or color channel (e.g., any of Y / R / G / B, etc.) of some or all dimensions or color channels (e.g., Y / R / G / B, etc.) for which local reshaping is to be performed may utilize its own separate HSSL LUT, different from other HSSL LUTs used for other dimensions or color channels.
[0098] In some operating scenarios, each dimension or color channel (e.g., either R / G / B, either Y / R / G / B, etc.) of some or all dimensions or color channels for which local reshaping is to be performed may utilize or share the same HSSL LUT as used by the other dimensions or color channels.
[0099] As shown in FIG. 2C , first, an input HDR image (denoted as “HDR RGB”) may be converted to an HDR image in the YCbCr domain or color space using an RGB-to-YCbCr color space conversion. The Y or luma codeword of the HDR image in the Y channel may be processed (e.g., a local mean Y value or a mid-L value) and used to generate a local reshaping function (LRF) index map, which may include, for example, the local mean Y value or the mid-L value (e.g., per pixel, per pixel block, etc.) as an LRF index (e.g., per pixel, per pixel block, etc.).
[0100] 2C , a luma reshaping module (labeled "Y local reshaping") receives or obtains a Y channel codeword value for each pixel in the HDR image in the YCbCr color space, uses an LRF index in the LRF index map corresponding to the pixel to find or select an LUT applicable to the pixel from among a family of Y channel HSSL LUTs, and uses the selected LUT to map the Y channel codeword value to an enhanced or locally reshaped Y channel codeword value. In a YCbCr-to-RGB color space conversion back to the RGB codeword value of the pixel in the luma-enhanced HDR image in the RGB domain or color space, the enhanced Y channel codeword value may be combined with the (original, or unenhanced, or unlocally reshaped) Cb and Cr codeword values in the HDR image in the YCbCr color space.
[0101] A Y channel enhancement operation performed using the Y channel local reshaping described above may be followed by an R / G / B enhancement operation performed using R / G / B local reshaping.
[0102] For each color channel in the RGB color space, an R / G / B reshaping module as shown in FIG. 2C (denoted “RGB local reshaping”) receives or obtains per-channel (R or G or B) luminance-enhancement codeword values for each pixel of each channel in the HDR image in the RGB color space, uses the per-channel LRF index in the LRF index map corresponding to the pixel and channel to find or select an LUT applicable to the pixel from among a family of per-channel HSSL LUTs for the channel (denoted RGB HSSL LUTs for simplicity), and uses the selected LUT to map the per-channel luminance-enhancement codeword values to final enhanced or locally reshaped per-channel codeword values for the channel in the enhanced HDR image or output HDR image (denoted “enhanced HDR RGB”).
[0103] For purposes of example only, it has been described that RGB local reshaping may be followed by Y channel local reshaping to generate an enhanced HDR image or an output HDR image. It should be noted that in other embodiments, Y channel local reshaping may be followed by RGB local reshaping to generate an enhanced HDR image or an output HDR image. Additionally, optionally, or alternatively, in some operating scenarios, Y channel-only local reshaping may be performed without separate RGB local reshaping to generate a final enhanced HDR image or an output HDR image. Additionally, optionally, or alternatively, in some operating scenarios, RGB-only local reshaping may be performed without separate Y channel local reshaping to generate a final enhanced HDR image or an output HDR image. HSSL(Hybrid Shifted Sigmoid with Linear) LUT
[0104] A family of HSSL LUTs may be constructed for a single color channel (denoted as ch) or for several or all color channels. A family of HSSL LUTs may include a total number K (e.g., 1024, 4096, etc.) of different (HSSL) LUTs or local reshaping functions. These different (HSSL) LUTs or local reshaping functions may be indexed by an LRF index k ranging from 0 to K-1. Each LUT is constructed by merging a shifted sigmoid function and a linear function and may be characterized by a set of parameters as shown in Table 1 below. [Table 1]
[0105] Figure 3H shows an example curve used to generate a LUT within the family of HSSL LUTs. As shown, the curve giving rise to the LUT is a function of the parameter c (ch,k)The left and right sigmoid function segments may each have the parameters
number
number
[0106] A sigmoid function, used to represent a sigmoid function segment as described herein, may be expressed as a function with a slope denoted m as follows:
number
[0107] Different slope values of m result in different maximum and minimum values of the sigmoid function on the support of x in the value range [0,1].
[0108] The maximum and minimum values of the left and right sigmoidal functions used to specify the left and right sigmoidal function segments of FIG. 3H can be expressed as:
number
[0109] Using the maximum and minimum values in equation (20) above as normalization factors, one can define left and right scaled sigmoid functions as follows:
number
[0110] The left and right scaled sigmoid functions in equation (21) are centered on c (ch,k) to yield left and right shifted sigmoid functions as follows:
number
[0111] The k-th local reshaping function (which may be used to generate the corresponding HSSL LUT) may initially be set to a default 1:1 mapping, i.e., LRF (ch,k) (x)=x. In addition, the center c of the k-th local reshaping function (ch,k) teeth,
number
[0112] center c (ch,k) The default function LRF around (ch,k) The left and right regions of (x) can be replaced by left and right shifted scaled sigmoid functions in equation (22) above. More specifically, the left region or segment
number
number
number
number
number
[0113] These maximum and minimum values are used in normalizing or scaling, as follows:
number
number
[0114] The final (output) local reshaping function value in the left region or segment may be given as:
number
[0115] Similarly, the default local reshaping function LRF (ch,k) (x) is the right region or segment
number
number
number
number
number
[0116] These maximum and minimum values are used in normalizing or scaling, as follows:
number
number
[0117] The final (output) local reshaping function value in the right region or segment may be given as:
number
[0118] FIG. 3I shows exemplary local reshaping functions (belonging to a family of K=4096 different local reshaping functions) with k=512, 2048, and 3584, respectively, where:
number
number
number
number
number
[0119] FIG. 3J shows exemplary local reshaping functions (belonging to a family of K=4096 different local reshaping functions) with k=512, 2048, and 3584, respectively, where:
number
number
number
number
number
[0120] HSSL LUT or Local Reshaping Function LRF (ch,k) Once the family of (x) is generated, a local reshaping function (LRF) index map can be generated from the input HDR image and used to select different HSSL LUTs or different local reshaping functions for reshaping different pixels in the input HDR image into a final enhanced or locally reshaped HDR image.
[0121] By way of illustration and not limitation, take Y channel local reshaping as an example: for each pixel in the input HDR image, the luma or luminance of the pixel may be locally reshaped or enhanced using an applicable HSSL LUT or local reshaping function specifically selected for that pixel, thereby increasing the saturation or contrast (luminance difference) between pixels in the local neighborhood around each pixel.
[0122] As mentioned above, the curves that give rise to each HSSL LUT or local reshaping function within a family of HSSL LUTs or local reshaping functions are more or less linear, except for the regions / segments replaced by shifted / scaled sigmoid curves. The center connecting the left and right sigmoid curve segments may be located differently for each such curve, depending on a parameter k corresponding to an LRF index, such as the local mean (luma value) or mean L value. Thus, a local mean may be calculated for each pixel's local neighborhood and used to select a particular local reshaping function having a center given as the local mean. Pixel or codeword values within the local neighborhood may use a particular local reshaping function to boost or enhance luminance differences, contrast, or saturation levels or differences, thereby enhancing the visual appearance of the final enhanced image.
[0123] The challenge to be addressed in using the local mean to select the local reshaping function is how to calculate the local mean while avoiding the occurrence of visual artifacts such as halos near edges that separate different image details with different average luminance levels. As mentioned above, guided filters are edge-aware (or edge-preserving) filters and therefore can be applied to select the local mean to be used to calculate the local mean used to select the local reshaping function.
[0124] Let the i-th pixel of the input HDR image (of size WxH, where W is a non-zero positive integer representing the total number of pixels in a row and H is a non-zero positive integer representing the total number of pixels in a column) be v.i It is written as follows.
[0125] Several different types of filters can be used to generate an LRF index map, including Gaussian filters, denoted as follows: Y=GAUSS(X,σ g ) (29) where Y is the pre-filtered image / map X with standard deviation σ g represents the Gaussian blurred or filtered image / map by applying Gaussian filtering with
[0126] In some operating scenarios, Gaussian blurring or filtering may be accelerated, for example, by using an iterated box filter to approximate a Gaussian filter. Exemplary approximations of Gaussian filters using iterated box filters are described in Pascal Getreuer, “A Survey of Gaussian Convolution Algorithms” Image Processing On Line, vol. 3, pp. 286-310 (2013); William M. Wells, “Efficient Synthesis of Gaussian Filters by Cascaded Uniform Filters” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 8, no. 2, pp. 234-239 (1986), the contents of which are incorporated herein by reference in their entirety.
[0127] A gradient filter can be used to calculate the local gradient Y of an image or map X as follows: Y=GRAD(X) (30)
[0128] In some operational scenarios, a Sobel filter may be used as a gradient filter to generate local gradients within an image or map. For example, exemplary local reshaping operations related to Gaussian blur, Sobel, and multi-level guided filters are described in U.S. Provisional Patent Application No. 63 / 086,699, "ADAPTIVE LOCAL RESHAPING FOR SDR-TO-HDR UP-CONVERSION," filed October 2, 2020, by Tsung-Wei Huang et al., the contents of which are incorporated herein by reference in their entirety.
[0129] A guided filter can be used to generate a guided filtered image or map Y of an image or map X as follows: Y=GUIDE(X,I,B g ) (31) where I represents the guided image and B g represents the dimension of the local neighborhood used in guided filtering.
[0130] In some operating scenarios, multi-level (T levels in total) filtering may be performed to generate the LRF index map.
[0131] First, a common feature map (or image) can be created for all levels in the multi-level filtering as follows:
number
[0132] This common feature map (or image) may be used by each level of the multi-level filtering to generate the respective filtered maps according to the exemplary procedure shown in Table 2 below. [Table 2]
[0133] All these multi-level filtered images
number
number
[0134] The LRF map index for the t-th pixel may be calculated as follows:
number
[0135] As a non-limiting example:
number
[0136] The family of HSSL LUTs or local reshaping functions {LRF} to be used for the color channel ch and LRF index map. (ch,k) (x)} and the LRF index map {f i}, local reshaping may be performed on an input HDR codeword of an input HDR image to generate corresponding locally reshaped HDR codewords for the color channels. The locally reshaped HDR codeword for a color channel ch may be combined or further enhanced with other locally reshaped HDR codewords in one or more other color channels to generate final or enhanced HDR codewords in some or all color channels of a final or enhanced HDR image.
[0137] More specifically, for a color channel ch, the input HDR pixel or codeword value v for the i-th pixel i Given the LRF index map {f i LRF index f in i is a family of HSSL LUTs or local reshaping functions {LRF (ch,k) (x)},
number
number
number
[0138] As previously mentioned, the locally reshaped HDR pixel or codeword value may be further locally reshaped or enhanced in other color channel(s). The operational parameters for controlling the amount of local reshaping enhancement (e.g., slope, region / segment size, etc.) for a given channel ch are set as follows for each LRF index k supported in the family of HSSL LUTs or local reshaping functions, as described herein:
number
[0139] In some operational scenarios, a computing device that performs image transformation and / or enhancement operations as described herein, such as an image acquisition device, a mobile device, a camera, etc., may be configured or pre-configured with operational or functional parameters that specify each of some or all of the local reshaping functions in a family.
[0140] Additionally, optionally, or alternatively, in some operating scenarios, each of some or all of the local reshaping functions in the family may be represented using a corresponding (HSSL) LUT generated using the respective local reshaping function in the family of local reshaping functions, thereby resulting in a corresponding family of HSSL LUTs generated from the corresponding family of local reshaping functions. A computing device that performs image transformation and / or enhancement operations as described herein, such as an image acquisition device, a mobile device, a camera, or the like, may be configured or pre-configured with operational or functional parameters that specify each of some or all of the HSSL LUTs in the family of HSSL LUTs generated from the corresponding family of local reshaping functions. Exemplary Process Flow
[0141] 4B shows an exemplary process flow according to one embodiment. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. At block 452, an image processing system applies guided filtering to a first image in a first dynamic range using the raw camera image as a guide image to generate an intermediate image in a first dynamic range, the first image in the first dynamic range being generated from the raw camera image using an image signal processor of an image capture device.
[0142] At block 454, the image processing system performs dynamic range mapping on the intermediate image in the first dynamic range to generate a second image in a second dynamic range different from the first dynamic range.
[0143] At block 456, the image processing system uses the second image in the second dynamic range to generate a specific local reshaping function index value for selecting a specific local reshaping function for the second image in the second dynamic range.
[0144] At block 458, the image processing system applies a particular local reshaping function to the second image in the second dynamic range to generate a locally reshaped image in the second dynamic range.
[0145] In one embodiment, the particular local reshaping function index values collectively form a local reshaping function index value map.
[0146] In one embodiment, the first dynamic range represents a standard dynamic range and the second dynamic range represents a high dynamic range.
[0147] In one embodiment, each local reshaping function index value of a particular local reshaping function index value is used to look up a corresponding local reshaping function within a family of local reshaping functions.
[0148] In one embodiment, each local reshaping function in a particular local reshaping function is represented by a corresponding lookup table.
[0149] In one embodiment, each local reshaping function in a particular local reshaping function represents a hybrid shifted sigmoid function.
[0150] In one embodiment, each particular local reshaping function reshapes component codewords in a color channel obtained from the second image into locally reshaped component codewords in the color channel.
[0151] In one embodiment, the guided filtering includes a first guided filtering with a relatively small spatial kernel and a second guided filtering with a relatively large spatial kernel.
[0152] In one embodiment, a first guided filtering with a relatively small spatial kernel is used to recover image details associated with edges that are present in the raw camera image but lost in the first image, and a second guided filtering with a relatively large spatial kernel is used to recover image details associated with dark and light regions that are present in the raw camera image but lost in the second image.
[0153] In one embodiment, the intermediate image is fused from a first guided filtered image generated from a first guided filtering using a relatively small spatial kernel and a second guided filtered image generated from a second guided filtering using a relatively large spatial kernel.
[0154] In one embodiment, a smoothed binarized local standard deviation map is generated from the first image, and weighting factors used to fuse the first guided filtered image and the second guided filtered image are determined based at least in part on the smoothed binarized local standard deviation map.
[0155] In one embodiment, the dynamic range mapping represents a tensor product B-spline (TPB) mapping.
[0156] In one embodiment, the dynamic range mapping represents a static mapping specified with optimized operational parameters generated in a model training phase using training images in the first and second dynamic ranges obtained by spatially aligning captured images in the first and second dynamic ranges.
[0157] In one embodiment, a computing device, such as a display device, a mobile device, a set-top box, a multimedia device, or the like, is configured to perform any of the aforementioned methods. In one embodiment, an apparatus comprises a processor and is configured to perform any of the aforementioned methods. In one embodiment, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause the performance of any of the aforementioned methods.
[0158] In one embodiment, a computing device comprises one or more processors and one or more storage media storing a set of instructions that, when executed by the one or more processors, cause the performance of any of the aforementioned methods.
[0159] It should be noted that although separate embodiments are described herein, any combination of the embodiments and / or sub-embodiments described herein may be combined to form further embodiments. Computer system implementation example
[0160] Embodiments of the present invention may be implemented in a computer system, a system comprised of electronic circuits and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA) or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or IC may implement, control, or execute instructions related to adaptive perceptual quantization of images with extended dynamic range, such as those described herein. The computer and / or IC may calculate any of the various parameters or values related to the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0161] Certain implementations of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors in a display, an encoder, a set-top box, a transcoder, etc., may implement the method for adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may be provided in the form of a program product. The program product may include any non-transitory medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention may be in any of a wide variety of forms. The program product may include physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, and electronic data storage media including flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0162] Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise specified, reference to that component (including reference to "means") should be interpreted as including, as an equivalent of that component, any component that performs the function of the described component (e.g., is functionally equivalent), including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the invention.
[0163] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hardwired to perform the techniques, may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) permanently programmed to perform the techniques, or may include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices incorporating hardwired and / or program logic to implement the techniques.
[0164] 5 is a block diagram illustrating a computer system 500 in which embodiments of the present invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled to bus 502 for processing information. Hardware processor 504 may be, for example, a general-purpose microprocessor.
[0165] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions executed by processor 504. Main memory 506 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored on a non-transitory storage medium accessible to processor 504, render computer system 500 a specialized machine customized to perform the operations specified in the instructions.
[0166] Computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic or optical disk, is provided and coupled to bus 502 for storing information and instructions.
[0167] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), allowing the device to specify a position in a plane.
[0168] Computer system 500 may implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, renders or programs computer system 500 into a dedicated machine. According to one embodiment, the techniques described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0169] The term "storage medium" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific fashion. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, or any other memory chip or cartridge.
[0170] Storage media is distinct from but can be used in conjunction with transmission media. Transmission media involves transferring information to and from storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0171] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.
[0172] Computer system 500 also includes a communication interface 518 coupled to bus 502. The communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a LAN card to provide a data communication connection to a compatible LAN local area network (LAN). Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0173] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 528. Local network 522 and Internet 528 both use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are exemplary forms of transmission media.
[0174] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.
[0175] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution. Equivalents, Extensions, Substitutes and Others
[0176] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the claimed embodiments of the invention are and are intended by the applicant to be is the set of claims issuing from this application, in the particular form in which such claims are issued, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall control the meaning of such terms as used in the claims. Accordingly, any limitations, elements, properties, features, advantages, or attributes not expressly recited in a claim should in no way limit the scope of such a claim. Accordingly, the specification and drawings are to be regarded in an illustrative and not a restrictive sense. Enumerated exemplary embodiments
[0177] The present invention may be embodied in any of the forms described herein, including but not limited to the following enumerated exemplary embodiments (EEE), which illustrate the structure, features, and functionality of some portions of embodiments of the present invention.
[0178] EEE1. A method comprising: applying guided filtering to a first image of a first dynamic range using a raw camera image as a guide image to generate an intermediate image of a first dynamic range, the first image of the first dynamic range being generated from the raw camera image using an image signal processor of an image capture device; performing dynamic range mapping on the intermediate image in the first dynamic range to generate a second image in a second dynamic range different from the first dynamic range; generating a specific local reshaping function index value using the second image in the second dynamic range to select a specific local reshaping function for the second image in the second dynamic range; applying the particular local reshaping function to the second image in the second dynamic range to generate a locally reshaped image in the second dynamic range; A method comprising:
[0179] EEE2. The method of EEE1, wherein the particular local reshaping function index values collectively form a local reshaping function index value map.
[0180] EEE3. The method of EEE1 or EEE2, wherein the first dynamic range represents a standard dynamic range and the second dynamic range represents a high dynamic range.
[0181] EEE4. The method of any of EEE1-EEE3, wherein each local reshaping function index value of a particular local reshaping function index value is used to look up a corresponding local reshaping function within a family of local reshaping functions.
[0182] EEE5. The method of any one of EEE1-EEE4, wherein each local reshaping function in a particular local reshaping function group is represented by a corresponding lookup table.
[0183] EEE6. The method of any of EEE1-EEE5, wherein each local reshaping function in a particular local reshaping function range represents a hybrid shifted sigmoid function.
[0184] EEE7. The method of EEE6, wherein the hybrid shifted sigmoid function is constructed by combining, at the midpoint, a first sigmoid function of a first segment length with a first slope located to the left of the midpoint, and a second sigmoid function of a second segment length with a second slope located to the right of the midpoint.
[0185] EEE8. The method of EEE7, wherein each of the particular local reshaping functions is selected from a family of local reshaping functions, wherein each local reshaping function within the family of local reshaping functions is indexed with a respective local reshaping function index value within an index value range, and wherein each local reshaping function within the family of local reshaping functions is generated from a hybrid shifted sigmoid function based on a distinct combination of values of the midpoint, the first segment length, the first slope, the second segment length, and the second slope.
[0186] EEE9. The method of any of EEE1-EEE8, wherein each particular local reshaping function reshapes component codewords in a color channel obtained from the second image into locally reshaped component codewords in the color channel.
[0187] EEE10. The method of any of EEE1-EEE9, wherein the guided filtering includes a first guided filtering with a relatively small spatial kernel and a second guided filtering with a relatively large spatial kernel.
[0188] EEE11. The method of EEE10, wherein a first guided filtering using a relatively small spatial kernel is used to recover image details associated with edges that are present in the raw camera image but lost in the first image, and a second guided filtering using a relatively large spatial kernel is used to recover image details associated with dark and light regions that are present in the raw camera image but lost in the second image.
[0189] EEE12. The method of EEE10 or EEE11, wherein the intermediate image is fused from a first guided filtered image generated from a first guided filtering using a relatively small spatial kernel and a second guided filtered image generated from a second guided filtering using a relatively large spatial kernel.
[0190] EEE13. The method of claim 12, wherein a smoothed binarized local standard deviation map is generated from the first image, and wherein weighting factors used to fuse the first guided filtered image and the second guided filtered image are determined based at least in part on the smoothed binarized local standard deviation map.
[0191] EEE14. The method of any of EEE1 to EEE13, wherein the dynamic range mapping represents a tensor product B-spline (TPB) mapping.
[0192] EEE15. The method of any of EEE1-EEE14, wherein the dynamic range mapping represents a static mapping specified with optimized operating parameters generated in a model training phase using training images in the first dynamic range and the second dynamic range obtained by spatially aligning captured images in the first dynamic range and the second dynamic range.
[0193] EEE16. The dynamic range mapping is represented by a three-dimensional look-up table (3D-LUT) including a plurality of entries that map an input three-dimensional codeword to a predicted three-dimensional codeword in a mapped RGB color space, and the method comprises: converting the predicted RGB codewords in the plurality of entries in the 3D-LUT to corresponding YCbCr codewords in a YCbCr color space; dividing the corresponding YCbCr codewords along a Y dimension of a YCbCr color space into a plurality of Y-dimension partitions; identifying a subset of entries in the 3D-LUT having dark luminance levels based on a plurality of Y-dimension partitions of corresponding YCbCr codewords; adjusting the predicted 3-D codewords in the subset of entries according to the dark intensity level; The method according to any one of EEE1 to EEE15, further comprising:
[0194] EEE17. An apparatus comprising a processor, the apparatus being configured to perform any one of the methods described in EEE1-EEE16.
[0195] EEE18. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing, with one or more processors, a method according to any of the methods set forth in EEE1-EEE16.
[0196] EEE19. A computer system configured to perform any one of the methods described in EEE1 to EEE16.
Claims
1. 1. A method comprising: applying guided filtering to a first image in a first dynamic range using a raw camera image as a guide image to generate an intermediate image in the first dynamic range, the first image in the first dynamic range having been generated from the raw camera image using an image signal processor of an image capture device; performing dynamic range mapping on the intermediate image in the first dynamic range to generate a second image in a second dynamic range different from the first dynamic range; using the second image in the second dynamic range to generate a specific local reshaping function index value for selecting a specific local reshaping function for the second image in the second dynamic range; applying the particular local reshaping function to the second image in the second dynamic range to generate a locally reshaped image in the second dynamic range; A method comprising:
2. The method of claim 1 , wherein the particular local reshaping function index values collectively form a local reshaping function index value map.
3. The method of claim 1 , wherein the first dynamic range represents a standard dynamic range and the second dynamic range represents a high dynamic range.
4. The method of claim 1 , wherein each local reshaping function in the particular local reshaping function represents a hybrid shifted sigmoid function.
5. 5. The method of claim 4, wherein the hybrid shifted sigmoid function is constructed by combining, at a midpoint, a first sigmoid function of a first segment length with a first slope located to the left of the midpoint, and a second sigmoid function of a second segment length with a second slope located to the right of the midpoint.
6. 6. The method of claim 5, wherein each of the particular local reshaping functions is selected from a family of local reshaping functions, each local reshaping function within the family of local reshaping functions being indexed with a respective local reshaping function index value within an index value range, and each local reshaping function within the family of local reshaping functions is generated from the hybrid shifted sigmoid function based on a distinct combination of values of the midpoint, the first segment length, the first slope, the second segment length, and the second slope.
7. 2. The method of claim 1 , wherein the guided filtering comprises first guided filtering with a first spatial kernel having a first kernel size and second guided filtering with a second spatial kernel having a second kernel size larger than the first kernel size.
8. 8. The method of claim 7, wherein the first guided filtering with the first spatial kernel is configured to restore image details associated with edges that are present in the raw camera image but lost in the first image, and the second guided filtering with the second spatial kernel is configured to restore image details associated with dark and light regions that are present in the raw camera image but lost in the second image.
9. 8. The method of claim 7, wherein the intermediate image is fused from a first guided filtered image generated from the first guided filtering using the first spatial kernel and a second guided filtered image generated from the second guided filtering using the second spatial kernel.
10. 10. The method of claim 9, wherein a smoothed binarized local standard deviation map is generated from the first image, and weighting factors used to fuse the first guided filtered image and the second guided filtered image are determined based at least in part on the smoothed binarized local standard deviation map.
11. The method of claim 1 , wherein the dynamic range mapping represents a tensor product B-spline (TPB) mapping.
12. 2. The method of claim 1 , wherein the dynamic range mapping represents a static mapping specified with optimized operational parameters generated in a model training phase using training images in the first and second dynamic ranges obtained by spatially aligning captured images in the first and second dynamic ranges.
13. An apparatus comprising a processor, the apparatus being configured to perform the method of any one of claims 1 to 12.
14. 13. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed on one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 12.
Citation Information
Patent Citations
Adaptive local reshaping for SDR-to-HDR up-conversion
WO2022072884A1