Methods, apparatuses, media for adaptive local reshaping for sdr to hdr upconversion
By employing local shaping techniques, multi-level edge-preserving filtering, and local index value selection of local shaping functions, low dynamic range image data is transformed into high dynamic range image data, solving the problems of insufficient dynamic range and color saturation, and achieving higher dynamic range and local contrast.
Patent Information
- Application Number
- CN202180071928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-02
- Filing Date
- 2021-10-01
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2041-10-01
AI Technical Summary
Existing technologies struggle to effectively convert low dynamic range image data into high dynamic range image data, resulting in insufficient dynamic range and color saturation in the output image.
Local shaping techniques are employed, which estimate local brightness levels through multi-level edge-preserving filtering and select specific local shaping functions using local index values to locally shape the input image, generating high dynamic range image data.
It improves the dynamic range and local contrast of the output image, enhances the vividness of the colors, and maintains the integrity of the visual semantic content.
Smart Images

Figure CN116508324B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 086,699 and European Patent Application No. 20199785.5, both filed on October 2, 2020, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure generally relates to image processing operations. More specifically, embodiments of this disclosure relate to video codecs. Background Technology
[0004] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, brightness) in an image, such as from the darkest black (darkness) to the brightest white (highlight). In this sense, DR relates to the intensity "scene-referred". DR can also refer to the ability of a display device to fully or approximately render a specific breadth of intensity range. In this sense, DR relates to the intensity "display-referred". Unless a particular meaning is explicitly specified to have a specific implication at any point in the description herein, it should be inferred that the terms can be used interchangeably in either sense, for example.
[0005] As used herein, the term "high dynamic range (HDR)" refers to a DR width spanning approximately 14 to 15 or more orders of magnitude across the human visual system (HVS). In practice, the DR, which represents a broad range of intensity that humans can simultaneously perceive relative to HDR, may be slightly truncated. As used herein, the terms "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)" can be associated, individually or interchangeably, with this type of DR: DR that can be perceived within a scene or image by the human visual system (HVS), including eye movements, allowing for some changes in light adaptability on the scene or image. As used herein, EDR can refer to a DR spanning 5 to 6 orders of magnitude. While this may be slightly narrower than HDR relative to a reference real-world scene, EDR indicates a wide DR width and can also be referred to as HDR.
[0006] In fact, an image comprises one or more color components in a color space (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by a per pixel. n Bit precision representation (e.g., n = 8). Using non-linear light intensity coding (e.g., gamma coding), where... nImages with a dynamic range ≤ 8 (e.g., a color 24-bit JPEG image) are considered to have a standard dynamic range, where... n Images with a dynamic range greater than 8 can be considered images with enhanced dynamic range.
[0007] A reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values (e.g., luminance) of the input video signal and the color values (e.g., screen luminance) of the output screen generated by the display. For example, ITU-R BT.1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011), defines a reference EOTF for flat panel displays, the contents of which are incorporated herein by reference in their entirety. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term “metadata” in this document refers to any auxiliary information transmitted as part of the encoded bitstream and used to assist the decoder in rendering the decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as described herein.
[0008] As used herein, the term "PQ" refers to Perceived Luminance Amplitude Quantization. The human visual system responds to increasing light levels in a highly nonlinear manner. The human ability to perceive stimuli is influenced by factors such as the luminance of the stimulus, the size of the stimulus, the spatial frequency constituting the stimulus, and the luminance level to which the eye adapts at a particular moment of viewing the stimulus. In some embodiments, the perceived quantizer function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter referred to as "SMPTE"), which is incorporated herein by reference in its entirety, wherein, given a fixed stimulus size, for each luminance level (e.g., stimulus level, etc.), the minimum visible contrast step size at said luminance level is selected based on the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).
[0009] Supports 200 to 1,000 cd / m³ 2A display with a brightness of nits or nits represents a lower dynamic range (LDR) associated with EDR (or HDR), also known as standard dynamic range (SDR). EDR content can be displayed on an EDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or higher). Such a display can be defined using an alternative EOTF that supports high brightness capabilities (e.g., 0 to 10,000 nits or higher). In SMPTE 2084 and Rec. ITU-R BT. 2100, “ Image parameter values for high dynamic range television for use in production and international program exchange An example of such an EOTF is defined in "[Image Parameter Values for High Dynamic Range Television Used in Production and International Program Exchange]" (06 / 2017). As the inventors understand it herein, an improved technique is desired for converting input video content data into output video content with high dynamic range, high local contrast, and vivid colors.
[0010] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Summary of the Invention
[0011] A method for generating an image of the first dynamic range from an input image of a second dynamic range below a first dynamic range, the method comprising: generating a global index value for selecting a global shaping function for the input image of the second dynamic range using luminance codewords in the input image, wherein the global index value is an L1 median representing an average luminance level predicted by using a multinomial regression model based on the average of the luminance SDR codewords in the input image; applying an image filtering operation to the input image to generate a filtered image, the input image comprising a plurality of regions with up to per-pixel precision, and generating a filtered value of the filtered image for each of the plurality of regions, providing a measure of the local luminance level in that region; and using the global index value and the filtered value of the filtered image to generate a global index value for selecting a global shaping function for each of the plurality of regions of the input image. Determine local index values for a local shaping function, wherein each local index value is a local L1 median representing the average luminance level of the region, wherein generating a local index value for each region includes: determining the difference between a local luminance level estimated based on a filtered value generated for the region and the luminance level of an individual pixel in the region; estimating a local L1 median adjustment value based on the determined difference; predicting a local L1 median value based on the global index value and the local L1 median adjustment value, the local L1 median value representing the local index value; such that at least in part, a shaped image of the first dynamic range is generated by shaping the input image using the specific local shaping function selected using the local index value, wherein the specific local shaping function selected using the local index value shapes the luminance SDR codewords in the input image into luminance HDR codewords in the shaped image. Attached Figure Description
[0012] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:
[0013] Figure 1 An example process of a video transmission pipeline is described;
[0014] Figure 2A The diagram illustrates an example workflow for applying local shaping operations; Figure 2B The diagram illustrates an example framework or architecture for the upconversion process, which converts an SDR image to an HDR image through local shaping operations. Figure 2C The diagram illustrates an example flow for applying multi-level edge-preserving filtering;
[0015] Figure 3A The diagram illustrates an example backward integer function; Figure 3B The diagram illustrates the example basic integer function after adjustments and modifications; Figure 3CThe example least squares solution is illustrated. Figure 3D The diagram illustrates example global and local integer functions; Figure 3E and Figure 3F The illustration shows an example of a local integer function; Figure 3G The illustration shows an example nonlinear function used to adjust a linear regression model for predicting L1 median;
[0016] Figure 4 The example process flow is illustrated; and
[0017] Figure 5 A simplified block diagram of an example hardware platform is shown, on which a computer or computing device as described herein can be implemented. Detailed Implementation
[0018] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent that this disclosure may be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing this disclosure.
[0019] The local shaping techniques described herein can be implemented to (backward) shape or upconvert image data with a relatively narrow dynamic range, such as SDR image data, into image data with a higher dynamic range, such as HDR image data with enhanced local contrast and color saturation.
[0020] As used in this article, "upconversion" or "(backward) shaping" refers to the transformation of image data with a lower dynamic range into image data with a higher dynamic range through shaping operations such as local shaping operations under the techniques described in this article, global shaping operations under some other methods, etc.
[0021] Global reshaping refers to an up-conversion or backward reshaping operation that applies the same global reshaping function / mapping to all pixels of an input image (such as an input SDR image) to generate a corresponding output image—such as a reshaped HDR image—that depicts the same visual semantic content as the input image.
[0022] For example, the HDR luminance or luminance codeword in a shaped HDR image can be constructed or up-transformed by applying the same global shaping function—such as the same 8-segment second-order polynomial or the same backward lookup table (BLUT)—to the SDR luminance or luminance codeword of all pixels in the input SDR image.
[0023] Similarly, HDR chroma or chroma codewords in a shaped HDR image can be constructed or upconverted by applying the same global shaped mapping—such as the same backward multivariate multiple regression (backward MMR or BMMR) mapping specified by a set of MMR coefficients—to the SDR codewords of all pixels in the input SDR image (in both the luminance and chroma channels).
[0024] Example backward reshaping operations are described in U.S. Provisional Patent Application Serial No. 62 / 136,402, filed March 20, 2015 (also published January 18, 2018, as U.S. Patent Application Publication Serial No. 2018 / 0020224) and PCT Application Serial No. PCT / US2019 / 031620, filed May 9, 2019, the entire contents of which are incorporated herein by reference as fully set forth herein.
[0025] Local shaping, compared to global shaping which applies the same shaping function or mapping to all pixels of an input image, refers to applying different shaping functions or mappings to different pixels of the input image, performing up-conversion or backward shaping operations. Therefore, in local shaping, the first shaping function applied to the first pixel of the input image can be a different function than the second shaping function applied to a second different pixel of the input image.
[0026] Specific shaping functions can be selected or identified for specific pixels in an input image that have local brightness levels in local regions containing specific pixels. Local brightness levels can be estimated using image filtering, such as multi-level edge-preserving filtering of a guide image. Under the techniques described herein, image filtering for estimating or predicting local brightness levels can be performed in a manner that minimizes the leakage of information (e.g., pixel values, codewords, etc.) between different visual objects / characters / regions / fragments, with the aim of reducing or preventing visual artifacts such as halo artifacts.
[0027] Local reshaping, as described in this paper, can take into account the local image characteristics of the input image. Different reshaping functions or mappings can be located at each pixel of the input image to enhance the local contrast and color saturation (level) in the output image, and make the overall output image have a higher local contrast ratio, better viewer-perceived image detail, more vibrant colors, etc.
[0028] The example embodiments described herein relate to generating an image of the first dynamic range from an input image of a second dynamic range, which is below the first dynamic range. A global index value is generated for selecting a global shaping function for the input image of the second dynamic range. The global index value is generated using luminance codewords in the input image. Image filtering is applied to the input image to generate a filtered image. The filtered values of the filtered image provide a measure of local luminance levels in the input image. Local index values are generated for selecting a specific local shaping function for the input image. The local index values are generated using the global index value and the filtered values of the filtered image. The generated image of the first dynamic range is at least partially achieved by shaping the input image using a specific local shaping function selected with the local index values.
[0029] Example video transmission and processing pipeline
[0030] Figure 1 An example process of a video transmission pipeline (100) is depicted, illustrating the various stages from video capture / generation to HDR or SDR display. The example HDR display may include, but is not limited to, image displays operating in conjunction with TVs, mobile devices, home theaters, etc. The example SDR display may include, but is not limited to, SDR TVs, mobile devices, home theater displays, head-mounted displays, wearable displays, etc. It should be noted that the SDR to HDR upconversion can be performed on the encoder / server side (before video compression) or the decoder / playback side (after video decompression). To support SDR to HDR upconversion on the playback side, alternatives to [other methods] can be used. Figure 1 Different system configurations besides the system configuration described. In these different system configurations, additional systems may be used besides... Figure 1 The processing unit described uses a different image metadata format than the image metadata format it employs to transmit image metadata.
[0031] In a preferred embodiment of the invention, the image metadata includes L1 metadata. As used herein, the term "L1 metadata" refers to one or more of the minimum (L1 minimum) luminance value, intermediate (L1 intermediate) luminance value, and maximum (L1 maximum) luminance value associated with a specific portion of the video content (e.g., an input frame or image). The L1 metadata is associated with the video signal. To generate the L1 metadata, pixel-level frame-by-frame analysis of the video content is preferably performed at the encoding end. Alternatively, the analysis can be performed at the decoding end. The analysis describes the distribution of luminance values over a defined portion of the video content covered by the analysis process (e.g., a single frame or a series of frames such as a scene). The L1 metadata can be calculated during analysis covering a single video frame and / or a series of frames such as a scene. The L1 metadata may include various values obtained during the analysis process that together form the L1 metadata, which are associated with the corresponding portion of the video content from which the L1 metadata is calculated and are associated with the video signal. L1 metadata includes at least one of the following: (i) an L1 minimum representing the lowest black level in a corresponding portion of the video content, (ii) an L1 median representing the average luminance level in a corresponding portion of the video content, and (iii) an L1 maximum representing the highest luminance level in a corresponding portion of the video content. Preferably, the L1 metadata is generated and appended to each video frame and / or each scene encoded in the video signal. L1 metadata can also be generated for regions of the image referred to as local L1 values. L1 metadata can be calculated by converting RGB data to a luminance-chrominance format (e.g., YCbCr) and then calculating at least one or more of the minimum, median (average), and maximum values in the Y plane, or L1 metadata can be calculated directly in the RGB space.
[0032] In some embodiments, the minimum L1 value represents the PQ encoding of a corresponding portion of the video content (e.g., a video frame or image). min The minimum (RGB) value is considered, while only considering the active area (e.g., by excluding gray or black bars, black borders in video, etc.). min (RGB) represents the minimum value of the pixel's color component values {R, G, B}. The L1 median and L1 maximum values can also be calculated in the same way. Specifically, in this embodiment, the L1 median represents the PQ encoding of the image. max The average of the (RGB) values, and the L1 maximum value represents the PQ encoding of the image. max The maximum value of (RGB) values, where, max (RGB) represents the maximum value of the pixel's color component values {R, G, B}. In some embodiments, L1 metadata can be normalized to [0, 1].
[0033] For illustrative purposes only, Figure 1Used to illustrate or depict the SDR-HDR upconversion process performed on the server side using local backward shaping techniques as described herein. Figure 1 The encoder / server-side SDR-to-HDR upconversion illustrated in the diagram uses the input SDR image to generate a locally shaped HDR image. The combination of the input SDR image and the locally shaped HDR image can be used by either a backward-compatible or non-backward-compatible codec to generate a backward-compatible or non-backward-compatible SDR video signal. In some operational scenarios, such as... Figure 1 As illustrated, the video signal can be encoded using a shaped SDR image generated by forward-shaping a locally shaped HDR image.
[0034] HDR image generation block 105 can receive video frames such as a sequence of continuously input SDR images 102. These SDR images (102) can be received from a video source or obtained from video data storage. Some or all of the SDR images (102) can be generated from source images, for example, through video editing or transformation operations (e.g., automatic without human input, manual, automatic with human input, etc.), color grading operations, etc. The source images can be captured digitally (e.g., by a digital camera), generated by converting analog camera images captured on film into digital format, generated by a computer (e.g., using computer animation, image rendering, etc.), etc. The SDR images (102) can be images associated with one or more of the following: film releases, archived media programs, media program libraries, video recording / editing, media programs, TV programs, user-generated video content, etc.
[0035] The HDR image generation block (105) applies a local shaping operation to each SDR image in the sequence of consecutive SDR images (102) to generate a corresponding (shaped) HDR image in the corresponding consecutive (shaped) HDR image sequence. The HDR image depicts the same visual semantic content as the SDR image (102), but has a higher dynamic range, more vivid colors, etc. compared to the SDR image (102).
[0036] The parameter generation block 142 generates specific values for at least some of the operational parameters used in the local shaping operation based on a predictive model used to predict operational parameter values. The predictive model can be trained using training images such as HDR-SDR image pairs from a training dataset and real data associated with the training images or image pairs.
[0037] The families of shaping functions for the luminance or Y channel and the families of shaping maps for the chrominance channel can be generated by the shaping map generation block 146, and are preloaded into the image generation block 105 by the shaping map generation block 146, for example, during the system startup period before the SDR image (102) is processed to generate the shaped HDR image. The families of shaping functions for the luminance or Y channel may include multiple BLUTs for multiple different L1 medians (or indexed by them). The families of shaping functions for the chrominance channel may include multiple BMMR maps for the same multiple different L1 medians (or indexed by them).
[0038] For each input SDR image, the HDR image generation block (105) generates or computes a local luminance level up to per-pixel precision based on the luminance or Y-channel codeword in the input SDR image. In some operational scenarios, a global filtered image can be generated from the luminance or Y-channel codeword of the input SDR image, for example, as a weighted sum of filtered images generated by multi-level edge-preserving filtering. The filtered values in the global filtered image can be used to estimate or approximate the local luminance level, and then the local luminance level is used as part of the input to estimate or predict the local L1 median up to per-pixel precision. The local L1 median can be represented in an L1 median map and used as an index to select a local shaping function or BLUT for the luminance channel from a family of BLUTs with up to per-pixel precision. These local shaping functions or BLUTs selected with the local L1 median provide a higher local slope or a higher local contrast ratio. Additionally, optionally, or alternatively, the local L1 median represented in the L1 median map can be used as an index to select a local shaping map or BMMR for the chrominance channel from a family of BMMR maps with up to per-pixel precision.
[0039] Therefore, based on the individual local median mappings generated for the input SDR image (102) and the preloaded BLUT and BMMR families, the HDR image generation block (105) can perform local shaping operations on the input SDR image (102) to generate a corresponding shaped HDR image with higher dynamic range, higher local contrast, more vivid colors, etc. compared to the SDR image (102).
[0040] Some or all of the SDR image (102) and the shaped HDR image can be provided to the synthesizer metadata generation block 115 to generate a shaped SDR image 112 that can be encoded more efficiently than the input SDR image (102) by forward shaping the shaped HDR image and to generate image metadata 177 (e.g., synthesizer metadata, etc.). The image metadata (177) may include synthesizer data for generating a backward-shaping map (e.g., BLUT, backward-shaping function / curve or polynomial set, MMR coefficients, etc.), which generates the corresponding HDR image when applied to the input SDR image.
[0041] The shaped SDR image (112) and image metadata (177) can be encoded by a block of code 120 (e.g., a coded bitstream, etc.) or a set of consecutive video segments in the video signal 122. Given the video signal (122), as part of the internal processing or post-processing of the video signal (122) on the device, a receiving device such as a mobile phone can decide to use the metadata with the SDR image data to generate and render an image with a higher dynamic range (such as HDR) and more vibrant colors within the display capabilities of the receiving device. Additionally, optionally, or alternatively, the video signal (122) or video segments allow backward compatibility with conventional SDR displays that can ignore the image metadata (177) and simply display the SDR image.
[0042] Example video signals or video clips may include, but are not limited to, single-layer video signals / segments. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray and other transmission formats, to generate video signals (122) or video clips.
[0043] The video signal (122) or video segment is then transmitted downstream to receivers such as mobile devices, tablets, decoding and playback devices, media source devices, media streaming client devices, televisions (e.g., smart TVs), set-top boxes, cinemas, etc. In the downstream device, the video signal (122) or video segment is decoded by the decoding block (130) to generate a decoded image 182, which may be similar to or identical to a shaped SDR image (112) subjected to quantization errors and / or transmission errors and / or synchronization errors and / or errors caused by packet loss during compression performed by the encoding block (120) and decompression performed by the decoding block (130).
[0044] In a non-limiting example, the video signal (122) (or video clip) may be a backward-compatible SDR video signal (or video clip). Here, "backward-compatible" means a video signal or video clip carrying an SDR image optimized for SDR displays (e.g., preserving specific artistic intent).
[0045] The decoding block (130) can also obtain or decode image metadata (177) from the video signal (122) or video segment. The image metadata (177) specifies a back-shaping map, which can be used by a downstream decoder to perform back-shaping on the decoded SDR image (182) to generate a back-shaped HDR image for rendering on an HDR (e.g., target, reference, etc.) display. The back-shaping map represented in the image metadata (177) can be generated by the compositor metadata generation block (115) by minimizing the error or difference between the back-shaped HDR image generated using the image metadata (177) and the shaped HDR image generated using local shaping operations. Therefore, the image metadata (177) helps ensure that the back-shaped HDR image generated by the receiver using the image metadata (177) is relatively close to and accurately approximates the shaped HDR image generated using local shaping operations.
[0046] Additionally, optionally, or alternatively, the image metadata (177) may include display management (DM) metadata, which may be used by a downstream decoder to perform display management operations on the back-shaped image to generate a display image optimized for rendering on an HDR display device (e.g., an HDR display image, etc.).
[0047] In an operational scenario where the receiver operates with (or is attached to) an SDR display 140 that supports a standard dynamic range or a relatively narrow dynamic range, the receiver can render a decoded SDR image directly or indirectly on the target display (140).
[0048] In an operational scenario where the receiver operates with (or is attached to) an HDR display 140-1 that supports high dynamic range (e.g., 400 nits, 1000 nits, 4000 nits, 10000 nits or higher), the receiver can extract compositor metadata from the video signal (122) or video segment (e.g., metadata containers therein) and use the compositor metadata to synthesize an HDR image (132). The HDR image can be a backward-shaped image generated by backward-shaping an SDR image based on the compositor metadata. Alternatively, the receiver can extract DM metadata from the video signal (122) or video segment and apply DM operations (135) to the HDR image (132) based on the DM metadata to generate a display image (137) optimized for rendering on the HDR display device (140-1), and render the display image (137) on the HDR display device (140-1).
[0049] For illustrative purposes only, it has been described that the local shaping operations described herein can be performed by an upstream device, such as a video encoder, to generate shaped HDR images from SDR images. These shaped HDR images are then used by the video encoder as target or reference HDR images to generate back-shaped metadata, which helps the receiving device generate a back-shaped HDR image that is relatively close to or accurately approximates the shaped HDR image generated by the local shaping operations.
[0050] It should be noted that in various embodiments, some or all of the local shaping operations may be performed by a separate video encoder, a separate video decoder, a separate video transcoder, or a combination thereof. For example, a video encoder may generate a local L1 median map including an index of a shaping function / mapping for an SDR image. Additionally, optionally, or alternatively, a playback-side video decoder or a video transcoder between the video encoder and video decoder may generate a local L1 median map including an index of a shaping function / mapping for an SDR image. The video encoder may not apply local backward shaping operations to generate an HDR image. The video encoder may defer local backward shaping operations so that the HDR image is generated by the video transcoder or video decoder at a later time. The local L1 median map may be included by the video encoder as part of the image metadata encoded with the SDR image in the video signal or video segment. The video transcoder or video decoder may preload using the BLUT family and / or BMMR mapping family. The local L1 median map of the SDR image may be used by the video transcoder or decoder to search for or find a specific shaping function or mapping in the BLUT family and / or BMMR mapping family. The video transcoder or decoder can then perform local shaping operations on the SDR image with up to per-pixel precision to generate a shaped HDR image, at least in part, based on a specific shaping function / mapping that utilizes an index lookup in the local L1 median map. In some operational scenarios, once the HDR image is generated through local shaping, the video encoder can encode the HDR image or a version of the HDR image derived from the locally shaped HDR image into the video signal (e.g., the underlying layer of the video signal), instead of encoding the SDR image into the video signal (e.g., the underlying layer of the video signal). The HDR image decoded from this video signal can then be viewed directly on an HDR display.
[0051] Local plastic surgery
[0052] Figure 2A The illustration shows an example flow for applying local shaping operations (e.g., 2^12, etc.) to generate a corresponding HDR image from an input SDR image. The image processing system, and the coded blocks therein (e.g., Figure 1 (e.g., 105, etc.) or the decoding blocks therein (e.g., Figure 1The process flow can be implemented or executed by functions such as 130, etc. In some operating scenarios, given an input SDR image 202, a global shaping function 204 is selected to determine the main HDR appearance of the output HDR image 214 to be generated from the input SDR image (202) through local shaping operations.
[0053] Multiple fundamental shaping functions, such as an 8-segment second-order polynomial for the luminance channel and the MMR for the chrominance channel, can be determined at least partially (e.g., for codewords in the luminance or luminance channels) via a multinomial regression model and an MMR framework based on a training dataset consisting of a population of image pairs including training HDR images and training SDR images.
[0054] In some operational scenarios, the (backward) shaping function used to generate an HDR image from an SDR image can be specified by L1 metadata. L1 metadata can include three parameters, such as the L1 maximum, L1 median, and L1 minimum, which can be obtained from the HDR codewords of the HDR image (e.g., RGB codewords, YUV codewords, YCbCr codewords, etc.). The L1 maximum represents the highest brightness level in the video frame. The L1 median represents the average brightness level on the video frame. The L1 minimum represents the lowest darkness level in the video frame. Alternatively or additionally, (backward) shaping can be performed based on L1 metadata associated with the scene to which the current frame belongs. One or more of these parameters specify the shaping function. For example, the L1 maximum and L1 minimum may be disregarded, but the L1 median identifies or specifies the shaping function. According to the invention, one or more parameters from the L1 metadata—preferably the L1 median—are used to identify the global shaping function and are referred to as the global index value. Examples of the construction of forward and backward reshaping functions are described in U.S. Provisional Patent Application Serial No. 63 / 013,063, filed April 21, 2020, “Reshaping functions for HDR imaging with continuity and reversibility constraints” by GM. Su and U.S. Provisional Patent Application Serial No. 63 / 013,807, filed April 22, 2020, “Iterative optimization of reshaping functions in single-layer HDR image codec” by GM. Su and H. Kadu, the contents of which are fully set forth herein and are incorporated herein by reference in their entirety.
[0055] A specific basic shaping function can be selected as the global shaping function (204) from multiple basic shaping functions using the global (e.g., global, per image / frame, etc.) L1 median determined, estimated, or predicted based on the input SDR image (202).
[0056] The goal of applying local shaping operations at the per-pixel level to generate an HDR image (214) from an input SDR image (202) is to enhance the local contrast ratio in local regions of the HDR image (214) without changing the local brightness levels of these local regions of the HDR image (214), thereby maintaining the main HDR appearance of the HDR image (214) determined by the global shaping function (204).
[0057] To help prevent or reduce common artifacts such as halo artifacts that may arise from altering the local brightness levels of edges / boundaries of neighboring visual objects / characters, a high-precision filter, such as a multi-level edge-preserving filter 206, can be applied to the input SDR image (202) to form a filtered image. Alternatively, a global shaping function can be selected using the overall (e.g., global, per-image / frame, etc.) L1 median estimated or predicted based on the filtered input SDR image, rather than based on the L1 median obtained from the (unfiltered) input SDR image. The filtered image can be used to obtain or estimate the local brightness levels (or region-specific brightness levels) in different local regions (up to per-pixel precision or up to the local region around each pixel) of the input SDR image (202). The estimated local brightness levels in different local regions (up to per-pixel precision or up to the local region around each pixel) of the input SDR image (202) can be used to estimate or predict the local L1 median (up to per-pixel precision or up to the local region around each pixel) in the HDR image (214). Local L1 metadata describes the distribution of luminance values over a region surrounding a pixel. A region can be defined as a single pixel, thus supporting per-pixel precision when applying local shaping. The local L1 maximum value represents the highest luminance level in the region. The local L1 median value represents the average luminance level in the region. The local L1 minimum value represents the lowest darkness level in the region.
[0058] The high-precision filter (206) used to estimate the local brightness level of the SDR image (202)—which can then be used as input to a prediction model to predict the local L1 median in the HDR image (214)—can be specifically selected or tuned to avoid or reduce information leakage (or pixel value or codeword information diffusion) of edges / boundaries of visual objects / characters between different visual objects / characters, between visual objects / characters and background / foreground, and adjacent to those depicted in the input SDR image (202) and / or to be depicted in the output HDR image (214).
[0059] To improve the efficiency or response time of local shaping operations, a family of (multiple) local shaping functions 208 can be loaded or constructed in an image processing system as described herein (e.g., initially, pre-processed, etc.) during the system startup period. The loading or construction of the family of (multiple) local shaping functions (208) can be performed before shaping or up-converting a sequence of consecutive input SDR images (e.g., including an input SDR image (202), etc.) into a corresponding sequence of consecutive output HDR images (e.g., including an HDR image (214), etc.) based at least in part on the family of (multiple) local shaping functions (208). The family of (multiple) local shaping functions (208) can be generated, but is not limited to, from a plurality of basic shaping functions used to select a global shaping function (204) (e.g., by extrapolation and / or interpolation). Each local integer function in a family of (multiple) local integer functions (208) can be indexed or identified, in whole or in part, using a corresponding value (preferably a corresponding local L1 median) from the local L1 metadata, as in the case of the global integer function (204) or each of a plurality of basic integer functions selected from the global integer function (204). According to the invention, one or more parameters (preferably local L1 medians) from the L1 metadata are used to identify the local integer function and are referred to as local index values.
[0060] As mentioned, the filtered image generated by applying multi-level edge-preserving filtering (206) to the input SDR image (202) can be used to generate or estimate local brightness levels, which can be used in combination with or referenced to global shaping functions in the prediction model to generate or predict local L1 medians. These local L1 medians form an index of the local L1 median map 210, which can be used to look up pixel-specific local shaping functions in the family of (multiple) local shaping functions (208).
[0061] Pixel-specific local shaping functions, found by indexing in the local L1 median map (210), can be applied per-pixel level to the input SDR image (202) via local shaping operations (represented as 212) to back-shape the SDR codewords in the luminance and chrominance channels into shaped HDR codewords for the HDR image (214) in the luminance and chrominance channels. Each of these pixel-specific local shaping functions, which are fine-tuned (up to per-pixel precision) nonlinear functions, can be used to enhance local contrast and / or saturation in the HDR image (214).
[0062] Figure 2BAn example framework or architecture is illustrated for an upconversion process that transforms an input SDR image into an output or shaped HDR image via a local shaping operation (e.g., 212, etc.). The framework or architecture fuses multiple levels of filtered images generated by multi-level edge-preserving filtering of the input SDR image (e.g., 202, etc.) into a global filtered image. The filtered values in the global filtered image are used as predictions, estimates, and / or proxies for local brightness levels. These filtered values can then be used as part of the input to a prediction model to generate or predict local L1 medians with up to per-pixel precision. The local L1 medians serve as indices for selecting or finding specific local shaping functions with up to per-pixel precision. For the purpose of generating the corresponding HDR image (e.g., 214, etc.), the selected local shaping function can be applied to the input SDR image via the local shaping operation (212). The framework or architecture may employ a small number of (e.g., main, etc.) components to perform image processing operations related to the local shaping operation (212).
[0063] More specifically, multiple integer functions, such as basic integer functions, can be constructed first. In some operational scenarios, these basic integer functions can correspond to different L1 medians and can be indexed using L1 medians, such as twelve different L1 medians that are uniformly or non-uniformly distributed in some or all of the entire HDR codeword space or range (e.g., for a 12-bit HDR codeword space or range of 4096, etc.).
[0064] Given an input SDR image (202), a multinomial regression model (e.g., expression (3) below, etc.) can be used to generate or predict a global L1 median from the average value of the luminance or Y-channel SDR codewords in the input SDR image (202). This global L1 median can be used to generate (e.g., using approximation, interpolation, etc.) or to select a specific shaping function (e.g., 204, etc.) from a plurality of basic shaping functions as the global shaping function for the input SDR image (202). This global shaping function provides or represents the primary HDR appearance of the HDR image (e.g., 214, etc.). This primary HDR appearance can be preserved in a locally shaped HDR image (214) generated from the input SDR image (202) through a local shaping operation (212).
[0065] Additionally, multi-level edge-preserving filtering (206) can be applied to generate multiple filtered images (e.g., 206-1 to 206-4) for multiple different levels (e.g., levels 1 to 4). A guide image filter can be used to perform or implement edge-preserving filtering at each of the multiple levels, using a guide image of a single dimension / channel or multiple dimensions / channels to guide the filtering of the input SDR image (202). A total filtered image as a weighted sum 224 can be generated from the multiple filtered images generated for multiple levels through edge-preserving filtering (e.g., 206-1 to 206-4, etc.).
[0066] Given a local luminance level estimated or approximated by filtered values in the overall filtered image with up to per-pixel accuracy, the difference between the local luminance level and the luminance or brightness value of an individual pixel can be determined with up to per-pixel accuracy using the subtraction / difference operator 222.
[0067] Differences up to per pixel precision (or individual differences for each pixel in the input SDR image (202)) can be used to estimate, for example, a desired local L1 median adjustment 218 with enhancement level 216 to generate L1 median adjustments up to per pixel precision (or individual L1 median adjustments for each pixel in the input SDR image (202)). The L1 median adjustment can be further modified using non-linear activation functions to generate modified L1 median adjustments up to per pixel precision (or individual modified L1 median adjustments for each pixel in the input SDR image (202)).
[0068] The prediction model for local L1 median generation / prediction (e.g., represented by the following expression (6) etc.) can be used to generate local L1 medians (or single local L1 medians for each pixel in the input SDR image (202) with up to per-pixel precision based on the modified L1 median adjustment and global L1 median (204).
[0069] These local L1 medians (210) can be collectively represented as an L1 median map. Each of the local L1 medians (210) can be used as an index value for selecting a local shaping function / map from the family of shaping functions / maps for the corresponding pixel in the input SDR image (202). Thus, multiple local shaping functions / maps can be selected using these local L1 medians (210) as index values. The local shaping operation (212) can be performed by applying local shaping functions / maps to the input SDR image (202) to generate a shaped HDR image (214).
[0070] Basic and non-basic integer functions
[0071] Figure 3AFor use in setting L1 median for multiple targets The form of the backward shaping function (also known as a backward lookup table or BLUT) illustrates an example basic backward shaping function used for SDR to HDR upconversion as described herein. As mentioned, the basic backward function can be constructed or built using a training dataset with different target L1 medians / settings. The global shaping function (204) used to inform or represent the HDR appearance of the input SDR image (202) can be selected from or linearly interpolated from these basic shaping functions based on the global L1 median determined or predicted for the input SDR image (202).
[0072] In some operational scenarios, these basic shaping functions used for multiple target L1 medians / settings can be used to further construct other shaping functions for other L1 medians / settings. Among these shaping functions, given the same input SDR codeword values, the higher the L1 median / setting of the basic shaping function, the higher the mapped HDR codeword value generated by the basic shaping function from the input SDR codeword value. This translates to a brighter output HDR appearance in the HDR image that includes the mapped HDR codeword value.
[0073] The input SDR image (202) and the globally shaped HDR image are represented as follows: and The globally shaped HDR image is globally shaped from the input SDR image (202), and represents the HDR appearance of the HDR image (214) locally shaped from the input SDR image (202). The luminance or Y channel of the input SDR image and the globally shaped HDR image are represented as follows: and In the given input SDR image (202), the first... The SDR luminance or Y codeword value of each pixel (represented as...) In the case of ), and given the corresponding representation as The L1 median / set or the backward integer function indexed by it (represented as In the case of (), the following can be given the (global reshaping) HDR image of the th The corresponding output HDR luminance or Y codeword value for each pixel (represented as...) ):
[0074] (1)
[0075] Backward Integer Function (or basic integer functions) can be pre-computed and stored as backward lookup tables (BLUTs) in image processing systems as described in this paper. For bit depths of... Input SDR codeword, Having 0, 1, ... indivual Entries. In In operational scenarios, basic integer functions It has 1024 entries.
[0076] As mentioned, the image processing system can be preloaded using multiple basic integer functions. In operation scenarios where the L1 median / settings corresponding to the multiple basic integer functions have 12-bit precision, these L1 median / settings occupy multiple different values distributed within the 12-bit range of
[04095] .
[0077] By way of examples rather than limitations, several basic integer functions include those for 12 L1 medians / sets. The 12 basic integer functions. 12 L1 median / settings. Each L1 median / setting corresponds to a specific basic shaping function or curve among the 12 basic shaping functions. In this example, initially after system startup, 12 basic 1024-entry tone curves are available (represented as follows). ()).
[0078] Shape functions, such as forward and backward shaping functions, can be trained or obtained based on content mapping (CM), tone mapping (TM), and / or display management (DM) algorithms / operations performed on training images in a training dataset that includes HDR-SDR training image pairs. CM / TM / DM algorithms or operations may or may not provide disjoint shaping functions. Figure 3A As illustrated in the figure, the basic integer functions can include a number of integer functions in the low L1 median range, which may intersect between or within different integer functions.
[0079] The basic shaping function can be adjusted or modified so that local shaping (or local tone mapping) can achieve higher local contrast and a better, more consistent HDR look than can be achieved using global shaping. Figure 3B The illustration shows the example basic integer function after adjustments and modifications.
[0080] In some operational scenarios, extrapolation can be used to adjust or modify the basic integer functions to ensure that they do not intersect with each other. To ensure that the basic integer functions do not intersect with each other, for any given SDR codeword... The basic integer function can be adjusted or modified to satisfy the constraint of monotonically increasing L1 median / set for the same given SDR codeword, as shown below:
[0081] (2)
[0082] To achieve the monotonically increasing property represented in expression (2) above, the pre-adjusted or pre-modified basic integer function of the intersection of the low L1 medians can be replaced by the adjusted or modified basic integer function generated by extrapolating the pre-adjusted or pre-modified basic integer function.
[0083] For illustration purposes only, assume L1 median First The basic shaping curves (intersecting and) will be replaced. The next two basic tone curves ( and This can be used to perform extrapolation (e.g., linear, etc.) using the L1 median distance ratio (or the ratio of the two differences / distances of the L1 medians). Example procedures for such extrapolation are illustrated in Table 1 below.
[0084] Table 1
[0085]
[0086] like Figure 3A As illustrated, when the L1 median is <= 1280, the original basic shaping function / curve or BLUT trained from the training HDR-SDR images does not monotonically increase relative to the L1 median. Therefore, the first four basic shaping functions / curves or BLUTs (where the L1 medians are 512, 768, 1024, and 1280, respectively) can be replaced by adjusted or modified basic shaping functions / curves through linear extrapolation of the two subsequent basic shaping functions / curves with L1 medians of 1536 and 1792.
[0087] like Figure 3B The basic shape function / curve illustrated in the figure can be extrapolated. Figure 3A It is generated using basic integer functions / curves. For example, in Figure 3B As can be seen, given the SDR codeword, Figure 3B The basic integer functions satisfy the condition or constraint that the mapped or shaped HDR codewords are monotonically increasing relative to the L1 median. It should be noted that a more general cleanup procedure can also be used or implemented compared to the procedure illustrated in Table 1 above. This cleanup procedure can implement program logic to detect intersecting (or non-monotonically increasing relative to the L1 median) basic integer functions, curves, or BLUTs in all input basic integer functions, curves, or BLUTs, and replace those intersecting basic integer functions, curves, or BLUTs by extrapolation and / or interpolation from adjacent or encoded nearest-neighbor basic integer functions, curves, or BLUTs.
[0088] If applicable, non-basic integer functions, curves, or BLUTs can be generated from adjusted or modified basic integer functions, curves, or BLUTs. In some operational scenarios, non-basic integer functions, curves, or BLUTs can be generated from basic integer functions, curves, or BLUTs by performing bilinear interpolation from the most recent basic integer function, curve, or BLUT. Since the basic integer function, curve, or BLUT is already monotonically increasing, the interpolated non-basic integer function, curve, or BLUT also inherits this monotonically increasing property. Table 2 below illustrates example programs for generating non-basic integer functions, curves, or BLUTs from basic integer functions, curves, or BLUTs using bilinear interpolation.
[0089] Table 2
[0090]
[0091] In many operational scenarios, the basic and non-basic shaping functions generated by extrapolation and interpolation as described herein can enable local shaping operations (212) to provide higher local contrast in (locally shaped) HDR images (214) than those provided in HDR images using pre-adjusted or pre-modified shaping functions. Using adjusted or modified shaping functions, dark areas in locally shaped HDR images (214) have lower codeword values, indicating or achieving a higher contrast ratio in locally shaped SDR images (214).
[0092] Global Integer Function Selection
[0093] To determine or achieve the overall (e.g., best, desired, target, etc.) HDR appearance of the HDR image (214) to be generated from the SDR image (202) through local shaping, the global L1 median (denoted as ) can first be predicted using SDR features (one or more feature types) extracted from the input SDR image (202) based on a multinomial regression model for global L1 median prediction. Global L1 median This can then be used to search for or identify the corresponding global integer function representing the overall HDR appearance (denoted as...). ).
[0094] In some operational scenarios, for the sake of model efficiency and robustness, the multinomial regression model used for global L1 median prediction can be a first-order multinomial (linear) model, as shown below:
[0095] (3)
[0096] in, The average value of the SDR codewords of the input SDR image (202) in the luminance or Y channel, wherein the SDR codewords can be normalized within a specific value range such as [0, 1]. and This represents the model parameters of the multinomial regression model.
[0097] The multinomial regression model for global L1 median prediction can be trained by extracting SDR features (of the same type as those used in the actual prediction operation) from a training dataset consisting of multiple HDR-SDR image pairs, as described in this paper, to obtain the model parameters. and The optimal value.
[0098] In some operational scenarios, to avoid any deviations caused by letter boxes that may exist in the visual semantic content depicted in the SDR image (202)—which could ultimately affect the overall HDR appearance—when acquiring or calculating the luminance or Y channel average value, You can exclude such letter boxes (if they exist).
[0099] The exclusion of letter boxes as described in this article can be performed during the training operation / process of the multinomial regression model and during the actual prediction operation / process based on the multinomial regression model.
[0100] When training a regression model with HDR-SDR image pairs, the training SDR image in each HDR-SDR image pair can be used to extract SDR features, while the training HDR image in the HDR-SDR image pair can be used as the target. The shaped HDR image generated or predicted by back-shaping the SDR image using a global shaping function predicted / estimated by the regression model is compared with the target.
[0101] Similarity metrics (e.g., cost, error, quality metrics, etc.) can be used to compare shaped HDR images with training HDR images. In some operational scenarios, peak signal-to-noise ratio (PSNR), which may empirically have a relatively strong correlation with the L1 median, can be used as a similarity metric.
[0102] Represent the HDR codewords of the (original) training HDR image in the luminance or Y channel as follows: Further, the HDR codewords of the shaped HDR image in the luminance or Y channel are represented as follows: The PSNR of HDR codewords in the luminance or Y channel between the (original) training HDR image and the shaped HDR image can be used to improve or optimize model parameters. and The purpose is to determine this, as illustrated in Table 3 below.
[0103] Table 3
[0104]
[0105] For a given HDR-SDR (training) image pair and its brightness or Y channel codeword The "real data" used for model training can be constructed in the form of the best L1 median (denoted as...). In some operational scenarios, the optimal L1 median can be determined by applying brute-force computation to generate a back-shaped image and calculate the corresponding PSNR for all (candidate) L1 medians. In other scenarios, to avoid or reduce the relatively high computational cost of this brute-force method, a two-step approach can be implemented or performed to find the optimal L1 median. The two-step method includes: (1) a coarse search to find the representation as (2) Iterate through a fine search to find the optimal L1 median. The solution (e.g., a local optimum).
[0106] In the first step, the initial L1 median can be determined by exhaustive search of the L1 medians {512, 768, ..., 3328} that correspond to or are indexed by the basic integer function. Among the (candidate) L1 values {512, 768, ..., 3328}, the one with the highest PSNR (denoted as ) can be selected. The initial L1 median is obtained by using the L1 median corresponding to one of the basic shaping functions in the basic shaping function of the HDR image after shaping. .
[0107] In the second step, assuming that the PSNR, as a function of L1 median, has a unique maximum point (e.g., global, local, convergent, etc.), then gradient information can be obtained from the current best L1 median. initial L1 median Begin iteratively or gradually refining the sample point grid to find the (final) optimal L1 median. In each iteration, a search can be performed in both the left and right directions for items with relatively small step sizes. The (current) best L1 median (For example, it can initially be set to an initial value, such as...) (etc.). (Current) best L1 median Move to the position or value that gives the maximum PSNR. If neither the left nor right direction gives the maximum or greater PSNR, the step size can be reduced to half of the previous value. This iterative process can be repeated until the step size is less than a threshold (e.g., (etc.) and / or until the maximum number of iterations is reached (e.g., (etc.). Due to (the (final) optimal L1 median) The "real data" representing the L1 median is used to train the regression model, therefore the (final) optimal L1 median is... It can be a real number, but is not necessarily limited to integers.
[0108] For any (final) best L1 median determined for a given pair of HDR-SDR training images that does not correspond to (e.g., cached, readily available, stored, preloaded, etc.) an integer function, an integer function can be generated by interpolating the nearest available L1 median (e.g., relative to the (final) best L1 median determined for a given pair of HDR-SDR training images) with a basic or non-basic integer function (e.g., tone curve, etc.).
[0109] Tables 4 and 5 below illustrate the example initial coarse search procedure and the example iterative search procedure.
[0110] Table 4
[0111]
[0112] Table 5
[0113]
[0114] A two-step search method can be performed on each HDR-SDR training image pair in the training dataset to obtain all the "real data" for each such image pair in the form of the optimal L1 median. The regression problem for global L1 median prediction can be formulated to be at least partially based on the predicted L1 median (in... Figure 3C The data is denoted as "pred" and the actual data (best L1 median; in Figure 3C The difference between (denoted as "gt") is used to optimize model parameters (or polynomial coefficients). and It can be used as follows: Figure 3C The diagram illustrates a relatively simple least-squares solution to obtain the model parameters (or polynomial coefficients). and . Figure 3CEach data point corresponds to an HDR-SDR training image pair and includes the average SDR codeword of the training SDR image (in the luminance or Y channel) and the true data or (final) best L1 median of the HDR-SDR training image pair. Example values for the model parameters of the regression model can be (but are not limited to): and .
[0115] Selection of local brightness shaping function
[0116] To increase or enhance the local contrast ratio in the HDR image (214) up to per pixel, a local shaping function can be created or selected to have a higher slope (corresponding to a higher contrast ratio) than the global shaping function, which is selected based on the global L1 median predicted using a regression model from the average of the SDR codewords in the luminance or Y channel of the input SDR image (202). As before, the regression model can be trained using HDR-SDR training image pairs.
[0117] In some operational scenarios, for each pixel represented in the input SDR image (202), a specific local shaping function may be selected based on an estimate, approximation, or proxy of the local brightness level in the local region surrounding the pixel or in the form of a filtered brightness value obtained through multi-level edge-preserving filtering, and may differ from the global shaping function selected based on the global L1 median. This local shaping function can be applied in conjunction with the local shaping operation (212) to achieve relatively optimal performance in local contrast enhancement, while having relatively low computational cost and little or no halo artifacts.
[0118] Instead of designing unique local shaping functions from scratch for up to each pixel at relatively high computational cost and time, a family of pre-computed (or pre-loaded) BLUTs, including some or all of the basic and non-basic shaping functions as discussed earlier, can be used as a set of candidate local shaping functions. For each pixel, a local shaping function (or BLUT) can be specifically selected from the BLUT family to obtain or achieve a higher slope or local contrast ratio for the pixel.
[0119] The construction, selection, and application of local shaping functions result in an increase in local contrast in local regions around a pixel, while maintaining the overall local brightness in those regions.
[0120] Conceptually, given the input SDR image (202) in the first... i The corresponding pixel (or the corresponding pixel in the corresponding HDR image (214)) i In the case of (a number of pixels), the local region of that pixel can be a pixel. The set of, where, Indicates the first i 1 pixel Neighborhood; Indicates the first i The local brightness or local brightness level around each pixel; and A decimal representing the local brightness deviation / change in a local area.
[0121] The goal is to find local integer functions. Locally oriented functions in The following has a higher value than the global integer function. A higher slope (and therefore a higher contrast ratio), and satisfies the condition / constraint: in the... i Under each pixel (or provide the same local brightness), such as Figure 3D As shown in the diagram.
[0122] It should be noted that some or all of the techniques described herein can be implemented without requiring a strict definition of the local regions of a pixel. The local regions described herein can be viewer-dependent (e.g., different viewers may perceive brightness at different spatial resolutions) and content-dependent (e.g., the local brightness of different content or image features in an image may be perceived in different ways).
[0123] In some operational scenarios, including but not limited to spatial filtering with multi-level edge-preserving filtering of an applicable spatial kernel size (e.g., determined empirically), spatial filtering can be used to generate a measure or proxy of the local brightness around an approximate pixel.
[0124] Compared to methods that use a strict, inflexible definition of local regions around a pixel to calculate local brightness, image processing systems as described in this paper can use filtering to better consider viewer relevance (e.g., spatial size through multiple levels in the filter) and content relevance (e.g., edge preservation through the filter) of how local brightness in an image can be perceived.
[0125] Existing global integer functions It can be used to construct the first i A local shaping function for the nth pixel to ensure that the result comes from a function with the function for the nth pixel. i The locally shaped HDR codeword value of the local shaping function of the pixel is within a reasonable range, or is consistent with the value obtained from the local shaping function of the pixel. i The predicted HDR codeword values predicted by the global shaping function for each pixel showed no significant deviation.
[0126] The predicted HDR codeword values can be generated from a global integer function, which uses the first... iInput in the form of global L1 median and SDR luminance or Y codeword values for each pixel.
[0127] Similarly, locally shaped HDR codeword values can be generated from a global shaping function, which uses the first... i Local L1 median of 1 pixel and SDR luminance or Y codeword value Input in the form shown below:
[0128] (4)
[0129] It should be noted that if become If the function is integer, then the integer function becomes the first... i A local function for each pixel. More specifically, to obtain a higher slope for the local shaping function, one can utilize a value that depends on the SDR luminance or Y codeword value. The BLUT index is used to select the local integer function in the expression (4) above, as shown below:
[0130] (5a)
[0131] Or equivalently:
[0132] (5b)
[0133] in,
[0134] (6)
[0135] By calculating the local and global integer functions, we can better understand their relationship. derivative under This can show the local shaping function in the first... i It has a higher slope at a pixel than the global integer function.
[0136] For global integer functions Therefore, the slope of the global integer function is given as follows:
[0137] (7)
[0138] For local shaping, the slope is given as follows:
[0139] (8)
[0140] As discussed, the basic and non-basic integer functions in the BLUT family can satisfy the condition / constraint of being monotonically increasing with respect to the L1 median with the corresponding property, as follows:
[0141] (9)
[0142] Therefore, assuming in non-flat regions and ,exist The slope of the local shaping function will be greater than that of the global shaping function, resulting in a larger local contrast ratio for the local shaping function. Therefore, expressions (5) and (6) can be used to provide a higher or better local contrast ratio.
[0143] In some operational scenarios, the following relationships can be set: ,in, Same as in expression (3) above, and This indicates the level of local augmentation. Any, some, or all of these parameters can be obtained from simulation, training, or experience. Local Augmentation Level Example values for can be, but are not necessarily limited to: a constant between 0 and 3 for the (selected) intensity based on the local contrast ratio. When the local integer function is equivalent to the global integer function, the local integer function is equivalent to the global integer function.
[0144] Figure 3E The illustrations show example local integer functions selected using the techniques described in this paper with reference to global integer functions. These example local integer functions are obtained by fixing... and And the differences in the above expressions (5) and (6) It is created by switching between these functions. Each local shaping function corresponds to a specific local brightness value. The local luminance value is the SDR value at the intersection with the global shaping function. For example... Figure 3E As shown, at any intersection of the local and global integer functions, the slope of the local integer function is greater than that of the global integer function.
[0145] Figure 3F The illustration shows an example local integer function selected under the technique described in this paper with reference to multiple (e.g., actual, candidate, possible, etc.) global integer functions. As shown, the slope of the local integer function is larger at the intersection of the local integer function and any global integer function compared to all the depicted global integer functions.
[0146] Additionally, alternatively, or as a further alternative, nonlinear adjustments may be made as part of determining or selecting the local shaping function for the local shaping operation (212).
[0147] As previously mentioned, local shaping functions with up to per-pixel precision can be determined or selected from a pre-computed family of (basic or non-basic) shaping functions in the BLUT family for all integer values in L1, such as 0 to 4095 (inclusive) (0, 1, 2, ..., 4095), where integer values cover some or all of the entire HDR codeword space in the luminance or Y channel. Extrapolation and / or interpolation can be used to obtain some or all of these (candidate) shaping functions from the pre-computed family of BLUTs. This avoids, for example, explicitly calculating the exact local shaping function for each pixel and / or each image / frame while the pixel or image / frame is being processed at runtime.
[0148] Using a family of pre-computed (basic or non-basic) shaping functions, for each image / frame, such as the input SDR image (202), the local luminance can first be calculated for each pixel in the image / frame. Image filtering can be used to generate or estimate local brightness. (e.g., its proxy or approximation, etc.). Local brightness ( This can then be used to generate or predict the corresponding L1 value for each such pixel in the image / frame based on the above expressions (5) and (6). This produces the L1 midpoint value for each pixel in the image / frame. L1 median mapping of ).
[0149] In some operational scenarios, to prevent the L1 median from becoming too low and giving an unnatural appearance on the shaped HDR image (214), a nonlinear function can be introduced. To further adjust the L1 values obtained from the above expressions (5) or (6), as shown below:
[0150] (10a)
[0151] Or equivalently:
[0152] (10b)
[0153] By using examples rather than limitations, nonlinear functions Alternatively, the soft-limiting function can be an S-shaped linear function (e.g.) Figure 3G (As shown in the diagram), it sets the minimum soft limit as follows: Equal offset values, as shown below.
[0154] (11)
[0155] in, This indicates an S-shaped offset. For (or applicable to) the entire L1 median range of [0, 4095], d Example values can be, but are not necessarily limited to, 200.
[0156] As shown in expression (11) above, in nonlinear functions The maximum value can be left unadjusted, allowing brighter pixels to gain or retain relatively sufficient highlights. Note the non-linear function. exist The slope at a point can be set to equal one (1). Table 6 below illustrates an example procedure for adjusting the L1 median of a localized shape using a sigmoid linear shape function.
[0157] Table 6
[0158]
[0159] In some operational scenarios, sigmoid linear functions / activations can help avoid or reduce over-enhancement (e.g., further darkening) in dark areas compared to (e.g., original linear functions / activations). Dark areas with sigmoid linear functions or activations may appear brighter than dark areas with linear functions or activations.
[0160] Guided filtering
[0161] Although local shaping functions provide a larger contrast ratio in local regions, the local brightness level should be set appropriately for each pixel. This makes the enhancement of the local shaping function appear natural. Like local regions, local brightness is subjective (viewer-dependent) and content-dependent. Intuitively, the local brightness level of a pixel... This should be set to the average brightness of pixels near or around the same visual object / character or pixels sharing the same light source, such that pixels belonging to the same visual object / character and / or sharing the same light source can be shaped using (multiple) identical or similar local shaping functions. However, finding or identifying visual objects / characters and light sources is not an easy task. In some operational scenarios, for example, where high computational costs are not required to (e.g., explicitly) find or identify visual objects / characters and light sources in an image / frame, edge-preserving image filtering can be used to approximate or achieve the effect that pixels belonging to the same visual object / character and / or sharing the same light source can be shaped using (multiple) identical or similar local shaping functions.
[0162] In some operational scenarios, guided image filters can be used to perform edge-preserving image filtering, which can be performed efficiently with relatively low or moderate computational cost. An example of guided image filtering is described in the following literature: Kaiming He, Jian Sun, and Xiaoou Tang, “Guided Image Filtering”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 35, No. 6, pp. 1397-1409 (2013), the contents of which are incorporated herein by reference in their entirety.
[0163] Given an input image such as an input SDR image (202). Guide image Kernel size And expected / target smoothness In this case, the guide image filter can perform a locally weighted average of the codeword values in the input image, similar to a bilateral filter. The weights can be calculated or determined in such a way that pixels from the same side of an edge in the guide image (e.g., similar colors, the same visual object / character, etc.) contribute more to the filtered value than pixels from the opposite side of the edge (e.g., different colors, different visual objects / characters, etc.). Therefore, pixels from the same side of an edge in the guide image are averaged together and dominate the filtered value, resulting in an edge-preserving property that efficiently approximates the effect obtained by calling computationally expensive methods (e.g., explicitly finding or identifying visual objects / characters and light sources in an image / frame).
[0164] The guided filter imposes a constraint that the filtered output, or filtered image, is a linear function relative to the noisy guided image. Therefore, a constraint optimization problem can be formulated to minimize the prediction error in each local region of the input image. This optimization problem can be solved using ridge regression in each local region. By implementing an efficient solution for pixel-by-pixel ridge regression, the guided filter-based filtering process can be accelerated.
[0165] Table 6 below illustrates the application of guided image filters to the input image. P Example program for filtering.
[0166] Table 6
[0167]
[0168] The mean filter in Table 6 above ( It can be at least partially achieved through calculation.O ( N Time complexity from the guide image I Generate an integral image to obtain or implement, wherein... N It is the number of pixels in the input image.
[0169] Given a size of Guide image In the case of its integral image It is the size of The image, where each pixel value is The sum of all pixels above and to the left of the center is shown below:
[0170] (12)
[0171] Note the integral image. It may include an input guiding image from which an integral image is obtained. The value in [the text] is a higher precision value. For Bit depth and The size of the (input) guiding image, whose integral image can include images with precision. The value. For example, for a 10-bit 1080p input image, its integral image can include values with precision. The value of .
[0172] Table 7 below illustrates an example procedure for mean filtering using an integral image. Table 8 below illustrates an example procedure for generating an integral image. Table 9 below illustrates an example auxiliary procedure used in the mean filtering illustrated in Table 7.
[0173] Table 7
[0174]
[0175]
[0176] Table 8
[0177]
[0178] Table 9
[0179]
[0180] When the input image is used as the guide image, the filtering result is similar to that generated by the bilateral filter, and pixels with similar colors are averaged together.
[0181] In some operational scenarios, a faster version of the guided image filter can be implemented or executed to speed up the filtering process for performing regression in the subsampled image.
[0182] Table 10 below illustrates the process of resampling an input image using a guided image filter. P Example program for filtering.
[0183] Table 10
[0184]
[0185] Tables 11 and 12 below illustrate the execution of the functions mentioned in Table 10 above. and .
[0186] Table 11
[0187]
[0188] Table 12
[0189]
[0190]
[0191] The subsampling factors in Tables 11 and 12 above can be selected without sacrificing the final quality of the HDR image (214), as will be discussed in more detail later.
[0192] The parameters used in the guided image filter can be compared with those used in other image filters. The radius associated with the kernel can be defined as... And with the smoothness in the brightness or Y channel. The relevant radius can be defined as These parameters are similar to the standard deviation / radius in a Gaussian or bilateral filter, where pixels within a radius can be filtered or averaged. Therefore, the parameters of a guided image filter can be interpreted or specified as follows:
[0193] (13)
[0194] Image filtering as described in this article can handle inputs with different spatial resolutions. Therefore, it can be adjusted proportionally to the image size. Parameters such as these. For example, to achieve similar visual effects at different spatial resolutions, for Images of resolution, It can be set to equal to 50, and for For the same content at the same resolution, it can be set to 25. For smoothness... The following settings can be configured for normalized SDR images: and (For example, codewords normalized within the value range of [0, 1]).
[0195] Additionally, optionally, or alternatively, the guide image, as described herein, can have multiple dimensions associated with multiple features. These dimensions of the guide image are not necessarily associated with the color channels YUV or RGB. Instead of simply calculating a single (e.g., scalar, etc.) covariance matrix across multiple dimensions, it is possible to compute the covariance matrix across multiple dimensions.
[0196] Table 13 below illustrates the application of guided image filters to an input image using multiple dimensions (e.g., three dimensions). P Example program for filtering.
[0197] Table 13
[0198]
[0199]
[0200] In Table 13 above, express Identity matrix. Parameters of the 3D guiding image. It can be scaled proportionally based on the total number of dimensions, as shown below:
[0201] (14)
[0202] Reduced halo artifacts
[0203] A major artifact that can result from image filtering across a large area in an image is halo artifacts. Halo artifacts can be caused by local filtering, which introduces unwanted or uniformly smoothed gradients at the edges of visual objects / characters. For example, such local image filtering can introduce unwanted "enhancements" (bright halos) caused by unwanted or uniformly smoothed gradients at the edges of foreground visual objects / characters such as a girl against a background such as the sky. Therefore, in an output image generated at least partially by local image filtering, a bright halo can be generated around a foreground visual object / character such as a girl.
[0204] Although guided image filters, as described in this paper, can be designed to preserve edges, halo artifacts can or may not be completely avoided, given that guided image filters impose a local linear model within a non-zero radius and operate with an imperfect (or approximate) ridge regression solution.
[0205] In some operational scenarios, image gradients can be used as weighting factors to adjust the strength or degree of local image filtering. When artifacts occur in textureless and smooth areas such as the sky, viewers may be more likely to notice halo artifacts. Therefore, the strength or degree of local image filtering can be reduced in these areas to make any halo artifacts less noticeable or less obvious. By comparison, it is not necessary to reduce local image filtering in textured areas.
[0206] It should be noted that the guiding filter attempts to preserve edges to avoid strong smoothing along them. If the guiding image has relatively large values (which can be positive or negative) near sharp edges, the guiding image filtering or smoothing along the edges becomes relatively weak, thereby avoiding blurred edges and reducing halo artifacts near the edges. To achieve this smoothing result, high-frequency features extracted from the input image can be incorporated into one dimension of the guiding image and used to provide additional guidance regarding filtering / smoothing. The high-frequency features can be weighted by the inverse of the image gradient, as follows:
[0207] (15)
[0208] in, This indicates a pixel-by-pixel Gaussian blur. This indicates, for example, the pixel-by-pixel image gradient magnitude obtained or acquired using a Sobel filter; This represents the regularization constant (e.g., for an input image of size ). The normalized SDR image is 0.01. The value can be inversely proportional to the image resolution of the input image.
[0209] The molecule in the above expression (15) This represents the pixel-wise high-frequency components of the input image. The goal of replacing the (original) input image with high-frequency (component) features is to guide filtering / smoothing by utilizing value differences (such as those measured from the high-frequency components extracted from the input image based on the difference between the input image and the Gaussian blurred image) rather than input values at or near edges in the input image (e.g., the original, etc.).
[0210] The denominator in the above expression (15) This represents the texture intensity of the input image. The weighted high-frequency features in expression (15) above... It can be used as one of the dimensions of the guiding image. If these locations or regions have relatively high-frequency components ( ) and relatively low texture intensity ( If the guide image's dimension provides a relatively high value near edges / boundaries, then the intensity or degree of smoothing or enhancement near these locations or regions will be relatively low. On the other hand, if the location or region has relatively high-frequency components and relatively high texture intensity, then the weighted high-frequency features of the guide image will be higher. The dimension provides a relatively low value near or at a location or region, resulting in a relatively high intensity or degree of smoothing or enhancement near such a location or region. This utilizes high-frequency characteristics. The form of this guided image implements an enhancement or smoothing scaling strategy, which can either amplify or preserve the differences between edges if the location or region near the edge is a non-textured and smooth area. This applies to higher dimensions at given locations or regions. In the case of a certain value, the guided image filter will reduce unwanted smoothing or enhancement.
[0211] Table 14 below illustrates how to obtain pixel-wise image gradients using a Sobel filter. (or Example programs for pixel-by-pixel Gaussian blur are shown in Table 15 below.
[0212] Table 14
[0213]
[0214] Table 15
[0215]
[0216] As used in this article, the clipping function in the expression b = clip(a, x, y) in Table 14 above performs the following operation: if (a <x ),则b="x;否则如果(" a>If y), then b = y; otherwise, b = a.
[0217] In some operational scenarios, the high-frequency features in the above expression (15) The third dimension of the guide image can be used, while the input image can be used as the first and second dimensions of the guide image. Alternatively, the three dimensions in the guide image can be adjusted to the same dynamic range. The third dimension in the guide image can be constant. (e.g., 0.3, etc.) multiplied. Therefore, the three dimensions of the guide image can be specified or defined as follows:
[0218] (16)
[0219] Multi-stage filtering
[0220] An important aspect of image filtering is the kernel size used for filtering. A smaller kernel size means a smaller local region is considered during filtering to produce the filtered result (and therefore fewer pixels are used or shaped within that local region). Conversely, a larger kernel size means a larger local region is considered during filtering to produce the filtered result (and therefore more pixels are used or shaped within that local region). For a relatively small kernel (size), relatively fine image details and textures corresponding to a relatively small local region are enhanced or altered, while relatively coarse image details or textures remain similar to or almost unchanged. Conversely, for a relatively large kernel (size), relatively coarse image details and textures corresponding to a relatively large local region are enhanced or altered, while relatively fine image details or textures remain similar to or almost unchanged.
[0221] The techniques described herein can be implemented to utilize different kernel sizes (e.g., to better illustrate how viewer relevance or content relevance is perceived in local regions or local brightness levels). For example, multi-level filtering can be implemented under these techniques to generate a filtered image as a whole. As in multiple different Multiple filtered images generated by kernel size The weighted sum of these factors results in the final shaped HDR image (214) being compared with multiple different... The kernel size appears to be enhanced across all levels (multi-stage filtering), as shown below:
[0222] (17)
[0223] in, Indicates the first The first at the level n The weights of the filtered image (or the weights of the first filtered image) The first kernel size at the [number]th n (kernel size).
[0224] Filtered values from the filtered image derived from the combination of the above expression (17) (e.g., used as local brightness levels in the input SDR image (202)). An estimate or proxy (such as an estimate or proxy) can be used as input in expression (10a). This is used for selecting local shaping functions.
[0225] In some operational scenarios, in order to help provide a better appearance for shaped HDR images (e.g., Figure 2A or Figure 2B (e.g., 214), four levels can be used (or = 4) or kernel-size image filtering, and combined with the corresponding weight factors. The example kernel sizes for these 4 levels can be, but are not necessarily limited to: for The image size is (12, 25, 50, 100). Example values for the weighting factor can be, but are not limited to: (0.3, 0.3, 0.2, 0.2).
[0226] While computational cost and number of iterations may increase with the total number of levels of the filtered image used to obtain the overall composition, this cost and number can be reduced by reusing or sharing some variables or quantities across different levels. Additionally, alternatively, or as a substitute, approximations can be applied where applicable to help prevent computational cost from increasing with kernel size.
[0227] For the purpose of providing (local contrast) enhancement in the corresponding shaped HDR image with little or no halo artifacts, HDR-SDR training image pairs in the training dataset can be used to help find optimal values for some or all operating parameters. Video annotation tools can be applied to these images to label or identify foreground, background, image objects / characters, etc., in the image pairs. Example video annotations are described in U.S. Provisional Patent Application Serial No. 62 / 944,847, "User guided image segmentation methods and products," filed December 6, 2019, by A. Khalilian-Gourtani et al., which is incorporated herein by reference in its entirety. A halo artifact mask (denoted as HA) can be used to define or identify background pixels, image objects / characters, etc., around the foreground. For illustrative purposes only, these background pixels may be located within a certain distance range (e.g., 5 to 30 pixels) of the foreground, image object, or character, for example, using a distance transformation of the foreground, image object / character, etc., as identified in or together with the foreground mask (denoted as FG) generated by a video annotation tool or object segmentation technique.
[0228] To quantitatively represent enhancement, a local foreground region can be defined. Region-based texture metrics (represented as) Local foreground area Having located in the n The first image i The center at 1 pixel 。 Region-based texture measurement It can be defined or specified as a local foreground region. HDR images in the image (e.g., Figure 2A or Figure 2B (e.g., 214) and input SDR image (e.g., Figure 2A or Figure 2B The ratio of the variances between (e.g., 202, etc.) is shown below:
[0229] (18)
[0230] Region-based texture measurement It can be used to measure local shaping operations (e.g., Figure 2A or Figure 2B Increased high-frequency components in (e.g., 212, etc.). Region-based texture measurement. A higher value indicates a greater increase in local contrast ratio from local shaping of the input SDR to the generation of a shaped HDR image.
[0231] Additionally, halo measurement can be defined as the local halo artifact region identified within a halo artifact mask. The correlation coefficients between the shaped HDR image and the input SDR image are shown below:
[0232] (19)
[0233] Halo measurement Halo This can be used to measure the inconsistency between the shaped HDR image and the input SDR introduced through local shaping. When halo artifacts are present in the shaped HDR image, the shaped HDR image shows a different trend from the input SDR in the halo artifact mask (e.g., background pixels around the foreground, visual objects / characters, etc. as defined in the expression (19) above). The correlation coefficient in the expression (19) above can be used to measure the inconsistency between the two variables. and The linear correlation between them is shown below:
[0234] (20a)
[0235] Among them, covariance The definition is as follows:
[0236] (20b)
[0237] Here, This represents the average value.
[0238] The halo measurement in the above expression (19) Halo The relatively small value of the second term implies a relatively weak linear relationship (or relatively large inconsistency) between the input SDR image and the shaped HDR image, thus indicating that the halo artifact is relatively strong.
[0239] The goal of finding optimal values for the operational parameters described in this article is to provide or generate large values (e.g., maximum values, etc.) for texture metrics and small values (e.g., minimum values, etc.) for halo metrics. Example operational parameters for which optimized values can be generated may include, but are not limited to, some or all of the following: , , , wait.
[0240] In some operational scenarios, (For example, (etc.) The SDR image in the HDR-SDR training image pair can be randomly sampled to a size of (For example, (etc.) (For example, Within a square region (e.g.), a foreground mask can be generated for that square region. and halo mask The optimization problem can be formulated as follows to obtain the optimal values of the operating parameters:
[0241] (twenty one)
[0242] In a non-restricted example, the parameter space used to search for the optimal value of the operation parameter can be defined as follows: , , and The optimal values for the operating parameters can be selected, at least in part, based on halo and texture metrics, as follows: , , and , used to normalize the input SDR input image (e.g., normalized SDR codeword values in the normalized value range of [0, 1], etc.).
[0243] In some operational scenarios, to speed up image filtering, results from some or all of the computations can be shared or reused in the local shaping operation (212), for example in the following two aspects. First, two of the three dimensions of the guide image used to guide the filtering of the input SDR image can be set to the input SDR image: ,in, P This represents the input SDR image. Therefore, images from the input SDR image and / or guide image can be shared or reused. , , and Related calculations and / or calculation results. Second, in multi-level filtering, variables (or the number of variables) that are independent of or unaffected by kernel size can be shared between different levels. Table 16 below illustrates the simplification or reuse of calculations and / or calculation results at different levels (e.g., , Example programs (etc.).
[0244] Table 16
[0245]
[0246]
[0247] Secondly, Gaussian blurring can be accelerated, for example, by using an iterative box filter to approximate a Gaussian filter. Example approximations of Gaussian filters and iterative box filters are described in the following literature: Pascal Getreuer, "A Survey of Gaussian Convolution Algorithms," Image Processing On Line, Vol. 3, pp. 286-310 (2013); William M. Wells, "Efficient Synthesis of Gaussian Filters by Cascaded Uniform Filters," IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 8, No. 2, pp. 234-239 (1986), the contents of which are incorporated herein by reference in their entirety. The total number of iterations for the iterative box filter... It can be (but is not limited to) set to three (3) times. Table 17 below illustrates an example program for approximating a Gaussian filter using an iterative box filter.
[0248] Table 17
[0249]
[0250] The effectiveness of subsampling in fast-guided image filters can be analyzed. The computational cost / number of iterations and visual quality of shaped HDR images generated with different subsampling factor values can be compared to determine the optimal value of the subsampling factor. In some operational scenarios, the subsampling factor... A value of 2 provides a significant speedup (e.g., approximately 17%, etc.) with no or almost no visible degradation in reshaped HDR images, while the subsampling factor... A value of 4 indicates a visible degradation or difference. Therefore, in some operational scenarios where computational resources may be limited, a quadratic sampling factor can be used as described in this paper. The value 2 is used to implement or perform secondary sampling in multi-stage filtering.
[0251] Figure 2C The illustration depicts an example flow for applying multi-level edge-preserving filtering. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform some or all of the operations in this flow.
[0252] Box 250 includes a receiving input SDR image (e.g., Figure 2A or Figure 2B Box 252 includes using the input SDR image (202) and the image gradient to generate a Gaussian blurred image gradient and a Gaussian blurred SDR image, for example, for the current level in multiple levels (or multiple iterations), wherein the current level is set as the first level in multiple levels in a specific processing order. Box 254 includes generating a guide image from the Gaussian blurred image gradient and the Gaussian blurred SDR image. Box 256 includes using the input SDR image (202) to calculate an integral image (of the input SDR image (202)). Box 258 includes applying a fast guide image filter using the guide image and the integral image to generate a filtered SDR image 260 of the input SDR image (202) at the current level. Box 262 includes adding the filtered SDR image (260) to a weighted sum of filtered SDR images at multiple levels (which may be initialized to zero or empty). Box 264 includes determining whether the maximum iteration or level has been reached. If the maximum iteration or level has been reached ("Yes"), the weighted sum is used as the multi-level filtered SDR image (266). On the other hand, if the maximum iteration or level has not been reached ("No"), the current level is incremented and the process returns to box 252.
[0253] MMR-based shaping / mapping can be used to perform or carry out chroma shaping, which is based on the input SDR image (e.g., Figure 2A or Figure 2B The luminance and chrominance SDR codewords in (e.g., 202, etc.) generate a shaped HDR image in the chrominance channel. Figure 2A or Figure 2B HDR codewords (e.g., 214, etc.) after reshaping.
[0254] Similar to the BLUT family of basic or non-basic (BLUT) shaping functions with multiple different L1 medians for local luminance shaping (or generating shaped luminance or Y-channel HDR codewords in a shaped HDR image from locally shaped luminance or Y-channel SDR codewords in an input SDR image), the backward MMR family (BMMR) of basic or non-basic (BMMR) shaping maps with multiple different L1 medians for local chrominance shaping.
[0255] Basic or non-basic (BMMR) shaping maps can be specified by multiple sets of MMR coefficients, respectively. Each (BMMR) shaping map in the BMMR family can be specified by a corresponding set of MMR coefficients from the multiple sets of MMR coefficients. These sets of MMR coefficients can be pre-computed and / or pre-loaded during the system initialization / startup cycle of the image processing system described herein. Therefore, when an input SDR image is locally shaped into an HDR image, it is not necessary to generate MMR coefficients for local chroma shaping at runtime. Extrapolation and / or interpolation can be performed to extend the basic BMMR coefficient set of the chroma channels to the set of non-basic BMMR coefficients with L1 midpoints in other BMMR families.
[0256] Since L1 median mapping can be generated through local luminance shaping, the same L1 median mapping can be used to select or look up a specific set of BMMR coefficients for a specific L1 median used for local chrominance shaping. The specific set of BMMR coefficients for a specific L1 median (which represents a specific BMMR shaping / mapping for that specific L1 median) combined with a specific BLUT for a specific L1 median selected or looked up in a BLUT family can be applied to an input SDR image to generate shaped colors in the shaped HDR image, i.e., increasing color saturation while matching the overall HDR hue / appearance (e.g., determined by training an HDR-SDR image, etc.), as indicated by the global luminance and chrominance shaping functions / mappings corresponding to the global L1 median determined from the input SDR image.
[0257] The set of MMR coefficients can be selected or looked up with up to per-pixel precision based on the L1 median / index indicated in the L1 median map. Chromaticity local shaping can be performed on the chroma channels Cb and Cr up to per pixel, as follows:
[0258] (twenty two)
[0259] In many operating scenarios, when the L1 median becomes larger, especially for relatively low input SDR luminance or Y codewords, the sculpted Cb and Cr HDR codewords cause the corresponding sculpted colors to become more saturated (e.g., as by...). and The distance between the measured corresponding shaped color and the neutral or grayscale color becomes larger. Therefore, in local shaped color using local L1 median mapping, pixels with brighter brightness or luminance levels become more saturated, while pixels with darker brightness or luminance levels become less saturated.
[0260] In some operational scenarios, a linear segment-based structure can be used when locally reshaping an SDR image into an HDR image. An example of a linear segment-based structure is described in U.S. Patent 10,397,576, "Reshaping curve optimization in HDR coding," the entire contents of which are fully set forth herein and incorporated by reference.
[0261] Some or all of the techniques described herein may be implemented and / or performed as part of real-time operation in broadcast video applications, real-time streaming applications, etc. Additionally, alternatively, or alternatively, some or all of the techniques described herein may be implemented and / or performed as part of delayed or offline operation in non-real-time streaming applications, cinema applications, etc.
[0262] Example process flow
[0263] Figure 4 An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding device / module, transcoding device / module, decoding device / module, inverse tone mapping device / module, tone mapping device / module, media device / module, inverse mapping generation and application system, etc.) may perform this process flow. In block 402, the image processing system generates a global index value for selecting a global shaping function for the input image of a second dynamic range. The global index value is generated using luminance codewords in the input image.
[0264] In box 404, the image processing system applies image filtering to the input image to generate a filtered image. The filtered values of the filtered image provide a measure of the local brightness level in the input image.
[0265] In box 406, the image processing system generates local index values for selecting a specific local shaping function for the input image. The local index values are generated using the global index value and the filtered values of the filtered image.
[0266] In box 408, the image processing system causes at least in part to generate a shaped image of a first dynamic range by shaping the input image using a specific local shaping function selected with local index values.
[0267] In this embodiment, local index values are represented in a local index value mapping.
[0268] In the embodiments, the first dynamic range represents the high dynamic range; the second dynamic range represents the standard dynamic range, which is lower than the high dynamic range.
[0269] In an embodiment, a specific local shaping function selected using a local index value shapes the luminance codewords in the input image into luminance codewords in the HDR image; the local index value is further used to select a specific local shaping map, which maps the luminance codewords and chrominance codewords in the input image into chrominance codewords in the shaped image.
[0270] In the embodiments, a specific local shaping function is represented by a backward shaping lookup table; a specific local shaping mapping is represented by a backward multivariate multiple regression mapping.
[0271] In this embodiment, the local index value is generated from the global index value and the filtered value using a prediction model trained on training image pairs in the training dataset; each of the training image pairs includes a training image of a first dynamic range and a training image of a second dynamic range; the training image of the first dynamic range and the training image of the second dynamic range depict the same visual semantic content.
[0272] In this embodiment, each of the specific local shaping functions is selected from a plurality of preloaded shaping functions based on the corresponding local index value in the local index value.
[0273] In an embodiment, multiple preloaded shaping functions are generated from a base shaping function by one of the following: interpolation, extrapolation, or a combination of interpolation and extrapolation; the base shaping function is determined using a training dataset comprising multiple pairs of training images; each of the training image pairs comprises a training image of a first dynamic range and a training image of a second dynamic range; the training image of the first dynamic range and the training image of the second dynamic range depict the same visual semantic content.
[0274] In this embodiment, the shaped image is generated by a video encoder that has already generated global and local index values.
[0275] In this embodiment, the shaped image is generated by a receiving device other than a video encoder that has already generated global and local index values.
[0276] In this embodiment, the filtered image is a weighted sum of multiple individually filtered images generated by image filtering at multiple kernel size levels.
[0277] In this embodiment, image filtering is applied to the input image with the guide image to reduce halo artifacts.
[0278] In the embodiments, image filtering refers to one of the following: guided image filtering, unguided image filtering, multi-level filtering, edge-preserving filtering, multi-level edge-preserving filtering, non-multi-level edge-preserving filtering, multi-level non-edge-preserving filtering, non-multi-level non-edge-preserving filtering, etc.
[0279] In this embodiment, the local index value is generated at least in part based on a soft-limiting function; the soft-limiting function takes as input the scaled difference between the luminance codeword and the filtered value; the soft-limiting function is specifically chosen to generate an output value that is not less than the minimum value of the dark pixels of the input image.
[0280] In an embodiment, image filtering refers to guide image filtering performed with a guide image; the guide image includes high-frequency feature values calculated for multiple pixel locations, each high-frequency feature value being a weighted inverse of an image gradient calculated for multiple pixel locations in one or more channels of the guide image.
[0281] In an embodiment, the guiding image includes guiding image values at least partially based on a set of halo reduction operation parameters; the image processing system is further configured to perform the following operations: for each of a plurality of image pairs, using a plurality of candidate value sets of the halo reduction operation parameters, calculate a region-based texture metric of a local foreground region in one or more local foreground regions as a variance ratio of a training image of a first dynamic range and a corresponding training image of a second dynamic range in the image pair, wherein the one or more local foreground regions are sampled from a foreground mask identified in the training image of the first dynamic range in the image pair; for each of the plurality of image pairs, using a plurality of candidate value sets of the halo reduction operation parameters, calculate a halo metric of a local halo artifact region in one or more local halo artifact regions as a correlation coefficient between the training image of the first dynamic range and the corresponding training image of the second dynamic range in the image pair, wherein the one or more local halo artifact regions are sampled from a halo mask identified in the corresponding training image of the second dynamic range; and for a plurality of candidate value sets of the halo reduction operation parameters, calculate the following corresponding weighted sums: (a) (a) All region-based texture measures of all local foreground regions sampled from the foreground mask identified in the training image of the first dynamic range in multiple image pairs, and (b) All halo measures of all local halo artifact regions sampled from the halo mask identified in the corresponding training image of the second dynamic range in multiple image pairs; Determine an optimized set of halo reduction operation parameters, the optimized set of values being used to generate a minimized weighted sum of: (a) all region-based texture measures of all local foreground regions sampled from the foreground mask identified in the training image of the first dynamic range in multiple image pairs, and (b) all halo measures of all local halo artifact regions sampled from the halo mask identified in the corresponding training image of the second dynamic range in multiple image pairs; Apply image filtering at multiple spatial kernel size levels to the input image having the optimized set of halo reduction operation parameters.
[0282] In embodiments, computing devices such as display devices, mobile devices, set-top boxes, and multimedia devices are configured to perform any of the methods described above. In embodiments, an apparatus includes a processor and is configured to perform any of the methods described above. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause any of the methods described above to be performed.
[0283] In one embodiment, a computing device includes one or more processors and one or more storage media, the storage media storing an instruction set that, when executed by the one or more processors, causes any of the methods described above to be performed.
[0284] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.
[0285] Example computer system implementation
[0286] Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions relating to adaptive perceptual quantization of images with enhanced dynamic range, as described herein. The computer and / or IC may calculate any of the various parameters or values relating to the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0287] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement the methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a processor-accessible program memory. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention may take any of a variety of forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0288] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to components (including references to "devices") should be interpreted as including any component that performs the function of the described component as an equivalent of the component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions in the illustrated exemplary embodiments of the invention.
[0289] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that incorporates hardwired and / or program logic to implement the techniques.
[0290] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.
[0291] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in non-transitory storage media accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.
[0292] Computer system 500 further includes read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions of processor 504. Storage device 510 (such as a magnetic disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.
[0293] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.
[0294] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the instruction sequence contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used instead of or in combination with software instructions.
[0295] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.
[0296] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0297] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 fetches and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.
[0298] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0299] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the "Internet" 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry streams of digital data. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.
[0300] Computer system 500 can send messages and receive data, including program code, through multiple networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.
[0301] The received code may be executed by processor 504 and / or stored in storage device 510 or other non-volatile storage device for later execution upon receipt.
[0302] Equivalents, extensions, alternatives and miscellaneous
[0303] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the claimed embodiments of the invention, and of the applicant's view, is the set of claims published in specific form according to this application, wherein such claim publication includes any subsequent corrections. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and the drawings should be viewed in an illustrative rather than restrictive sense.
[0304] Exemplary examples of enumeration
[0305] This invention may be practiced in any of the forms described herein, including but not limited to the enumerated example embodiments (EEE) that describe some parts of the structure, features and functions of embodiments of the invention.
[0306] EEE1. A method for generating an image with a first dynamic range from an input image with a second dynamic range lower than a first dynamic range, the method comprising:
[0307] Use the luminance codewords in the input image to generate a global index value for selecting a global shaping function for the input image in the second dynamic range;
[0308] Image filtering is applied to the input image to generate a filtered image. The filtered values of the filtered image provide a measure of the local brightness level in the input image.
[0309] The global index value and the filtered value of the filtered image are used to generate local index values for selecting a specific local shaping function for the input image;
[0310] This allows for the generation of a first dynamic range shaped image by at least partially shaping the input image using a specific local shaping function selected with local index values.
[0311] EEE2. Similar to the method in EEE1, where local index values are represented in a local index value mapping.
[0312] EEE3. The method of EEE1 or EEE2, wherein the first dynamic range represents the high dynamic range; wherein the second dynamic range represents the standard dynamic range below the high dynamic range.
[0313] EEE4. A method as described in any of EEE1 to EEE3, wherein a specific local shaping function selected using a local index value shapes the luminance codewords in the input image into luminance codewords in the HDR image; wherein the local index value is further used to select a specific local shaping map that maps the luminance codewords and chrominance codewords in the input image into chrominance codewords in the shaped image.
[0314] EEE5. Similar to the method in EEE4, where a specific local integer function is represented by a backward integer lookup table; where a specific local integer mapping is represented by a backward multivariate multiple regression mapping.
[0315] EEE6. The method of any one of EEE1 to EEE5, wherein the local index value is generated from the global index value and the filtered value using a prediction model trained on training image pairs in the training dataset; wherein each of the training image pairs includes a training image of a first dynamic range and a training image of a second dynamic range; wherein the training image of the first dynamic range and the training image of the second dynamic range depict the same visual semantic content.
[0316] EEE7. The method of any of EEE1 to EEE6, wherein each of a particular local integer function is selected from a plurality of preloaded integer functions based on the corresponding local index value in the local index value.
[0317] EEE8. As with the method in EEE7, wherein multiple preloaded shaping functions are generated from a base shaping function by one of the following: interpolation, extrapolation, or a combination of interpolation and extrapolation; wherein the base shaping function is determined using a training dataset comprising multiple pairs of training images; wherein each of the pairs of training images comprises a training image of a first dynamic range and a training image of a second dynamic range; wherein the training image of the first dynamic range and the training image of the second dynamic range depict the same visual semantic content.
[0318] EEE9. The method of any of EEE1 to EEE8, wherein the shaped image is generated by a video encoder that has already generated global and local index values.
[0319] EEE10. The method of any of EEE1 to EEE9, wherein the shaped image is generated by a receiving device other than a video encoder that has already generated global and local index values.
[0320] EEE11. The method of any of EEE1 to EEE10, wherein the filtered image is a weighted sum of multiple individually filtered images generated by image filtering at multiple kernel size levels.
[0321] EEE12. The method of any of EEE1 to EEE11, wherein image filtering is applied to the input image having a guide image to reduce halo artifacts.
[0322] EEE13. The method of any one of EEE1 to EEE10, wherein image filtering means one of the following: guided image filtering, unguided image filtering, multi-level filtering, edge-preserving filtering, multi-level edge-preserving filtering, non-multi-level edge-preserving filtering, multi-level non-edge-preserving filtering, or non-multi-level non-edge-preserving filtering.
[0323] EEE14. The method of any one of EEE1 to EEE13, wherein the local index value is generated at least in part based on a soft limiting function; wherein the soft limiting function accepts the scaled difference between the luminance codeword and the filtered value as input; wherein the soft limiting function is specifically chosen to generate an output value that is not less than the minimum value of the dark pixels of the input image.
[0324] EEE15. The method of any one of EEE1 to EEE14, wherein the image filtering represents guided image filtering performed with a guided image; wherein the guided image includes high-frequency feature values calculated for multiple pixel locations, the high-frequency feature values being weighted by the inverse of the image gradient calculated for multiple pixel locations in one or more channels of the guided image.
[0325] EEE16. The method of any one of EEE1 to EEE15, wherein the guide image comprises guide image values obtained at least in part based on a set of halo reduction operation parameters; the method further comprises:
[0326] For each of a plurality of image pairs, a region-based texture metric of a local foreground region in one or more local foreground regions is calculated as the variance ratio of the training image of the first dynamic range and the corresponding training image of the second dynamic range in the image pair using a plurality of candidate value sets of the halo reduction operation parameter set, wherein the one or more local foreground regions are sampled from a foreground mask identified in the training image of the first dynamic range in the image pair.
[0327] For each of the multiple image pairs, the halo metric of one or more local halo artifact regions is calculated as the correlation coefficient between the training image of the first dynamic range and the corresponding training image of the second dynamic range in the image pair using multiple candidate value sets of the halo reduction operation parameter set. The one or more local halo artifact regions are sampled from the halo mask identified in the corresponding training image of the second dynamic range.
[0328] For multiple candidate sets of halo reduction operation parameters, the following weighted sums are calculated: (a) all region-based texture measures of all local foreground regions sampled from the foreground mask identified in the training image of the first dynamic range in multiple image pairs, and (b) all halo measures of all local halo artifact regions sampled from the halo mask identified in the corresponding training image of the second dynamic range in multiple image pairs.
[0329] Determine an optimized set of halo reduction operation parameters, wherein the optimized set of parameters is used to generate a weighted sum that minimizes the following: (a) all region-based texture measures of all local foreground regions sampled from a foreground mask identified in a training image of a first dynamic range in a plurality of image pairs, and (b) all halo measures of all local halo artifact regions sampled from a halo mask identified in a corresponding training image of a second dynamic range in a plurality of image pairs;
[0330] Multiple image filters at spatial kernel size levels are applied to the input image, which has an optimized set of parameters for halo reduction.
[0331] EEE17. A computer system configured to perform any of the methods described in EEE1 through EEE16.
[0332] EEE18. An apparatus including a processor and configured to perform any of the methods described in EEE1 to EEE16.
[0333] EEE19. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for performing any of the methods such as EEE1 to EEE16.< / x>
Claims
1. A method for generating an image of a first dynamic range from an input image of a second dynamic range lower than the first dynamic range, the method comprising: generating, using luminance codewords in the input image, a global index value for selecting a global reshaping function for the input image of the second dynamic range, wherein the global index value is a local L1 median representing an average light brightness level predicted from an average of luminance SDR codewords in the input image using a polynomial regression model; applying an image filtering operation to the input image to generate a filtered image, the input image comprising a plurality of regions up to per-pixel precision, and generating, for each region of the plurality of regions, a filtered value of the filtered image providing a measure of a local brightness level in that region; generating, using the global index value and the filtered values of the filtered image, a local index value for selecting a particular local reshaping function for each region of the plurality of regions of the input image, wherein each local index value is a local L1 median representing an average light brightness level of that region, wherein generating the local index value for each region comprises: determining a difference between a local brightness level estimated from the filtered value generated for that region and light brightness levels of individual pixels of that region; estimating a local L1 median adjustment value from the determined difference; predicting a local L1 median from the global index value and the local L1 median adjustment value, the local L1 median representing the local index value; generating a reshaped image of the first dynamic range at least partly by reshaping the input image using the particular local reshaping function selected using the local index value, wherein the particular local reshaping function selected using the local index value reshapes luminance SDR codewords in the input image to luminance HDR codewords in the reshaped image.
2. The method of claim 1, wherein, the local index value is represented in a local index value map.
3. The method of claim 1, wherein, the first dynamic range represents a high dynamic range; wherein the second dynamic range represents a standard dynamic range lower than the high dynamic range.
4. The method of any one of claims 1 to 3, wherein, the local index value is further used to select a particular local reshaping mapping that maps luminance codewords and chrominance codewords in the input image to chrominance codewords in the reshaped image.
5. The method of any one of claims 1 to 3, wherein, the local index value is generated from the global index value and the filtered values using a prediction model trained using pairs of training images in a training data set; wherein each of the pairs of training images comprises a training image of the first dynamic range and a training image of the second dynamic range; wherein the training image of the first dynamic range and the training image of the second dynamic range depict the same visual semantic content.
6. The method of any one of claims 1 to 3, wherein, each of the particular local reshaping functions is selected from a plurality of preloaded reshaping functions based on a respective one of the local index values.
7. The method of any one of claims 1 to 3, wherein, the reshaped image is generated by a video encoder that has generated the global index value and the local index values.
8. The method of any one of claims 1 to 3, wherein, The reshaped image is generated by a receiving device other than a video encoder that has generated the global index value and the local index value.
9. The method of any one of claims 1 to 3, wherein, The filtered image is a weighted sum of a plurality of separate filtered images generated by an image filtering operation performed at a plurality of levels, wherein the plurality of levels use different kernel sizes.
10. The method of any one of claims 1 to 3, wherein, The image filtering operation is applied to an input image using a guide image to reduce halo artifacts.
11. The method of any one of claims 1 to 3, wherein, The image filtering operation represents one of: guided image filtering, unguided image filtering, multi-level filtering, edge-preserving filtering, multi-level edge-preserving filtering, unmulti-level edge-preserving filtering, multi-level unedge-preserving filtering, or unmulti-level unedge-preserving filtering.
12. The method of any one of claims 1 to 3, wherein, The local index value is generated based at least in part on a soft-clipping function; wherein the soft-clipping function accepts as input a scaled difference between the luminance codeword and the filtered value; wherein the soft-clipping function is specifically selected to generate an output value that is not less than a minimum value of dark pixels of the input image.
13. The method of any one of claims 1 to 3, wherein, The image filtering operation represents guided image filtering with a guide image; wherein the guide image includes a plurality of high-frequency feature values computed for a plurality of pixel locations, the plurality of high-frequency feature values being respectively weighted by inverses of image gradients computed for the plurality of pixel locations in one of one or more channels of the guide image.
14. An apparatus for adaptive local reshaping comprising a processor and configured to perform any of the methods of claims 1-13.
15. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for causing one or more processors to perform any of the methods of claims 1-13.
16. An apparatus for adaptive local reshaping comprising means for performing operations of any of the methods of claims 1-13.
17. A computer program product comprising computer-executable instructions for causing one or more processors to perform any of the methods of claims 1-13.
Citation Information
Patent Citations
Reshaping curve optimization in HDR coding
US10397576B2
Signal reshaping approximation
US20180020224A1
Systems and methods for optimizing video coding based on luminance transfer function or video color component values
CN107852512A
License plate image brightness processing method, device and equipment
CN110473158A