Method, apparatus, medium, and computer program product for generating integer functions

By generating and optimizing the polynomial representation of forward and backward shaping functions, the computational efficiency and quality problems in SDR to HDR image mapping are solved, and an efficient and reversible image shaping process is achieved, improving image quality and consistency.

CN115699734BActive Publication Date: 2025-07-11DOLBY LABORATORIES LICENSING CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202180037249.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-21
Filing Date
2021-04-20
Publication Date
2025-07-11
Estimated Expiration
2041-04-20

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently realize image shaping from standard dynamic range (SDR) to high dynamic range (HDR) in display devices, especially in devices with limited computing capabilities, and the generation efficiency and quality of dynamic shaping functions are difficult to take into account.

Method used

By generating a set of forward and backward shaping functions, using polynomial representation and optimization algorithms, segment parameters are adjusted to minimize the gap between consecutive segments, and dynamic shaping functions are generated based on input image characteristics to ensure image quality and reversibility.

Benefits of technology

It realizes efficient mapping of SDR images to HDR images with limited computing power, reduces visual artifacts, improves image quality, and ensures reversibility and consistency of the plastic shaping process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699734B_ABST
    Figure CN115699734B_ABST
Patent Text Reader

Abstract

Methods, apparatuses, media, and computer program products for generating shaping functions are provided. Given an initial set of forward shaping functions, an output forward shaping function is constructed by: a) using the forward shaping functions to generate a first set of corresponding backward shaping functions, b) using a multi-segment polynomial representation with a set of common pivot points to generate a second set of backward shaping functions, c) generating an output set of backward shaping functions by optimizing the polynomial representation of the second set of backward shaping functions to minimize the gap value between consecutive segments, and d) using the output set of backward shaping functions to generate an output set of forward shaping functions by minimizing the distance between the original input HDR codewords and the reconstructed HDR codewords.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 013,063 and European Patent Application No. 20170567.0, both filed on April 21, 2020, each of which is incorporated herein by reference in its entirety. Technical field

[0003] The present invention generally relates to images. More specifically, embodiments of the present invention relate to generating a series of shaping functions for HDR imaging that satisfy both continuity constraints and reversibility constraints. Background art

[0004] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, luma) in an image, such as the range from the darkest gray (black) to the brightest white (highlights). In this sense, DR is related to "scene - referred" intensities. DR can also refer to the ability of a display device to adequately or approximately render a particular breadth of intensity range. In this sense, DR is related to "display - referred" intensities. Unless a specific meaning is explicitly specified at any point in the description herein, it should be inferred that the term can be used in either sense, e.g., interchangeably.

[0005] As used herein, the term "high - dynamic range (HDR)" refers to a DR breadth that spans 14 to 15 orders of magnitude of the human visual system (HVS). In practice, the DR breadth that humans can simultaneously perceive may be slightly truncated relative to HDR.

[0006] In fact, an image includes one or more color components (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented with a precision of n bits per pixel (e.g., n = 8). Using linear or gamma luminance encoding, an image where n ≤ 8 (e.g., a 24 - bit color JPEG image) is considered a standard - dynamic - range image, while an image where n>8 can be considered an enhanced or high - dynamic - range image. High - precision (e.g., 16 - bit) floating - point formats can also be used to store and distribute HDR images, such as the OpenEXR document format developed by Industrial Light and Magic.

[0007] Most consumer desktop displays currently support 200 to 300 cd / m 2or the luminance of the nit. Most consumer HDTVs range from 300 to 500 nits, where new models reach 1000 nits (cd / m 2 ). Thus, such traditional displays represent a lower dynamic range (LDR) associated with HDR, also known as standard dynamic range (SDR). As the availability of HDR content increases due to the development of both capture devices (e.g., cameras) and HDR displays (e.g., Dolby Laboratories' PRM-4200 professional reference monitor), HDR content can be color graded and displayed on HDR displays that support a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or higher).

[0008] In a traditional image pipeline, a non-linear opto-electronic transfer function (OETF) is used to quantify the captured image, which converts linear scene light into a non-linear video signal (e.g., gamma-encoded RGB or YCbCr). Then, the signal is processed at the receiver by an electro-optical transfer function (EOTF) before being displayed on a display, which converts the video signal values into output screen color values. Such non-linear functions include the traditional "gamma" curves documented in ITU-R Rec.BT.709 and BT.2020, the "PQ" (Perceptual Quantization) curve described in SMPTE ST 2084, and the "Hybrid Log-gamma" or "HLG" curve described in Rec.ITU-R BT.2100.

[0009] As used herein, the terms "reshaping" or "remapping" refer to the process of sample-to-sample or codeword-to-codeword mapping of a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma, PQ, or HLG, etc.) to an image of the same or different bit depth and different codeword distribution or representation. Reshaping allows for improved compressibility or improved image quality at a fixed bit rate. For example, without limitation, forward reshaping can be applied to 10-bit or 12-bit PQ-encoded HDR video to improve the encoding efficiency in a 10-bit video coding architecture. At the receiver, after decompressing the received signal (which may or may not involve reshaping), the receiver can apply an inverse (or backward) reshaping function to restore the signal to its original codeword distribution and / or achieve a higher dynamic range.

[0010] Shaping can be static or dynamic. In static shaping, a single shaping function is generated and used for a single stream or across multiple streams. In dynamic shaping, the shaping function can be customized based on input video stream characteristics, which can change at the stream level, scene level, or even frame level. Dynamic shaping is preferred; however, some devices may not have sufficient computing power to support it. As understood by the present inventors herein, improved techniques are desired for efficient image shaping when displaying video content, particularly HDR content.

[0011] The methods described in this section are methods that can be pursued, but not necessarily methods that have been previously envisioned or pursued. Thus, unless otherwise indicated, no method described in this section should be considered to be prior art merely by virtue of its inclusion in this section. Similarly, unless otherwise indicated, problems identified with respect to one or more methods should not be considered to be identified in any prior art based on this section. Summary of the Invention

[0012] Example embodiments described herein relate to image shaping. In an embodiment, in a device including one or more processors, the processor receives an input pair of reference images in HDR and SDR. Given an initial set of forward shaping functions generated using the reference HDR and SDR images, a set of output forward shaping functions and output backward shaping functions is constructed by: a) using the initial set of forward shaping functions to generate a first set of corresponding backward shaping functions, b) generating a second set of backward shaping functions, where each backward shaping function is represented using a multi-segment polynomial representation with a common set of pivot points, c) generating an output set of backward shaping functions by optimizing the polynomial representation of the second set of backward shaping functions to minimize the gap output values between consecutive segments, and d) using the output set of backward shaping functions to generate an output set of forward shaping functions by minimizing the distance between the reference HDR values and the reconstructed HDR values.

[0013] In an embodiment where an output backward shaping function is generated in which the gap values between segments are minimized, the processor receives a first set of input images in a first dynamic range (e.g., SDR) and a second set of input images in a second dynamic range (e.g., HDR), where corresponding pairs between the first set of input images and the second set of input images represent the same scene. The processor:

[0014] Access a first set of backward shaping functions generated based on a first set of input images and a second set of input images, where the backward shaping functions map pixel codewords from a first codeword representation in a first dynamic range to a second codeword representation in a second dynamic range, and each backward shaping function is characterized by a shaping index parameter and a segment parameter of a segment-based representation of the backward shaping function having a set of common pivots;

[0015] Adjust the segment parameters in the first set of backward shaping functions to generate an output set of backward shaping functions having the same set of common pivots as the first set of backward shaping functions but having an adjusted segment-based polynomial representation, where adjusting the segment parameters in the backward shaping functions to generate updated polynomial coefficients includes:

[0016] For a segment of a backward shaping function between a first pivot and a second pivot, where the segment is represented by three original polynomial coefficients in an original polynomial representation:

[0017] Access the three original polynomial coefficients of the segment; and

[0018] For one or more new values of a first updated polynomial coefficient of the segment:

[0019] Based on the first pivot and the second pivot, the new value of the first updated polynomial coefficient, and the corresponding outputs of the first pivot and the second pivot according to the original polynomial representation of the segment, solve for a second updated polynomial coefficient and a third updated polynomial coefficient of the segment; and

[0020] When a distortion criterion is satisfied, update an optimal set of polynomial coefficients of the segment with the updated polynomial coefficients.

[0021] According to some embodiments, a method for generating a shaping function using a processor is provided, the method including:

[0022] Access a first set of input images in a first dynamic range and a second set of input images in a second dynamic range, where corresponding pairs between the first set of input images and the second set of input images represent the same scene;

[0023] Access a first set of backward shaping functions generated based on the first set of input images and the second set of input images, where the backward shaping functions map pixel codewords from a first codeword representation in a first dynamic range to a second codeword representation in a second dynamic range, and each backward shaping function is characterized by a shaping index parameter and a segment parameter of a segment-based representation of the backward shaping function having a set of common pivots (e.g., the set of common pivots may be common among the backward shaping functions in the first set of backward shaping functions); and

[0024] Adjust the segment parameters in the first set of backward shaping functions to generate an output set of backward shaping functions that have a common set of pivots with the first set of backward shaping functions but have an adjusted segment-based polynomial representation, wherein adjusting the segment parameters in the backward shaping functions to generate updated polynomial coefficients includes:

[0025] For a segment of the backward shaping function between a first pivot and a second pivot, where the segment is represented by three original polynomial coefficients in the original polynomial representation:

[0026] Access the three original polynomial coefficients of the segment;

[0027] Initialize a distortion parameter to a first distortion value; and

[0028] For one or more new values of a first updated polynomial coefficient of the segment:

[0029] Based on the first pivot and the second pivot, the new values of the first updated polynomial coefficient, and the corresponding outputs of the first pivot and the second pivot according to the original polynomial representation of the segment, solve for the second updated polynomial coefficient and the third updated polynomial coefficient of the segment; and

[0030] When a distortion criterion is satisfied, update the optimal set of polynomial coefficients of the segment with the updated polynomial coefficients.

[0031] The first set of backward shaping functions can be generated from an initial set of forward shaping functions generated using a first set of input images and a second set of input images.

[0032] In the sense defined by a shaping index parameter and segment parameters, each backward shaping function (e.g., in the first set of backward shaping functions) can be characterized by the shaping index parameter and the segment parameters.

[0033] Each backward shaping function can be characterized by a different shaping index parameter.

[0034] Each backward shaping function and / or forward shaping function can be calculated for characteristics that differ between the input images in the second set of input images. The characteristics can include a measurement or metric of the average luminance of the input images in a second dynamic range. For example, each backward shaping function (and forward shaping function) can be calculated for different average luminances of the input images in the second dynamic range.

[0035] According to some embodiments, the method may further include: generating a set of output forward shaping functions based on a set of output backward shaping functions, wherein the forward shaping functions map pixel codewords from a second codeword representation in a second dynamic range to a first codeword representation in a first dynamic range. Each output forward shaping function in the set of output forward shaping functions may be characterized by a shaping index parameter (e.g., different shaping index parameters) and a segment parameter of a segment-based representation of the output forward shaping function having a set of common pivots. Generating the output forward shaping functions corresponding to the output backward shaping functions may include: for each input codeword in the second dynamic range, identifying the codeword index of the output backward shaping function that minimizes the difference between the codeword generated by the backward shaping function and the input codeword.

[0036] According to some embodiments, the method may further include: generating a new forward shaping function by interpolating between two forward shaping functions in the set of output forward shaping functions. Thus, a new forward shaping function may be generated based on two forward shaping functions characterized by different shaping index parameters.

[0037] According to some embodiments, the method may further (or alternatively) include: generating a new backward shaping function by interpolating between two backward shaping functions in the output set of backward shaping functions. Thus, a new backward shaping function may be generated based on two backward shaping functions characterized by different shaping index parameters.

[0038] Interpolating between two backward shaping functions or two forward shaping functions may include: based on the respective set of polynomial coefficients of the two backward or forward shaping functions, generating a set of interpolated polynomial coefficients for each segment of the new backward or forward shaping function. The set of interpolated polynomial coefficients may be generated using linear interpolation, e.g., between the corresponding polynomial coefficients of the set of polynomial coefficients of the two backward or forward shaping functions.

[0039] The method may further include: selecting the two backward shaping functions or forward shaping functions based on a measurement of the average luminance of the input HDR image to be encoded (e.g., in the second dynamic range). Each output backward shaping function and / or each output forward shaping function may be calculated for different average luminances of the input image in the second dynamic range. The two backward shaping functions or forward shaping functions may be selected such that the average luminance of the input HDR image to be encoded lies between the respective measurements of the average luminances for which the selected two backward or forward shaping functions were calculated. The method may further include encoding the input HDR image using the new forward shaping function. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Embodiments of the present invention are illustrated in the accompanying drawings by way of example and not limitation, and like reference numerals refer to like elements, and in the drawings:

[0041] FIG. 1A depicts an example single-layer encoder for HDR data using a shaping function according to the prior art;

[0042] FIG. 1B depicts an example HDR decoder corresponding to the encoder of FIG. 1A according to the prior art;

[0043] Figure 2 An example process for designing a shaping function according to an embodiment of the present invention is depicted;

[0044] Figure 3 An example process for reducing the gap between two second-order polynomial segments in a multi-segment representation of a backward shaping function according to an embodiment of the present invention is depicted; and

[0045] Figure 4 An example encoder using a shaping function according to an embodiment is depicted. Detailed Description

[0046] Image shaping techniques for efficient encoding of images are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it is apparent that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in detail so as not to unnecessarily obscure, obfuscate or confound the present invention.

[0047] In the present disclosure, an "output forward shaping function" refers to a forward shaping function in the output set of a forward shaping function, an "output backward shaping function" refers to a backward shaping function in the output set of a backward shaping function, an "output set of forward shaping functions" refers to the output set of a forward shaping function, and an "output set of backward shaping functions" refers to the output set of a backward shaping function.

[0048] Example HDR Encoding and Decoding System

[0049] As described in References [1] and [2], FIGS. 1A and 1B illustrate an example single-layer backward-compatible codec framework using image shaping. More specifically, FIG. 1A illustrates an example encoder-side codec architecture that may be implemented with one or more computing processors in an upstream video encoder. FIG. 1B illustrates an example decoder-side codec architecture that may also be implemented with one or more computing processors in one or more downstream video decoders.

[0050] Under this framework, given a reference HDR content (120), the corresponding SDR content (134) (also referred to as the reshaped content) is encoded and transmitted in a single layer of the encoded video signal (144) by an upstream encoding device implementing an encoder-side codec architecture. The SDR content is received and decoded in a single layer of the video signal by a downstream decoding device implementing a decoder-side codec architecture. The backward reshaping metadata (152) is also encoded and transmitted in the video signal together with the SDR content, such that an HDR display device can reconstruct the HDR content based on the SDR content and the backward reshaping metadata.

[0051] As shown in FIG. 1A, a backward-compatible SDR image, such as SDR image (134), is generated using a forward reshaping map (132). Herein, a "backward-compatible SDR image" may refer to an SDR image that is specifically optimized or color-graded for an SDR display. A compression block 142 (e.g., an encoder implemented according to any known video coding algorithm such as AVC, HEVC, AV1, etc.) compresses / encodes the SDR image (134) into a single layer 144 of the video signal.

[0052] A forward reshaping function in 132 is generated using a forward reshaping function generator 130 based on a reference HDR image (120). Given the forward reshaping function, the forward reshaping map (132) is applied to the HDR image (120) to generate a reshaped SDR base layer 134. Additionally, a backward reshaping function generator 150 can generate a backward reshaping function, which can be transmitted as metadata 152 to the decoder.

[0053] Examples of backward reshaping metadata representing / specifying an optimal backward reshaping function may include but are not necessarily limited to any one of the following: an inverse tone mapping function, an inverse luminance mapping function, an inverse chrominance mapping function, a look-up table (LUT), a polynomial, inverse display management coefficients or parameters, etc. In various embodiments, the luminance backward reshaping function and the chrominance backward reshaping function can be obtained / optimized jointly or separately, and various techniques described in reference [2] can be used to obtain them.

[0054] The backward reshaping metadata (152) generated by the backward reshaping function generator (150) based on the SDR image (134) and the target HDR image (120) can be multiplexed as part of the video signal 144, e.g., passed as a supplementary enhancement information (SEI) message.

[0055] In some embodiments, the backward reshaping metadata (152) is carried in the video signal as part of the overall image metadata, which is carried in the video signal separately from the single layer in which the SDR image is encoded in the video signal. For example, the backward reshaping metadata (152) can be encoded in a component stream in the encoded bitstream, which may or may not be separate from the single layer (of the encoded bitstream) in which the SDR image (134) is encoded.

[0056] Thus, the backward reshaping metadata (152) can be generated or pre-generated on the encoder side to leverage the powerful computational resources and offline encoding processes available on the encoder side (including but not limited to content-adaptive multi-pass, look-ahead operations, inverse luminance mapping, inverse chrominance mapping, CDF-based histogram approximation, and / or transfer, etc.).

[0057] The encoder-side architecture of FIG. 1A can be used to avoid directly encoding the target HDR image (120) as the encoded / compressed HDR image in the video signal; instead, the backward reshaping metadata (152) in the video signal can be used to enable a downstream decoding device to backward reshape the SDR image (134) (encoded in the video signal) into a reconstructed image that is the same as or a close / best approximation of the reference HDR image (120).

[0058] In some embodiments, as illustrated in FIG. 1B, a video signal encoded with an SDR image in a single layer (144) and the backward reshaping metadata (152) as part of the overall image metadata are received as inputs on the decoder side of the codec framework. The decompression block 154 decompresses / decodes the compressed video data in the single layer (144) of the video signal into a decoded SDR image (156). The decompression 154 typically corresponds to the inverse process of the compression 142. The decoded SDR image (156) can be the same as the SDR image (134), but has undergone quantization errors in the compression block (142) and the decompression block (154), which have been optimized for the SDR display device. The decoded SDR image (156) can be output in the output SDR video signal (e.g., via an HDMI interface, via a video link, etc.) for presentation on an SDR display device.

[0059] In addition, a backward shaping block 158 extracts backward shaping metadata (152) from the input video signal, constructs an optimal backward shaping function based on the backward shaping metadata (152), and performs a backward shaping operation on the decoded SDR image (156) based on the optimal backward shaping function to generate a backward shaped image (160) (or a reconstructed HDR image). In some embodiments, the backward shaped image represents a production quality or near-production quality HDR image that is the same as or close to / best approximates the reference HDR image (120). The backward shaped image (160) can be output in the form of an output HDR video signal (e.g., via an HDMI interface, via a video link, etc.) for presentation on an HDR display device.

[0060] In some embodiments, as part of an HDR image presentation operation for presenting the backward shaped image (160) on an HDR display device, display management operations specific to the HDR display device can be performed on the backward shaped image (160).

[0061] Example Systems for Adaptive Shaping

[0062] Introduction

[0063] Shaping can be static or dynamic. In static shaping, a single shaping function is generated and used for a single stream or across multiple streams. In dynamic shaping, the shaping function can be customized based on input video stream characteristics, which can change at the stream level, scene level, or even frame level. For example, in an embodiment, non-limitingly, a shaping function can be generated based on a measure of the average luminance value in a frame or scene (referred to as L1-mid). For example, non-limitingly, in an embodiment with PQ-encoded RGB data, L1-mid can represent the average of the max(R, G, B) values among all RGB pixels in the region of interest in a frame. In another embodiment, for YCbCr or ICtCp-encoded data, L1-mid can represent the average of all Y or I values in the region of interest in a frame (e.g., calculating the average may exclude the letterbox or sidebar regions in the frame).

[0064] Thus, for a 12-bit system, all 4,096 possible functions can be pre-constructed; however, due to the large memory requirements in a practical system, this approach is both time-consuming and impractical. In an embodiment, a smaller L (L < 2 位深度) A set of curves is used as basis shaping functions, stored in memory, and then additional functions are generated for missing intermediate luminance values by interpolating among the available L functions during runtime. This can be referred to as the "scalable-static mode" as it combines a small set of statically generated shaping curves to generate the complete set. For example, in an embodiment, for a 10-bit signal, 13 basis forward functions can be generated for the average luminance values in the set including {768 1024 1280 1536 1792 2048 2304 2560 2816 3072 3328 3584 3840}.

[0065] In an embodiment, to save bandwidth, the backward shaping function can be approximated using piecewise linear or non-linear polynomials, where in such a representation, the polynomial segments are separated by pivot points. To facilitate interpolation of the polynomial coefficients (and avoid additional computations), in an embodiment, the pivot points in all pre-computed shaping functions should be aligned (e.g., common). This allows for simpler interpolation between polynomials without worrying about their pivot points.

[0066] Nomenclature

[0067] Consider a database containing a reference (or "master") HDR set and multiple SDR sets of images or video clips generated for different L1-mid values, i.e., for each HDR image, there is a set of corresponding SDR images using Yc0c1 color data (e.g., YCbCr data where y = Y, c0 = Cb, and c1 = Cr). The SDR images can be generated from the HDR images manually, with the help of a color grader, automatically, using an automatic color mapping algorithm, or by using a combination of computer tools and human interaction.

[0068] Let and represent the data value of the i-th pixel of the j-th frame or picture in the l-th SDR database set generated from the L1-mid mapping. Denote the number of pixels in each frame as P. Let the bit depth of the HDR image be represented as HDR_bitdepth, and the bit depth in the SDR image be represented as SDR_bitdepth. Then the number of possible codeword values in the HDR and SDR signals are given by N V = 2 HDR_bitdepth and N S = 2 SDR_bitdepth respectively.

[0069] As used herein, the term "unnormalized pixel value" represents [0, 2 BValues in [-1], where B represents the bit depth of the pixel values (e.g., B = 8, 10, or 12 bits). As used herein, the term "normalized pixel value" represents a pixel value in [0, 1).

[0070] In some embodiments, rather than operating at the pixel level, it is possible to operate with average pixel values. For example, the input signal codewords can be divided into M non-overlapping bins with equal intervals w b (e.g., for 16-bit input data, w b = 65,536 / M) (e.g., M = 16, 32, or 64) to cover the entire normalized dynamic range (e.g., (0, 1]). Then, rather than operating with pixel values, it is possible to operate with the average pixel value within each such bin. Let the number of bins in the SDR and HDR signals be represented as M S and M V , and let their corresponding intervals be represented as w bS and w bV . When operating at the pixel level, then M S = N S and / or M V = N V .

[0071] Represent the minimum luminance value and the maximum luminance value within the j-th frame or picture in the HDR database as and Represent the minimum luminance value and the maximum luminance value within the j-th frame in the SDR database as and Construct a separate forward shaping function

[0072] Before constructing a series of basis shaping functions with a common pivot point, a separate shaping function needs to be constructed. In an embodiment and without limitation, such a function is constructed using the histogram or cumulative density function (CDF) matching method used in references [1-3]. For completeness, the algorithm is also described herein. The key steps include: a) collecting the statistics (histogram) of the images in the database, b) generating the cumulative density function (CDF) for each set, c) applying CDF matching (reference [3]) to generate the shaping function, and d) cropping and smoothing the shaping function. These steps are depicted in the pseudocode in Tables 1 and 2.

[0073] Table 1: Example steps for collecting statistics

[0074]

[0075] Table 2: Example of constructing a forward shaping function using CDF matching

[0076]

[0077]

[0078] In Table 2, the function y = clip3(x, Min, Max) is defined as:

[0079]

[0080] In Table 2, the CDF matching step (Step 4) can be simply explained as follows. Considering the SDR codeword x s corresponding to a specific CDF value c in CDFc s,(l) and the HDR codeword x v also corresponding to the same specific CDF value c in CDFc v,(l) , then the SDR value s = x s is determined to be mapped to the HDR value x v . Alternatively, in Step 4, for each SDR value (x s ), the corresponding SDR CDF value (say c) is calculated and then an attempt is made to identify from the existing HDR CDF values the HDR value (x v,(l) ) for which c v = c via simple linear interpolation.

[0081] By repeating the steps in Table 1 and Table 2 for each L1 - mid value of interest, a first set of forward shaping functions can be generated. The first set of forward shaping functions maps the HDR input image (or frame) to the corresponding SDR output according to its L1 - mid value or any other characteristic used to generate the l - th SDR database.

[0082] Construct a separate backward shaping function

[0083] Given an original set of separate forward shaping functions a corresponding set of separate backward shaping functions can now be generated where the SDR codewords are mapped to the reconstructed HDR codewords. An example process is depicted in Table 3.

[0084] Table 3: Example process for generating the backward shaping function

[0085]

[0086] Construct a series of backward shaping functions with a common pivot point

[0087] At this stage, there is a set of backward shaping functions and a set of forward shaping functions set. In an embodiment, the luminance backward shaping function can be represented / approximated by a multi-segment polynomial (e.g., using eight second-order polynomials), where the polynomial segments are separated by pivot points. To implement interpolation in the existing L1-mid function, it is desirable that for all functions have a common pivot point. In an embodiment, the set of common pivot points is generated based on the technique described in reference [6] summarized below.

[0088] For each SDR codeword b, its normalized value is calculated

[0089]

[0090] Represent the pivot points as {λ m}, m = 0, 1, …, K, where K represents the total number of segments. For example, [λ m , λ m+1 represents the m-th polynomial segment. The m-th polynomial is selected when the input value b is within λ m and λ m+1 . In an embodiment, both the values of b and λ m are integers in [0, 2 B )), where B represents the SDR bit depth. The m-th second-order polynomial of the l-th L1-mid shaping function is used to approximate the previously generated input backward shaping function, as follows:

[0091] When λ m ≤ b < λ m+1 ,

[0092] To reduce the gap between two adjacent polynomials, an overlap constraint can be applied to smooth the transition around the pivot points. The overlap window size is represented as W m . To achieve the minimum gap, the design optimization objective can be formulated as

[0093]

[0094] where and 0 ≤ l ≤ L - 1.

[0095] In equation (3), the pivot points can be bounded by specific communication constraints (e.g., the valid range of SMPTE 274M codeword values), and the communication constraints can be represented as a lower limit or can also be represented as an upper limit where values below and above will be clipped. For example, for an 8-bit signal, the broadcast safe area is [16, 235] instead of [0, 255].

[0096] Using overlapping windows, the expansion points are represented as

[0097]

[0098]

[0099] Thus, in the matrix representation,

[0100]

[0101] and Equation (3) can be represented as

[0102]

[0103] Under least squares optimization, the solution is given by

[0104]

[0105] Thus, the overall problem to be solved is given by: under the combination of optimization criteria

[0106] · For all m and all l, solve

[0107] · For all m, solve for {λ m} using the set of windowing parameters {W m}.

[0108] The set of overlapping window sizes {W m} can be varied to aid in the search for the best result. Table 4 depicts an example process for solving a set of multi-segment polynomials with a common pivot, where Monte Carlo simulation can be applied by performing U iterations (e.g., U = 10,000) of randomized pivot points and overlapping windows

[0109] Table 4: Example process for identifying common pivots and polynomial segments in a set of backward shaping functions

[0110]

[0111] In step 2, regarding checking implementation constraints, some embodiments can limit the accuracy of polynomial coefficients (e.g., limit to signed 7-bit or 8-bit integer values, etc.).

[0112] As discussed, the above joint optimization may not output a smooth backward shaping function without jumps near the pivot points. In an embodiment, another round of optimization may need to be performed as follows to minimize the gap.

[0113] Consider the pivot point λ m+1, where the endpoints of the two polynomials (the m-th and m+1 polynomials) should be joined together. When using a second-order polynomial and predicting the HDR value using the SDR codeword received at the position of λ m , the equation for applying the m-th polynomial is:

[0114]

[0115] For the remaining codewords in the m-th segment, the predicted HDR value is calculated as:

[0116] When λ m < b < λ m+1

[0117] To reduce the gap between the pivots, in an embodiment, the goal is to adjust the polynomial coefficients of the m-th segment such that the revised m-th polynomial satisfies the output values of the original polynomial at and , i.e., equations (7) and

[0118]

[0119] A second-order polynomial is completely defined by three coefficients. Thus, given only these two equations (7) and (9), solving for three unknowns is an underdetermined problem. Assuming that one of the revised polynomial coefficients is known (e.g., ), the remaining polynomial coefficients (e.g., and ) can be obtained from equations (7) and (9) as a closed-form solution of a system of equations with two equations and two unknowns. For example, assuming is known, the solutions for the other two are given by

[0120]

[0121]

[0122] Since is actually unknown, the problem can be formulated as an optimization problem of finding the best coefficients such that the sum of the differences between the original HDR values and the new HDR values using the new polynomial coefficients is minimized, i.e.:

[0123]

[0124] There is no need to include the point λ m in the optimization because the equation has already passed this point. In an embodiment, without limitation, the optimal solution for can be searched for in a series of values.

[0125] For the last segment, the (m + 1)-th segment is outside the last pivot point. In an embodiment, the curve after the last point can be a constant extending from the last output value of the previous polynomial. Thus, for m = K, one can set

[0126]

[0127] For the first segment (e.g., m = 0), if the video signal range is within the full range (i.e., within [0, 1]), the pivot point λ0 will be 0. In this case, must be equal to the original coefficient. That is, for m = 0:

[0128]

[0129] Assume is known, can be obtained from Equation (9) as follows

[0130]

[0131] Example implementations of the gap reduction algorithm are depicted in Table 5 (m = 0) and Table 6 (m > 0). When m = K, the constraints of Equation (12) also apply.

[0132] Table 5: Example process for reducing the gap between pivot points, m = 0 (first segment)

[0133] Table 6: Example process for reducing the gap between pivot points, m > 0

[0134]

[0135] Adjust the forward shaping function to achieve reversibility

[0136] After reducing the gap between all consecutive segments, determine the set of final backward shaping functions. Since the gap reduction algorithm changes the backward shaping functions, the forward shaping functions need to be reconstructed to ensure proper reversibility. This can be done by backtracking the backward shaping functions.

[0137] For the l-th polynomial, the calculated polynomial coefficients can be used to reconstruct the output backward shaping function for all m

[0138] When λ m ≤ b < λ m+1 then,

[0139] The corresponding updated forward shaping function can be constructed by searching for the codeword index that minimizes the difference between the codewords generated by the backward shaping function and the original HDR codewords. For each input HDR codeword b, for optimal reconstruction, the ideal backward shaping output should be as close to b as possible. Given the value of b, assume that the value mapped from the forward shaping will be k, where k can be within the entire valid range of SDR codewords. By performing backward shaping on each valid SDR codeword the corresponding backward mapped HDR value b can be found. Among all these backward mapped HDR values the one with the smallest difference from the original input HDR codeword b can be found In other words, for input b, the forward shaping mapping should map b to Table 7 depicts an example process.

[0140] Table 7: Example process for reconstructing the forward shaping function based on an optimized backward shaping function

[0141]

[0142] Figure 2 An example process pipeline is depicted, which summarizes the process of generating the forward and backward shaping functions based on the continuity constraints and reversibility constraints between pivots described previously. Given a reference HDR and SDR image, in step 205, a first set of forward shaping functions is constructed, where each function is optimized for specific characteristics of the input HDR image (e.g., a measure of the average luminance per frame). For example, in an embodiment, non - restrictively, the CDF matching criteria described in Tables 1 and 2 can be used to generate such shaping functions.

[0143] In step 210, a first set of backward shaping functions is constructed using the first set of forward shaping functions, for example, using the process depicted in Table 3. Next, in step 215, a second set of backward shaping functions is constructed by: a) applying piecewise approximation to the first set of backward shaping functions under the constraint that all functions in the first set of forward shaping functions share a common set of pivots. An example process is provided in Table 4.

[0144] To improve image quality and reduce visual artifacts, in step 220, the polynomial representation of the second set of forward shaping functions is further optimized to reduce the distance between values in adjacent pivot points, thereby generating an output set of backward shaping functions. Example processes are provided in Tables 5 and 6.

[0145] Given a set of forward shaping functions (with optimized gaps), step 225 generates a set of outputs of the forward shaping functions under the constraint that the distance between the codewords of the reference HDR input and the codewords of the reconstructed HDR input (using the output backward shaping function) is minimized. An example process is described in Table 7.

[0146] Return to processing block 220, Figure 3 An example process for reducing the gap between two second-order polynomial segments in the multi-segment representation of the backward shaping function according to an embodiment is depicted. As described in Table 6, for segment m > 0, reducing the gap between two segments includes the following steps:

[0147] Step 305: Initialization. This step initializes the distortion (D) to a very large number and, given the original polynomial coefficients of the m-th and m+1-th segments of the l-th shaping function, sets how to calculate the HDR prediction value for the codewords within the m-th segment and the pivot point λ, e.g.: m+1 When λ

[0148] When λ m ≤b<λ m+1 When,

[0149]

[0150] Step 310 starts an iterative process that iterates over values in the range [A, B] in increments of C (e.g., without limitation, and and

[0151] · In step 315, it calculates the other two coefficients and (see Equation (10))

[0152] · In step 320, it calculates the total distortion (see Table 6)

[0153]

[0154] · In step 325, if D’ < D, the set of calculated updated polynomial coefficients replaces the previous coefficients as the new set of optimal polynomial coefficients

[0155] The iterative process repeats until all values within [A, B] have been processed.

[0156] · In step 340, at the end of the iterative process, the final set of optimal polynomial coefficients is output.

[0157] Repeat for all segments and all L functions in the set of the backward shaping functions. Figure 3 The process in

[0158] For the first segment (e.g., m = 0), the same iterative process can still be applied, but with some minor changes. After initialization (see Equation (15), but with m = 0 here), since must be equal to the original coefficient, the iterative process in step 310 now changes with a step size C0 within the range [A0, B0] where, in the embodiment (see Table 5), and Given and Equations (11) and (16) can be applied in steps 315 and 320 to calculate and the distortion D'.

[0159] For the last segment (m = K), The values of (j = 0, 1, and 2 (see Equation (15))) are calculated using Equation (12).

[0160] Shaping function interpolation

[0161] Given a set of L pre-computed forward shaping functions, each function is calculated for the values in a set of different adaptive control signals {r (0) , r (1) , … r (L-1)}. For example, if r represents the average luminance value, the corresponding forward shaping functions for the three color planes (e.g., YCbCr) can be expressed as

[0162]

[0163] Given a new control signal value (r) between two pre-computed values, i.e., r (l) ≤ r < r (l+1) , and wanting to obtain a new forward shaping function from the existing pre-computed functions. Assume that the new reconstructed HDR sample can be interpolated from two neighboring samples. Let Then, using linear interpolation, the HDR interpolated value can be expressed as

[0164]

[0165] Next, check the interpolation of the luminance and chrominance shaping functions.

[0166] Interpolation of the luminance interpolation function

[0167] According to Equations (15) and (16)

[0168]

[0169] Alternatively, the functional form can be expressed as

[0170]

[0171] Assuming that all pivot points are aligned, in an embodiment, an interpolation function can be generated by directly generating an interpolation set of polynomial coefficients for each segment. For example and without limitation, consider a multi-segment polynomial format with a second-order polynomial. Suppose the HDR range under consideration is within the m-th sub-segment / segment, then the corresponding SDR shaping values using the l-th and l+1-th forward shaping functions can be expressed as:

[0172]

[0173] After polynomial interpolation, the new polynomial coefficients are expressed as

[0174]

[0175] Then, the interpolated and shaped SDR value can be expressed as

[0176]

[0177] The backward shaping function can be constructed in the same way Interpolation of the chroma shaping function

[0178] In an embodiment, and without loss of generality, instead of expressing the shaping function as a multi-segment polynomial (e.g., as discussed previously for the luminance component), alternative schemes such as multi-color channel, multiple regression prediction, etc. discussed in References [4] and [6] can be used to represent shaping, where the chroma value is predicted based on a combination of the luminance value and the chroma value.

[0179] In Reference [6], it has been shown that, given r (l) ≤r<r (l+1) the MMR coefficient can also be a linear combination of two adjacent MMR coefficients, or:

[0180]

[0181] where represents the set of MMR coefficients of the l-th shaping function. Alternatively,

[0182]

[0183] where C can be color component c0 or c1 (e.g., Cb or Cr).

[0184] Figure 4 depicts an example of shaping based on interpolation of basis forward and backward shaping functions in an encoder according to an embodiment. As Figure 4 depicted, the forward shaping stage may include: a set of basis forward shaping functions (405); a function interpolation unit (410) that may generate a new forward shaping function (412) by interpolating from two basis forward shaping functions; and a forward shaping unit (415) that will apply the generated forward function (412) to generate a shaped signal (417), such as an SDR signal.

[0185] Given the forward shaping function (412), the encoder may generate the parameters of the reverse or backward shaping function (e.g., 150) (e.g., see reference [5]), which may be transmitted to the decoder, as shown in FIG. 1. Alternatively, as Figure 4 shown, the encoder may include a separate backward shaping stage that may include: a set of basis backward shaping functions (420); and a second function interpolation unit (425) that may generate a new backward shaping function (427) by interpolating from two basis backward shaping functions. The parameters of the backward shaping function may be conveyed as metadata. As previously discussed, the shaping functions may be selected based on a measure of the average luminance of the input HDR signal.

[0186] For decoding, the decoder may use the system functionality of FIG. 1B. In another embodiment, to reduce the amount of metadata transmitted, given the L1-mid value used by the encoder, function interpolation may also be performed at the decoder side (e.g., see reference [6]).

[0187] References

[0188] Each of these references is hereby incorporated by reference in its entirety.

[0189] 1. G-M, Su et al., “Encoding and decoding reversible, production-quality single-layer video signals,” PCT application serial number PCT / US 2017 / 023543, filed Mar. 22, 2017, WIPO publication number WO 2017 / 165494.

[0190] 2. Q. Song et al., "High-fidelity full-reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline", PCT Application Serial No. PCT / US 2019 / 031620, filed on May 9, 2019, WIPO Publication No. WO 2019 / 217751.

[0191] 3. B. Wen et al., "Inverse luma / chroma mappings with histogram transfer and approximation", U.S. Patent 10,264,287, issued on April 16, 2019.

[0192] 4. G-M. Su et al., "Multiple color channel multiple regression predictor", U.S. Patent 8,811,490.

[0193] 5. A. Kheradmand et al., "Block-based content-adaptive reshaping for high-dynamic range", U.S. Patent 10,032,262.

[0194] 6. H. Kadu et al., "Interpolation of reshaping functions", PCT Application Serial No. PCT / US 2019 / 063796, filed on November 27, 2019.

[0195] Example computer system implementations

[0196] Embodiments of the present invention can be implemented using a computer system, a system configured with electronic circuits and components, an integrated circuit (IC) device (such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete-time or digital signal processor (DSP), an application-specific IC (ASIC)), and / or an apparatus including one or more of such systems, devices, or components. The computer and / or the IC can execute, control, or implement instructions related to the generation of the shaping function, such as those described herein. The computer and / or the IC can calculate any of the various parameters or values related to the generation of the shaping function described herein. The image and video dynamic range extension embodiments can be implemented in hardware, software, firmware, and their various combinations.

[0197] Certain embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, an encoder, a set-top box, a transcoder, etc. can implement the method of generating the shaping function as described above by executing software instructions in a program memory accessible to the processor. The present invention can also be provided in the form of a program product. The program product can include any non-transitory and tangible medium that bears a set of computer-readable signals, the set of computer-readable signals including instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. The program product according to the present invention can take any of various non-transitory and tangible forms. The program product can include, for example, a physical medium, such as a magnetic data storage medium including a floppy disk, a hard disk drive, an optical data storage medium including a CD ROM, a DVD, an electronic data storage medium including a ROM, a flash RAM, etc. The computer-readable signals on the program product can optionally be compressed or encrypted.

[0198] In the case of the above-mentioned components (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, a reference to such a component (including a reference to an "apparatus") should be construed to include any component that performs the function of the described component as an equivalent (e.g., functionally equivalent) of the component, including a component that is not structurally equivalent to the disclosed structure that performs the function in the illustrated exemplary embodiments of the present invention.

[0199] Equivalents, Extensions, Alternatives, and Miscellaneous

[0200] Accordingly, example embodiments related to the generation of shaping functions for HDR images are described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary according to the implementation. Therefore, the only and exclusive indication of the present invention and the applicant's inventive intent is the set of claims issued in a specific form according to this application, where such claim issuance includes any subsequent amendments. Any definition explicitly set forth herein for terms contained in such claims should govern the meaning of such terms as used in the claims. Accordingly, limitations, elements, characteristics, features, advantages, or attributes not explicitly recited in the claims should not in any way limit the scope of such claims. Therefore, the present specification and the drawings should be regarded in an illustrative rather than a restrictive sense.

Claims

1. A method for generating a shaping function using a processor, the method comprising: Accessing a first set of input images in a first dynamic range and a second set of input images in a second dynamic range, wherein a corresponding pair between the first set of input images and the second set of input images represents the same scene; Accessing a first set of backward shaping functions generated based on the first set of input images and the second set of input images, wherein each backward shaping function in the first set of backward shaping functions maps a pixel codeword from a first codeword representation in the first dynamic range to a second codeword representation in the second dynamic range, and each backward shaping function is characterized by a shaping index parameter and segment parameters of a segment-based representation of the backward shaping function having a common set of pivots; and Adjusting the segment parameters in the first set of backward shaping functions to generate an output set of backward shaping functions having the same common set of pivots as the first set of backward shaping functions but having an adjusted segment-based polynomial representation, wherein adjusting the segment parameters in the backward shaping function to generate updated polynomial coefficients includes: For a segment of a backward shaping function between a first pivot and a second pivot, wherein the segment is represented by three original polynomial coefficients in an original polynomial representation: i. Accessing the three original polynomial coefficients of the segment; ii. Initializing a distortion parameter to a first distortion value; and iii. For one or more new values of a first updated polynomial coefficient of the segment: a. Solving a system of equations for a second updated polynomial coefficient and a third updated polynomial coefficient of the segment based on: The first pivot and the second pivot, The new value of the first updated polynomial coefficient, and The corresponding outputs of the first pivot and the second pivot according to the original polynomial representation of the segment; and b. Updating an optimal set of polynomial coefficients of the segment with the updated polynomial coefficients when a distortion criterion is met.

2. The method according to claim 1, wherein, Meeting the distortion criterion includes: After solving for the updated polynomial coefficients, Calculating a new distortion parameter value based on the updated polynomial coefficients and the original polynomial coefficients; Comparing the first distortion value with the new distortion parameter value, and If the new distortion parameter value is less than the first distortion value, the distortion criterion is met, and the first distortion value is updated with the new distortion parameter value.

3. The method according to claim 2, wherein, Calculating the new distortion parameter value includes: calculating a cumulative error between outputs generated using the original polynomial representation and the updated polynomial representation for all codewords within the segment.

4. The method of claim 1, further comprising Generate an output set of the forward shaping function based on the output set of the backward shaping function, where, A forward shaping function maps a pixel codeword from a second codeword representation in the second dynamic range to the first codeword representation in the first dynamic range.

5. The method according to claim 4, wherein Generating an output set of forward shaping functions in the output set of forward shaping functions corresponding to the backward shaping functions in the output set of backward shaping functions includes: For each input codeword in the second dynamic range, identify the codeword index in the output set of the backward shaping function that minimizes the difference between the codeword generated by the backward shaping function and the input codeword.

6. The method according to claim 1, wherein, The shaping index parameter includes a measurement of the average luminance in the input image in the second dynamic range.

7. The method according to claim 1, wherein The first dynamic range includes a standard dynamic range, and the second dynamic range includes a high dynamic range.

8. The method according to claim 2, wherein, For the first segment in the segment-based representation of a backward shaping function that includes two or more segments m = 0, the original polynomial representation includes wherein, , and represent the original polynomial coefficients, l represents the shaping index parameter of the backward shaping function in the first set of the backward shaping functions, and represents the input codeword in the first dynamic range between the first pivot ( ) and the second pivot ( ), and given the first updated polynomial coefficient , solving for the second updated polynomial coefficient and the third updated polynomial coefficient includes calculating: Among them, represents the optimal update coefficient of represents the updated coefficient, and Represents the output of the original polynomial representation of the second segment in the backward shaping function for the second pivot.

9. The method according to claim 8, wherein, The first updated polynomial coefficient is iterated within the range of values of , where n is an integer value less than 5.

10. The method according to claim 8, wherein, Calculating the new distortion parameter value ( D’ ) includes calculating 。 11. The method according to claim 2, wherein, For segment m , m > 0, in the segment-based representation of the backward shaping function, the original polynomial representation includes Among them, , and represent the original polynomial coefficients, l represents the shaping index parameter of the backward shaping function in the set of the backward shaping functions, and represents the input codeword in the first dynamic range between the first pivot ( ) and the second pivot ), and given the first updated polynomial coefficient , solving for the second updated polynomial coefficient ( ) and the third updated polynomial coefficient ( ) includes calculating: Wherein, Represents the output of the original polynomial representation of the segment in the backward shaping function for the second pivot m + 1.

12. The method according to claim 11, wherein, The first updated coefficient is iterated within the range of , where n is an integer value less than 5.

13. The method according to claim 11, wherein, Calculating the new distortion parameter value ( D’ ) includes calculating 。 14. The method according to claim 1, further comprising: A new backward shaping function is generated by interpolating between two backward shaping functions in the output set of the backward shaping function, the two backward shaping functions being characterized by different shaping index parameters.

15. The method according to claim 4, further comprising: A new forward shaping function is generated by interpolating between two forward shaping functions in the output set of the forward shaping function, the two forward shaping functions being characterized by different shaping index parameters.

16. The method according to claim 14, wherein, Interpolating between two backward or forward shaping functions includes: generating a set of interpolated polynomial coefficients for each segment of the new backward or forward shaping function based on the set of polynomial coefficients of the two backward or forward shaping functions.

17. The method according to claim 16, further comprising: Based on a measurement of the average luminance of the input HDR image to be encoded, the two backward or forward shaping functions are selected.

18. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing the method according to any one of claims 1 to 17 with one or more processors.

19. An apparatus for generating a shaping function, comprising a processor and configured to perform the method according to any one of claims 1 to 17.

20. A computer program product, the computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1 - 17.

Citation Information

Patent Citations

  • Block-based content-adaptive reshaping for high dynamic range images

    US10032262B2

  • Inverse luma / chroma mappings with histogram transfer and approximation

    US10264287B2

  • Multiple color channel multiple regression predictor

    US8811490B2

  • Encoding and decoding reversible production-quality single-layer video signals

    WO2017165494A2

  • High-fidelity full reference and high-efficiency reduced reference encoding in end-to-end single-layer backward compatible encoding pipeline

    WO2019217751A1