Image enhancement through global and local reshaping.

Combining global and local reshaping operations addresses the challenge of enhancing dynamic range and local contrast in video content, ensuring improved viewing experiences and metadata preservation.

JP7831903B2Active Publication Date: 2026-03-17DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to effectively enhance dynamic range and local contrast in video content, leading to subdued viewing experiences, especially in mobile device applications, and fail to preserve metadata integrity during reshaping operations.

Method used

Implementing a combination of global and local reshaping operations, including forward and backward reshaping functions, to enhance dynamic range and local contrast, while preserving metadata integrity by minimizing changes in codeword statistics.

Benefits of technology

Enhances video content to provide improved dynamic range and local contrast, resulting in more appealing images on various displays, and maintains metadata accuracy through targeted reshaping techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007831903000118
    Figure 0007831903000118
  • Figure 0007831903000119
    Figure 0007831903000119
  • Figure 0007831903000120
    Figure 0007831903000120
Patent Text Reader

Abstract

A first re-shaping mapping is performed on the first image represented in the first domain to generate a second image represented in a second domain, the first domain having a first dynamic range different from a second dynamic range of the second domain. A second re-shaping mapping is performed on the second image represented in the second domain to generate a third image represented in the first domain, the third image perceptually different from the first image in at least one of global contrast, global saturation, local contrast, local saturation, etc. A display image is derived from the third image and rendered on a display device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Application No. 63 / 142,270 and European Patent Application No. 21153722.0, both filed on 27 January 2021, each of which is invoked by reference in whole.

[0002] technology This disclosure relates, in general terms, to image processing operations. More specifically, some embodiments of this disclosure relate to video codecs. [Background technology]

[0003] As used herein, the term “dynamic range” (DR) may relate to the human visual system’s (HVS) ability to perceive a range of intensity (e.g., luminance, luma) within an image, from, for example, the darkest black (sunset) to the brightest white (highlight). In this sense, DR relates to scene-referred intensity. DR may also relate to the ability of a display device to render a given range of intensity well or approximately. In this sense, DR relates to display-referred intensity. Unless a particular meaning is expressly designated as having a specific significance in any aspect of this description, it should be inferred that the term may be used interchangeably, for example, in either sense.

[0004] As used herein, the term “High Dynamic Range (HDR)” refers to a DR width spanning approximately 14 to 15 orders of magnitude or more within the human visual system (HVS). In practice, the DR over which humans can simultaneously perceive a wide range of intensity may be somewhat truncated compared to HDR. As used herein, the terms “Enhanced Dynamic Range (EDR)” or “Visual Dynamic Range (VDR)” refer to the DR perceptible within a scene or image by the human visual system (HVS), including eye movements, taking into account any changes in light adaptation across the scene or image, individually or interchangeably. As used herein, EDR may refer to a DR spanning 5 to 6 orders of magnitude. While perhaps somewhat narrower in relation to true scene-based HDR, EDR still represents a wide DR width and is sometimes referred to as HDR.

[0005] In practice, an image contains one or more color components in a color space (e.g., lumens Y and chromins Cb and Cr), and each color component is represented by n bits per pixel (e.g., n=8). Using nonlinear luminance coding (e.g., gamma coding), an image with n≦8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range image, while an image with n>8 is considered an improved dynamic range image.

[0006] A reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​(e.g., luminance) of an input video signal and the output screen color values ​​(e.g., screen luminance) produced by the display. For example, Non-Patent Literature 1, which is incorporated herein by reference in its entirety, defines a reference EOTF for a flat-panel display. Given a video stream, information regarding its EOTF can be embedded in the bitstream as (image) metadata. The term “metadata” as used herein refers to any auxiliary information transmitted as part of an encoded bitstream that assists the decoder in rendering the decoded image. Such metadata may include, but is not limited to, color space or color gamut information, reference display parameters, and auxiliary signal parameters, as described herein. [Non-Patent Document 1] ITU Rec. ITU-R BT.1886 "Reference electro-optical transfer function for flat panel displays used in HDTV studio production" (March 2011)

[0007] As used herein, the term “PQ” refers to perceptual luminance amplitude quantization. The human visual system responds very nonlinearly to increases in light levels. A person’s ability to see a stimulus is influenced by the stimulus’s luminance, magnitude, the spatial frequencies that make up the stimulus, and the luminance level to which the eye is adapted at a particular moment in time when the stimulus is being viewed. In some embodiments, a perceptual quantization function maps a linear input gray level to an output gray level that better matches the contrast sensitivity thresholds in the human visual system. An exemplary PQ mapping function is described in Non-Patent Literature 2 (hereinafter “SMPTE”), which is incorporated herein by reference in its entirety. Here, given a fixed stimulus size, for all luminance levels (e.g., stimulus levels), the smallest visible contrast step at that luminance level is selected according to the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model). [Non-Patent Document 2] SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays"

[0008] 200-1,000 cd / m² 2 Alternatively, displays that support luminance of nits represent low dynamic range (LDR), also known as standard dynamic range (SDR), in contrast to EDR (or HDR). EDR content may be displayed on an EDR display that supports a higher dynamic range (e.g., 1,000 nits to 5,000 nits or more). Such displays may be defined using alternative EOTFs that support high luminance capabilities (e.g., 0 to 10,000 nits or more). Examples of such EOTFs are defined in SMPTE 2084 and Non-Patent Literature 3. As acknowledged herein by the inventors, improved techniques are desired for converting input video content data into output video content having high dynamic range, high local contrast, and vivid colors. [Non-Patent Document 3] Rec.ITU-R BT.2100,"Image parameter values ​​for high dynamic range television for use in production and international program exchange,"(06 / 2017)

[0009] The approaches described in this section are approaches that can be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section is eligible as prior art simply because it is included in this section. Similarly, unless otherwise indicated, it should not be assumed that any problem identified with respect to one or more approaches was recognized in any prior art based on this section. [Brief explanation of the drawing]

[0010] Embodiments of the present invention are illustrated in the drawings of the accompanying drawings, not as limitations but as examples, and similar reference numerals refer to similar elements.

[0011] [Figure 1] This illustrates an exemplary process for a video distribution pipeline. [Figure 2A] This provides an exemplary architecture for applying global and / or local reshaping operations. [Figure 2B] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2C] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2D] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2E]This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2F] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2G] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2H] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2I] This document presents an exemplary architecture for generating an output HDR video signal from an input HDR video signal. [Figure 2J] This illustrates an exemplary video application in which a source image is captured as an SDR image by a mobile device in an input SDR video signal. [Figure 2K] This example flow demonstrates using unpaired forward and backward reshaping functions to modify the brightness and / or saturation of an input video signal. [Figure 2L] This section demonstrates exemplary luma-local and chroma-local reshaping that can be performed on an input image to generate a locally reshaped output image. [Figure 2M] This provides an illustrative flow for generating a self-derived single-channel local function. [Figure 2N] This shows an exemplary flow for generating cross-channel local functions. [Figure 2O] This shows an exemplary flow for generating cross-channel local functions. [Figure 2P] This shows an exemplary flow for generating cross-channel local functions. [Figure 2Q] This shows an illustrative flow for local reshaping function selection dithering. [Figure 2R] An exemplary mapping flow for generating coefficients for the chroma-local reshaping function is shown. [Figure 2S]An exemplary mapping flow for generating coefficients for the chroma-local reshaping function is shown. [Figure 2T] An exemplary mapping flow for generating coefficients for the chroma-local reshaping function is shown.

[0012] [Figure 3A] An example of an indexed luminance-global forward reshaping function is shown. [Figure 3B] An example of an indexed Luma global back reshaping function is shown. [Figure 3C] 3C and 3F demonstrate exemplary ruma codeword mappings using a pair of global forward and backward reshaping functions, each with the same index value. [Figure 3D] 3D and 3E illustrate exemplary chroma codeword mapping using a pair of global forward and backward reshaping functions, each with the same index value. [Figure 3E] 3D and 3E illustrate exemplary chroma codeword mapping using a pair of global forward and backward reshaping functions, each with the same index value. [Figure 3F] 3C and 3F demonstrate exemplary ruma codeword mappings using a pair of global forward and backward reshaping functions, each with the same index value. [Figure 3G] An example of a lumer backward function with different index values ​​is shown. [Figure 3H] An exemplary lumer-forward function with different index values ​​is shown. [Figure 3I] An exemplary nonlinear function used to achieve a change in saturation is shown.

[0013] [Figure 4] An exemplary process flow is shown.

[0014] [Figure 5]A simplified block diagram of an exemplary hardware platform on which the computers or computing devices described herein may be implemented is shown. [Modes for carrying out the invention]

[0015] The following description includes many specific details to provide a full understanding of the disclosure for illustrative purposes. However, it will be apparent that the disclosure may be carried out without these specific details. On the other hand, well-known structures and devices are not described in exhaustive detail to avoid unnecessarily obscuring, burying, or obscuring the disclosure.

[0016] overview To reshape or transform input image data into output image data, forward or backward reshaping techniques described herein can be employed. Exemplary reshaping or transformation operations may include one or more combinations of local reshaping operations, global reshaping operations, local forward reshaping operations, global forward reshaping operations, local backward reshaping operations, global backward reshaping operations, and combinations thereof.

[0017] Global reshaping refers to a transformation or reshaping operation that applies the same global (forward and / or backward) reshaping function / mapping to every pixel of an input image to produce a corresponding output image—a reshaped image (e.g., forward, backward, circular, etc.)—that depicts the same visual semantic content as the input image.

[0018] In contrast to global reshaping, which applies the same reshaping function or mapping to all pixels of an input image, local reshaping refers to a transformation or reshaping operation that applies different (forward and / or backward) reshaping functions or mappings to different pixels of an input image. Thus, in local reshaping, the first reshaping function applied to a first pixel of an input image may be a different function from the second reshaping function applied to a second different pixel of the input image.

[0019] Exemplary global reshaping operations are described in U.S. Provisional Patent Application No. 62 / 136,402, filed March 20, 2015 (also published as U.S. Patent Application Publication No. 2018 / 0020224 on January 18, 2018), and PCT Application No. PCT / US2019 / 031620, filed May 9, 2019, PCT Application No. PCT / US2019 / 063796, filed November 27, 2019, “Interpolation of Reshaping Functions,” also published as WO2020 / 117603 on June 11, 2020; and U.S. Provisional Patent Application No. 63 / 013,063, “Reshaping functions for HDR imaging with continuity and reversibility,” filed April 21, 2020 by GM. Su. The constraints described are included in U.S. Provisional Patent Application No. 63 / 013,807, “Iterative optimization of reshaping functions in single-layer HDR image codec,” filed April 22, 2020, and the entire contents thereof are incorporated herein by reference as if they were entirely contained herein. An exemplary local reshaping operation is described in U.S. Provisional Patent Application No. 63 / 086,699, filed October 2, 2020, and the entire contents thereof are incorporated herein by reference as if they were entirely contained herein.

[0020] The reshaping techniques described herein can be used to implement a unified image enhancement method for producing reshaped or reconstructed images that have better image quality compared to the input images used to generate the reshaped or reconstructed images, by selecting, combining, and / or mixing various global and / or local reshaping operations.

[0021] For example, HDR video may provide a significantly better viewing experience than its corresponding SDR video in TV viewing applications, but HDR video may exhibit a much more subdued viewing experience in mobile device viewing applications, comparable to or similar to its corresponding SDR video.

[0022] Under the technologies described herein, HDR images within HDR video can be further enhanced beyond the existing SDR-HDR dynamic range improvements. These further enhanced HDR images may appear far more appealing in TV and mobile device viewing applications than the SDR images in the corresponding SDR video. Additionally, optionally, or alternatively, even for SDR video carried in the base layer of an SDR backward-compatible video signal or a layered video signal, the SDR images within SDR video can be enhanced with better local contrast and / or saturation to enrich or improve the viewing experience of the base layer.

[0023] As used herein, “circular reshaping” or “circularly reshaped” refers to a combination of anterior and posterior reshaping.

[0024] Circular reshaping, which consists of global reshaping operations, can be classified into paired forward / backward reshaping and unpaired forward / backward reshaping. An input image received by a paired forward / backward global reshaping operation can be reconstructed (for example, completely, faithfully, with quantization errors, etc.) using the output image generated by the paired forward / backward global reshaping operation. The luminance, hue [tone], and / or saturation of an input image received by an unpaired forward / backward global reshaping operation can be adjusted using the output image generated by the unpaired forward / backward global reshaping operation.

[0025] Local reshaping operations, such as forward, backward, and / or circular local reshaping, may not be categorized as paired or unpaired. Regardless of whether a correspondence or relationship exists between a local forward reshaping operation and a local backward reshaping operation, the local forward reshaping operation and the local backward reshaping operation do not have to be used as a pair to produce the same reconstructed image (e.g., exactly, faithfully, subject to quantization errors, etc.) as the input image used to generate the reconstructed image by the local forward reshaping operation and the local backward reshaping operation.

[0026] Local reshaping operations, whether paired or unpaired, can be used to increase local contrast ratio and saturation and to achieve better local appearance in the reshaped and / or reconstructed images produced by the local reshaping operations.

[0027] In some operational scenarios, incoming image metadata received along with the input image from an image source (e.g., an upstream encoder, media content server, media streaming server, production studio, etc.) is preserved or nearly undisturbed in the outgoing image metadata of the output video signal encoded with the reshaped or reconstructed image generated by the reshaping operations described herein.

[0028] Unpaired global reshaping operations can excessively alter the codeword statistics or distribution in the output image relative to the input image, thereby partially or entirely invalidating the incoming image metadata, or at least introducing inaccuracies in the outgoing image metadata due to the changes in codeword statistics or distribution caused by the unpaired global reshaping operation. In some operating scenarios, local reshaping operations can be applied whole or partially instead of unpaired global reshaping operations to minimize or reduce the changes in codeword statistics / distribution in the output image relative to the input image, thereby maintaining the validity and / or accuracy of the incoming image metadata. As a result, some or all of the incoming image metadata can be preserved or not corrupted in the outgoing image metadata of the output video signal encoded with the output image.

[0029] Depending on the use case, different combinations of global / local reshaping functions / operations to improve an image can, for example, produce an output image that looks better in terms of contrast ratio, hue, and / or saturation than the input image received by the reshaping function / operation.

[0030] A variety of methods can be used to generate reshaping functions as described herein. For example, local reshaping functions can be generated without limitation by using one or more of the following: (1) self-derivation methods, (2) pre-construction methods, or (3) hybrid methods combining self-derivation and pre-construction methods. Some or all of these methods can be deployed in online / real-time reshaping use cases and offline / non-real-time reshaping use cases. For example, in some operating scenarios, local reshaping functions can be generated in real time from a family of global reshaping functions trained offline.

[0031] In some operating scenarios, local reshaping can cause false contouring or banding artifacts. Local reshaping increases the slope or contrast in the mapping function, thus widening the gaps between consecutive codewords in the codeword space for the available codewords to encode the image content within the image. Local reshaping function index dithering can be applied to mitigate or mask false contouring or banding artifacts. Additionally, optionally, or alternatively, film grain injection can be performed in the reshaped codewords to further mitigate or mask false contouring or banding artifacts.

[0032] The exemplary embodiments described herein relate to generating an output image from an input image. A first reshaping mapping is performed on a first image represented by a first domain to generate a second image represented by a second domain. The first domain has a first dynamic range that is different from the second dynamic range of the second domain. A second reshaping mapping is performed on the second image represented by the second domain to generate a third image represented by the first domain. The third image is perceptually different from the first image in at least one of the following: global contrast, global saturation, local contrast, local saturation, etc. A display image is derived from the third image to be rendered on a display device.

[0033] Exemplary video delivery processing pipeline Figure 1 illustrates an exemplary process of a video distribution pipeline (100) showing various stages from video capture / generation to an HDR or SDR display. The exemplary HDR display may include, but is not limited to, an image display operating in conjunction with a TV, mobile device, home theater, etc. The exemplary SDR display may include, but is not limited to, an SDR TV, mobile device, home theater display, head-mounted display device, wearable display device, etc.

[0034] It should be noted that image processing operations described herein, such as reshaping operations, can be performed either on the encoder / server side (before video compression) or the decoder / playback side (after video decompression), as well as in the video preprocessing system or block providing the input image for video encoding. To support playback-side reshaping operations, the same system configuration as shown in Figure 1 may be used, or different system configurations may be used in addition to those shown in Figure 1. In these system configurations, a wide variety of different image metadata formats other than those used by the processing components shown in Figure 1 may be used to transmit image metadata.

[0035] A video frame, such as a sequence of consecutive input images (e.g., SDR, HDR, etc.) 102, can be received by the image generation block 105. These images (102) may be received from a video source, generated by a video preprocessing system / block outside or inside the image generation block (105), or retrieved from a video data store. Some or all of the images (102) may be generated from the source images through video editing or conversion operations (e.g., automatic without human input, manual, automatic with human input, etc.), color grading operations, etc. The source images may be digitally captured (e.g., by a digital camera), generated by converting analog camera pictures captured on film into a digital format, generated by a computer (e.g., using computer animation, image rendering, etc.). The images (102) may be images associated with one or more of the following: a movie release, an archived media program, a media program library, a video recording / clip, a media program, a TV program, user-generated video content, etc.

[0036] The image generation block (105) applies a reshaping operation (e.g., local, global, local and global, etc.) to each input image in a sequence of consecutive input images (102) to produce each corresponding output image in a sequence of consecutive output images (e.g., reshaped, reconstructed, improved, further improved, etc.) that depicts the same visual semantic content as the input images (102), but has the same or different dynamic range, higher local contrast, sharper colors, etc., compared to the input images (102).

[0037] More specifically, based on the codewords present in the input image (102), the image generation block (105) can select or construct a global reshaping function 142 and / or a local reshaping function 146 that are general or specific to the input image. The image generation block (105) can reshape or transform the input image by performing a global reshaping operation based on the global reshaping function (142) and / or a local reshaping operation based on the local reshaping function (146) to produce a corresponding output image (e.g., reshaped, reconstructed, improved, further improved, etc.) that depicts the same visual semantic information as the input image. The image generation block (105) can perform only a global reshaping operation, only a local reshaping operation, or a combination of a global reshaping operation and / or a local reshaping operation to produce a corresponding output image from the input image. The output image may have the same or different dynamic range, higher local contrast, sharper colors, etc., compared to the input image.

[0038] In some operational scenarios, some reshaping operations can be performed offline to generate pre-built reshaping mappings / functions based on training image data such as training HDR images, training SDR images, or combinations / pairs of training HDR-SDR images. These pre-built reshaping mappings / functions can be used or adapted directly or indirectly for the purpose of performing dynamic online reshaping operations in the image generation block (105) at runtime.

[0039] In some operating scenarios, the output image generated by the image generation block (105) can be provided to the composer metadata generation block (115) to generate a forward-remodeled SDR image (112) and image metadata (177) (e.g., composer metadata, back-remodeling mapping, etc.). The image metadata (177) may include composer data for generating a back-remodeling mapping, which, when applied to the forward-remodeled SDR image, generates a corresponding HDR image that depicts the same visual semantic content as the forward-remodeled SDR image. Part or all of the input image (102) can be provided to the composer metadata generation block (115) to facilitate or assist the forward-remodeling operation that generates the forward-remodeled SDR image (112) and to facilitate or assist the generation of image metadata or the back-remodeling mapping within it.

[0040] Exemplary reshaping mappings / operations include, but are not limited to, backward mappings / operations, forward mappings / operations, and / or circular mappings / operations, but may include, but are not limited to, one or more of the following: backward lookup tables (BLUT), forward lookup tables (FLUT), backward and forward reshaping functions / curves or polynomial sets, multivariate multiple regression (MMR) coefficients, tensor product B-spline (TPB) coefficients, and any combination thereof. An example of MMR operation is described in U.S. Patent No. 8,811,490, and its entire content is incorporated by reference as if it were entirely contained herein. An example of TPB operation is described in U.S. Provisional Application No. 62 / 908,770, entitled "TENSOR-PRODUCT B-SPLINE PREDICTOR," filed on 1 October 2019 (Agent Reference Number 60175-0417), and its entire contents are incorporated herein by reference as if they were fully described herein.

[0041] The reshaped SDR image (112) and image metadata (177) can be encoded by a video signal 122 (e.g., an encoded bitstream) or an encoding block 120 within a set of consecutive video segments. Given a video signal (122), a receiving device such as a mobile phone may, as part of internally processing or post-processing the video signal (122) on the device, use the metadata along with the SDR image data to generate and render an image with a higher dynamic range and sharper colors, such as HDR, within the display capabilities of the receiving device. Additionally, optionally, or alternatively, the video signal (122) or video segments may ignore the image metadata (177) and simply allow backward compatibility with legacy SDR displays that can display the SDR image represented in the SDR image data.

[0042] Exemplary video signals or video segments may include, but are not limited to, single-layer video signals / segments. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray, and other distribution formats, to generate the video signal (122) or video segment.

[0043] The video signal (122) or video segment is delivered downstream to receivers such as mobile devices, tablet computers, decoding and playback devices, media source devices, media streaming client devices, television sets (e.g., smart TVs), set-top boxes, and cinemas. In the downstream devices, the video signal (122) or video segment is decoded by a decoding block (130) to produce a decoded image 182, which may be similar to or identical to the reshaped SDR image (112), although it is affected by errors caused by quantization errors and / or transmission errors and / or synchronization errors and / or packet loss that occur in the compression performed by the encoding block (120) and the decompression performed by the decoding block (130).

[0044] In a non-restrictive example, the video signal (122) (or video segment) may be a backward-compatible SDR video signal (or video segment), where “backward-compatible” means a video signal or video segment that carries an SDR image optimized for an SDR display (e.g., retaining certain artistic intent).

[0045] The decode block (130) can also obtain or decode image metadata (177) from a video signal (122) or video segment. The image metadata (177) specifies a back-reshaping mapping which can be used by a downstream decoder to perform back-reshaping on the decoded SDR image (182) to produce a back-reshaped HDR image for rendering on an HDR display (e.g., a target display, a reference display, etc.). The back-reshaping mapping represented in the image metadata (177) can be generated by the composer metadata generation block (115) by minimizing the error or difference between the back-reshaped HDR image generated using the image metadata (177) and the output image generated by the image generation block (115) using global and / or local reshaping operations. As a result, the image metadata (177) helps ensure that the back-reshaped HDR image generated by the receiver using the image metadata (177) approximates relatively well and accurately the output image generated by the image generation block (115) using global and / or local reshaping operations.

[0046] Additionally, optionally, or alternatively, the image metadata (177) may include display management (DM) metadata that can be used by downstream decoders to perform display management operations on the back-reshaped image to generate a display image optimized for rendering on an HDR display device (e.g., an HDR display image).

[0047] In operating scenarios where the receiver operates with (or is mounted on) an SDR display 140 that supports a standard or relatively narrow dynamic range, the receiver can render the decoded SDR image directly or indirectly onto the target display (140).

[0048] In an operating scenario in which the receiver operates with (or is mounted on) an HDR display 140-1 that supports a high dynamic range (e.g., 400 nits, 1000 nits, 4000 nits, 10000 nits or more), the receiver may extract composer metadata from a video signal (122) or video segment (e.g., from a metadata container such as one within it) and compose an HDR image (132) using the composer metadata, which may be a back-reshaped image produced by back-reshaping an SDR image based on the composer metadata. Furthermore, the receiver may extract DM metadata from a video signal (122) or video segment and apply DM operations (135) to the HDR image (132) based on the DM metadata to produce a display image (137) optimized for rendering on the HDR display device (140-1), which can then be rendered on the HDR display device (140-1).

[0049] For illustrative purposes only, it has been noted that the global and / or local reshaping operations described herein can be performed by upstream devices such as video encoders that generate output images from input images. These output images are then used by the video encoder as target or reference images to generate back-reshaping metadata, which helps a receiving device generate a back-reshaped HDR image that approximates the output images generated from the global and / or local reshaping operations relatively well or accurately.

[0050] It should be noted that in various embodiments, some or all global and / or local reshaping operations can be performed by the video encoder alone, the video decoder alone, the video transcoder alone, or a combination thereof.

[0051] Global and / or local reshaping options Figure 2A shows an exemplary architecture for applying global and / or local reshaping operations. Some of all reshaping operations can be performed or executed by an image processing system, such as one or more processing blocks from Figure 1 on the encoder side, one or more processing blocks from Figure 1 on the encoder side, or a combination of processing blocks from Figure 1 on the encoder and decoder sides.

[0052] As shown in Figure 2A, block 202 represents the HDR video signal input / output interface. In various operating scenarios, the HDR video signal input / output interface (202) can be the HDR video signal input interface alone, the HDR video signal output interface alone, or a combination of the HDR video signal input interface and the HDR video signal output interface. In some operating scenarios, the input HDR video signal or the input HDR image within it may be input or received using the HDR video input interface of the HDR video signal input / output interface (202). In some operating scenarios, the output HDR video signal or the output HDR image within it may be output or transmitted using the HDR video output interface of the HDR video signal input / output interface (202).

[0053] Similarly, block 208 represents an SDR video signal input / output interface. In various operating scenarios, the SDR video signal input / output interface (208) can be an SDR video signal input interface alone, an SDR video signal output interface alone, or a combination of an SDR video signal input interface and an SDR video signal output interface. In some operating scenarios, an input SDR video signal or an input SDR image within it may be input or received using the SDR video input interface of the SDR video signal input / output interface (208). In some operating scenarios, an output SDR video signal or an output SDR image within it may be output or transmitted using the SDR video output interface of the SDR video signal input / output interface (208).

[0054] Blocks 204 and 206 represent options for forward reshaping operations, methods, and / or functions in the forward reshaping path. Blocks 210 and 212 represent options for backward reshaping operations, methods, and / or functions in the backward reshaping path. Each of the forward and backward reshaping paths includes options for global reshaping operations, methods, and / or functions that may be designed to maintain the overall (SDR or HDR) appearance of the SDR or HDR image produced by these global reshaping operations, methods, and / or functions.

[0055] In some operating scenarios, a forward global reshaping operation, method, and / or function in the forward reshaping path forms a reversible reshaping pair (e.g., perfectly, faithfully, affected by quantization errors, etc.) with a backward global reshaping operation, method, and / or function in the backward reshaping path. In these operating scenarios, an HDR image (e.g., input, received, etc.) is forward reshaped by a forward global reshaping operation, method, and / or function in the forward reshaping path to produce a forward reshaped (e.g., SDR, etc.) image, and the reconstructed or backward reshaped image produced by then applying a backward global reshaping operation, method, and / or function in the backward reshaping path is the same as the HDR image (e.g., input, received, etc.) or, subject to possible quantization errors, closely approximates the HDR image (e.g., input, received, etc.).

[0056] Similarly, an SDR image (e.g., input, received, etc.) can be retrofitted by a retro-global retrofitting operation, method, and / or function in the retro-refitting path to produce a retro-refitted (e.g., HDR, etc.) image, and then the reconstructed or forward-refitted image produced by applying a forward-global retrofitting operation, method, and / or function in the forward-refitting path is either identical to the SDR image (e.g., input, received, etc.) or, subject to possible quantization errors, closely approximates the SDR image (e.g., input, received, etc.).

[0057] While local, visually perceptible properties can be improved, local reshaping operations, methods, and / or functions in forward or backward paths are, in contrast to global reshaping operations, methods, and / or functions, either revertible or incompatible with other global or local reshaping operations in backward or forward paths.

[0058] As shown in Figure 2A, in various operating scenarios, the input video signal or the input image within it can be HDR or SDR, depending on how specific reshaping operations, methods, and / or functions are used and how specific input and / or output image interfaces (e.g., input and / or output SDR image interfaces, input and / or output HDR image interfaces, etc.) are used. Similarly, as shown in Figure 2A, the output video signal or the output (e.g., reconstructed, etc.) image within it can be HDR or SDR, depending on how specific reshaping operations, methods, and / or functions are used and how specific input and / or output interfaces are used.

[0059] Several different combinations are shown below. Firstly, the input signal is an HDR input signal and the output signal is an HDR output signal. In this case, the input HDR image in the HDR input signal is first processed by forward reshaping and then by backward reshaping to produce the corresponding output HDR image in the HDR output signal (which depicts the same visual meaning as the input HDR image). There are four (2*2=4) combinations, formed by two options using global or local reshaping in the forward path, multiplied by two options using global or local reshaping in the backward path.

[0060] Secondly, the input signal is an HDR input signal, and the output signal is an SDR output signal. In this case, the input HDR image in the HDR input signal is processed by forward reshaping to produce the corresponding output SDR image in the SDR output signal (which depicts the same visual meaning as the input HDR image). There are two combinations formed by two options: using global or local reshaping in the forward path.

[0061] Thirdly, the input signal is an SDR input signal, and the output signal is an HDR output signal. In this case, the input SDR image in the SDR input signal is processed by back-reshaping to produce the corresponding output HDR image in the HDR output signal (which exhibits the same visual meaning as the input SDR image). There are two combinations formed by two options: using global or local reshaping in the back-path.

[0062] Fourth, the input signal is an SDR input signal, and the output signal is an SDR output signal. In this case, the input SDR image in the SDR input signal is first processed by back reshaping, and then by forward reshaping to produce the corresponding output SDR image in the SDR output signal (which represents the same visual meaning as the input SDR image). There are four (2*2=4) combinations formed by two options using global or local reshaping in the back path, multiplied by two options using global or local reshaping in the forward path.

[0063] Improved HDR Several combinations of the above can be used to support HDR enhancement in the output HDR video signal shaped from the input HDR video signal.

[0064] Figure 2B illustrates an example of generating an output HDR video signal from an input HDR video signal using global reshaping operations, methods, and / or functions in both the forward and backward paths. As illustrated, the input HDR image in the input HDR video signal received via the input HDR interface of the HDR input / output interface (202) may be transformed into the output HDR image in the output HDR video signal transmitted via the output HDR interface of the HDR input / output interface (202) using global forward reshaping (204) in the forward path and global backward reshaping (210) in the backward path.

[0065] In operating scenarios where global forward reshaping (204) in the forward path and global backward reshaping (210) in the backward path form a pair or are symmetrical to each other (e.g., revertible, mathematical inverse, mathematical inverse under quantization error), the output HDR image is identical (possibly within the range of quantization error) to the input HDR image (the original input HDR image without HDR enhancement). In some operating scenarios, the reshaped SDR image produced by performing global forward reshaping (204) on the input HDR image may be a watchable, unenhanced SDR image.

[0066] In operating scenarios where the global forward reshaping (204) in the forward path and the global backward reshaping (210) in the backward path do not form a pair or are asymmetrical to each other (e.g., non-recoverable, not mathematically inverse, not mathematically inverse due to differences not attributable to quantization errors, etc.), the output HDR image is modified from the (original) input HDR image, having a different HDR appearance. For example, the global forward reshaping (204) in the forward path and the global backward reshaping (210) in the backward path can generate different HDR appearances with different luminance and saturation than the (original) appearance of the input HDR image. Additionally, optionally, or alternatively, a reshaped SDR image generated from performing global forward reshaping (204) on the input HDR image may be a viewable, non-enhanced SDR image.

[0067] Figure 2C illustrates an example of generating an output HDR video signal from an input HDR video signal using a global forward reshaping operation, method, and / or function in the forward path and a local backward reshaping operation, method, and / or function in the backward path. As illustrated, the input HDR image in the input HDR video signal received via the input HDR interface of the HDR input / output interface (202) may be converted to an output HDR image in the output HDR video signal transmitted via the output HDR interface of the HDR input / output interface (202) using global forward reshaping (204) in the forward path and local backward reshaping (212) in the backward path.

[0068] The output HDR image is improved thanks to the use of local back reshaping. Additionally, optionally, or alternatively, a reshaped SDR image generated by performing global forward reshaping (204) on the input HDR image may be a non-locally improved SDR image.

[0069] Figure 2D illustrates an example of generating an output HDR video signal from an input HDR video signal using a local forward reshaping operation, method, and / or function in the forward path and a global backward reshaping operation, method, and / or function in the backward path. As illustrated, the input HDR image in the input HDR video signal received via the input HDR interface of the HDR input / output interface (202) may be converted to an output HDR image in the output HDR video signal transmitted via the output HDR interface of the HDR input / output interface (202) using local forward reshaping (206) in the forward path and global backward reshaping (210) in the backward path.

[0070] The output HDR image is improved thanks to the use of local forward reshaping. The reshaped SDR image generated by performing local forward reshaping(206) on the input HDR image is a locally improved SDR image.

[0071] Figure 2E illustrates an example of generating an output HDR video signal from an input HDR video signal using local forward reshaping operations, methods, and / or functions in the forward path and local backward reshaping operations, methods, and / or functions in the backward path. As illustrated, the input HDR image in the input HDR video signal received via the input HDR interface of the HDR input / output interface (202) may be converted to an output HDR image in the output HDR video signal transmitted via the output HDR interface of the HDR input / output interface (202) using local forward reshaping (206) in the forward path and local backward reshaping (212) in the backward path.

[0072] The output HDR image is doubly enhanced thanks to the use of local forward reshaping and local backward reshaping in both the forward and backward paths. The reshaped SDR image produced by performing local forward reshaping(206) on the input HDR image is an enhanced SDR. Additionally, optionally, or alternatively, the global settings in one or both of the forward and backward paths may be asymmetric, thereby further altering the (original) HDR appearance of the input HDR image to a different HDR appearance with different global luminance and saturation compared to the input HDR image.

[0073] Additionally, optionally, or alternatively, in some operating scenarios, some or all of the aforementioned operations in one or more combinations of the options described above may be performed sequentially and iteratively over multiple rounds of global / local forward and global / local backward reshaping to generate further HDR and / or SDR enhancements in the output HDR image and reshaped SDR image.

[0074] Improved SDR Several combinations of the above can be used to support SDR enhancement in the output SDR video signal reshaped from the input SDR video signal.

[0075] Figure 2F shows an example of generating an output SDR video signal from an input SDR video signal using global reshaping operations, methods, and / or functions in both the forward and backward paths. As illustrated, the input SDR image in the input SDR video signal received via the input SDR interface of the SDR input / output interface (208) may be transformed into the output SDR image in the output SDR video signal transmitted via the output SDR interface of the SDR input / output interface (208) using global forward reshaping (204) in the forward path and global backward reshaping (210) in the backward path.

[0076] In operating scenarios where global forward reshaping (204) in the forward path and global backward reshaping (210) in the backward path form a pair or are symmetrical to each other (e.g., recoverable, mathematical inverse, mathematical inverse under quantization error), the output SDR image is identical (possibly subject to quantization error) to the (original) input SDR image without SDR enhancement. In some operating scenarios, the reshaped HDR image produced by performing global backward reshaping (210) on the input SDR image may be a viewable unenhanced HDR image.

[0077] In operating scenarios where the global forward reshaping (204) in the forward path and the global backward reshaping (210) in the backward path do not form a pair or are asymmetrical to each other (e.g., non-recoverable, not mathematically inverse, not mathematically inverse due to differences not attributable to quantization errors, etc.), the output SDR image is modified from the (original) input SDR image to have a different SDR appearance. For example, using global forward reshaping (204) in the forward path and global backward reshaping (210) in the backward path can generate different SDR appearances with different luminance and saturation than the (original) appearance of the input SDR image. Additionally, optionally, or alternatively, a reshaped HDR image generated from performing global backward reshaping (210) on the input SDR image may be a viewable, non-enhanced HDR image.

[0078] Figure 2G illustrates an example of generating an output SDR video signal from an input SDR video signal using a global forward reshaping operation, method, and / or function in the forward path and a local backward reshaping operation, method, and / or function in the backward path. As illustrated, the input SDR image in the input SDR video signal received via the input SDR interface of the SDR input / output interface (208) may be converted to an output SDR image in the output SDR video signal transmitted via the output SDR interface of the SDR input / output interface (208) using global forward reshaping (204) in the forward path and local backward reshaping (212) in the backward path.

[0079] The output SDR image is improved thanks to the use of local back reshaping. Additionally, optionally, or alternatively, a reshaped HDR image produced by performing global back reshaping (212) on the input SDR image may be a locally improved HDR image.

[0080] Figure 2H illustrates an example of generating an output SDR video signal from an input SDR video signal using a local forward reshaping operation, method, and / or function in the forward path and a global backward reshaping operation, method, and / or function in the backward path. As illustrated, the input SDR image in the input SDR video signal received via the input SDR interface of the SDR input / output interface (208) can be converted to the output SDR image in the output SDR video signal transmitted via the output SDR interface of the SDR input / output interface (208) using local forward reshaping (206) in the forward path and global backward reshaping (210) in the backward path.

[0081] The output SDR image is improved thanks to the use of local forward reshaping. Additionally, optionally, or alternatively, a reshaped HDR image generated by performing global backward reshaping (210) on the input SDR image may be a non-locally improved HDR image.

[0082] Figure 2I illustrates an example of generating an output SDR video signal from an input SDR video signal using local forward reshaping operations, methods, and / or functions in the forward path and local backward reshaping operations, methods, and / or functions in the backward path. As illustrated, the input SDR image in the input SDR video signal received via the input SDR interface of the SDR input / output interface (208) can be converted into an output SDR image in the output SDR video signal transmitted via the output SDR interface of the SDR input / output interface (208) using local forward reshaping (206) in the forward path and local backward reshaping (212) in the backward path.

[0083] The output SDR image is doubly enhanced thanks to the use of local forward reshaping and local backward reshaping in both the forward and backward paths. The reshaped HDR image produced by performing local backward reshaping (212) on the input SDR image is an enhanced HDR. Additionally, optionally, or alternatively, the global settings in one or both of the forward and backward paths can be made asymmetric, thereby further altering the appearance of the input SDR image as an SDR with different global luminance and saturation compared to the input SDR image.

[0084] Additionally, optionally, or alternatively, in some operating scenarios, some or all of the aforementioned operations in one or more combinations of the options described above may be performed sequentially and iteratively in multiple rounds of global / local forward and global / local backward reshaping to generate further HDR and / or SDR enhancements in the output SDR image and the reshaped HDR image.

[0085] Restorable improvements Revertible enhancements, such as those based on global reshaping in both the forward and backward paths, can be applied in many video applications. Figure 2J shows an exemplary video application in which a source image is captured as an SDR image by a mobile device in an input SDR video signal. Upconversion from SDR to HDR can be performed using global backward reshaping (e.g., 210). An upconverted (or reshaped) HDR image with at least a dynamic range improvement can be produced by performing global backward reshaping (210) on the (source or input) SDR image. Compared to global reshaping, which can be performed with relatively little or modest computing power, local reshaping may be unaffordable or impractical due to the relatively limited computing power on mobile phones. Instead, global backward reshaping (210) can be used to produce an (upconverted or reshaped) HDR image. The HDR image can be encoded as the output version of an HDR encoded bitstream. This version of the HDR encoded bitstream can be shared, previewed, and / or displayed on the display screen of a mobile device.

[0086] The HDR image can be uploaded via an HDR encoded bitstream to another device, such as a more powerful machine (e.g., a PC, laptop, cloud-based storage and / or computing system, server, media or video sharing system, etc.). As shown in Figure 2C, the HDR encoded bitstream can be received as an input HDR video signal. This version of the HDR encoded bitstream—the HDR image—generated from a global back reshaping (210) performed by the mobile device on the (original) SDR image—can be further enhanced by first restoring the (original) SDR image generated by the mobile device (e.g., by a camera(s) operating in conjunction with the mobile device, etc.) using a restorable global forward reshaping (204)—which forms a restorable pair or corresponds to the global back reshaping (210) in Figure 2J. As shown in Figure 2C, local back reshaping (212) can be applied or performed on the recovered SDR image to generate a (new) HDR image with an improved HDR appearance compared to the HDR appearance of this version of the HDR image of the HDR encoded bitstream.

[0087] Therefore, in this video application involving a mobile device that acquires the original SDR image or input SDR image, the original SDR image or input SDR image can be restored from this version of the HDR image of the HDR encoded bitstream generated by the mobile device and viewed. Furthermore, an HDR image with at least dynamic range improvement generated by the mobile device can be received and viewed, for example, from the mobile device. In addition, an HDR image with further improvement can be generated and viewed by performing local reshaping as described herein, as shown in Figure 2C.

[0088] Similarly, in video applications where the input HDR image is received in an input HDR video signal (e.g., generated by a mobile device, a non-mobile device, or a camera), global forward reshaping can be performed on the HDR image to produce an output SDR image. The output SDR image can be restored to the (original) input HDR image by forming a reconstructible pair or by using global backward reshaping corresponding to the global forward reshaping. Additionally, optionally, or alternatively, local forward reshaping can be performed on the HDR image to produce an improved output SDR image, for example, with local enhancements.

[0089] Global reshaping In some operational scenarios, the global reshaping function or mapping used in global reshaping can be designed based on a scalable static framework for single-layer backward compatible (SLBC) video coding. Algorithms that guarantee complete reversibility from one domain, such as one of the SDR and HDR domains, to another domain, such as the other of the SDR and HDR domains, can be used in global reshaping as described herein. Different uses of global reshaping may exist depending on whether the forward and backward (global reshaping) functions used by global reshaping form a reversible pair.

[0090] In some operating scenarios, a pair of forward and backward global reshaping functions are used within a scalable static framework for SLBC video encoding. In these operating scenarios, there can be a total of G pairs of fully recoverable global forward and backward reshaping functions (e.g., G = 4096; covering all possible codeword values ​​in a 12-bit codeword space). Each fully recoverable global forward and backward reshaping function pair includes a global forward reshaping function and a corresponding global backward reshaping function.

[0091] The global forward reshaping function in the above recoverable pair is expressed as follows (each global forward reshaping function transforms or forward reshapes the codeword for the lumens and chroma channels of the first color space into a forward-reshaped codeword for the lumens and chroma channels of the second color space):

number

[0092] Similarly, the global back-reshaping function in the above recoverable pairs is expressed as follows (each of the global back-reshaping functions transforms or forward-reshapes the codewords for the lumens and chroma channels of the second color space into back-reshaped codewords for the lumens and chroma channels of the first color space):

number

[0093] As shown in Figure 3A and equation (1) above, the Luma-Global Forward Reshaping function is different for different index values ​​g F,Y Indexing or lookup may be performed using different index values ​​g. Similarly, as shown in Figure 3B and Equation (2) above, the Luma-Global Back Reshaping function can be used for different index values ​​g. B,Y It may be used for indexing or lookup.

[0094] In a plurality of recoverable pairs formed by the global forward reshaping function in equation (1) and the global backward reshaping function in equation (2), the recoverable pairs may be formed by the global forward reshaping function in the global forward reshaping function in equation (1) and the global backward reshaping function (in the global backward reshaping function in equation (2)) that corresponds to this global forward reshaping function.

[0095] Global forward reshaping and global backward reshaping functions in the same restorable pair may be indexed with the same index value. The same pair of global forward and global backward reshaping functions can be used to reconstruct the original input video signal or the input image within it in the output video signal or output image. For example, given the g-th pair in the plurality of restorable pairs, g F,Y =g B,Y =g F,Cx =g B,Cx If = g, the input HDR image may be forward reshaped into a reshaped SDR image as shown in the following equation, and then backward reshaped into the same output HDR as the input HDR image (e.g., perfectly, faithfully, under quantization error, etc.).

number

[0096] Given the same g-th pair, the input SDR image may be retrofitted to a refitted HDR image and then retrofitted to the same output SDR as the input SDR image (e.g., perfectly, faithfully, under quantization error, etc.).

number

[0097] When a reshaping function performs a bit depth conversion from the pre-reshaping bit depth (e.g., the total number of bits encoding each codeword in the pre-reshaping image within the pre-reshaping color space channels) to the post-reshaping bit depth (e.g., the total number of bits encoding each codeword in the post-reshaping image within the post-reshaping color space channels), for example, from a 16-bit pre-reshaping video signal / image to a 10-bit post-reshaping video signal / image, quantization errors or losses may occur. If the difference between the input image received by the recoverable pair and the output image produced by the recoverable pair is attributable to quantization errors or losses caused by the bit depth conversion (one or more), the input and output images can still be considered the same.

[0098] Figure 3C shows an exemplary ruma codeword mapping that uses the same (recoverable) pair of global forward and backward reshaping functions indexed by the same index value g to convert the input ruma codeword of an input HDR image (input bit depth = 16 bits) to the output ruma codeword of an output HDR image (output bit depth = 16 bits) using equation (3-1) above.

[0099] As illustrated, the codeword mapping from input to output in the luma channel is nearly linear (for example, within the range of SMPTE specifications that clip at both ends of a defined luma value range). This means that the luma codewords reconstructed or output in the output HDR image generated by the recoverable pair are the same as or very close to the luma codewords received in the input HDR image.

[0100] Figures 3D and 3E show an exemplary chroma codeword mapping using the same (recoverable) pair of global forward and backward reshaping functions indexed by the same index value g, to convert the input chroma codeword of an input HDR image (16 bits) to the output chroma codeword of an output HDR image (16 bits) using equation (3-2) above.

[0101] The input chroma codewords have fixed Cb (e.g., 35000) and Cr (e.g., 28000) values, but can form groups of colors with varying chroma codeword values.

[0102] As illustrated, the codeword mapping from input to output in the chroma channel for this color group is nearly linear. This means that the chroma codewords reconstructed or output in the output HDR image generated by the recoverable pair are the same as, or very close to, the chroma codewords received in the input HDR image with respect to this color group, for varying chroma codeword values.

[0103] Similarly, another group of colors, having a different set of fixed Cb and Cr codeword values ​​(e.g., Cb is 25000, Cr is 35000), but with varying rumor codeword values, can be used in chroma mapping using recoverable pairs of global forward and backward reshaping functions. The input-to-output codeword mapping in the chroma channel for such another group of colors can be shown to be nearly linear, as shown in Figures 3D and 3E. This means that the chroma codewords reconstructed or output in the output HDR image generated by the recoverable pairs are the same as, or very close to, the chroma codewords received in the input HDR image for the other group of colors, with respect to the varying rumor codeword values.

[0104] Figure 3F shows an exemplary ruma codeword mapping that uses the same (recoverable) pair of global forward and backward reshaping functions indexed by the same index value g to convert the input ruma codeword of an input SDR image (input bit depth = 10 bits) to the output ruma codeword of an output SDR image (output bit depth = 16 bits) using equation (4-1) above.

[0105] As illustrated, the codeword mapping from input to output in the lumar channel is nearly linear (for example, within the entire codeword range defined by industry standards, etc.). This means that the lumar codewords reconstructed or output in the output SDR image generated by the recoverable pair are the same as, or very close to, the lumar codewords received in the input SDR image.

[0106] Similarly, chroma codeword mapping can be performed on the recoverable pairs to map the input SDR chroma codewords in the input SDR image to the output SDR chroma codewords in the output SDR image.

[0107] In some operating scenarios, the global forward and backward reshaping functions in recoverable pairs may be, but are not limited to, MMR-based forward and backward mapping / functions for chroma mapping and FLUT / BLUT-based forward and backward mapping / functions for luma mapping.

[0108] An MMR-based backmapping / function in recoverable pairs may be used to map input lumer and chroma codewords in an input SDR image to back-remodeled chroma codewords in a back-remodeled HDR image. A BLUT-based backmapping / function in recoverable pairs may be used to map input lumer codewords in an input SDR image to back-remodeled lumer codewords in a back-remodeled HDR image.

[0109] An MMR-based forward mapping / function in recoverable pairs may be used to map the back-remodeled lumens and chroma codewords in the back-remodeled HDR image to the output chroma codeword in the output SDR image. A FLUT-based forward mapping / function in recoverable pairs may be used to map the back-remodeled lumens codeword in the back-remodeled HDR image to the output lumens codeword in the output SDR image.

[0110] Similar to the case of input HDR-output HDR conversion performed using a recoverable pair, in the input SDR-output SDR conversion performed using a recoverable pair, while accepting possible quantization errors, the input chroma codewords in the output SDR image can be recovered or can be the same as the output chroma codewords in the input SDR image.

[0111] Non-paired forward and backward global reshaping Reshaping functions from different recoverable pairs may not reproduce the input video signal (or the input image therein) in the output signal (or the output image therein) generated from the reshaping operations performed on the input video signal (or the input image therein) based on these functions.

[0112] For illustrative purposes only, the forward luma and chroma reshaping functions may be indexed by the first index values g F,Y and g F,Cx respectively. Here, the first index values may or may not be equal to each other. Similarly, the backward luma and chroma reshaping functions may be indexed by the second index values g B,C and g B,Cx respectively. Here, the second index values may or may not be equal to each other.

[0113] In operating scenarios where the first and second index values are different from each other, as follows, the input video signal (or the input image therein) may not be reproduced in the output signal (or the output image therein) generated from the reshaping operations performed on the input video signal (or the input image therein) based on these functions.

Number

[0114] However, in these operating scenarios, these reshaping functions, which may be called unpaired forward and backward reshaping functions, can be applied or used as a global image enhancement tool to adjust the input brightness and saturation of the input image to different brightness and saturation of the output image.

[0115] Figure 3G shows various different index values ​​g ranging from 512 to 3328. B,C This shows an exemplary lumens (or luminance) backward function having (for example, a fixed forward index g) F,Y (For example, =1792). Figure 3H shows different index values ​​g in the range of 512 to 3328. F,C An example of a forward luma (or luminance) function having (for example, a fixed forward index g) is shown. F,Y (For example, let's say =1792)

[0116] For lumen anterior and posterior reshaping, the index value is g B,Y >g F,Y If selected such that, the reconstructed HDR image produced by performing lumar forward and backward reshaping functions with these index values ​​on the input HDR image will be brighter than the input HDR image, at least partially. On the other hand, if the index value is g B,Y <g F,Y If selected as such, the reconstructed HDR image produced by performing lumar forward and backward reshaping functions with these index values ​​on the input HDR image will be darker than the input HDR image, at least in part. The gap Δg between the two index values. Y =g B,Y -g F,Y This affects the brightness change. In some operating scenarios, the larger the gap, the greater the brightness change between the input and output images.

[0117] Similarly, for chromatic anterior and posterior reshaping, g B,Cx >g F,CxIf the index values ​​are selected such that (where Cx can be Cb or Cr), the reconstructed HDR image produced by performing chroma forward and backward reshaping functions with these index values ​​on the input HDR image will have higher saturation than the input HDR image, at least in part. On the other hand, if the index value is g B,Cx <g F,Cx If selected in this manner, the reconstructed HDR image produced by performing chroma forward and backward reshaping functions with these index values ​​on the input HDR image will have lower saturation than the input HDR image, at least partially. The gap Δg between the two index values. Cx =g B,Cx -g F,Cx This affects the change in saturation. In some operating scenarios, the larger the gap, the greater the change in saturation between the input and output images.

[0118] Figure 2K shows an exemplary flow that uses unpaired forward and backward reshaping functions to change the luminance and / or saturation of an input video signal to different luminance and / or saturation of an output video signal. As with other flows shown herein, the flow in Figure 2K may be performed by one or more computing devices such as a video encoder, video transcoder, video decoder, or a combination thereof. For illustrative purposes only, the input video signal and output video signal are the input HDR video signal and the output HDR video signal, respectively.

[0119] Block 222 includes receiving an input HDR image in an input HDR video signal. Blocks 224-1 and 224-2 include selecting first luma and chroma index values ​​for a forward reshaping function. Blocks 226-1 and 226-2 include performing luma and chroma forward reshaping on the input HDR image based on a luma and chroma forward reshaping function indexed by the first index values ​​to generate a forward reshaped image. Blocks 228-1 and 228-2 include selecting second luma and chroma index values ​​for backward reshaping. Blocks 230-1 and 230-2 include performing luma and chroma backward reshaping on the forward reshaped image based on a luma and chroma backward reshaping function indexed by the second index values ​​to generate a reconstructed HDR image having different luminance and / or saturation than the input HDR image. Block 232 includes outputting a reconstructed HDR image.

[0120] Local reshaping and guided images Figure 2L shows an exemplary luma and chroma-local reshaping 246 that can be performed on an input image 242 to produce an output image 250 (locally reshaped). The processing block in Figure 2L (including, but not limited to, luma and chroma-local reshaping (246)) may be implemented or performed by one or more computing devices such as a video encoder, video transcoder, video decoder, or a combination thereof.

[0121] As shown in Figure 2L, two components may be used to assist or support the lumar and chromar-local reshaping (246).

[0122] The first component is F <l>< / l>This is a family of local reshaping functions 248, which includes a set (or total of L) of local reshaping functions denoted as ( ), where l is 0, 1, ..., L-1. These local reshaping functions can be used or called as part of ruma and chroma local reshaping operations. The family of local reshaping functions is F for ruma local reshaping. Y <l>< / l> A family of chroma-local reshaping functions, and F for chroma-local reshaping. Cx <l>< / l> This includes a family of chroma-local reshaping functions, which are described as follows.

[0123] The second component is guidance image generation, which generates a guided image (M) from an input image (V) using a guidance image generation operator (G) as follows:

number

[0124] The guided image M contains individual (reshaping function) index values ​​(within the range [0 L-1]) for each input pixel in part or all of the input image (242) to select which local reshaping function in the family of local reshaping functions will perform the local reshaping in order to perform local reshaping on the input pixels to produce the corresponding output pixels in the output image (250). The guided image M contains M for lumar and chroma local reshaping in different lumar and chroma channels. Y and M Cx It can include different guided images like the following.

[0125] More specifically, given a guided image generated by guided image generation (244), for each input pixel of the input image (242), local reshaping (246) can find a specific reshaping function index value(s) in the guided image that is stored in the same row and column as the image frame containing the input image (242). The specific reshaping function index value(s) can then be used to select or identify a specific local reshaping function(s) from among the reshaping functions that constitute a family of reshaping functions(248). The specific local reshaping function(s) can then perform lumen and chroma local reshaping on the input pixels to produce the corresponding output pixels in the output image(250).

[0126] For explanatory purposes only, local forward reformatting will be discussed in detail. Note that local backward reformatting can be similarly derived, implemented, or performed. For simplicity, superscripts such as "Y" and "Cx" may be removed in this discussion.

[0127] Single-channel lumar local reshaping In some operating scenarios, the output or reshaped rumor codeword of an output image can be generated by performing a single-channel rumor local reshaping on the input or pre-reshaping rumor codeword of the input image. More specifically, the input rumor codeword for each pixel of the input image (e.g., sufficiently, etc.) allows the local reshaping function to determine the mapped or reshaped rumor codeword for the corresponding pixel in the output image without using the input chroma codeword for that pixel in the input image.

[0128] A family of ruma-local reshaping functions for single-channel ruma-local reshaping can be generated in many different ways.

[0129] In the first example, a self-derived local reshaping may be used to generate a family of ruma-local reshaping functions. A self-derived local reshaping refers to an approach in which a family of ruma-local reshaping functions is derived from some ruma-global reshaping function. This family of ruma-local reshaping functions can be generated on the fly at runtime in response to each input image or a specific codeword distribution within it. The family of local reshaping functions may include customized ruma-local reshaping functions for each frame in different content.

[0130] In the second example, offline training may be used to generate a family of ruma-local reshaping functions that include a pre-built ruma-local reshaping function. This family of ruma-local reshaping functions can be applied to all input images in all content.

[0131] In the third example, hybrid offline and online operations may be used to generate a family of rumor-local reshaping functions, including rumor-local reshaping functions generated by performing some or all of the self-derived local reshaping from a pre-built global reshaping function. The benefit of this hybrid method is that it saves or reduces computational cost, as the rumor-local reshaping function can be constructed using a dynamic global function built at runtime in response to the input image and a pre-built local function generated from offline training.

[0132] Self-derived single-channel local reshaping The two goals of luma-local reshaping are (2) to increase the local contrast ratio while maintaining similar luminance in the output image generated by performing luma-local reshaping on the input image. The first goal can be achieved by increasing the slope of the local reshaping function, such as that represented by a tone curve. The second goal, as described above, can be achieved by intersecting the local reshaping function (e.g., a new one) with the global reshaping function at various positions. An exemplary generation of a local reshaping function given a global reshaping function is described in U.S. Provisional Application No. 63 / 004,609, “BLIND LOCAL RESHAPING IN HDR IMAGING,” filed 3 April 2020, and its entire content is incorporated herein by reference as if it were fully described herein.

[0133] Figure 2M shows an exemplary flow for generating a local function that self-derives from a global reshaping function represented as F(). In some operating scenarios, the global reshaping function can be computed in online operation by a video codec in response to receiving an input image for local reshaping, for example. Such a global reshaping function may be referred to as a dynamic global reshaping function (e.g., adaptive, runtime). In some operating scenarios, the global reshaping function can be computed or pre-built by an image processing system in offline operation, which is not necessarily in response to receiving an input image for local reshaping, for example. Such a global reshaping function may be referred to as a static global reshaping function (e.g., scalable, pre-configured).

[0134] Block 252 involves constructing a template reshaping function. To do so, flat regions in the global reshaping function F() may first be removed to generate a modified reshaping curve or function. These flat regions may correspond to subranges outside the full range of available codewords (e.g., valid, SMPTE, etc.) in the codeword space. Flat regions may be added back later (e.g., finally, etc.) to local reshaping functions that are constructed.

[0135] The modified reshaping curve or function can be shifted until it touches the y-axis. As used herein, plots or curves representing local or global reshaping functions (as shown, for example, in Figures 3A, 3B, 3G, and 3H) may be represented in a two-dimensional coordinate system where the y-axis represents the reshaped or output codeword and the x-axis represents the unshaped or input codeword. Shifting the modified reshaping curve or function can be used to avoid or eliminate the need to handle critical cases and to facilitate scaling and shifting operations that generate local reshaping functions. Thus, a shifted modified reshaping function (denoted F') can be constructed from the global reshaping function by removing flat regions and shifting to the y-axis as follows: F'=shift_to_Yaxis(remove_flat(F())) (7)

[0136] In order to achieve the higher local contrast ratio mentioned as the first objective above, α <l>< / l> The x-axis scaling factor can be used, as indicated by α. More specifically, the x-axis scaling factor is set to be greater than 1 for a given input codeword value, or α <l>< / l> If the x-scaling factor α is > 1, the local neighborhood around the input codeword value along the x-axis is expanded (for example, proportional to the x-axis scaling factor), thereby reducing or decreasing (or increasing the blur) the local contrast ratio at the x-axis location represented by that input codeword value. In other words, the x-axis scaling factor α <l>< / l>If the scaling factor is greater than 1, the curve is scaled along the x-axis. If the range along the x-axis (e.g., intervals or differences used for differential calculations) is increased while the mapping range along the y-axis (e.g., intervals or differences used for differential calculations) remains the same, the x-scaled curve will have a smaller slope than the curve before scaling, and therefore the contrast ratio of the x-scaled curve will be smaller than that of the curve before scaling. On the other hand, if the scaling factor for the x-axis in the input codeword value is set to less than 1, or α <l>< / l> If the x-scaling factor α is <1, the local neighborhood around the input codeword value along the x-axis is compressed (for example, proportional to the x-axis scaling factor), and thus the local contrast ratio (or sharpness) at the x-axis location represented by the input codeword value can be increased or improved. In other words, the x-axis scaling factor α <l>< / l> If the value is less than 1, the curve is compressed along the x-axis. Since the range along the x-axis (e.g., intervals or differences used for differential calculations) decreases while the mapped range along the y-axis (e.g., intervals or differences used for differential calculations) remains the same, the x-axis compressed curve has a steeper slope than the uncompressed curve, resulting in a higher contrast ratio in the x-axis compressed curve compared to the uncompressed curve.

[0137] x-axis scaling factor (α <l>< / l> <1) can be used to scale the shifted, modified, and reshaped functions at various input codeword values ​​along the x-axis, for example, as specified by the x-axis scaling factor, to generate a scaled (forward) function, which increases the local contrast ratio (or sharpness) in the scaled function at these input codeword values.

[0138] Template reshaping function (F T <l>< / l>()) can be constructed or configured by first transforming or shifting the template reshaping function into a scaled (or shifted) forward function, and then resampling the scaled function at various input codeword values ​​corresponding to the input codeword values ​​used to encode the input image: F T <l>< / l> ( )=resample(F'(),α <l>< / l> ) (8)

[0139] In some operating scenarios, unequal local contrast ratio improvements can be achieved between different local reshaping functions in the local reshaping family by using different values ​​of the x-axis scaling factor in equation (8) above for different l values ​​of the local reshaping function.

[0140] In some operating scenarios, equal local contrast ratio improvements can be achieved among some or all local reshaping functions in the local reshaping family by using the same value of the x-axis scaling factor in equation (8) above for the l-value of the local reshaping function.

[0141] For illustrative purposes only, the same x-axis scaling factor α is assigned or applied in the shift and resampling operations in equation (8) above. <l>< / l> Along with equal local contrast ratio enhancement, all local reshaping functions (F for all l) are used in the local reshaping family. T <l>< / l> ()=F T A template reshaping function can be generated for ()). Therefore, only one scaled and resampled function has the same x-axis scaling factor α in equation (8) above. <l>< / l> From a global reshaping function that has a certain property, it can be constructed as a template reshaping function for all local reshaping functions.

[0142] Block 254 includes shifting the template reshaping function in equation (8) above to generate a pre-fuse local reshaping function. This pre-fuse local reshaping function can be fused with the global reshaping function and placed in the l-th local reshaping function in the local reshaping family. This shift of the template reshaping function achieves a second objective, namely maintaining similar (e.g., locally averaged, locally filtered, global, etc.) luminance between the pre-fuse local reshaping function and the global reshaping function. As used throughout this disclosure, the term “fusion” refers to the interpolation or blending of two specified functions. In a preferred embodiment, the blending of functions involves computing a linear combination of the functions relating. The weights in the linear combination may be normalized to a given value. For example, the sum of the weights may be 1. The term “pre-fusing” as used throughout this disclosure is used in relation to reshaping functions. For example, a local reshaping function may be fused with a global reshaping function. The resulting fused reshaping function is still a local reshaping function despite the fusion with the global reshaping function. This is because the fused reshaping function still retains different functional characteristics at the pixel level from the pre-fusing local reshaping function. To distinguish between the pre-fusing local reshaping function and the local reshaping function after fusion with the global reshaping function, the initial pre-fusing local reshaping function is referred to as the “pre-fusing local reshaping function,” and the fused result is referred to as the “post-fusing local reshaping function,” “fused local reshaping function,” or simply “local reshaping function.” The term “pre-merging” may also apply to other reshaping functions. For example, when two global reshaping functions are merged, they remain even more global, and therefore, if a distinction is needed, one should refer to the pre-merging (“pre-merging”) global reshaping function or the post-merging global reshaping function.

[0143] More specifically, when an input pixel, such as the i-th input codeword in the input image, is reshaped into an output pixel, such as the i-th output codeword, the pre-fusion local reshaping function for the l-th local reshaping function is the i-th element (m) of the guided image M in the local reshaping family. i It is indexed by the corresponding local reshaping function index value l stored in (which is noted as ).

[0144] The bit depth of the input image is B v This is expressed as follows. The entire input codeword range with this bit depth can be divided into L uniform intervals with corresponding interval centers, as follows:

number

[0145] To maintain similar (e.g., locally averaged, locally filtered, global, etc.) luminance around the input pixel between the pre-merged local reshaping function and the l-th local reshaping function for the global reshaping function and the l-th local reshaping function, the (original) global reshaping function F(C) at the center of the l-th interval of the input codeword range is used. l v A map of the reshaped or globally reshaped values ​​and the pre-fused local reshaping function F for the l-th local reshaping function at the same center of the l-th interval of the input codeword range—generated by shifting the template reshaping function. <l>< / l> (C l v ) and can be constrained to satisfy the following conditions / constraints: F(C l v )=F <l>< / l> (C l v ) (11)

[0146] This constraint is the pre-fusion local reshaping function F <l>< / l> (C l vDetermine the shift amount or shift (value) for <l>< / l> (C l v ) and apply the shift (value) to a template reshaping function to obtain the pre-fusion local reshaping function F <l>< / l> (C l v ) for the l-th local reshaping function. The l-th local reshaping function or the pre-fusion local reshaping function F

[0147] Thus, given the input codeword C l v , the corresponding mapped (or globally reshaped) value F <l>< / l> (C l v ) can be obtained or looked up (for example, from a curve, function, and / or lookup table representing the global reshaping function F <l>< / l> (C l v )). Then, the corresponding mapped value F <l>< / l> (C l v ) can be used to determine the input codeword value φ l v such that the following equality condition (such as equal luminance) is satisfied: F <l>< / l> T (φ l v ) = F <l>< / l> (C l v ) (12)

[0148] Furthermore, the input codeword value φ l v can be used to determine the shift (value) as follows: a l v = C l v - φ l v (13)

[0149] Therefore, for all l between 0 and L - 1, the pre - fusion local reshaping function for the l - th local reshaping function is F <l>< / l> (C l v ) can be obtained as a shifted version of the template reshaping function as follows: F <l>< / l> (v)=F <l>< / l> T (v - a l v ) (14)

[0150] After scaling and shifting operations are performed to increase or change the slope of the global reshaping function to the slope of the template reshaping function, the template reshaping function or the local reshaping function generated from the template reshaping function may generate values outside the range that are hard - clipped within the range of acceptable codeword values, thereby generating visual artifacts. To avoid or reduce hard - clipping, soft - clipping may be used by fusing the global function and the local function with their respective weighting factors.

[0151] Block 256 soft - clips the pre - fusion local reshaping function F <l>< / l> (C l v ) by fusing it with the global reshaping function to obtain the pre - fusion local reshaping function F <l>< / l> (C l v ). Given the pre - fusion local reshaping function F <l>< / l> (C l v ) generated for the l - th local reshaping function in block 254, the global reshaping function weighting factor represented as θ <l>< / l> Y,G and the local reshaping function weighting factor represented as θ <l>< / l> Y,L are respectively for the global reshaping function and the pre - fusion local reshaping function F <l>< / l> (C lv They may be assigned to each of the following. In some operating scenarios, the two weighting factors satisfy the following constraints / conditions: θ <l>< / l> Y,G +θ <l>< / l> Y,L =1 (15)

[0152] The l-th local reshape function (after fusion) is given as follows: F <l>< / l> ( )=θ <l>< / l> Y,L ·F <l>< / l> ( ) + θ <l>< / l> Y,G ·F( ) (16)

[0153] It should be noted that the above operations can be implemented or performed to generate both local forward reshaping functions and local backward reshaping functions.

[0154] Pre-built single-channel local reshaping In some operating scenarios, in addition to the self-derived single-channel local reshaping described above, or instead, there are multiple pre-constructed single-channel global reshaping functions F, each indexed by a global reshaping function index value represented by g. <g>< / g> Y ( ) can be obtained through offline training. These pre-built global reshaping functions can be used, for example, in a forward path to generate local reshaping functions, each indexed by a local reshaping function index value l, as follows: F <l>< / l> ( )=F <g>< / g> Y () (17-1) Here, g=l (17-2)

[0155] Similarly, in the rear route, B <l>< / l> The local reshaping function represented by ( ) is as follows: B <g>< / g>Y It can be generated from a pre-built global reshaping function represented as ( ): B <l>< / l> ( )=B <g>< / g> Y () (18)

[0156] An exemplary derivation of a local reshaping function from a global reshaping function is described in U.S. Provisional Application No. 63 / 086,699, “ADAPTIVE LOCAL RESHAPING FOR SDR-TO-HDR UP-CONVERSION,” filed 2 October 2020, and its entirety is incorporated by reference as if it were entirely contained herein.

[0157] Hybrid of predefined and self-derived single-channel local reshaping In some operational scenarios, a hybrid approach combining predefined single-channel local reshaping and self-derived single-channel local reshaping may be used to generate local reshaping functions. More specifically, the self-derived local reshaping described above can be performed on an existing pre-built global reshaping function instead of a global reshaping function that is constructed in response to receiving an input image (e.g., dynamic). For example, a particular pre-built global reshaping function may be selected from among several pre-built global reshaping functions (e.g., those obtained through offline training) based on the codeword distribution of the input image. The particular pre-built global reshaping function may be used in the self-derived local reshaping described above instead of a global reshaping function that is constructed in response to receiving an input image.

[0158] Cross-channel lumar-local reshaping In some operating scenarios, the output or reshaped rumor codewords of the output image can be generated by performing a cross-channel rumor local reshaping on the input or pre-reshaping rumor and chroma codewords of the input image. More specifically, the input rumor and chroma codewords for the pixels of the input image (e.g., suffix, etc.) collectively enable a local reshaping mapping (e.g., TPB-based, MMR-based, etc.) to determine the mapped or reshaped rumor codewords of the corresponding pixels in the output image.

[0159] The family of ruma-local reshaping (functions) for cross-channel ruma-local reshaping can be generated in many different ways. Similar to single-channel ruma-local reshaping, cross-channel ruma-local reshaping includes at least (1) using self-derived local reshaping functions, (2) using pre-constructed reshaping functions, and (3) hybrids that use both self-derived and pre-constructed reshaping functions.

[0160] Self-derived cross-channel luma-local reshaping In various operating scenarios, any of the diverse and different types of cross-channel lumer reshaping functions (e.g., TPB, MMR, etc.) can be used. For illustrative purposes only, the cross-channel lumer reshaping function may be Tensor-Product B-Spline (TPB) based. TPB-based reshaping functions can capture a wide range of nonlinearities in lumer reshaping.

[0161] Figure 2N shows an exemplary flow for generating a self-derived cross-channel local function based on an image pair including a first image and a second image. The flow can be used to generate a TPB-based local reshaping function to map the input codeword of the first image to a locally reshaped codeword that approximates or reconstructs the codeword of the reference image represented by the second image. For illustrative purposes only, the first image in the image pair may be an HDR image, while the second image (or reference image) may be an SDR image. In one example, the SDR image may have been used to generate the HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations). In another example, the SDR image may have been generated from the HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations).

[0162] Block 262 involves constructing or building a 3D mapping table (3DMT) from an image pair. Each pixel of an HDR image corresponds to a color channel (e.g., three, etc.) represented as y, c0, c1 (or alternatively y, c0, c1) in the HDR color space or domain, v i =[v i y v i c0 v i c1 ] T The rumor and chroma codewords may be represented as follows: Each pixel of the SDR image is represented as y, c0, c1, and so on, in the SDR color space or domain, in the color channels (for example, three, etc.) represented as y, c0, c1 respectively. i =[s i y s i c0 s i c1 ] T It may include rumor and chroma codewords represented as follows.

[0163] The HDR color space, which includes the codewords available in each of the three channels, has a corresponding fixed number—Q for each component. y Q c0 Q c1 It can be quantized or partitioned using one-dimensional bins such as (for example, uniformly, etc.). As a result, the three-dimensional histogram is (Q y ×Q c0 ×Q c1 It can be initialized to contain ) cubes or 3D histogram bins.

[0164] 3D histogram Ω Q,v This is expressed as follows: Here, Q = [Q y Q c0 Q c1 ]. 3D histogram Ω Q,v is the total number (Q y Q c0 Q c1 Includes ) 3D histogram bins. 3D histogram Ω Q,v The 3D histogram bins within are represented by the 3D bin index q = (q y ,q c0 ,q c1 ) may be specified or indexed using, where for each ch={y,C0,C1}, q ch =0,…,Q ch-1 These may be used to represent pixels having quantized values ​​for these three channels within a 3D histogram bin—or to store a count of such pixels (e.g., total number).

[0165] The sum of the codeword values ​​(e.g., reference, mapped, etc.) in the SDR image may be calculated for each 3D histogram bin in the 3D histogram. Ψ y Q,s Ψ C0 Q,s Ψ C1 Q,s Let these be the sums of the codewords in the lumen channel and chroma channel {y,C0,C1} of the output domain, respectively.

[0166] Assume that each of the HDR image and the SDR image contains P pixels. An exemplary procedure for calculating the count of HDR pixels in a 3D histogram bin of the 3D histogram and the sum of the SDR codeword values for those 3D histogram bins is shown in Table 1 below. [Table 1]

[0167] (v q y,(B) ,v q C0,(B) , v q C1,(B) ) is assumed to represent the center of the q-th (HDR) 3D histogram bin in the 3D histogram. An exemplary procedure for pre-calculating the centers of some or all of the 3D histogram bins in the 3D histogram is illustrated in Table 2 below. [Table 2] <##**##>

[0168] Next, among the 3D histogram bins in the 3D histogram, 3D histogram bins having a non-zero total number of pixels can be identified. All other 3D histogram bins having no pixels - or in some other operating scenarios having a total number of pixels less than a minimum pixel count threshold - can be discarded from further processing.

[0169] q0, q1,..., q k-1 are assumed to be k 3D histogram bins where the count of HDR pixels is Ω q Q,v ≠0. For these k 3D histogram bins, the average value [Equation] of the SDR codewords is the sum [Equation] This can be calculated based on the total number of HDR pixels in the bin. An exemplary procedure for such a calculation is shown in Table 3 below. [Table 3]

[0170] As a result, there are multiple mapping pairs from the first image (an HDR image in this example) to the second image (an SDR image in this example). Each such mapping pair may include the center of a 3D histogram bin indexed by a valid q having a non-zero HDR pixel count, and the average of the SDR codeword values ​​for the SDR pixels mapped from the HDR pixels counted in the 3D histogram bin, as shown below:

number

[0171] In some operating scenarios,

number

[0172] Block 264 involves constructing a modified 3DMT for each local reshaping function. To achieve the change in local contrast ratio, h Y A function denoted as () is designed to calculate the average SDR or mapped value in each 3D histogram bin (indexed by a valid q with a non-zero HDR pixel count) of the 3D histogram.

number

number

number

[0173] The point where the global reshaping function and the l-th local reshaping function intersect (e.g., the center) is given, for example, in a local neighborhood where the local contrast ratio is increased, to maintain similar brightness between the global reshaping function and the l-th local reshaping function, as follows: C l v = l / L (21)

[0174] In some operating scenarios, the above function h is used to realize changes in the local contrast ratio. Y () is for the l-th local reshaping function, α <l>< / l> It may also be defined using a linear scaling factor expressed as (α to increase the local contrast ratio) <l>< / l> <1. To reduce the local contrast ratio, α <l>< / l> >1). Therefore, equation (21) above can be rewritten as follows:

number

[0175] Block 266 involves constructing the TPB coefficients for the global reshaping function and each of the local reshaping functions (for example, the l-th local reshaping function).

[0176] Based on the (original) 3DMT pair defined in equation (19) above, the first TPB coefficient for the global reshaping function

number

number

number

[0177] Vector

number

number

number

[0178] Optimized global reshaping solution or first TPB coefficient

number

number

[0179] First TPB coefficient

number

number

number

number

number

[0180] Similar to the global reshaping function, for each of the L local reshaping functions (l=0,…,L-1), the optimized local reshaping solution or the second TPB coefficient

number

number

[0181] Block 268 involves constructing a TPB-based local reshaping function (e.g., a fused local reshaping function). Similar to single-channel local reshaping / prediction, local reshaping functions with higher local contrast may generate out-of-range values ​​that are hard-clipped within an acceptable range of codeword values, thereby generating visual artifacts. To avoid or reduce hard clipping, soft clipping may be used by fusing the global and local functions together with their respective weighting factors.

[0182] Given the l-th local reshaping function generated in block 266, θ <l>< / l> Y,G The global reshaping function weighting factors and θ are expressed as follows: <l>< / l> Y,L Local reshaping function weighting factors expressed as may be assigned to the global reshaping function and the l-th local reshaping function, respectively. In some operating scenarios, the two weighting factors satisfy the constraints / conditions shown in equation (15) above. Thus, the local reshaping described herein may be performed using a fused local reshaping function generated by fusing the first TPB coefficient for the global reshaping function and the second TPB coefficient for the local reshaping function, as follows:

number

[0183] Pre-built cross-channel luma-local reshaping In some operating scenarios, in addition to the self-derived cross-channel local reshaping described above, or instead, there are multiple pre-constructed cross-channel global reshaping functions F, each indexed by a global reshaping function index value represented by g. <g>< / g> Y() can be obtained through offline training. These pre-built global reshaping functions can be used, for example, in a forward path to generate local reshaping functions, each indexed by the local reshaping function index value l, as shown below: F <l>< / l> ( )=F <g>< / g> Y ( ) (29-1) Here, g=l (29-2)

[0184] Similarly, in the rear route, B <l>< / l> The local reshaping function represented by () is B <g>< / g> Y From a pre-built global reshaping function represented as (), it can be generated as follows: B <l>< / l> ( )=B <g>< / g> Y ( ) (30)

[0185] A hybrid of predefined cross-channel lumar-local reshaping and self-derived cross-channel lumar-local reshaping. In some operational scenarios, a hybrid approach combining predefined single-channel local reshaping and self-derived single-channel local reshaping can be used to generate local reshaping functions.

[0186] In some embodiments, one or more 3DMTs can be generated in offline training using one or more sets of training image pairs in the training data. Each training image pair in each of the one or more sets of training image pairs may include an input HDR image and a corresponding input SDR image that exhibits the same visual semantic content as the input HDR image. In one example, the corresponding input SDR image may be the one used to generate the input HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations). In another example, the corresponding input SDR image may be generated from the input HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations).

[0187] Each of the one or more 3DMTs can be generated using each set of training image pairs in the training data. Image data such as rumor and chroma codewords in each set of training image pairs can be used to generate statistics collected in the corresponding 3DMT. These statistics can be used to form multiple 3DMT mapping pairs for the corresponding 3DMT. Some or all of the one or more 3DMTs, the statistics in the one or more 3DMTs, and / or the 3DMT mapping pairs can be stored, cached, and / or loaded at system startup.

[0188] Each of the one or more 3DMTs can be indexed by its own index value (e.g., g) for lookup purposes. In some operating scenarios, the index value (e.g., g) that uniquely identifies a corresponding 3DMT in the one or more 3DMTs may be an L1 mid value. An example of a mid L value can be found in the aforementioned U.S. Provisional Patent Application No. 63 / 086,699.

[0189] Figure 2O shows an exemplary flow for combining self-derived cross-channel local function generation with pre-built cross-channel local function generation.

[0190] Block 262' involves reusing or selecting a 3DMT generated using a corresponding set of corresponding training image pairs. In some operating scenarios, the distribution of input codeword values ​​determined from the input HDR images can be used to calculate or predict index values ​​(represented as g), such as the mid-L value. The index values ​​can be used to look up, identify, or select one or more of the 3DMTs. The selected 3DMT may be generated based on image data in the g-th set of training image pairs containing the g-th SDR data.

[0191] As a result, from the aforementioned 3DMT, multiple 3DMT mapping pairs

number

number

[0192] Block 264' involves constructing individual modified 3DMTs for each local reshaping function (e.g., the l-th local reshaping function).

[0193] To achieve a change in local contrast, h Y The function denoted as () can be used to correct the average of SDR codeword values ​​in 3DMT mapping pairs, as follows:

number

[0194] The point where the global reshaping function and the l-th local reshaping function intersect may be given, for example, as follows, in order to maintain similar brightness between the global reshaping function and the l-th local reshaping function in a local neighborhood where the local contrast ratio is increased: C l v = l / L (33)

[0195] For the l-th local reshaping function, α <l>< / l> Linear scaling expressed as (to increase the local contrast ratio, α <l>< / l> <1. To reduce the local contrast ratio, α <l>< / l> >1) may be used in the following formula to change the average or mapped value of the SDR codeword in the 3DMT mapping pair:

number

[0196] Block 266' involves constructing the first TPB coefficient for the global reshaping function and the second TPB coefficient for the l-th local reshaping function.

[0197] More specifically, based on the original 3DMT mapping pairs generated using the training data, the first TPB coefficient is calculated using the following design matrix S Y and vectors

number

number

number

[0198] Optimized global reshaping solution or optimized value for the first TPB coefficient

number

number

[0199] First TPB coefficient

number

number

number

number

number

[0200] Similar to the global reshaping function, the optimized local reshaping solution or the optimized value for the second TPB coefficient for each of the L local reshaping functions (l=0,…,L-1)

number

number

[0201] Block 268' involves constructing a TPB-based local reshaping function (e.g., a fused local reshaping function). Similar to single-channel local reshaping / prediction, local reshaping functions with higher local contrast may produce out-of-range values ​​that are hard-clipped within an acceptable range of codeword values, thereby potentially generating visual artifacts. To avoid or reduce hard clipping, soft clipping may be used by fusing the global and local functions together with their respective weighting factors.

[0202] Given the l-th local reshaping function generated in block 266', θ <l>< / l> Y,G The global reshaping function weighting factors are expressed as follows, and θ <l>< / l> Y,L The local reshaping function weighting factors, expressed as , may be assigned to the global reshaping function and the l-th local reshaping function, respectively. In some operating scenarios, the two weighting factors satisfy the constraints / conditions shown in equation (15) above. Thus, the local reshaping described herein can be performed using a fused local reshaping function generated by fusing the first TPB coefficient for the global reshaping function and the second TPB coefficient for the local reshaping function, as shown below.

number

[0203] Figure 2P shows an exemplary flow using 3DMT to generate multiple cross-channel local reshaping functions for performing local reshaping (e.g., forward) on an input HDR image 272. The cross-channel local reshaping functions are at least partially alpha for achieving local contrast changes and soft clipping. <l>< / l> , θ <l>< / l> Y,G , θ <l>< / l> Y,L It can be generated based on improvement parameters such as these.

[0204] Block 276 includes constructing a 3DMT using a 3D histogram having a mapping pair (or item) that stores the count of HDR pixels in the input HDR image (272) and the mean of the corresponding input SDR image in the image pair formed by the input HDR image and the input SDR image.

[0205] In one example, the corresponding input SDR image may be one that was previously used to generate the HDR image (272) (for example, through previously performed content mapping, tone mapping, and / or reshaping operations). In another example, the corresponding input SDR image may be one that was generated from the HDR image (272) (for example, through previously performed content mapping, tone mapping, and / or reshaping operations).

[0206] Block 278 involves using 3DMT to construct a global reshaping function. For example, the first TPB coefficients can be generated to specify the global reshaping function using a complete set of tensor product B-spline basis functions.

[0207] Block 280 improves the parameter α for local contrast changes. <l>< / l> This includes modifying the 3DMT using such a method to obtain each modified 3DMT represented by a modified 3D histogram for each cross-channel local reshaping function in a plurality of cross-channel local reshaping functions used for local reshaping.

[0208] Block 282 includes constructing a pre-fusion local reshaping function using a modified 3DMT. The pre-fusion local reshaping function can be fused with a global reshaping function to become the cross-channel local reshaping function in the plurality of cross-channel local reshaping functions used for local reshaping. For example, a second TPB coefficient may be generated to specify the pre-fusion local reshaping function using tensor product B-spline basis functions.

[0209] Block 284 is an improvement parameter θ for soft clipping. <l>< / l> Y,G and θ <l>< / l> Y,L This includes using a weighted fusion with both a global reshaping function and a pre-fusion local reshaping function to generate a corresponding fused cross-channel local reshaping function. Block 286 includes outputting multiple (fused) cross-channel local reshaping functions.

[0210] Guide image for Ruma Local reshaping The guide image M=G(V) can be generated in one or more different ways to provide local reshaping function index values ​​for the pixels of the input image. In some operating scenarios, a multilevel edge-preserving filter may be applied to generate the guide image to avoid or reduce visual artifacts such as halo artifacts (e.g., near the edges of image features / objects, near or around the background-foreground boundary). An exemplary guide image generation using multilevel edge-preserving filtering can be found in the aforementioned U.S. Patent Provisional Application 63 / 086,699.

[0211] Local reshaping can increase the local contrast ratio, but it can also increase the likelihood of introducing false contour generation or banding artifacts, especially when local reshaping is performed from a high-bit-depth input domain (or input color space) to a low-bit-depth domain (or output color space) and the output image has a lower bit depth than the input image.

[0212] The injection of film grain into high-bit-depth video signals encoded with input images, as described in U.S. Provisional Application No. 63 / 061,937, “ADAPTIVE STREAMING WITH FALSE CONTOURING ALLEVIATION,” filed August 6, 2020 (the entire content of which is incorporated herein by reference as if it were fully described herein), may not be sufficient to avoid or significantly reduce banding artifacts because the output domain has fewer available codewords than the input domain. This banding artifact problem may be exacerbated when local reshaping increases the local contrast ratio. Additionally, optionally, or alternatively, film grain noise may be significantly introduced at the cost of resulting in an unpleasant visual appearance.

[0213] To overcome or mitigate the banding artifact problem, local reshaping function selection dithering may be implemented or performed to avoid or reduce banding artifacts.

[0214] Figure 2Q shows an exemplary flow for local reshaping function selection dithering. Block 288 includes performing local reshaping index noise injection. For example, Gaussian noise is injected into the local reshaping function index m in the guided image M. i Injected into the following noise-injected local reshaping function index

number

number

[0215] Block 290 is the noise-injected guided image or the noise-injected local reshaping function index within it.

number

[0216] The locally (e.g., forward) reshaped codeword value is the corresponding noise-injected index value in the noise-injected guided image for each pixel of the input image (e.g., the i-th pixel).

number

number

number

[0217] In some operating scenarios, the lumen modulation function can be based on a global reshaping function. In some operating scenarios, the lumen modulation function can be based on a local reshaping function. The entire lumen range of the input domain (or color space) is Δ v It can be divided into multiple non-overlapping bins having intervals represented as follows. The discrete slope in interval k is given as follows: Regarding (or based on) global reshaping

number

number

[0218] The maximum or best slope among all k

number

number

[0219] The slope for all bins can be normalized by the maximum slope as follows:

number

number

number

[0220] Value for a certain bin

number

number

[0221] In some operating scenarios, the normalized slope

number

[0222] The maximum and minimum noise intensities to be added to equation (42) above (e.g., user-specified, pre-configured, dynamically configurable, etc.) are specified by the user ψ max and ψ min This is expressed as follows: Each input codeword v in the input image i Regarding this, the applicable noise intensity represented by the luma modulation function can be designed or derived as follows:

number

number

[0223] To avoid or reduce computational costs, the lumar modulation function can be derived in equation (46) above, depending on the slope calculated using the global reshaping function.

[0224] In some operating scenarios, a combination of local reshaping function index dithering and film grain noise injection into locally reshaped codeword values ​​can effectively avoid or reduce banding artifacts, especially when the output image is an 8-bit image.

[0225] Chroma-Local Reshaping In some operational scenarios, the output or reshaped chroma codeword of the output image can be generated by performing a chroma-local reshaping (e.g., cross-channel) on the input or pre-reshaping chroma and chroma codeword of the input image, based on the input or pre-reshaping chroma and chroma codeword of the input image.

[0226] The family of chroma-local reshaping (functions) for chroma-local reshaping can be generated in many different ways, such as self-derived local reshaping, pre-constructed local reshaping, and combinations or hybrids of the aforementioned.

[0227] To construct chroma-local reshaping mappings / functions, MMR or TPB-based techniques can be implemented or performed. The methods and / or processes used to construct or generate operating parameters for chroma-local reshaping, such as coefficients for both MMR-based and TPB-based local reshaping functions, are the same or similar to each other, with the difference (e.g., only difference, main difference, etc.) being the basis functions used in the optimization problem to obtain optimized values ​​for the coefficients. For MMR, polynomial basis functions / terms are used as basis functions in relation to the MMR coefficients. For TPB, tensor product B-spline basis functions are used as basis functions in relation to the TPB coefficients. Both MMR and TPB reshaping functions are involved in the construction and modification of 3DMTs.

[0228] Self-derived chroma-local reshaping It may not be straightforward to directly derive a local chroma reshaping function from a global chroma reshaping function (e.g., free-form). This is because chroma reshaping may use cross-color channel predictors (e.g., MMR-based, TPB-based). Similar to cross-channel chroma local reshaping, cross-channel chroma local reshaping on an input image can be performed by returning to the original 3DMT related to the global reshaping function or mapping, and modifying the original 3DMT to obtain a local reshaping function or mapping for local variations (e.g., local enhanced saturation, local reduced saturation, etc.). The 3DMT can be constructed online and / or offline. In some operating scenarios, two global chroma reshaping functions can be generated for each of the two chroma channels Cb and Cr, and these can be merged to generate the respective local reshaping functions for the two local reshaping functions (e.g., after merging) for each of the two chroma channels Cb and Cr.

[0229] Figure 2R shows an exemplary flow for generating coefficients (e.g., MMR-based, TPB-based, etc.) for chroma-local reshaping functions / mappings using self-derived chroma-local reshaping.

[0230] Block 292 includes constructing a 3DMT for a global reshaping function or mapping based on, for example, an image pair (e.g., an HDR image to be locally reshaped and a corresponding SDR image). The 3DMT can be constructed using the same or similar operations as those used for cross-channel lumar-local reshaping, as described in Tables (1) to (3) above.

[0231] As a result, for 3DMT, multiple 3DMT mapping pairs can be generated as shown in equation (19). These multiple 3DMT mapping pairs can then be used to derive a global reshaping function or mapping, and to modify or improve it for the purpose of generating a second global chroma reshaping function or mapping.

[0232] Block 294 involves modifying the 3DMT or mapped values ​​in the mapping pair. To achieve a change in saturation, h C A function represented as () is designed, which maps the values ​​in each 3D histogram bin of the 3D histogram (indexed by a valid q with a non-zero HDR pixel count).

number

number

number

[0233] function h C The first step involves transforming the Cartesian coordinate system of Cb (denoted as C0) and Cr (denoted as C1) (for example, with respect to the neutral color point (0.5,0.5) instead of the non-neutral color origin (0,0)) into a polar coordinate system of polar radius ρ and polar angle θ as follows:

number

[0234] function h C () represents the polar diameter ρ as follows: C,q Q,SThis includes further applying a nonlinear function or mapping to the saturation represented by:

number

[0235] As shown in equation (49) above, a color close to neutral (for example, ρ C,q Q,s <δ C For example, linear mapping is used to avoid or prevent changes in the hue of colors that are close to neutral.

[0236] Figure 3I shows an exemplary nonlinear function h used to realize the change in saturation as specified by equation (49) above. C (This indicates)

[0237] Nonlinear function h C The modified and shifted Cb / Cr values ​​generated from () can be converted from polar coordinates to the original Cartesian coordinates. The modified and shifted Cb / Cr values ​​converted to Cartesian coordinates can then be shifted in reverse to generate the modified Cb / Cr values ​​by adding an inverse offset of 0.5, as shown below:

number

[0238] The modified Cb / Cr value in equation (50) above is used in part to form multiple modified 3DMT mapping pairs for the modified 3DMT, as shown below. It is possible.

number

[0239] Block 296 includes constructing or generating a first coefficient (e.g., MMR, TPB, etc.) for a global reshaping function from multiple (original) 3DMT mapping pairs, and constructing or generating a second coefficient (e.g., MMR, TPB, etc.) for a second global reshaping function from multiple modified 3DMT mapping pairs within the modified 3DMT.

[0240] For example, based on the original 3DMT pair constructed using the original colors, the first coefficient is the design matrix S Cx and vectors

number

number

number

[0241] Optimized global reshaping solution or optimized value for the first coefficient

number

number

[0242] Optimal value for the first coefficient of the global reshaping function

number

number

[0243] These extended operations involve the vector in equation (54) above for the global reshaping function.

number

number

number

[0244] Similar to the case of the global reshaping function, the second coefficient for the second global reshaping function

number

number

[0245] Block 298 involves constructing local reshaping functions (e.g., fused, MMR-based, TPB-based, etc.).

[0246] Given the first and second global reshaping functions generated in block 296, θ <l>< / l> Cx,G1 The first global reshaping function weighting factor is expressed as, and θ <l>< / l> Cx,G2 A second global reshaping function weighting factor, expressed as , may be assigned to the first global reshaping function and the second global reshaping function, respectively. In some operating scenarios, the two weighting factors satisfy the following constraints / conditions:

number

[0247] Therefore, the local reshaping described herein may be performed using a fused local reshaping function generated by fusing a first coefficient for a global reshaping function with a second coefficient for a second global reshaping function, as follows:

number

[0248] Pre-built cross-channel chroma-local reshaping In some operating scenarios, in addition to the self-derived cross-channel local reshaping described above, or instead, there are multiple pre-constructed cross-channel global reshaping functions F, each indexed by a global reshaping function index value represented by g. <g>< / g> Cx () is obtained through offline training using the g-th training dataset, which contains each SDR-HDR image pair.

[0249] For example, multiple 3DMT mapping pairs

number

number

[0250] The g-th global reshaping function can be obtained, for example, using the least-squares approach described above, as follows:

number

[0251] These pre-built global reshaping functions may be used, for example, in a forward path to generate local reshaping functions, each indexed by a local reshaping function index value l, as follows: F <l>< / l> Cx ()=F <g>< / g> Cx () (62-1) Here, g=l (62-2)

[0252] Similarly, in the rear route, B <l>< / l> Cx The local reshaping function, expressed as follows, B <g>< / g> Cx It can be generated from a pre-built global reshaping function represented as: B <l>< / l> Cx ()=B <g>< / g> Cx () (63)

[0253] A hybrid of predefined chroma-local reshaping and self-derived chroma-local reshaping. In some operational scenarios, a hybrid approach combining predefined and self-derived cross-channel chroma-local reshaping may be used to generate local chroma reshaping functions.

[0254] In some embodiments, two or more 3DMTs can be generated in offline training, each using one or more sets of training image pairs within the training data. Each of the two or more sets of training image pairs may include an input HDR image and a corresponding input SDR image that exhibits the same visual semantic content as the input HDR image. In one example, the corresponding input SDR image may be the one used to generate the input HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations). In another example, the corresponding input SDR image may be generated from the input HDR image (e.g., through previously performed content mapping, tone mapping, and / or reshaping operations).

[0255] Each of the two or more 3DMTs can be generated using each set of training image pairs in the training data. Image data such as lumens and chroma codewords in each set of training image pairs can be used to generate statistics collected in the corresponding 3DMT. These statistics can be used to form multiple 3DMT mapping pairs for the corresponding 3DMT.

[0256] Each of the two or more 3DMTs can be indexed by its respective index value (e.g., g) for lookup purposes. In some operating scenarios, the index value (e.g., g) that uniquely identifies a corresponding 3DMT in the two or more 3DMTs may be the L1 mid value.

[0257] Two or more pre-constructed global chroma reshaping functions can each be constructed from the two or more (original, unmodified) 3DMTs using the self-derived local reshaping approach described above. The two or more 3DMTs can be improved or modified using the same saturation factor or two or more different saturation factors to generate two or more modified 3DMTs. From the two or more modified 3DMTs, two or more pre-fusion second global chroma reshaping functions can each be constructed using the self-derived local reshaping approach described above. Furthermore, the two or more global chroma reshaping functions from the two or more (original, unmodified) 3DMTs can each be fused with the two or more pre-fusion second global chroma reshaping functions generated from the two or more modified 3DMTs to generate two or more pre-constructed local chroma reshaping functions using the self-derived local reshaping approach described above.

[0258] Figure 2S shows an exemplary flow for generating coefficients (e.g., MMR-based, TPB-based, etc.) for chroma-local reshaping functions / mappings using predefined and self-derived chroma-local reshaping combinations.

[0259] Block 2092 involves constructing a 3DMT for a global reshaping function or mapping based on an image data subset, such as the g-th set of training image pairs selected from multiple sets of training image pairs. The 3DMT can be constructed using the same or similar behavior as that used for cross-channel lumar-local reshaping, as described in Tables (1) to (3) above.

[0260] As a result, multiple 3DMT mapping pairs, as shown in equation (19), can be generated for 3DMT as follows:

number

[0261] Next, multiple 3DMT mapping pairs in equation (64) can be used to derive a global reshaping function or mapping and to modify or extend it for the purpose of generating a second chroma reshaping function or mapping. Block 2094 includes modifying the 3DMT or mapped values ​​in the mapping pairs to generate modified 3DMTs.

[0262] To achieve a change in saturation, h C A function represented as () is designed, which maps the values ​​in each 3D histogram bin (indexed by a valid q with a non-zero HDR pixel count) of the 3D histogram as follows:

number

number

number

[0263] function h C The first step involves transforming the shifted Cb / Cr in the Cartesian coordinate system of Cb (represented as C0) and Cr (represented as C1) (for example, using the neutral color point (0.5,0.5) instead of the non-neutral color origin (0,0)) into polar coordinates of polar radius ρ and polar angle θ, as shown below:

number

[0264] function h C ( ) represents the polar diameter ρ as shown below. C,q Q,s <g>< / g>This includes further applying a nonlinear function or mapping to the saturation represented by:

number

[0265] As shown in equation (67) above, a color close to neutral (for example, ρ C,q Q,S, <g>< / g> <δ C Linear mapping is used to avoid or prevent the alteration of near-neutral colors (more specifically, their hue) for things like this.

[0266] Nonlinear function h C The modified and shifted Cb / Cr values ​​generated from () can be converted from polar coordinates to the original Cartesian coordinates. The modified and shifted Cb / Cr values ​​converted to Cartesian coordinates can then be reverse-shifted by adding an offset of 0.5 in reverse, as shown below, to generate the modified Cb / Cr values:

number

[0267] The modified Cb / Cr value in equation (68) above can be used to form multiple modified 3DMT mapping pairs within the modified 3DMT, as follows:

number

[0268] Block 2096 includes constructing or generating first (e.g., MMR, TPB, etc.) coefficients for a global reshaping function from the plurality of (original) 3DMT mapping pairs, and constructing or generating second (e.g., MMR, TPB, etc.) coefficients for a second global reshaping function from the plurality of modified 3DMT mapping pairs.

[0269] For example, based on the original 3DMT pair constructed using the original colors, the first coefficient is the design matrix S Cx and vectors

number

number

number

[0270] Optimized global reshaping solution or optimized value for the first coefficient

number

number

[0271] Optimal value for the first coefficient of the global reshaping function

number

number

[0272] These extended operations involve the vector in equation (24) for the global reshaping function.

number

number

number

[0273] Optimized local reshaping solution or optimized value for the second coefficient

number

number

[0274] Block 2098 involves constructing local reshaping functions (e.g., fused, MMR-based, TPB-based, etc.).

[0275] Given two global reshaping functions generated in block 2096, θ <l>< / l> Cx,G1 The first global reshaping function weighting factor and θ are expressed as follows: <l>< / l> Cx,G2A second global reshaping function weighting factor, expressed as , may be assigned to the first global reshaping function and the second global reshaping function, respectively. In some operating scenarios, the two weighting factors satisfy the following constraints / conditions:

number

[0276] Therefore, the local reshaping described herein may be performed using a fused local reshaping function generated by fusing a first coefficient for a global reshaping function with a second coefficient for a second global reshaping function, as follows:

number

[0277] For illustrative purposes only, the mid-L1 values ​​predicted or estimated from 12-bit (offline) training input images in equations (64) through (76) above. <g>It is used as an index value represented as follows. Given a bit depth of 12 bits, the index value <g>The value can be selected from the range [0, 4095]. Similarly, given a bit depth of 10 bits, the index value <g>This can be a value selected from the range [0, 1023].

[0278] Similarly, the mid-L1 values ​​predicted or estimated from 12-bit input images (e.g., actual, untrained, to be improved, online processed, etc.) in equations (64) through (76) above <l>It is used as an index value represented as follows. Given a bit depth of 12 bits, the index value <l>The value can be selected from the range [0, 4095]. Similarly, given a bit depth of 10 bits, the index value <l>This can be a value selected from the range [0, 1023].

[0279] Index value <g>and / or <l>The memory space, volatile or non-volatile storage space, etc., used to store operating parameters such as MMR or TPB coefficients for all possible combinations of reshaping functions / mappings can be relatively large and expensive.

[0280] In some operating scenarios, the index value <g>and / or <l>Only operating parameters such as MMR or TPB coefficients for a true subset of index values ​​(e.g., representative, etc.) in all possible combinations are calculated or generated offline and stored / cached in memory space or storage. For example, the entire range of values ​​for all possible combinations can be partitioned or divided into multiple subranges (e.g., 16, 64, 128, 256, etc.). Representative index values <g>or <l>However, for each sub-range among multiple sub-ranges (for example, evenly, at specific positions, every 16, every 64, every 128, every 256, etc.), a representative index value is selected. <g>(and corresponding <l>Operating parameters such as MMR or TPB coefficients (values)

number

[0281] During the system boot-up of an image processing system that performs a local reshaping operation on an input image, a representative index value <g>(and corresponding <l>Operating parameters such as MMR or TPB coefficients (values)

number

[0282] At runtime, local reshaping as described herein applies to any index value within a relatively wide range of values, such as [0,4095], [0,1023], etc. <l>This can be performed on input images that have the following pre-generated index values. <g>(and corresponding <l>For index values ​​not covered by the (value), operating parameters such as MMR or TPB coefficients may not be readily available from memory space or storage at the start of system boot-up. Index values ​​not included in the representative index values. <l>For this, pre-generated index values <g>and / or <l>The interpolation operation may be performed after the available operating parameters, such as MMR or TPB coefficients, for the index values ​​covered by (for example, the two nearest ones) have been retrieved, loaded, or otherwise made available.

[0283] Figure 2T shows an exemplary flow for generating a 3DMT-based chroma-local reshaping function or mapping. For illustrative purposes, an MMR or TPB-based chroma-local reshaping function is generated to enhance the local saturation of the input HDR image 2192. The chroma-local reshaping function (or mapping) is δ to achieve local saturation changes and soft clipping. C M C , β, θ <l>< / l> Cx,G1 It can be generated based at least partially on an improvement parameter 2194 such as the following.

[0284] Block 2196 includes constructing a 3DMT using a 3D histogram having bins that store the count of HDR pixels in the input HDR image (2192) and the mean of the corresponding input SDR image in the image pair formed by the input HDR image and the input SDR image.

[0285] In one example, the corresponding input SDR image may be one that was previously used to generate an HDR image (2192) (for example, through content mapping, tone mapping, and / or reshaping operations). In another example, the corresponding input SDR image may be one that was previously generated from an HDR image (2192) (for example, through content mapping, tone mapping, and / or reshaping operations).

[0286] Block 2198 involves constructing a global chroma reshaping function using 3DMT. For example, a first MMR or TPB coefficient may be generated to specify the global chroma reshaping function using an MMR term or a complete set of tensor product B-spline basis functions.

[0287] Block 2200 improves the parameter δ for local saturation changes. C M C Then, using β, the 3DMT is made into each modified 3DMT, including the modified mapping pair.

[0288] Block 2202 involves constructing a second global chroma reshaping function using the modified 3DMT. The second global chroma reshaping function can be merged with the first global reshaping function. For example, a second MMR or TPB coefficient may be generated to specify the global reshaping function using an MMR term or a tensor product B-spline basis function.

[0289] Block 2204 is the improvement parameter θ <l>< / l> Cx,G1 , θ <l>< / l> Cx,G2 This involves using a weighted fusion with both global chroma reshaping functions derived from the original 3DMT and the modified 3DMT to generate the corresponding fused cross-channel local reshaping function. Block 2206 includes outputting the (fused) cross-channel local reshaping function.

[0290] Similar to local lumar reshaping, local chromar reshaping can use a guidance image with index values ​​to identify specific local chromar reshaping functions / mappings for specific input pixels of the input image. There are several different ways to construct a guidance image for local chromar reshaping. In the first example, the guidance image may be a lumar-independent guidance image constructed based on chromar codewords in the input image, or it may be a saturation image calculated from the input image. In the second example, the guidance image may be constructed based on lumar and chromar codewords in the input image. In the third example, the same guidance image for local lumar reshaping can be used as a guidance image for local chromar reshaping, or to derive it. Depending on the color sampling format, such as 444 or 420, the guidance image for local lumar reshaping may be downsampled as appropriate to fit the chroma channel image size. In many operational scenarios, chroma guide images can be reused for chroma reshaping without introducing visual artifacts, while providing a better appearance, including improved local saturation, in the locally reshaped image.

[0291] Exemplary process flow Figure 4 shows an exemplary process flow according to one embodiment. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) can perform this process flow. In block 402, the image processing system performs a first reshaping mapping on a first image represented by a first domain to generate a second image represented by a second domain. The first domain has a first dynamic range that is different from the second dynamic range of the second domain.

[0292] In block 404, the image processing system performs a second reshaping mapping on the second image represented in the second domain to generate a third image represented in the first domain. The third image is perceptually different from the first image in at least one of the following: global contrast, global saturation, local contrast, local saturation, etc.

[0293] In block 406, the image processing system renders the display image derived from the third image onto the display device.

[0294] In one embodiment, the first reshaping mapping and the second reshaping mapping form one of the following: (a) a combination of global forward reshaping mapping and global backward reshaping mapping; (b) a combination of global backward reshaping mapping and global forward reshaping mapping; (c) a combination of global forward reshaping mapping and local backward reshaping mapping; (d) a combination of global backward reshaping mapping and local forward reshaping mapping; (e) a combination of local backward reshaping mapping and global forward reshaping mapping; (f) a combination of local forward reshaping mapping and global backward reshaping mapping; (g) a combination of local forward reshaping mapping and local backward reshaping mapping; (h) a combination of local backward reshaping mapping and local forward reshaping mapping; and so on.

[0295] In one embodiment, the first dynamic range for the first image and the second dynamic range for the second image form one of the following: a combination of high dynamic range (HDR) for the first image and standard dynamic range (SDR) for the second image; a combination of SDR for the first image and HDR for the second image, and so on.

[0296] In one embodiment, an encoded image is generated from a third reshaping mapping performed on a second image; the encoded image is encoded in a video signal received by a receiving device; and the receiving device generates a display image from a decoded version of the encoded image received along with the video signal.

[0297] In one embodiment, the first image represented by the first domain is generated and uploaded by a mobile device.

[0298] In one embodiment, at least one of the first reshaping mapping and the second reshaping mapping includes a ruma-local reshaping mapping.

[0299] In one embodiment, the ruma-local reshaping mapping represents one of the following: (a) a single-channel ruma-local reshaping mapping that generates an output ruma codeword from an input ruma codeword independently of the input chroma codeword; (b) a cross-channel ruma-local reshaping mapping that generates an output ruma codeword from both the input ruma codeword and the input chroma codeword.

[0300] In one embodiment, the luma-local reshaping mapping represents a cross-channel luma-local reshaping mapping; the cross-channel luma-local reshaping mapping is generated by fusing a cross-channel luma-global reshaping mapping with a pre-fusion cross-channel luma-local reshaping mapping; the cross-channel luma-global reshaping is generated using a three-dimensional mapping table (3DMT) calculated using codewords in the image pair; and the pre-fusion cross-channel luma-local reshaping mapping is generated using a modified 3DMT obtained by modifying the 3DMT with a local contrast enhancement function.

[0301] In one embodiment, at least one of the first reshaping mapping and the second reshaping mapping includes a chroma-local reshaping mapping.

[0302] In one embodiment, the chroma-local reshaping mapping represents one of the following: (a) a cross-channel MMR chroma-local reshaping mapping that generates an output chroma codeword from an input lumer codeword and an input chroma codeword; or (b) a cross-channel TPB chroma-local reshaping mapping that generates an output chroma codeword from both an input lumer codeword and an input chroma codeword.

[0303] In one embodiment, a chroma-local reshaping mapping is generated by fusing a first cross-channel chroma-global reshaping mapping and a second cross-channel chroma-global reshaping mapping; the first cross-channel chroma-global reshaping is generated using a three-dimensional mapping table (3DMT) calculated using codewords in the image pair; and the second cross-channel chroma-global reshaping mapping is generated using a modified 3DMT derived by modifying the 3DMT with a saturation enhancement function.

[0304] In one embodiment, to reduce banding artifacts, image filtering is applied to at least one of a first image, a second image, or an image derived from the second image using a noise-injected guide image; the noise-injected guide image includes a per-pixel noise-injected local reshaping function index; the filtered image includes codewords generated from the image filtering using the noise-injected guide image; and the codewords in the filtered image are applied with further noise injection to generate noise-injected codewords.

[0305] In one embodiment, a second image is received by a video encoder as an input image in a sequence of input images; the sequence of input images received by the video encoder is encoded into a video signal by the video encoder.

[0306] In one embodiment, a computing device such as a display device, a mobile device, a set-top box, or a multimedia device is configured to perform one of the methods described above. In one embodiment, the device includes a processor and is configured to perform one of the methods described above. In one embodiment, a non-temporary computer-readable storage medium storing software instructions, which, when executed by one or more processors, causes the execution of one of the methods described above.

[0307] In one embodiment, the computing device comprises one or more processors and one or more storage media storing a set of instructions that, when executed by the one or more processors, cause the execution of any of the methods described above.

[0308] While separate embodiments are discussed herein, it should be noted that any combination of the embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.

[0309] Exemplary implementation of a computer system Embodiments of the present invention may be implemented by computer systems, systems configured in electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field-programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or devices comprising one or more such systems, devices, or components. Computers and / or ICs can execute, control, or run instructions relating to the adaptive perceptual quantization of images with an enhanced dynamic range, as described herein. Computers and / or ICs can compute any of the various parameters or values ​​related to the adaptive perceptual quantization process described herein. Embodiments of images and videos can be implemented in hardware, software, firmware, and various combinations thereof.

[0310] Certain implementations of the present invention include a computer processor that executes software instructions causing the processor to perform the method of the present disclosure. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., may implement the method relating to the adaptive perceptual quantization of HDR images as described above by executing software instructions in program memory accessible to the processor. Embodiments of the present invention may be provided in the form of a program product. A program product may include any non-temporary medium that carries a set of computer-readable signals, which, when executed by a data processor, causes the data processor to perform the method of the embodiment of the present invention. A program product according to an embodiment of the present invention may be any of a wide variety of forms. A program product may include physical media such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, and electronic data storage media including ROMs and flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0311] Where components (e.g., software modules, processors, assemblies, devices, circuits, etc.) are referred to above, unless otherwise indicated, references to such components (including references to “means”) should be interpreted as including any components that perform the function of the described component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structures that perform the function of the described component (e.g., functionally equivalent).

[0312] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. A special-purpose computing device may be fixedly configured to perform those techniques, or may include one or more digital electronic devices such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) permanently programmed to perform those techniques, or may include one or more general-purpose hardware processors programmed to perform those techniques according to program instructions in firmware, memory, other storage, or combination. Such a special-purpose computing device may also achieve those techniques by combining custom fixed-configuration logic, ASICs, or FPGAs with custom programming. A special-purpose computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device incorporating fixed configuration and / or programmed logic to implement these techniques.

[0313] For example, Figure 5 is a block diagram showing a computer system 500 in which an exemplary embodiment of the present disclosure may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for communicating information and a hardware processor 504 coupled to the bus 502 for processing information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0314] The computer system 500 also includes main memory 506 coupled to bus 502 for storing information and instructions to be executed by processor 504, such as random access memory (RAM) or other dynamic storage devices. Main memory 506 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When such instructions are stored in a non-temporary storage medium accessible to processor 504, the computer system 500 becomes a special-purpose machine customized to perform the operations specified in the instructions.

[0315] The computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to the bus 502 for storing static information and instructions for the processor 504. A storage device 510, such as a magnetic disk or optical disk, or semiconductor RAM, is provided and coupled to the bus 502 for storing information and instructions.

[0316] The computer system 500 may be coupled via a bus 502 to a display 512, such as a liquid crystal display, for displaying information to the computer user. An input device 514, including alphanumeric and other keys, is coupled to the bus 502 to transmit information and command selections to the processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor directional keys, for transmitting directional information and command selections to the processor 504 and for controlling cursor movement on the display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), so that the device can specify a position in a plane.

[0317] The computer system 500 may use customized fixed-configuration logic, one or more ASICs or FPGAs, firmware and / or programmed logic that, in combination with the computer system, make or program the computer system 500 a special-purpose machine. According to one embodiment, the technique described herein is executed by the computer system 500 in response to the processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. By executing the sequence of instructions contained in main memory 506, the processor 504 performs the process steps described herein. In alternative embodiments, fixed-configuration circuitry may be used instead of or in combination with software instructions.

[0318] As used in this paper, the term “storage medium” refers to any non-temporary medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage mediums may include non-volatile and / or volatile media. Non-volatile media include, for example, optical or magnetic disks such as storage device 510. Volatile media include dynamic memory such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, semiconductor drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, and any other memory chips or cartridges.

[0319] A storage medium is distinct from a transmission medium, but may be used in conjunction with a transmission medium. A transmission medium participates in transferring information between storage mediums. For example, a transmission medium includes coaxial cables, copper wires, and optical fibers, and includes wires forming a bus 502. A transmission medium may also take the form of acoustic or optical waves, such as those generated during radio and infrared data communications.

[0320] Various forms of media may be involved in transporting one or more sequences of one or more instructions to the processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or semiconductor drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send them over a telephone line using a modem. A modem local to the computer system 500 can receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data onto the bus 502. The bus 502 transports the data to the main memory 506, from which the processor 504 retrieves and executes the instructions. The instructions received by the main memory 506 may optionally be stored on the storage device 510 before or after execution by the processor 504.

[0321] The computer system 500 also includes a communication interface 518 coupled to bus 502. The communication interface 518 provides bidirectional data communication coupling to a network link 520 connected to a local network 522. For example, the communication interface 518 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or modem for providing data communication connectivity to a corresponding type of telephone line. As another example, the communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity to a compatible LAN. A wireless link may also be implemented. In any such implementation, the communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0322] Network link 520 typically provides data communication to other data devices through one or more networks. For example, network link 520 may provide connection to a data facility operated by a host computer 524 or an Internet service provider (ISP) 526 through a local network 522. ISP 526 provides data communication services through a global packet data network now commonly referred to as the “Internet” 528. Both the local network 522 and the Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals across various networks and signals over network link 520 and through communication interface 518 that carry digital data to and from computer system 500 are exemplary forms of transmission media.

[0323] The computer system 500 can send messages and receive data, including program code, through a network(s), network link 520, and communication interface 518. In the internet example, server 530 may send requested code for an application program through the internet 528, ISP 526, local network 522, and communication interface 518.

[0324] The received code may be executed by the processor 504 upon receipt, and / or stored in the memory device 510 or other non-volatile memory for later execution.

[0325] 7. Equivalents, extensions, substitutes, etc. The above specification has described embodiments of the disclosure, referring to numerous specific details that may vary depending on the implementation. Thus, the sole and exclusive indicator of what constitutes an embodiment of the present invention, and what is intended by the applicant to be an embodiment claimed by the present invention, is the specific form in which such claims are patented, including any subsequent amendments. If there are any definitions of terms included in such claims that are expressly provided in this document, those definitions govern the meaning of such terms as used in the claims. Therefore, no limitations, elements, characteristics, features, advantages, or attributes not expressly provided in the claims should in any way limit the scope of such claims. Accordingly, the specification and drawings should be considered illustrative and not restrictive. Several aspects are described below. [Aspect 1] A way to improve the image: A step of performing a first reshaping mapping on a first image represented by a first domain to generate a second image represented by a second domain, wherein the first domain has a first dynamic range different from the second dynamic range of the second domain; A step of performing a second reshaping mapping on the second image represented by the second domain to generate a third image represented by the first domain, wherein the third image differs from the first image in at least one of global contrast, global saturation, local contrast, or local saturation, the difference in at least one of global contrast and global saturation being caused by at least one of first and second reshaping mappings represented by the same function applied to all pixels of each image, and the difference in at least one of local contrast and local saturation being caused by at least one of first and second reshaping mappings represented by a function selectable at the pixel level of each image; The process includes the step of rendering the display image derived from the third image onto a display device, method. [Aspect 2] At least one of the first reshaping mapping and the second reshaping mapping includes a ruma-local reshaping mapping; the ruma-local reshaping mapping represents a cross-channel ruma-local reshaping mapping that generates an output ruma codeword from both an input ruma codeword and an input chroma codeword; the cross-channel ruma-local reshaping mapping is generated by fusing a cross-channel ruma-global reshaping mapping with a pre-fuse cross-channel ruma-local reshaping mapping; the cross-channel ruma-global reshaping mapping and the pre-fuse cross-channel ruma-local The method according to aspect 1, wherein merging the reshaping mappings comprises calculating a linear combination of the cross-channel lumar-global reshaping mapping and the cross-channel lumar-local reshaping mapping before merging; the cross-channel lumar-global reshaping mapping is generated using a three-dimensional mapping table (3DMT) calculated using codewords in the image pair represented by the first and second images; and the cross-channel lumar-local reshaping mapping before merging is generated using a modified 3DMT derived from modifying the 3DMT with a local contrast enhancement function. [Aspect 3] The process further includes the step of constructing a three-dimensional mapping table (3DMT) from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, and the 3DMT stores a plurality of mapping pairs, each of which corresponds to a respective bin in the plurality of bins, and includes (a) the count value of the first codeword located in each of the bins in the one or more first images, and (b) the average second luma codeword value of the second codeword corresponding to the first codeword in the one or more second images. The method described in Embodiment 2. [Aspect 4] The method according to any one of embodiments 1 to 3, wherein at least one of the first reshaping mapping and the second reshaping mapping includes a chroma-local reshaping mapping; the chroma-local reshaping mapping is generated by fusing a first cross-channel chroma-global reshaping mapping and a second cross-channel chroma-global reshaping mapping; fusing the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping includes computing a linear combination of the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping; the first cross-channel chroma-global reshaping mapping is generated using a three-dimensional mapping table (3DMT) computed using codewords in an image pair represented by the first and second images; and the second cross-channel chroma-global reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT with a local saturation enhancement function. [Aspect 5] The process further includes the step of constructing a three-dimensional mapping table (3DMT) from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, and the 3DMT stores a plurality of mapping pairs, each of which corresponds to a respective bin in the plurality of bins, and includes (a) the count value of the first codeword located in each of the bins in the one or more first images, and (b) the average chroma codeword value of the second codeword corresponding to the first codeword in the one or more second images. The method according to aspect 4. [Aspect 6] The method according to aspect 4 or 5, wherein the chroma-local reshaping mapping represents one of the following: (a) a cross-channel multivariate multiple regression (MMR) chroma-local reshaping mapping that generates an output chroma codeword from an input lumer codeword and an input chroma codeword; or (b) a cross-channel tensor product B-spline (TPB) chroma-local reshaping mapping that generates an output chroma codeword from both an input lumer codeword and an input chroma codeword. [Aspect 7] The method according to embodiment 1, wherein at least one of the first reshaping mapping and the second reshaping mapping includes a ruma-local reshaping mapping; the ruma-local reshaping mapping represents a single-channel ruma-local reshaping mapping that generates an output ruma codeword from an input ruma codeword independently of the input chroma codeword. [Aspect 8] The method according to any one of embodiments 1 to 7, wherein the first reshaping mapping and the second reshaping mapping form one of the following: (a) a combination of a global forward reshaping mapping and a global backward reshaping mapping; (b) a combination of a global backward reshaping mapping and a global forward reshaping mapping; (c) a combination of a global forward reshaping mapping and a local backward reshaping mapping; (d) a combination of a global backward reshaping mapping and a local forward reshaping mapping; (e) a combination of a local backward reshaping mapping and a global forward reshaping mapping; (f) a combination of a local forward reshaping mapping and a global backward reshaping mapping; (g) a combination of a local forward reshaping mapping and a local backward reshaping mapping; (h) a combination of a local backward reshaping mapping and a local forward reshaping mapping. [Aspect 9] The method according to any one of embodiments 1 to 8, wherein the first dynamic range for the first image and the second dynamic range for the second image form one of the following: a combination of high dynamic range (HDR) for the first image and standard dynamic range (SDR) for the second image; or a combination of SDR for the first image and HDR for the second image. [Aspect 10] The method according to any one of embodiments 1 to 9, wherein an encoded image is generated from a third reshaping mapping performed on the second image; the encoded image is encoded in a video signal received by a receiving device; and the receiving device generates the display image from a decoded version of the encoded image received together with the video signal. [Aspect 11] The method according to any one of embodiments 1 to 10, wherein the first image represented in the first domain is generated and uploaded by a mobile device. [Aspect 12] The method according to any one of embodiments 1 to 11, wherein image filtering is applied to at least one of the first image, the second image, or an image derived from the second image using a noise-injected guide image to reduce banding artifacts; the image filtering includes generating a guide image from the image to be filtered, the guide image providing local reshaping function index values ​​for each pixel of the image to be filtered; the noise-injected guide image includes a local reshaping function index calculated from the guide image and with per-pixel noise injected; the filtered image includes a codeword generated from the image filtering using the noise-injected guide image; the codeword in the filtered image is applied together with further noise injection to generate a noise-injected codeword. [Aspect 13] The method according to any one of embodiments 1 to 12, wherein the second image is received by a video encoder as an input image in a sequence of input images; and the sequence of input images received by the video encoder is encoded into a video signal by the video encoder. [Aspect 14] An apparatus having one or more processors and configured to perform the method described in any one of embodiments 1 to 13. [Aspect 15] A non-temporary computer-readable storage medium storing computer-executable instructions for performing the method described in any one of embodiments 1 to 13 on one or more processors.

[0326] Bulleted Exemplary Embodiments The present invention includes, but is not limited to, the following Enumerated Example Embodiments (EEEs) describing the structure, features, and functions of some parts of the embodiments of the present invention, and can be embodied in any of the forms described herein.

[0327] [EEE1] A way to improve the image: A step of performing a first reshaping mapping on a first image represented by a first domain to generate a second image represented by a second domain, wherein the first domain has a first dynamic range different from the second dynamic range of the second domain; A step of performing a second reshaping mapping on the second image represented by the second domain to generate a third image represented by the first domain, wherein the third image is perceptually different from the first image in at least one of global contrast, global saturation, local contrast, or local saturation; The process includes the step of rendering the display image derived from the third image onto a display device, method.

[0328] [EEE2] The method according to EEE1, wherein the first reshaping mapping and the second reshaping mapping form one of the following: (a) a combination of a global forward reshaping mapping and a global backward reshaping mapping; (b) a combination of a global backward reshaping mapping and a global forward reshaping mapping; (c) a combination of a global forward reshaping mapping and a local backward reshaping mapping; (d) a combination of a global backward reshaping mapping and a local forward reshaping mapping; (e) a combination of a local backward reshaping mapping and a global forward reshaping mapping; (f) a combination of a local forward reshaping mapping and a global backward reshaping mapping; (g) a combination of a local forward reshaping mapping and a local backward reshaping mapping; (h) a combination of a local backward reshaping mapping and a local forward reshaping mapping.

[0329] [EEE3] The method according to EEE1 or 2, wherein the first dynamic range for the first image and the second dynamic range for the second image form one of the following: a combination of high dynamic range (HDR) for the first image and standard dynamic range (SDR) for the second image; or a combination of SDR for the first image and HDR for the second image.

[0330] [EEE4] The method according to any one of EEE1 to 3, wherein an encoded image is generated from a third reshaping mapping performed on the second image; the encoded image is encoded in a video signal received by a receiving device; and the receiving device generates the display image from a decoded version of the encoded image received together with the video signal.

[0331] [EEE5] The first image represented in the first domain is generated and uploaded by a mobile device, according to the method described in any one of EEE1 to 4.

[0332] [EEE6] The method according to any one of EEE1 to 5, wherein at least one of the first reshaping mapping and the second reshaping mapping includes a ruma-local reshaping mapping.

[0333] [EEE7] The method according to EEE6, wherein the ruma-local reshaping mapping represents one of the following: (a) a single-channel ruma-local reshaping mapping that generates an output ruma codeword from an input ruma codeword independently of the input chroma codeword; or (b) a cross-channel ruma-local reshaping mapping that generates an output ruma codeword from both the input ruma codeword and the input chroma codeword.

[0334] [EEE8] The method according to EEE7, wherein the luma-local reshaping mapping represents the cross-channel luma-local reshaping mapping; the cross-channel luma-local reshaping mapping is generated by fusing a cross-channel luma-global reshaping mapping with a pre-fusion cross-channel luma-local reshaping mapping; the cross-channel luma-global reshaping is generated using a three-dimensional mapping table (3DMT) calculated using codewords in an image pair; and the pre-fusion cross-channel luma-local reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT with a local contrast enhancement function.

[0335] [EEE9] The method according to any one of EEE1 to 8, wherein at least one of the first reshaping mapping and the second reshaping mapping includes a chroma-local reshaping mapping.

[0336] [EEE10] The method according to EEE9, wherein the chroma-local reshaping mapping represents one of the following: (a) a cross-channel multivariate multiple regression (MMR) chroma-local reshaping mapping that generates an output chroma codeword from an input lumer codeword and an input chroma codeword; or (b) a cross-channel tensor product B-spline (TPB) chroma-local reshaping mapping that generates an output chroma codeword from both an input lumer codeword and an input chroma codeword.

[0337] [EEE11] The method according to EEE9 or 10, wherein the chroma-local reshaping mapping is generated by fusing a first cross-channel chroma-global reshaping mapping and a second cross-channel chroma-global reshaping mapping; the first cross-channel chroma-global reshaping is generated using a three-dimensional mapping table (3DMT) computed using codewords in an image pair; and the second cross-channel chroma-global reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT using a local saturation enhancement function.

[0338] [EEE12] The method according to any one of EEE1 to 11, wherein image filtering is applied to at least one of the first image, the second image, or an image derived from the second image using a noise-injected guide image to reduce banding artifacts; the noise guide image includes a pixel-by-pixel noise-injected local reshaping function index; the filtered image includes a codeword generated from the image filtering using the noise-injected guide image; and the codeword in the filtered image is applied along with further noise injection to generate a noise-injected codeword.

[0339] [EEE13] The method according to any one of EEE1 to 12, wherein the second image is received by a video encoder as an input image in a sequence of input images; and the sequence of input images received by the video encoder is encoded into a video signal by the video encoder.

[0340] [EEE14] A method for improving an image: A step of constructing a three-dimensional mapping table (3DMT) from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, the 3DMT stores a plurality of mapping pairs, each of the mapping pairs corresponds to a respective bin in the plurality of bins, and includes (a) the count value of the first codeword located in each of the bins in the one or more first images, and (b) the average second luma codeword value of the second codeword corresponding to the first codeword in the one or more second images; A step of generating a modified 3DMT by applying a local contrast enhancement operation to the average second luma codeword value in each of the plurality of mapping pairs; The steps include generating a global rumor reshaping function from the aforementioned 3DMT, and generating a local rumor reshaping function from the aforementioned modified 3DMT; The steps include generating a final local rumor reshaping function from the global rumor reshaping function and the local rumor reshaping function; The steps include applying the final local lumens reshaping function to the input image to generate a local contrast enhancement image, wherein a display image is derived from the local contrast enhancement image and rendered on a display device, method.

[0341] [EEE14] A method for improving an image: A step of constructing a three-dimensional mapping table (3DMT) from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, the 3DMT stores a plurality of mapping pairs, each of the mapping pairs corresponds to a respective bin in the plurality of bins, and includes (a) the count value of the first codeword located in each of the bins in the one or more first images, and (b) the average chroma codeword value of the second codeword corresponding to the first codeword in the one or more second images; A step of generating a modified 3DMT by applying one or more local saturation enhancement operations to the average chroma codeword value in each of the plurality of mapping pairs; For each chroma channel, the process involves generating a global chroma reshaping function from the 3DMT, and generating a second global chroma reshaping function from the modified 3DMT; For each chroma channel, the process involves generating a final local chroma reshaping function from the global chroma reshaping function and the second global chroma reshaping function; The process includes the step of applying the final local chroma reshaping function to the input image for each chroma channel to generate a locally saturated image, wherein a display image is derived from the locally saturated image and rendered on a display device. method.

[0342] [EEE15] The method according to any one of EEE1 to 14, wherein image filtering is represented as guide image filtering applied using a guide image, and the guide image includes high-frequency feature values ​​calculated for a plurality of pixel positions, each weighted by the inverse of an image gradient calculated for a plurality of pixel positions in one or more channels of the guide image.

[0343] [EEE16] The method according to any one of EEE1 to 15, wherein the guide image includes guide image values ​​derived at least partially based on a set of halo reduction operation parameters, and the method further: For each of the multiple image pairs, a region-based texture measure in one or more local foreground areas is calculated as the ratio of the variances of the training image for the first dynamic range and the corresponding training image for the second dynamic range in the image pair, using multiple sets of candidate values ​​for the set of halo reduction operation parameters, wherein the one or more local foreground areas are sampled from a foreground mask identified in the training image for the first dynamic range in the image pair; For each of the plurality of image pairs, the step of calculating a halo measure in one or more local halo artifact areas as a correlation coefficient between the training image of the first dynamic range and the corresponding training image of the second dynamic range in the image pair, using the plurality of sets of candidate values ​​for the set of halo reduction operation parameters, wherein the one or more local halo artifact areas are sampled from the halo mask identified in the corresponding training image of the second dynamic range; For the set of candidate values ​​for the set of halo reduction operation parameters, the steps include: calculating a weighted sum of (a) all region-based texture measures for all local foreground areas sampled from the foreground mask identified in the training images of the first dynamic range in the set of image pairs, and (b) all halo measures for all local halo artifact areas sampled from the halo mask identified in the corresponding training images of the second dynamic range in the set of image pairs; A step of determining an optimized set of values ​​for the set of halo reduction operation parameters, wherein the optimized set of values ​​is used to generate a minimized weighted sum of (a) all of the region-based texture measures for all local foreground areas sampled from the foreground mask identified in the training images of the first dynamic range in the plurality of image pairs, and (b) all of the halo measures for all local halo artifact areas sampled from the halo mask identified in the corresponding training images of the second dynamic range in the plurality of image pairs; The step includes applying the image filtering at multiple spatial kernel size levels to the input image using the set of optimized values ​​for the set of halo reduction operating parameters, method.

[0344] [EEE16] A computer system configured to perform any one of the methods described in any one of the EEE1 through EEE15.

[0345] [EEE17] A device having a processor and configured to perform the method described in any one of EEE1 to 15.

[0346] [EEE18] A non-temporary computer-readable storage medium storing computer-executable instructions for performing the method described in any one of the EEE1 to EEE15.< / l> < / g> < / l> < / l> < / g> < / l> < / l> < / g> < / l> < / g> < / l> < / g> < / l> < / g> < / l> < / g> < / l> < / l> < / l> < / g> < / g> < / g>

Claims

1. Here's a way to improve the image: A step (402) of performing a first reshaping mapping on a first image represented by a first domain to generate a second image represented by a second domain, wherein the first domain has a first dynamic range different from the second dynamic range of the second domain; Step (404) of performing a second reshaping mapping on the second image represented by the second domain to generate a third image represented by the first domain, wherein the third image differs from the first image in at least one of local contrast or local saturation, and the difference in at least one of local contrast and local saturation is caused by at least one of the first and second reshaping mappings represented by a pixel-level selectable function; The process includes the step (406) of rendering the display image derived from the third image onto a display device (140, 140-1), At least one of the first reshaping mapping and the second reshaping mapping includes a ruma-local reshaping mapping; the ruma-local reshaping mapping represents a cross-channel ruma-local reshaping mapping that generates an output ruma codeword from both the input ruma codeword and the input chroma codeword; the cross-channel ruma-local reshaping mapping is generated by mixing a cross-channel ruma-global reshaping mapping with an initial cross-channel ruma-local reshaping mapping (268, 268', 284); the cross-channel ruma-global reshaping mapping and the initial channel Mixing the cross-channel lumar-local reshaping mappings involves calculating a linear combination of the cross-channel lumar-global reshaping mapping and the initial cross-channel lumar-local reshaping mapping; the cross-channel lumar-global reshaping mapping is generated using a three-dimensional mapping table (3DMT) calculated using codewords in the image pair represented by the first and second images; and the initial cross-channel lumar-local reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT with a local contrast enhancement function. method.

2. At least one of the first reshaping mapping and the second reshaping mapping includes a chroma-local reshaping mapping; the chroma-local reshaping mapping is generated by mixing a first cross-channel chroma-global reshaping mapping and a second cross-channel chroma-global reshaping mapping (298, 2098, 2204); Mixing the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping involves calculating a linear combination of the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping; The method according to claim 1, wherein the first cross-channel chroma-global reshaping mapping is generated using a 3DMT calculated using codewords in the image pair represented by the first and second images; and the second cross-channel chroma-global reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT using a local saturation enhancement function.

3. Here's a way to improve the image: A step (402) of performing a first reshaping mapping on a first image represented by a first domain to generate a second image represented by a second domain, wherein the first domain has a first dynamic range different from the second dynamic range of the second domain; Step (404) of performing a second reshaping mapping on the second image represented by the second domain to generate a third image represented by the first domain, wherein the third image differs from the first image in at least one of local contrast or local saturation, and the difference in at least one of local contrast and local saturation is caused by at least one of the first and second reshaping mappings represented by a pixel-level selectable function; The process includes the step (406) of rendering the display image derived from the third image onto a display device (140, 140-1), At least one of the first reshaping mapping and the second reshaping mapping includes a chroma-local reshaping mapping; the chroma-local reshaping mapping is generated by mixing a first cross-channel chroma-global reshaping mapping and a second cross-channel chroma-global reshaping mapping (298, 2098, 2204); Mixing the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping involves calculating a linear combination of the first cross-channel chroma-global reshaping mapping and the second cross-channel chroma-global reshaping mapping; The first cross-channel chroma global reshaping mapping is generated using a 3DMT calculated using codewords in the image pair represented by the first and second images; the second cross-channel chroma global reshaping mapping is generated using a modified 3DMT derived from modifying the 3DMT using a local saturation enhancement function. method.

4. The step (262) further includes constructing the 3DMT used to generate the cross-channel lumer global reshaping mapping from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, and the 3DMT stores a plurality of mapping pairs, each of which corresponds to a respective bin in the plurality of bins, and includes (a) a count value of the first codeword located in each of the bins in the one or more first images, and (b) an average second lumer codeword value of the second codeword to which the first codeword corresponds in the one or more second images. The method according to claim 1 or 2.

5. The process further includes the step of constructing the 3DMT used to generate the first cross-channel chroma global reshaping mapping from one or more first images encoded with a first codeword in a first codeword space and one or more second images encoded with a second codeword in a second codeword space, wherein the one or more first images correspond to the one or more second images, the first codeword space is divided into a plurality of bins, and the 3DMT stores a plurality of mapping pairs, each of which corresponds to a respective bin in the plurality of bins, and includes (a) a count value of the first codeword located in each of the bins in the one or more first images, and (b) an average chroma codeword value of the second codeword corresponding to the first codeword in the one or more second images. The method according to claim 2 or 3.

6. The method according to claim 5, wherein the chroma-local reshaping mapping represents one of the following: (a) a cross-channel multivariate multiple regression (MMR) chroma-local reshaping mapping that generates an output chroma codeword from an input lumer codeword and an input chroma codeword; or (b) a cross-channel tensor product B-spline (TPB) chroma-local reshaping mapping that generates an output chroma codeword from both an input lumer codeword and an input chroma codeword.

7. The first reshaping mapping and the second reshaping mapping are: (a) a combination of global forward reshaping mapping and global backward reshaping mapping; (b) a combination of global backward reshaping mapping and global forward reshaping mapping; (c) a combination of global forward reshaping mapping and local backward reshaping mapping; (d) A combination of global back reshaping mapping and local forward reshaping mapping; (e) A combination of local back reshaping mapping and global forward reshaping mapping; (f) A combination of local forward reshaping mapping and global backward reshaping mapping; The method according to any one of claims 1 to 6, which forms one of the following: (g) a combination of a local forward reshaping mapping and a local backward reshaping mapping; or (h) a combination of a local backward reshaping mapping and a local forward reshaping mapping.

8. The method according to any one of claims 1 to 7, wherein the first dynamic range for the first image and the second dynamic range for the second image form one of the following: a combination of high dynamic range (HDR) for the first image and standard dynamic range (SDR) for the second image; or a combination of SDR for the first image and HDR for the second image.

9. The method according to any one of claims 1 to 8, wherein an encoded image is generated from a third reshaping mapping performed on the second image; the encoded image is encoded in a video signal received by a receiving device; and the receiving device generates the display image from a decoded version of the encoded image received together with the video signal.

10. The method according to any one of claims 1 to 9, wherein the first image represented in the first domain is generated and uploaded by a mobile device.

11. The method according to any one of claims 1 to 10, wherein image filtering is applied to at least one of the first image, the second image, or an image derived from the second image using a noise-injected guide image to reduce banding artifacts; the image filtering includes generating a guide image from the image to be filtered, the guide image providing local reshaping function index values ​​for the pixels of the image to be filtered; the noise-injected guide image includes a local reshaping function index calculated from the guide image and with per-pixel noise injected; the filtered image includes a codeword generated from the image filtering using the noise-injected guide image; the codeword in the filtered image is applied together with further noise injection to generate a noise-injected codeword.

12. The method according to any one of claims 1 to 11, wherein the second image is received by a video encoder as an input image in a sequence of input images; and the sequence of input images received by the video encoder is encoded into a video signal by the video encoder.

13. An apparatus (500) having one or more processors (504) and configured to perform the method described in any one of claims 1 to 12.

14. Non-temporary computer-readable storage media (506, 510) storing computer-executable instructions for performing the method according to any one of claims 1 to 12 on one or more processors (504).

Citation Information

Patent Citations

  • How to inverse tone map an image

    JP2017502602A

  • Chroma Reshaping for High Dynamic Range Images

    US20190320191A1

  • Reducing banding artifacts in backward-compatible HDR imaging

    WO2020072651A1

  • Interpolation of reshaping functions

    WO2020117603A1