Image optimization in mobile capture and editing applications

TPB-based image reshaping solutions in video capture applications address the limitations of SDR-HDR transitions by optimizing forward and reverse mappings and using boundary clipping, ensuring high-quality image reconstruction and compatibility across devices.

JP7797682B2Active Publication Date: 2026-01-13DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024554121
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-18
Filing Date
2023-03-17
Publication Date
2026-01-13
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

Existing video editing operations in SDR and HDR domains often result in adverse effects such as color space limitations and clipping, impairing the reversibility of edited images, particularly on mobile devices with limited computational resources.

Method used

Implementing tensor-product B-spline (TPB)-based image reshaping solutions in video capture applications, including mobile capture, to generate optimized forward and reverse reshaping mappings, and using boundary clipping techniques to preserve color space and maintain image quality.

Benefits of technology

The TPB-based solutions provide accurate reconstructed images with reduced power consumption, minimize data overhead, and ensure compatibility across various devices, while preserving image quality and enabling seamless transitions between SDR and HDR formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797682000179
    Figure 0007797682000179
  • Figure 0007797682000180
    Figure 0007797682000180
  • Figure 0007797682000181
    Figure 0007797682000181
Patent Text Reader

Abstract

The HDR color patch is sampled over an HDR color space parameterized by the parameters. From the sampled HDR color patch, a reference SDR color patch, an input HDR color patch, and a reference HDR color patch are generated. An optimization algorithm is performed to generate an optimized forward reshaping mapping and an optimized inverse reshaping mapping. The optimized forward reshaping mapping is used to forward reshape the input HDR image into a forward reshaped SDR image, and the optimized inverse reshaping mapping is used to inverse reshape the forward reshaped SDR image into a inverse reshaped HDR image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 321,363, filed March 18, 2022, and European Patent Application No. 22162984.3, filed March 18, 2022, the entire contents of each of which are incorporated by reference.

[0002] [Technical field] FIELD OF THE DISCLOSURE The present disclosure relates generally to images. More particularly, embodiments of the present disclosure relate to video codecs used to process images. [Background technology]

[0003] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive a range of intensities (e.g., luminance, luma) in an image, e.g., from darkest black (dark) to brightest white (highlight). In this sense, DR relates to "scene-referred" intensity. DR may also relate to the ability of a display device to properly or approximately render an intensity range of a particular width. In this sense, DR relates to "display-referred" intensity. Unless a particular meaning is explicitly specified as having particular significance at any point in the description herein, it should be presumed that the terms may be used in either sense, e.g., interchangeably.

[0004] As used herein, the term high dynamic range (HDR) refers to a DR width that spans approximately 14 to 15 orders of magnitude or more of the human visual system (HVS). In practice, the DR that humans can simultaneously perceive across a wide range of intensities may be somewhat truncated relative to HDR. As used herein, the terms enhanced dynamic range (EDR) or visual dynamic range (VDR) may individually or interchangeably refer to the DR perceivable within a scene or image by the human visual system (HVS), including eye movements, that allows for some light-adaptive changes across the scene or image. As used herein, EDR may refer to a DR that spans 5 to 6 orders of magnitude. Thus, while perhaps somewhat narrower relative to true scene-referential HDR, EDR still represents a wide DR width and may also be referred to as HDR.

[0005] In practice, an image comprises one or more color components of a color space (e.g., luma Y and chroma Cb and Cr), each color component represented with n bits of precision per pixel (e.g., n=8). Using non-linear luminance coding (e.g., gamma coding), images with n≦8 (e.g., color 24-bit JPEG images) are considered standard dynamic range images, while images with n>8 may be considered extended dynamic range images.

[0006] The reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​(e.g., luminance) of an input video signal and the output screen color values ​​(e.g., screen luminance) produced by the display. For example, ITU Rec. ITU-R BT.1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011), the entire contents of which are incorporated herein by reference, defines a reference EOTF for flat panel displays. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term “metadata” herein refers to any auxiliary information transmitted as part of a coded bitstream to assist a decoder in rendering a decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters, as described herein.

[0007] As used herein, the term "PQ" refers to perceptual luminance amplitude quantization. The human visual system responds to increases in light levels in a highly nonlinear manner. A person's ability to see a stimulus is affected by the luminance of the stimulus, its size, the spatial frequencies that make up the stimulus, and the luminance level to which the eye is adapted at the particular moment the stimulus is viewed. In some embodiments, a perceptual quantization function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An exemplary PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter "SMPTE"), the entire contents of which are incorporated herein by reference, where, given a fixed stimulus size, for each luminance level (e.g., stimulus level), the minimum visible contrast step at that luminance level is selected according to the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).

[0008] 200-1,000 cd / m 2 Displays supporting 100 nits of brightness represent a lower dynamic range (LDR), also called standard dynamic range (SDR), as opposed to EDR (or HDR). EDR content can be displayed on EDR displays that support a higher dynamic range (e.g., 1,000 nits to 5,000 nits or more). Such displays can be defined using an alternate EOTF that supports high brightness capabilities (e.g., 0 to 10,000 nits or more). Examples of such EOTFs are SMPTE 2084 and Rec. It is defined in ITU-R BT.2100, "Image parameter values ​​for high dynamic range television for use in production and international program exchange," (06 / 2017). As recognized by the inventors, improved techniques for structuring video content data that can be used to support the display capabilities of a wide variety of SDR and HDR display devices are desirable.

[0009] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Likewise, it should not be assumed that problems identified with one or more approaches have been recognized by virtue of this section in any prior art, unless otherwise indicated. [Brief explanation of the drawings]

[0010] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference symbols refer to similar elements and in which: [Figure 1A] 1 illustrates an exemplary process flow for applying an optimized reshaping operation. [Figure 1B] 1 illustrates an exemplary process flow for applying an optimized reshaping operation. [Figure 1C] 1 illustrates an exemplary process flow for applying an optimized reshaping operation. [Figure 1D] 1 illustrates an exemplary process flow for applying an optimized reshaping operation. [Figure 1E] 1 illustrates an exemplary process flow for applying an optimized reshaping operation. [Figure 2A]1 shows examples of exemplary irregularly shaped color distributions associated with HDR and SDR color spaces such as the R.2020, P3, and R.709 color spaces. [Figure 2B] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2C] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2D] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2E] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2F] 10 shows exemplary percentile Cb and Cr values ​​in a forward-reshaped SDR color space corresponding to different parameterized HDR color spaces. [Figure 2G] 10 shows exemplary percentile Cb and Cr values ​​in a forward-reshaped SDR color space corresponding to different parameterized HDR color spaces. [Figure 2H] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2I] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2J] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2K] 1 illustrates an example parameterized color space in the R.2020 color space. [Figure 2L] 1 shows an exemplary color chart image. [Figure 2M] 1 shows an exemplary checkerboard image. [Figure 2N] 1 shows an exemplary checkerboard image. [Figure 2O] 1 shows an exemplary original color chart image displayed on a reference image display, an exemplary captured image from the color chart image, and a modified captured image generated from the captured image. [Figure 2P]1 shows an exemplary distribution of RGB values. [Figure 2Q] 1 shows an exemplary alpha shape. [Figure 2R] 1 shows an exemplary alpha shape. [Figure 2S] 1 shows an exemplary alpha shape. [Figure 2T] 1 shows an exemplary alpha shape. [Figure 3A] 1 illustrates an exemplary process flow for finding a color space optimized for representing HDR colors. [Figure 3B] 1 illustrates an exemplary process flow for determining optimized values ​​for programmable parameters in an ISP pipeline. [Figure 3C] 1 illustrates an exemplary process flow for finding a color space optimized for representing HDR colors. [Figure 3D] 1 illustrates an exemplary process flow for generating optimized reshaping operating parameters. [Figure 3E] 1 illustrates an exemplary process flow for generating multiple different color charts. [Figure 3F] 1 illustrates an exemplary process flow for matching SDR and HDR colors between an image pair of captured SDR and HDR images. [Figure 3G] 1 illustrates exemplary encoder-side and decoder-side image / video editing operations. [Figure 3H] 1 illustrates exemplary encoder-side and decoder-side image / video editing operations. [Figure 3I] 1 illustrates an exemplary solution for clipping out-of-range codewords in image / video editing applications. [Figure 3J] 1 illustrates an exemplary solution for clipping out-of-range codewords in image / video editing applications. [Figure 3K] 1 illustrates an exemplary solution for clipping out-of-range codewords in image / video editing applications. [Figure 3L]1 illustrates an exemplary solution for clipping out-of-range codewords in image / video editing applications. [Figure 3M] 1 illustrates an exemplary solution for clipping out-of-range codewords in image / video editing applications. [Figure 3N] 10 illustrates an exemplary process flow for constructing a boundary clipping 3D-LUT. [Figure 4A] 1 illustrates an exemplary process flow. [Figure 4B] 1 illustrates an exemplary process flow. [Figure 4C] 1 illustrates an exemplary process flow. [Figure 4D] 1 illustrates an exemplary process flow. [Figure 4E] 1 illustrates an exemplary process flow. [Figure 5] FIG. 1 illustrates a simplified block diagram of an exemplary hardware platform on which a computer or computing device described herein may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, obscuring, or confusing the present disclosure.

[0012] [overview] Described herein are image optimizations, such as those associated with tensor-product B-spline (TPB)-based image reshaping solutions, for video capture applications, including but not limited to mobile (video) capture applications. TPB-based solutions can provide or generate relatively accurate reconstructed images using robust predictors implemented in video codecs, such as backward-compatible video codecs. These predictors can operate with static mappings to gracefully handle or recover from transmission errors, such as frame dropping, minimize power consumption, such as battery power consumption, reduce data overhead and / or processing / transmission latency, etc. In some operating scenarios, some or all of the TPB-based solutions described herein can be implemented in mobile applications that must comply with relatively stringent battery or power constraints.

[0013] The TPB-based image reshaping solution can be adapted to work with different use cases. Some of the use cases allow flexibility to design and / or apply HDR-SDR mapping to generate an SDR image to be encoded into the (base layer of) a video signal as described herein. Some of the use cases rely on a mobile image signal processor (ISP) with relatively limited programmable registers to generate an SDR image to be encoded into the (base layer of) a video signal. Different architectures may be used to implement the TPB-based solution described herein. Additionally, optionally or alternatively, these architectures or TPB-based solutions can be used to provide or support backward compatibility.

[0014] Given a reference SDR image represented in an SDR color space or a color space within the SDR domain, a TPB-based solution can use a TPB optimization process to achieve or determine the largest or widest HDR color space or color space within the HDR domain to represent the corresponding reconstructed HDR image predicted from the SDR image. The wider the HDR color space, the more color deviations will be introduced into the SDR color space. The TPB optimization process can be performed to achieve or determine an optimized balance point or trade-off between achieving the largest or widest HDR color space, on the one hand, and introducing SDR color deviations, on the other hand. Additionally, optionally or alternatively, the TPB optimization process described herein may be implemented to incorporate neutral color processing or perform constrained optimization to help preserve gray levels represented in the SDR image in the reconstructed HDR image predicted from the SDR image.

[0015] As more and more video is captured in practice by various camera-equipped computing devices or mobile devices, video editing is also becoming increasingly popular. Users operating these devices may be enabled to manually adjust the appearance of the captured video or simply apply a default theme template to the captured video. Depending on the available computational resources and / or the preferences of the video editing tool designer / provider, video editing may be performed either in a source domain, such as a source HDR domain in which the source image is represented, or in a non-source domain, such as a generated SDR domain in which pre-edited images may be converted from the source image in the source domain.

[0016] In an operating scenario where an SDR image is encoded into a video signal, video editing operations in the HDR domain or source domain do not adversely affect devices downstream of the video signal to reconstruct an HDR image from an SDR image decoded from the video signal at the decoder side, because these video editing operations do not interfere with the HDR-SDR mapping at the encoder side to generate an SDR image from an HDR image or source image in the HDR domain, and do not interfere with the SDR-HDR mapping at the decoder side to map back to the HDR image or reconstruct an approximation thereof.

[0017] However, in these operating scenarios, video editing operations in the SDR domain are likely to adversely affect restorability between the SDR and HDR domains. For example, the edited SDR image may not be able to be mapped back to the original or source HDR image. Furthermore, the color space representing the SDR image may be limited by a predefined range, such as the SMPTE range, that pixel values ​​or codewords cannot exceed. Color space conversions, such as YUV-RGB conversion, used in video editing operations may cause pixel values ​​or codewords to be clipped, thus impairing or harming reversibility. The clipping techniques described herein can be used to reduce or prevent the adverse effects of edits on video editing operations, including, but not limited to, those performed on mobile devices. Two-level TPB boundary clipping can be implemented in the RGB domain to support or preserve the maximum edited colors in the SDR domain that can be propagated to or mapped back to the reconstructed image in the HDR domain.

[0018] Exemplary embodiments described herein relate to image generation. Sampled HDR color space points distributed across an HDR color space are constructed. The HDR color space is parameterized by primary color scaling parameters having candidate values ​​selected from among a plurality of candidate values. The primary color scaling parameters are used to calculate color space coordinates of at least one of a plurality of primary colors describing the HDR color space. Reference SDR color space points represented in a reference SDR color space, input HDR color space points represented in an input HDR color space, and reference HDR color space points represented in a reference HDR color space are generated from the sampled HDR color space points in the HDR color space. A reshaping operation optimization algorithm is performed to generate a chain of optimized forward reshaping mappings and optimized reverse reshaping mappings. The reshaping operation optimization algorithm uses the reference SDR color space points, the input HDR color space points, and the reference HDR color space points as inputs. The optimized forward reshaping mapping is used to forward reshape an input HDR image in an input HDR color space into a forward reshaped SDR image in a forward reshaped SDR color space, while the optimized backward reshaping mapping is used to backward reshape a forward reshaped SDR image in the forward reshaped SDR color space into a backward reshaped HDR image.

[0019] Exemplary embodiments described herein relate to image generation. Sampled HDR color space points distributed across an HDR color space are constructed. The HDR color space is parameterized by primary color scaling parameters having candidate values ​​selected from among a plurality of candidate values. The primary color scaling parameters are used to calculate color space coordinates of at least one of a plurality of primary colors describing the HDR color space. Input SDR color space points represented in an input SDR color space and reference HDR color space points represented in a reference HDR color space are generated from the sampled HDR color space points in the HDR color space. A reshaping operation optimization algorithm is performed to generate an optimized inverse reshaping mapping. The reshaping operation optimization algorithm receives the input SDR color space points and the reference HDR color space points as inputs. The inverse reshaping mapping is used to inversely reshape an SDR image in the input SDR color space into an inversely reshaped HDR image.

[0020] The exemplary embodiments described herein relate to image generation. A set of SDR image feature points is extracted from a training SDR image, and a set of HDR image feature points is extracted from a training HDR image. A subset of one or more SDR image feature points in the set of SDR image feature points is matched with a subset of one or more HDR image feature points in the set of HDR image feature points. The subset of one or more SDR image feature points and the subset of one or more HDR image feature points are used to generate a geometric transformation for spatially aligning a set of SDR pixels in the training SDR image with a set of HDR pixels in the training HDR image. A set of SDR color patch and HDR color patch pairs is determined from the set of SDR pixels in the training SDR image and the set of HDR pixels in the training HDR image after the training SDR image and the training HDR image are spatially aligned by the geometric transformation. An optimized SDR-HDR mapping is generated based at least in part on the set of SDR color patch and HDR color patch pairs derived from the training SDR image and the training HDR image. The optimized SDR-HDR mapping is applied to one or more non-training SDR images to generate one or more corresponding non-training HDR images.

[0021] The exemplary embodiments described herein relate to image generation. A respective camera distortion correction operation is performed on each training image in a pair of training standard dynamic range (SDR) and high dynamic range (HDR) images to generate a respective undistorted image in a pair of undistorted training SDR and HDR images. Each projective transformation in the pair of SDR and HDR image projective transformations is generated using a corner pattern mark detected from each undistorted image in the pair of undistorted training SDR and HDR images. Each projective transformation in the pair of SDR and HDR image projective transformations is applied to each undistorted image in the pair of undistorted training SDR and HDR images to generate a respective rectified image in the pair of rectified training SDR and HDR images. A set of SDR color patches is extracted from the rectified training SDR image, and a set of HDR color patches is extracted from the rectified training HDR image. The optimized SDR-HDR mapping is generated based at least in part on the set of SDR color patches and the set of HDR color patches derived from the training SDR image and the training HDR image, and the optimized SDR-HDR mapping is applied to one or more non-training SDR images to generate one or more corresponding non-training HDR images.

[0022] Exemplary embodiments described herein relate to a clipping operation on an edited image. Sampled HDR color space points distributed across an HDR color space used to represent a reconstructed HDR image are constructed. The sampled HDR color space points are converted to SDR color space points in a first SDR color space representing an SDR image to be edited by an editing device. A bounding SDR color space rectangle is determined based on extreme SDR codeword values ​​of the SDR color space points in the first SDR color space. An irregular three-dimensional (3D) shape is determined from the distribution of the SDR color space points. Sampled SDR color space points distributed across the bounding SDR color space rectangle in the first SDR color space are constructed. The sampled SDR color space points and the irregular shape are used to generate a boundary clipping 3D lookup table (3D-LUT). The boundary clipping 3D-LUT uses the sampled SDR color space points as lookup keys. A clipping operation is performed on the edited SDR image in the first SDR color space based at least in part on the boundary clipping 3D-LUT to generate a boundary-clipped edited SDR image in the first SDR color space.

[0023] [Reshaping optimization in image / video capture applications] A reshaping optimization process, such as a TPB optimization process and / or a non-TPB optimization process, can be implemented in or incorporated into a video capture application executing on a computing device, such as a mobile device, in various operating scenarios. The video capture application having the reshaping optimization process can be used to generate or output a video signal, such as a base layer or SDR image, encoded therein. The reshaping optimization process can implement different solutions in different operating scenarios to generate or optimize reshaping operating parameters, such as TPB coefficients and / or non-TPB coefficients, used with a base layer or SDR image to generate, construct, or reconstruct a non-base layer or HDR image having optimized image quality.

[0024] For illustrative purposes, reshaping optimization processes, or solutions implemented therein, may be classified into different types based on the particular SDR generation process or sub-process employed by the video capture application, and based on the particular reshaping path that undergoes reshaping optimization, i.e., the forward (reshaping) path and / or the reverse (reshaping) path.

[0025] 1A shows a first exemplary reshaping optimization process that implements a white-box joint forward and reverse TPB optimization design / solution (referred to as "WFB") using a (single) video capture device. The white-box joint forward and reverse TPB optimization design / solution may be implemented for an operating scenario in which an SDR generation process or sub-process converts a reference HDR image to a reference SDR image using a white-box transformation operation, and both the forward and reverse (reshaping) paths can undergo TPB optimization.

[0026] As used herein, a "white-box transform" refers to an HDR-to-SDR mapping or conversion that uses a clearly defined transform or mapping function / formula, such as one specified or documented (e.g., publicly, etc.) in a standards-based or proprietary video coding specification. In comparison, a "black-box transform" refers to an HDR-to-SDR mapping or conversion that involves a transform or mapping operation that is not based on a clearly defined transform or mapping function / formula. For example, a black-box transform may be implemented as internal image signal processing performed by an image signal processor with little or no reliance on any clearly defined transform or mapping function or formula specified or documented (e.g., publicly, etc.) in a standards-based or proprietary video coding specification.

[0027] 1A, an input video signal including an HDR image may be represented in an input color space, such as the full R.2020 color space. A generated SDR signal including a reference SDR image may be represented in a reference SDR color space. The reference SDR image may be an image generated from the HDR image using known HDR-to-SDR mapping processes, functions, and / or formulas.

[0028] 1A, in the forward reshaping path, a reference HDR image may be forward reshaped into a (forward) reshaped SDR image that approximates the reference SDR image based at least in part on optimized forward reshaping operation parameters (denoted as "forward TPB optimization") generated from the joint forward and backward TPB optimization solution. The reshaped SDR image may be represented in a reshaped SDR color space and generated by an upstream device to be included in / encoded in a video signal output by the upstream device or its base layer (BL).

[0029] A downstream receiving device or video decoder of the video signal can decode the reshaped SDR image from the video signal. The decoded reshaped SDR image at the decoder side may be the same as the reshaped SDR image at the encoder side, subject to errors introduced in the compression / decompression, coding operations and / or data transmission.

[0030] In the inverse reshaping path as implemented by a downstream device, the reshaped SDR image may be inversely reshaped into a (inverse) reshaped HDR image based at least in part on optimized inverse reshaping operation parameters (denoted as "inverse TPB optimization") generated from the same joint forward and inverse TPB optimization solution. The reshaped HDR image generated using the optimized inverse reshaping operation parameters in the output HDR color space represents an approximation or reconstructed version of the reference HDR image.

[0031] The joint forward and backward TPB optimization solution can generate optimized forward and backward reshaping operation parameters to cover as wide as possible in the output HDR color space in which the reshaped or reconstructed HDR image generated at the decoder side is represented.

[0032] One or both of the optimized forward and backward reshaping operation parameters, such as the optimized forward and backward TPB coefficients, can be implemented or represented in a three-dimensional lookup table or 3D-LUT to reduce processing time in the reshaping operation.

[0033] Because forward and backward reshaping operation parameters, such as forward and backward TPB coefficients, are jointly or simultaneously designed or optimized in a joint forward and backward TPB optimization process, the supported reshaped SDR and HDR color spaces used to represent the reshaped SDR and HDR images can be jointly or simultaneously designed or optimized in the same process.

[0034] In some operating scenarios, the optimized forward and backward reshaping operating parameters generated from the joint forward and backward TPB optimization process may be applied in a static single-layer backward compatible (SLBC) framework, where it is not necessary to obtain (e.g., dynamically, etc.) optimized image-specific or image-dependent (or content-dependent) forward and backward operating parameters, such as image-specific or image-dependent (or content-dependent) forward and backward TPB coefficients, on the fly while images are being processed.

[0035] Rather, under a static SLBC framework, the same or static optimized forward and backward operation parameters, such as the same optimized forward and backward TPB coefficients, can be obtained or generated at once, for example, offline or before performing reshaping operations on either the (input) reference HDR image or the reshaped SDR image, for forward reshaping all (input) reference HDR images and backward reshaping all reshaped SDR images. In one example, a single set of static optimized forward and backward operation parameters can be generated offline by a system described herein and configured / arranged in or used by a capture device described herein to perform an image reshaping or reconstruction operation. In another example, multiple sets of static optimized forward and backward operation parameters can be generated offline by a system described herein and configured / arranged in or used by a capture device described herein to select a particular set of static optimized forward and backward operation parameters to perform an image reshaping or reconstruction operation.

[0036] The optimized static forward TPB coefficients can then be applied by an upstream device to forward reshape some or all of the (input) reference HDR images (e.g., a sequence of consecutive or sequential (input) reference HDR images, etc.) to generate a reshaped SDR image to be encoded into an (SLBC) video signal, while the optimized static backward TPB coefficients can be applied by a downstream receiving device of the video signal to some or all of the reshaped SDR images (e.g., a sequence of consecutive or sequential reshaped SDR images, etc.) decoded from the video signal to generate or reconstruct a reshaped HDR image.

[0037] 1B shows a second exemplary reshaping optimization process using a (single) video capture device to implement a white-box reverse-only TPB optimization design / solution (referred to as "WB"). The white-box reverse-only TPB optimization design / solution may be implemented for operating scenarios in which an SDR generation process or sub-process converts a reference HDR image to a reference SDR image using a white-box transformation operation, and only the reverse (reshaping) path undergoes TPB optimization.

[0038] 1B, the input video signal including the HDR image may be represented in an input color space, such as the P3 color space. The generated SDR signal including the reference SDR image may be represented in a reference SDR color space. The reference SDR image may be an image generated from the HDR image using a known HDR-to-SDR mapping process, function, and / or formula.

[0039] 1B, the reference HDR image may be processed by a programmable image signal processor or ISP (e.g., using a given image signal processing function / formula / operation, etc.) into an ISP SDR image that approximates the reference SDR image based at least in part on optimized ISP operating parameters (denoted as "ISP parameter optimization") generated from the ISP optimization solution. The ISP SDR image may be expressed in an ISP SDR color space and generated by an upstream device to be included in / encoded in a video signal (or its base layer (BL)) output by the upstream device.

[0040] A downstream receiving device or video decoder of the video signal can decode the ISP SDR image from the video signal. The decoded ISP SDR image at the decoder side may be the same as the ISP SDR image at the encoder side, subject to errors introduced in compression / decompression, coding operations and / or data transmission.

[0041] In the inverse reshaping path as implemented by a downstream device, the ISP SDR image may be inversely reshaped into a (inverse) reshaped HDR image based at least in part on optimized inverse reshaping operation parameters (denoted as "inverse TPB optimization") generated from the inverse-only TPB optimization solution. The reshaped HDR image generated using the optimized inverse reshaping operation parameters in the output HDR color space represents an approximation or reconstructed version of the reference HDR image.

[0042] The reverse-only TPB optimization solution can generate optimized reverse reshaping operation parameters to cover as wide as possible in the output HDR color space in which the reshaped or reconstructed HDR image generated at the decoder side is represented.

[0043] The optimized backward reshaping operation parameters, such as the optimized backward TPB coefficients, can be implemented or represented in a three-dimensional look-up table or 3D-LUT to reduce the processing time in the reshaping operation.

[0044] In some operating scenarios, the inverse reshaping operating parameters generated from the inverse-only TPB optimization process may be applied in a static single-layer inverse display mapping (SLiDM) framework, where optimized image-specific or image-dependent (or content-dependent) inverse operating parameters, such as image-specific or image-dependent (or content-dependent) inverse TPB coefficients, do not need to be obtained (e.g., dynamically).

[0045] Rather, under the static SLiDM framework, the same or static optimized inverse operation parameters, such as the same optimized inverse TPB coefficients, can be obtained or generated at one time, e.g., offline or before performing a reshaping operation on any of the reshaped SDR images, to inversely reshape an ISP SDR image. In one example, a single set of static optimized inverse operation parameters can be generated offline by a system described herein and configured / arranged in or used by a capture device described herein to perform an image reshaping or reconstruction operation. In another example, multiple sets of static optimized inverse operation parameters can be generated offline by a system described herein and configured / arranged in or used by a capture device described herein to select a particular set of static optimized inverse operation parameters to perform an image reshaping or reconstruction operation.

[0046] The optimized static backward TPB coefficients can then be applied by a downstream receiving device of the video signal to some or all of the ISP SDR images (e.g., a sequence of consecutive or sequential ISP SDR images) decoded from the video signal to generate or reconstruct a reshaped HDR image.

[0047] The ISP SDR image encoded in the video signal may or may not be identical to the desired SDR appearance as represented by the reference image. The operating parameters or their settings in the programmable ISP can be optimized to approximate the reference image. Thus, in the "WB" operating scenario of FIG. 1B, the SDR image for backward reshaping to the reshaped HDR image is provided as the ISP SDR image or is fixed. By comparison, in the WFB operating scenario of FIG. 1A, the SDR image for backward reshaping to the reshaped HDR image is a forward-reshaped SDR image, which can be optimized by applying optimized forward reshaping operating parameters generated in a joint forward and backward TPB optimization process that generates optimized forward and backward reshaping parameters. In other words, TPB optimization is not used in the forward path to generate the ISP SDR image with the aim of improving or enhancing both the desired appearance of the SDR image encoded in the video signal and the desired appearance of the reshaped HDR image. In these operating scenarios, relatively large deviations from the desired appearance can occur.

[0048] 1C illustrates a third exemplary reshaping optimization process using a (single) video capture device to implement a black-box reverse-only TPB optimization design / solution (referred to as "BB1"). The black-box reverse-only TPB optimization design / solution may be implemented for operating scenarios in which an SDR image is not generated from an HDR image using a white-box conversion operation, and only the reverse (reshaping) path undergoes TPB optimization.

[0049] The reshaping operating parameters for these operating scenarios can be generated using training SDR images and training HDR images generated or acquired by the same video capture device, such as the same mobile device, forming multiple SDR and HDR image pairs, each of which includes a training SDR image and a training HDR image corresponding to the training SDR image.

[0050] The training SDR image and the training HDR image in the same SDR and HDR image pair may be acquired at different time instances / points (e.g., a few milliseconds apart, a fraction of a second apart, a few seconds apart, etc.) using the same capture device with the same ISP. Because it is difficult, if not impossible, to maintain the same shooting position and process the training SDR image and the training HDR image in the same way at different time instances / points, the training SDR image and the training HDR image may not be precisely spatially or temporally aligned with each other. For example, the training SDR image may be locally tone mapped or locally enhanced, while the training HDR image may be generated or acquired from multiple camera exposures. Therefore, in the “BB1” operating scenario, the relationship between pixel or codeword values ​​in the training SDR image and the corresponding pixel or codeword values ​​in the training HDR image in the same image pair may be treated as or assumed to be a black box.

[0051] The training SDR image and the training HDR image in each image pair may first be spatially aligned. The spatially aligned training SDR image and the training HDR image in the image pair may be used to determine or find matching color pairs. These matched color pairs can then be used to generate optimized reshaping operation parameters, such as optimized TPB and / or non-TPB coefficients, for reshaping or mapping the (e.g., non-training, reference, etc.) SDR image to an inversely reshaped or reconstructed HDR image that approximates the (e.g., non-training, reference, etc.) HDR image.

[0052] 1C, an SDR image, such as a reference SDR image, may be generated by an upstream device, such as a capture device, to be included in / encoded in a video signal or its base layer (BL) output by the upstream device. Exemplary reference SDR images described herein may include, but are not limited to, an SDR image generated from a programmable image signal processor of the upstream device or a capture device operating in conjunction with the upstream device.

[0053] A downstream receiving device or video decoder of the video signal can decode a reference SDR image from the video signal. The decoded reference SDR image at the decoder side may be the same as the reference SDR image at the encoder side, and may be subject to errors introduced in compression / decompression, coding operations and / or data transmission.

[0054] In the inverse reshaping path as implemented by a downstream device, the reference SDR image may be inversely reshaped into a (inverse) reshaped HDR image based at least in part on optimized inverse reshaping operation parameters (denoted as "inverse TPB optimization") generated from a black-box inverse-only TPB optimization design / solution. The reshaped HDR image generated using the optimized inverse reshaping operation parameters in the output HDR color space represents an approximation or reconstructed version of the reference HDR image that can or may be generated by the same capture device.

[0055] The black-box backward-only TPB optimization design / solution can generate optimized backward reshaping operation parameters to cover as wide as possible in the output HDR color space in which the reshaped or reconstructed HDR image generated at the decoder side is represented.

[0056] The optimized backward reshaping operation parameters, such as the optimized backward TPB coefficients, can be implemented or represented in a three-dimensional look-up table or 3D-LUT to reduce the processing time in the reshaping operation.

[0057] In some operating scenarios, backward reshaping operating parameters generated from a black-box backward-only TPB optimization design / solution may be applied in the SLiDM framework, where there is no need to (e.g., dynamically) obtain optimized image-specific or image-dependent (or content-dependent) backward operating parameters, such as image-specific or image-dependent (or content-dependent) backward TPB coefficients.

[0058] Rather, under the static SLiDM framework, the same or static optimized inverse operation parameters, such as the same optimized inverse TPB coefficients, can be obtained or generated at once, e.g., offline or before performing the reshaping operation on any of the reshaped SDR images, to inversely reshape a spatially aligned training SDR image generated by an upstream capture device into a reshaped or reconstructed HDR image that approximates the spatially aligned training HDR image generated by the same capture device.

[0059] The optimized static backward TPB coefficients can then be applied by a downstream receiving device of the video signal to some or all of the (e.g., non-training, etc.) reference SDR images decoded from the video signal (e.g., a sequence of consecutive or sequential reference SDR images, etc.) to generate or reconstruct a reshaped HDR image that approximates a reference HDR image that can be generated or may be generated from the same upstream capture device that generates or captures the reference SDR image.

[0060] Because TPB optimization is used only in the reverse path, the resulting appearance of the reshaped HDR image generated from reverse reshaping the reference SDR image may have a relatively large deviation from the desired appearance of the reference HDR image in the "BB1" operating scenario.

[0061] In some "BB1" operating scenarios, as shown in FIG. 1D , instead of a reference SDR image generated from a programmable image signal processor, a camera raw image generated from a camera ISP may be directly encoded into a video signal output by an upstream (capture) device. During the training phase, the training camera raw image, which may or may not be an SDR image, can be spatially and / or temporally aligned with the training HDR image in the same camera raw image and HDR image pair. Matching color pairs from the spatially aligned camera raw image and HDR image pair can be used to optimize inverse reshaping operation parameters, such as inverse TPB coefficients. During the non-training or deployment phase, the optimized inverse reshaping operation parameters can be used by a downstream device of the video signal encoded with the non-training camera raw image to inversely reshape a decoded non-training camera raw image from the video signal into an inversely reshaped or reconstructed (non-training) HDR image that approximates a reference HDR image that can or may be generated by the same upstream (capture) device.

[0062] 1E illustrates a fourth exemplary reshaping optimization process implementing a black-box reverse-only reshaping optimization design / solution (referred to as "BB2") using different video capture devices. The "BB2" reshaping optimization design / solution may be implemented for an operating scenario in which an SDR image is generated by a first capture device, an HDR image generated from the SDR image is generated (e.g., intended to be generated) by a second, different capture device, and only the reverse (reshaping) path undergoes reshaping optimization. Furthermore, the reshaping optimization in the "BB2" operating scenario may or may not be a TPB optimization.

[0063] The reshaping operating parameters in the "BB2" operating scenario can be generated using training SDR images and training HDR images generated or acquired by two video capture devices, such as two mobile devices of different makes and / or models, respectively. These training SDR images and training HDR images form a plurality of SDR and HDR image pairs, each of which includes a training SDR image and a training HDR image corresponding to the training SDR image.

[0064] The training SDR image and the training HDR image in the same SDR and HDR image pair may be acquired using the same image, such as the same checkerboard image rendered on the same or similar image display. In the "BB1" operating scenario, the relationship between pixel or codeword values ​​in the training SDR image and the corresponding pixel or codeword values ​​in the training HDR image in the same image pair may be treated as or assumed to be a black box.

[0065] The training SDR image and the training HDR image in each image pair may first be spatially aligned. The spatially aligned training SDR image and the training HDR image in the image pair may be used to determine or find matching color pairs rendered with a test image, such as a checkerboard image, on the same image display or the same image display type. These matching color patches can then be used to construct a three-dimensional look-up table (3D-LUT) for mapping SDR pixel or codeword values ​​of color patches in the test image to corresponding HDR pixel or codeword values ​​of the same color patches.

[0066] A 3D-LUT derived partially or entirely based on the test image can be used to generate optimized static or dynamic reshaping operating parameters for reshaping or mapping a (e.g., non-training, reference, etc.) SDR image to an inversely reshaped or reconstructed HDR image that approximates the (e.g., non-training, reference, etc.) HDR image.

[0067] 1E, an SDR image, such as a reference SDR image, may be generated by an upstream device, such as a first capture device, to be included in / encoded in a video signal or its base layer (BL) output by the upstream device. Exemplary reference SDR images described herein may include, but are not limited to, an SDR image generated from a programmable image signal processor of the first capture device or an upstream device.

[0068] A downstream receiving device or video decoder of the video signal can decode from the video signal the reference SDR image captured using the first capture device. The decoded reference SDR image at the decoder side may be the same as the reference SDR image at the encoder side, and may be subject to errors introduced in compression, decompression, coding operations and / or data transmission.

[0069] In the inverse reshaping path as implemented by a downstream device, the reference SDR image may be inversely reshaped into a (inversely) reshaped HDR image based at least in part on optimized inverse reshaping operation parameters (denoted as "dynamic inverse function optimization") generated from the "BB2" reshaping optimization design / solution. The reshaped HDR image generated using the optimized inverse reshaping operation parameters in the output HDR color space represents an approximation or reconstructed version of the reference HDR image that can or may be generated by a second, different capture device.

[0070] The "BB2" reshaping design / solution can generate optimized inverse reshaping operation parameters to cover as wide as possible the output HDR color space in which the reshaped or reconstructed HDR image generated at the decoder side is represented.

[0071] In some operating scenarios, the reverse reshaping operating parameters generated from the "BB2" reshaping optimization design / solution may be applied in a static or dynamic SLiDM framework.

[0072] Because "BB2" reshaping is used only in the reverse path, the resulting appearance of a reshaped HDR image generated from reverse reshaping a reference SDR image may have a relatively large deviation from the desired appearance of the reference HDR image. The first and second capture devices may be mobile devices that can operate differently in their respective video capture applications, for example, using different exposure times, different video processing operations, different mapping functions / relationships, etc.

[0073] In particular, in operational scenarios where first and second capture devices operate with relatively large differences between the respective video / image capture applications running on the first and second capture devices and / or between the respective ISPs used by the first and second capture devices, a dynamic SLiDM frame may be used in which dynamic mapping using image-dependent or image-specific reshaping operation parameters is used to inversely reshape the SDR image of the first capture device into an HDR image that can or may be generated by the second capture device. For example, during a training phase, multiple sets of training SDR images and training HDR images or image pairs can be used to derive multiple sets of reshaping operation parameters optimized for each of the multiple sets of training SDR images and training HDR images or image pairs. During a deployment or application phase, specific image characteristics, such as the overall brightness of some or all regions of a specific image, can be dynamically evaluated as the specific image is being processed. A particular set of optimized reshaping operation parameters can be adaptively and / or dynamically selected for a particular image from among a plurality of sets of optimized reshaping operation parameters based on particular image characteristics of the particular image in relation to or in comparison with the respective image characteristics of a training SDR image and a training HDR image or a plurality of sets of image pairs.

[0074] [White-box TPB forward and backward joint optimization (WFB)] To help increase coverage of the supported color space in which the reshaped image is represented, several design factors for reshaping optimization may be considered. First, a tensor-product bi-spline (TPB) predictor with optimized TPB coefficients generated from the reshaping optimization may be used to obtain or achieve relatively higher prediction accuracy than other types of predictors, including, but not limited to, MMR predictors. The multi-knot and continuity properties of the tensor-product bi-spline predictor or prediction function can be utilized or used to cover a relatively wide color range or portion of the color space with relatively high accuracy. An exemplary TPB predictor is described in Guan-Ming Su, filed October 1, 2019, No. 62 / 908,770, entitled "Tensor-product B-spline predictor," by Harshad Kadu, Qing Song, and Neeraj J. Gadgil, the contents of which are fully incorporated herein by reference as if fully set forth herein.

[0075] In comparison, while a multi-part MMR predictor or prediction function may be better than a single-part MMR predictor or prediction function, discontinuities between different MMR parts are likely introduced by a multi-part MMR predictor or prediction function, thereby adversely affecting or even preventing the use of a multi-part MMR predictor or prediction function in many operational scenarios. An exemplary MMR-based operation is described in U.S. Patent No. 8,811,490, the entire contents of which are incorporated by reference as if fully set forth herein.

[0076] TPB predictors can be advantageously used in SLBC-based video codecs. More specifically, a forward TPB predictor can be used to generate a forward-reshaped SDR image that approximates a reference SDR image, while a backward TPB predictor can be used to generate a backward-reshaped or reconstructed HDR image that approximates a reference HDR image. An exemplary use of TPB predictors with SLBC-based video codecs can be found in U.S. Provisional Application No. 63 / 255,057, "Tensor-product B-spline prediction for HDR video in mobile applications," by H. Kadu et al., filed October 13, 2021, the contents of which are incorporated herein by reference in their entirety as if fully set forth herein. Furthermore, a BESA (Backward Error Subtraction for signal Adjustment) algorithm / method with modifications to support neutral color preservation can be used in a pipeline of chained reshaping functions to jointly optimize the forward and backward TPB predictors to achieve relatively high reversibility or relatively accurate approximation of the reference HDR image by the backward reshaping or reconstructed HDR image to the reference HDR image. Exemplary BESA algorithms / methods can be found in U.S. Provisional Patent Application No. 63 / 013,063, entitled "Reshaping functions for HDR imaging with continuity and reversibility constraints," filed April 21, 2020, by G.M. Su, U.S. Provisional Patent Application No. 63 / 013,807, entitled "Iterative optimization of reshaping functions in single-layer HDR image codec," filed April 22, 2020, by G.M. Su and H. Kadu, and PCT Application No. PCT / US2021 / 028475, filed April 21, 2021, the contents of which are hereby incorporated by reference in their entireties as if fully set forth herein.

[0077] Although the TPB predictor has better prediction accuracy than other types of predictors, it may incur a relatively large signal or bit rate overhead for transmitting the TPB coefficients within a video signal, and it may also incur a relatively large computational cost for constructing the TPB formulas or basis functions used in the TPB predictor.

[0078] In some operating scenarios, to avoid or reduce signal or bitrate overhead and computational costs, a built-in or static TPB predictor may be used to reshape some or all images of an entire video sequence without changing the TPB coefficients used in the static TPB predictor during playback of the video sequence. More specifically, some or all of the (built-in or static) TPB coefficients may be cached or stored in a video application, such as a video capture / editing application, without the need to explicitly transmit these TPB coefficients through or within an encoded video signal or bitstream along with the images to be reshaped by the built-in or static TPB predictor. In one example, some or all of the TPB coefficients may be pre-loaded or pre-configured in a downstream receiving device before the video signal or bitstream is received and processed by the downstream receiving device. Additionally, optionally or alternatively, multiple sets of TPB coefficients may be pre-loaded or pre-configured in a downstream receiving device before the video signal or bitstream is received and processed by the downstream receiving device. A video application may simply select or choose a particular set of TPB coefficients to use with a static TPB predictor from among multiple sets of TPB coefficients based on, for example, a simple indicator (e.g., a binary indicator, a multi-bit indicator, etc.) signaled or transmitted in the video signal or bitstream. Examples of static TPB predictors or prediction functions can be found in the above-referenced U.S. Provisional Patent Application No. 63 / 255,057.

[0079] In some operating scenarios, a mobile device may be used to host and execute video capture and / or editing applications. A user of the mobile device can edit captured images in these video capture and / or editing applications. Edited images, such as edited HDR images, may be intended to be displayed or viewed, for example, on an image display of a non-mobile device having a higher or larger display capability than the display capabilities of the image display of the mobile device, and may have a higher, larger, wider, and / or broader luminance range and / or color range or color distribution than the luminance range and / or color range or color distribution of an (original) captured image, such as an HDR image originally captured from a camera sensor of the mobile device. For example, a captured HDR image may often be limited by the mobile device's camera sensor and ISP output, such that it is represented in a relatively narrow color space or gamut, such as P3, while the edited HDR image may be represented in a relatively wide color space or gamut, such as the entire R.2020 color space.

[0080] To design a static mapping for the reshaping operation, it is highly desirable to optimize the forward / backward TPB coefficients to cover as high a dynamic range and as wide a color space or gamut as possible, while achieving as much bitrate and computational efficiency as possible.

[0081] Full HDR restoration between a reconstructed HDR image and a reference HDR image identical to the reconstructed HDR requires that colors in the original or reconstructed HDR domain (or color space) be distinguished or distinguishable in the reshaped SDR domain (or color space) so that distinguished colors in the SDR domain can be mapped back to distinguishable colors in the HDR domain. Thus, full restoration likely requires a one-to-one mapping from HDR to SDR and from SDR to HDR. The wider the width that an HDR domain or color space supports, the more codewords the SDR domain or color space needs to have. Given that the total number of available SDR pixel or codeword values ​​in the SDR domain (e.g., those corresponding to a bit depth lower than that of the HDR domain) is typically smaller than the total number of required HDR pixel or codeword values ​​in the HDR domain, full HDR restoration may not be possible in some operating scenarios. In fact, some captured or edited images may contain combinations of diffuse and / or specular colors that form color distributions such as those shown as irregular shapes in FIG. 2A that exceed the R.2020 color space (“BT.2020”) and therefore cannot be fully represented in smaller color spaces such as the R.709 color space (“BT.709”) or the (DCI) P3 color space or SDR color spaces.

[0082] In many operating scenarios, the nonlinear reshaping mapping or function can be used to distribute the available pixel or codeword values ​​relatively efficiently, generate a (forward) reshaped SDR that approximates the reference SDR image as closely as possible, and helps the reconstructed HDR image approximate the reference HDR image as closely as possible, even though the reshaped SDR generated using the nonlinear reshaping mapping or function may not be identical to the reference SDR image, but may instead contain some deviations from the reference SDR image.

[0083] Additionally, optionally or alternatively, in some operating scenarios, in order to maintain the same or similar appearance of the reference SDR and HDR images in the reshaped SDR and HDR images, the reshaping optimizations described herein can be performed with a neutral color constraint that preserves neutral colors in the color mapping performed in conjunction with the reshaping operation.

[0084] Given their ability to approximate functions with relatively high nonlinearity, TPB predictors or functions can be incorporated into reshaping operations to support or achieve a relatively wide HDR color space for representing inversely reshaped or reconstructed HDR images.

[0085] A potential drawback is that the TPB predictor may require a relatively large number of TPB coefficients to represent or approximate the nonlinear SDR-HDR and / or HDR-SDR mapping function. To prevent possible ill-defined conditions, numerical instability, slow convergence issues, etc. that may arise when solving the TPB coefficient optimization problem, a static TPB predictor may be pre-generated and deployed with a video codec, such as one used by an upstream device and / or a downstream device, before the video codec is used to process a video sequence, such as a captured and / or edited video sequence.

[0086] For illustrative purposes only, in an operating scenario such as that shown in FIG. 1A, the (forward reshaped) SDR domain or color space may be that of R.709, while the (original or inversely reshaped) HDR domain or color space may be that of R.2020. As shown in FIG. 2A, the R.2020 color space is much larger than the R.709 color space. Therefore, while it may be difficult to map the entire R.2020 to R.709 and then map the R.709 back to R.2020, the reshaping optimization techniques described herein can be used to optimize the coverage of the HDR color space for representing the inversely reshaped or reconstructed HDR image and to maintain the same or similar appearance of the reference SDR and HDR images in the reshaped SDR and HDR images. These techniques can be used to simultaneously optimize reshaping operation parameters, such as TPB coefficients, for both forward and inverse reshaping operations.

[0087] While it may not be possible to cover the entire R.2020 color space in all scenarios, a subset or subspace in the R.2020 color space, such as a particular color gamut depicted as a triangle formed by the primary colors in a color coordinate system, may be selected or chosen by the reshaping optimization operation to maintain HDR restorability in the reshaped HDR image and SDR image relative to the reference HDR image and an acceptable SDR approximation of the reference SDR image.

[0088] By way of example and not limitation, a subset or subspace in the R.2020 color space (or an HDR color space for representing a reshaped HDR image) may be defined or characterized by a specific white point and three specific primaries (red, green, and blue). The specific white point may be selected or fixed to be the D65 white point. There may be considerable design freedom for selecting three specific primaries from many possible primary combinations for a subset or subspace in R.2020 to be supported by a reshaping operation. The reshaping optimization techniques described herein can be used to determine or select the specific primaries for a subset or subspace in R.2020 to be optimized primaries to realize or achieve maximum perceptual quality and / or color coding efficiency and / or restorability between SDR and HDR images.

[0089] The MacAdam ellipse represents just perceptible color differences, in that the HVS may not be able to distinguish color differences within the same ellipse. The Pointer color gamut can represent all diffuse colors perceptible by the HVS. In the MacAdam ellipse and the Pointer color gamut, green is less important or less distinguishable / perceptible by the HVS than red and blue. Therefore, in some operating scenarios, especially when a one-to-one mapping relationship cannot be supported by the SDR-HDR mapping and HDR-SDR mapping in the reshaping operation, specific optimized primaries for a subset or subspace in R.2020 for representing a reshaped or reconstructed HDR image can be selected to cover more red and blue than green.

[0090] As used herein, primary colors may also be referred to as primary colors and may be used to define the corners of a polygon, such as a triangle, that represents a particular color space or gamut, such as a standard-specified color space, a color space supported by a display, a color space supported by a video signal, etc. For example, a standard color space with a standard-specified white point can be represented in the CIExy color space coordinate system or CIExy chromaticity diagram by a triangle whose corners are specified by the three standard-specified primary colors. The CIExy coordinates of the primary colors (red or R, green or G, blue or B) and white point that define the R.709, P3, and R.2020 color spaces, respectively, are specified in Table 1 below. [Table 1]

[0091] The CIExy coordinates of each primary color in the standard color space in Table 1 and the corresponding white point are (R x (c) ,R y (c) ), (G x (c) ,G y (c) ), (B x (c) )B y (c) ) and (W x (c) ,W y (c) ), where (c) is the standard color space.

[0092] Therefore, the P3 color space is (R x (P3) ),R y (P3) ), (G x (P3) ,G y (P3) ), (B x (P3) ,B y (P3) ), (W x (P3 ),W y (P3)) may be characterized by the CIExy coordinates of the P3 primaries and the P3 white point. Similarly, the R.2020 color space may be characterized by the CIExy coordinates of the P3 primaries and the P3 white point, such as (R x (R2020) ),R y (R2020) ), (G x (R2020) ,G y (R2020) ), (B x (R2020) ,B y (R2020) ), (W x (R2020) ,W y (R2020) ), it may be characterized by the CIExy coordinates of the R.2020 primaries and the R.2020 white point.

[0093] The inversely reshaped HDR color space for representing the inversely reshaped or reconstructed HDR image is shown as (a) color space. Therefore, the CIExy coordinates of the primary colors and white point that define the (a) color space are (R x (a) ,R y (a) ), (G x (a) ,G y (a) ), (B x (a) ,B y (a) ) and (W x (a) ,W y (a) )

[0094] As noted above, green is less important or distinguishable / perceptible by the HVS than red and blue. In some operating scenarios, the red and blue primaries of the (a) color space may be selected to match those of the R.2020 color space, as follows:

[0095] R x (a) =R x (R2020) (1-1) R y(a) =R y (R2020) (1-2) B x (a) =B x (R2020) (2-1) B y (a) =B y (R2020) (2-2)

[0096] Furthermore, like the R.709, P3 and R.2020 color spaces all being specified with a D65 white point, the (a) color space can also be specified with a D65 white point as follows: (W x (a) ,W y (a) )=(0.3127 0.3290) (3)

[0097] (a) The green primary color of the color space is the green primary color (G x (P3 ),G y (P3)) and the green primary (G x (R2020) ,G y (R2020) ) Any point along the line between two green primaries in the P3 and R.2020 color spaces can be expressed as a linear combination of these two green primaries with a weighting factor a as follows: G x (a) =aG x (P3) +(1-a)G x (R2020) (4-1) G y (a) =aG y (P3) +(1-a)G y (R2020) (4-2)

[0098] Therefore, the optimization problem of finding maximum support from the (a) color space for the R.2020 color space can be simplified to the problem of selecting a weighting factor a. When a = 0, the (a) color space becomes the entire R.2020 color space. When a = 1, as shown in Figure 2B (where the (a) color space is denoted as "TPB" or "Current TPB Covered Colors"), the red and blue corners or red and blue primaries of the (a) color space are the same as the red and blue corners or red and blue primaries of the R.2020 color space, while the green corner or green primaries of the (a) color space are the same as the green corner or green primaries in the P3 color space. Figures 2C-2E show three exemplary (a) color spaces for a = 0.9, 0.5, and 0.25, respectively.

[0099] [Joint color space and TPB optimization] Because the range or coverage by the inversely reshaped (a) color space on the R.2020 color space can be controlled by the parameter a, the overall TPB (based reshaping) optimization problem becomes how to optimize the TPB coefficients in both the forward and backward paths so that (i) the forward-reshaped SDR domain (or color space) for representing the reshaped SDR image is close to the reference SDR domain (or color space) for representing the reference SDR image, especially in the neural color or color space portions that are more sensitive to HVS than the non-neural color or color space portions, and (ii) the inverse-reshaped or reconstructed HDR domain (or color space) for representing the inverse-reshaped or reconstructed HDR image is as closely identical as possible to the reference HDR domain (or color space) for representing the reference HDR image. Ideally, if perfect reconstruction or perfect restorability can be achieved, the inverse-reshaped or reconstructed HDR image is identical to the reference HDR image.

[0100] As shown in Figures 2B-2E, in order to cover as much of the R.2020 color space as possible with the inversely reshaped or reconstructed HDR domain or color space, the TPB optimization problem reduces to finding or searching for the smallest possible value for the parameter a such that the above SDR and HDR quality conditions are met.

[0101] 3A shows an exemplary process flow for finding the smallest possible value for parameter a with the goal of maximizing SDR and HDR quality in reshaped SDR and HDR images. The process flow of FIG. 3A may iterate through multiple candidate values ​​for parameter a in an iterative order, which may be sequential or non-sequential.

[0102] Block 302 includes selecting the current (e.g., iterated, etc.) value of parameter a as the next candidate value (e.g., initially the first candidate value, etc.) among multiple candidate values ​​of parameter a. Given the candidate inversely reshaped HDR color space with the current value of parameter a, optimized reshaping operation parameters, such as optimized TPB coefficients, can be obtained in one or more subsequent process flow blocks of FIG. 3A.

[0103] Block 304 includes constructing sample points or preparing two sampled data sets in a candidate inversely reshaped HDR color space. By way of example and not limitation, the candidate inversely reshaped HDR color space may be a Hybrid Log Gamma (HLG) RGB color space (referred to as "(a) RGB color space HLG").

[0104] The first of the two sampled data sets is a uniformly sampled data set of color patches. Each color patch in the uniformly sampled data set of color patches is a uniformly sampled data set from the RGB color space HLG, which includes three dimensions for each RGB color (denoted as R, G, and B axes, respectively).

number

[0105] RGB color in a uniformly sampled dataset for each uniformly sampled data point or color patch

number

number

[0106] For simplicity, (i,j,k) in equation (5) above can be vectorized or simply denoted as p. Correspondingly, the uniformly sampled data points or RGB colors

number

number

number

[0107] The second of the two sampled data sets prepared or constructed in block 304 is a neutral color data set, which includes a plurality of neutral colors or neutral color patches (also called gray colors or gray color patches).

[0108] The second data set may be used to preserve input gray color patches in the input domain as output gray color patches in the output domain when the input gray color patches in the input domain are mapped or reshaped to output gray color patches in the reshaping operations described herein. The input gray color patches in the input domain (or input color space) may be given increased weighting in the optimization problem compared to other color patches to reduce the likelihood that these input gray color patches will be mapped to non-gray color patches in the output domain (or output color space) by the reshaping operations.

[0109] The second data set (gray color data set or gray color data set) may be prepared or constructed by uniformly sampling R, G, B values ​​along a line connecting between the first gray color (0,0,0) and the second gray color (1,1,1) in the RGB domain (e.g., (a) RGB color space HLG, etc.), as follows: n This results in nodes or gray color patches.

number

[0110] All N in the second data set n The nodes can be grouped or aggregated into a neutral color vector / matrix as follows:

number

[0111] The neutral color vector / matrix in equation (8) above is expressed as follows: t Repeat (a positive integer greater than or equal to 1) times to find N in the second data set (which is repeated here). n N t It can generate neutral color patches.

number

[0112] The repetition of neutral colors in the second data set increases the weighting of neutral or gray colors relative to other colors, and therefore neutral colors may be more preserved in the optimization problem than other colors.

[0113] The first data set of (all sampled) colors and the second (repeated) data set of neutral colors in equations (6) and (9) can be aggregated or placed together into a single combined vector / matrix as follows:

number

[0114] Combined Vector / Matrix

number

number

number

[0115] Block 306 is a combination vector / matrix in (a) RGB color space (or (b) RGB color space HLG).

number

[0116] The three red, green and blue primary colors in (a)RGB color space HLG, which correspond to the three corners or points of the triangle defining (a)RGB color space HLG in the CIExy chromaticity diagram, can be converted from CIE xy values ​​to CIE XYZ values ​​via the following formulas:

number

[0117] (R x (a) ,R y (a) ), (G x (a) ,G y (a) ), (B x (a) ,B y (a) ) as red, green, and blue primaries in (x,y) or CIE xy values, respectively. X (a) ,R Y (a) ,R Z (a) ), (G X (a) ,G Y (a) ,G Z (a) ), (B X (a) ,B Y (a) ,B Z (a) The same red, green and blue primaries in (X,Y,Z) or CIE XYZ values, denoted as (W), can be obtained using the conversion formula in equation (12) above. x (a) ,W y (a) ) as the white point in (x,y) or CIE xy values, then (W X (a) ,W Y (a) ,W Z (a) The same white point in (X,Y,Z) or CIE XYZ values, denoted as (X,Y,Z), can be obtained using the same transformation formula in equation (12) above.

[0118] P (a)→XYZ The 3x3 transformation matrix, denoted as

number

number

number

number

number

[0119] P (a)→R2020 The overall 3x3 transformation matrix, denoted as

number

number

[0120] Therefore, in block 306, a combination vector / matrix in RGB color space (or (a) RGB color space HLG)

number

number

[0121] Block 308 calculates the vector / matrix in the R.2020 RGB color space HLG as follows:

number

number

number

[0122]

number

number

[0123] Block 310 calculates the vector / matrix in the R.2020 RGB color space HLG as follows:

number

number

number

number

number

number

number

number

number

[0124] Block 312 calculates the vector / matrix in the R.2020 RGB color space HLG as follows:

number

number

number

number

number

number

number

[0125]

number

number

[0126] Block 314 uses the forward and reverse reshaping parameters, such as forward and reverse TPB coefficients, as inputs to the (TPB) BESA algorithm to obtain or generate the forward and reverse reshaping parameters.

number

[0127] The BESA algorithm is an iterative algorithm in which each current iteration may modify the reference SDR signal according to the backward prediction error measured or determined in the previous iteration.

[0128] The modified BESA algorithm described herein may implement neutral color preservation and avoid modifying neutral color patches. To do so, a set of neutral color patch indices that identify or correspond to neutral color patches may be generated as follows:

number

[0129] Let ch denote the channels in the forward reshaped SDR domain or color space (e.g., three channels, etc.). Let F denote forward reshaping or the forward path. Let B denote reverse reshaping or the reverse path.

[0130] S F ch The forward generator matrix (also called design matrix) for each channel, denoted as

number

number

number

number

number

[0131] The forward-reshaped SDR color patches may be predicted by multiplying the forward generator matrix in equation (28) with the forward TPB coefficients. Exemplary predicted reshaped codewords (equivalent or similar to color patches herein) with TPB coefficients are described in the above-mentioned U.S. Provisional Patent Application No. 62 / 908,770.

[0132] Forward reshaped SDR color patches or codewords can be inversely reshaped into inversely reshaped HDR color patches or codewords, for example, through a TPB-based inverse reshaping operation.

[0133] In the backward pass, at each iteration in the BESA algorithm (e.g., the kth iteration, etc.),

number

number

number

number

[0134] The inverse reshaped HDR color patches may be predicted by multiplying the inverse generator matrix in equation (29) with the inverse TPB coefficients.

[0135] The backward prediction error for each iteration in the BESA algorithm can be determined by comparing the backward reshaped HDR color patches or codewords aggregated into a vector / matrix with the reference HDR color patches or codewords in the per-channel backward observation vector / matrix derived from Equation (25). The per-channel backward observation vector / matrix can be generated or pre-computed from Equation (25), stored or cached in computer memory, and fixed at every iteration as follows:

number

[0136] A reference SDR signal or a reference SDR color patch or codeword may be used as a prediction target for a forward reshaping operation to generate a forward reshaped SDR color patch or codeword to approximate the reference SDR color patch or codeword. In the BESA algorithm, the reference SDR color patch or codeword

number

[0137] The reference SDR color patch or codeword for each channel at iteration k is given by the vector / matrix r F ch,(k) may be used to construct

number

[0138] Before the first iteration, the vector in equation (31) above is set to the original reference SDR signal represented by equations (22) and (23) as follows:

number

[0139] At iteration k, the optimized values ​​of the forward TPB coefficients (e.g., per channel) (m F ch,(k) ) can be generated via a least-squares solution to an optimization problem that minimizes the difference between the forward-reshaped SDR color patch or codeword and the reference SDR color patch or codeword determined for the iteration, as follows:

number

[0140] The predicted SDR color patch or codeword at iteration k for channel ch can be calculated from the optimized values ​​for the forward TPB coefficients (e.g., per channel) as follows:

number

[0141] At iteration k, the optimized values ​​of the backward TPB coefficients (e.g., per channel) (m B ch,(k) ) can be generated via a least-squares solution to an optimization problem that minimizes the difference between the inversely reshaped HDR color patch or codeword and the reference HDR color patch or codeword, where the latter is fixed for all iterations in the BESA algorithm, as follows:

number

[0142] The predicted HDR color patch or codeword at iteration k for channel ch can be calculated from the optimized values ​​for the backward TPB coefficients (e.g., per channel) as follows:

number

[0143] Calculate the backward prediction error and propagate the error back to the reference non-neutral SDR signal.

[0144] As mentioned above, the backward prediction error for each iteration in the BESA algorithm can be determined by comparing the backward reshaped HDR color patch or codeword with the reference HDR color patch or codeword derived from equation (25), which is fixed for all iterations in the BESA algorithm.

[0145] Among the backward prediction errors, the backward prediction errors for non-neutral colors (or non-gray colors) can be backpropagated to update or modify the reference SDR color patches or codewords for these non-neutral colors (or non-gray colors). The modified SDR color patches or codewords for non-neutral colors (or non-gray colors) can be combined with the (unmodified) reference SDR color patches or codewords for neutral colors (or gray colors) to serve as prediction targets for the forward pass in the next or subsequent iteration of the BESA algorithm.

[0146] In some operating scenarios, the backward prediction error can be calculated as the difference between the original HDR signal or reference HDR signal (or reference HDR color patch / codeword therein) and the predicted HDR signal (or backward reshaped HDR color patch / codeword therein) at iteration k as follows:

number

[0147] In response to determining that color patch i is in the neutral color set Φ, the reference SDR codeword of the color patch in the Cb and Cr channels of the reference SDR signal for the next iteration (k+1) can be set to a gray color value, such as 0.5, as follows:

number

[0148] Otherwise, in response to determining that color patch i is not in the neutral color set Φ, the reference SDR codeword of the color patch in the Cb and Cr channels of the reference SDR signal for the next iteration (k+1) can be set to an updated or modified value according to the backward prediction error as follows:

number

[0149] Neutral color preservation, as implemented using equations (38)-(40), can be used to ensure that neutral colors in the reference SDR signal for every iteration in the BESA algorithm deviate toward non-neutral colors, resulting in improved perceptual quality related to gray levels in the reshaped image.

[0150] The BESA algorithm may be iterated up to a total number of iterations. In some operating scenarios, the total number of iterations for the BESA algorithm may be specifically selected to balance color deviations in the reshaped SDR image and / or HDR image. For example, different total numbers of iterations in the BESA algorithm may generate forward and reverse TPB coefficients that produce reshaped SDR images with different SDR looks and reshaped HDR images with different HDR looks. A total number of iterations, such as 10, 15, etc., that produces reshaped SDR images and HDR images with relatively high quality looks can be selected for the BESA algorithm.

[0151] The BESA algorithm with neutral color preservation may be run for each candidate value of the parameter a. For example, in block 314, the BESA algorithm with neutral color preservation may be run for the current candidate value of the parameter a.

[0152] Block 316 involves determining whether the current candidate value for parameter a is the last candidate value of multiple candidate values ​​for parameter a. If so, process flow proceeds to block 318. If not, process flow returns to block 302.

[0153] Block 318 involves selecting an optimal or optimized value for parameter a, and calculating (a) optimized forward and backward TPB coefficients in RGB color space that correspond to the optimized value of parameter a (or simply selecting those that have already been calculated).

[0154] The optimized value of parameter a can be selected as follows: (1) (a) the RGB color space represents the maximum HDR color space (e.g., to cover the largest part of the R.2020 RGB color space);

number

number

number

[0155] This optimization problem is solved by solving the corresponding optimal TPB coefficients or optimized TPB coefficients m F,a ch ,m B,a ch This can be solved by examining the results from each a RGB color space using: Among all possible or candidate a values, the smallest a value (corresponding to the maximum color space coverage in the R.2020 color space) can achieve minimized HDR error with a relatively small SDR error.

number

number

[0156] 2F and 2G show example percentile Cb and Cr values ​​in the forward reshaped SDR domain or color space for different values ​​of parameter a. Different values ​​of parameter a determine different color space coverage of the backward reshaped HDR domain or color space in the R.2020 domain or color space, and therefore affect the reconstruction or generation of the backward reshaped HDR image in the R.2020 domain or color space differently, e.g., the distortion measure of the backward reshaped HDR image when compared to the reference HDR image that the backward reshaped HDR image should approximate.

number

[0157] The vertical axis of FIG. 2F represents the 3 percentile values ​​in the Cb and Cr channels in the forward-reshaped SDR domain for different values ​​of the parameter a (e.g., obtained using the BESA algorithm with neutral color preservation for 10 iterations) represented by the horizontal axis of FIG. 2F. The smaller the value of the parameter a, the larger the a color space covers in the R.2020 color space. The 3 percentile value may be used explicitly or implicitly to represent the minimum value in the Cb and Cr channels of the forward-reshaped SDR (e.g., YCbCr, etc.) domain or color space. The smaller the 3 percentile value, the smaller the minimum value, and therefore the more extended the codeword value range in the forward-reshaped SDR (YCbCr, etc.) domain or color space. As can be seen in FIG. 2F, the lower the value of the parameter a, the lower the 3 percentile value in the Cb and Cr channels of the forward-reshaped SDR (e.g., YCbCr, etc.) domain or color space.

[0158] The vertical axis of FIG. 2G represents the 97th percentile values ​​in the Cb and Cr channels in the forward-reshaped SDR domain for different values ​​of the parameter a (e.g., obtained using the BESA algorithm with neutral color preservation for 10 iterations) represented by the horizontal axis of FIG. 2G. The 97th percentile value may be used explicitly or implicitly to represent the maximum value in the Cb and Cr channels of the forward-reshaped SDR (e.g., YCbCr) domain or color space. The larger the 97th percentile value, the larger the maximum value, and therefore the more extended the codeword value range in the forward-reshaped SDR (YCbCr) domain or color space. As can be seen in FIG. 2G, the 97th percentile values ​​in both the Cb and Cr channels of the forward-reshaped SDR domain or color space remain approximately constant.

[0159] In some operating scenarios, the image dataset used in block 318 to determine distortion in the reshaped SDR and HDR images and select an optimized value for parameter a based partially or wholly on the distortion may include a video bitstream acquired using one or more specific video capture devices, such as a mobile phone (e.g., one that supports a video coding profile in SMPTE ST 2094, one that supports Dolby Vision Profile 8.4, etc.). A non-limiting example of an optimized value for parameter a may be, but is not limited to, 0.5, which determines the maximum color space coverage in the R.2020 color space for the inversely reshaped HDR domain or color space. When a wider inversely reshaped color space (e.g., one corresponding to a value smaller than the optimized value for parameter a) is used or supported, the reshaped SDR colors and codewords may begin to deviate from the reference SDR colors and codewords with relatively large distortion. This is because increasing the backward-reshaped HDR domain or color space means squeezing the differentiated forward-reshaped SDR colors (or codewords) more tightly into the forward-reshaped SDR domain or color space, resulting in the forward-reshaped SDR colors being displaced from their original 3D positions represented by the reference SDR colors (or codewords) that are approximated by the forward-reshaped SDR colors.

[0160] As noted above, the forward and backward TPB coefficients for a TPB-based reshaping operation can be determined for a given value, such as an optimized value for parameter a. These TPB coefficients can be mathematically multiplied with a forward or backward generator matrix constructed with TPB basis functions using the input codeword as an input parameter to generate a forward or reshaped codeword, which can involve performing multiple calculations that involve calculating a TPB basis function value for each pixel in a large number of pixels of the image.

[0161] In some operating scenarios, to speed up or reduce computation, forward and backward 3D-LUTs may be constructed from pre-computed TPB basis function values ​​from sampled values ​​and optimized forward and backward TPB coefficients generated for an optimized value for parameter a. The forward and backward 3D-LUTs or lookup nodes / entries therein may be pre-constructed and applied with relatively simple lookup operations in the forward and backward paths, or corresponding forward and backward reshaping operations performed therein on the input image at runtime, before being deployed at runtime to process the input image.

[0162] The optimized value of parameter a is a opt The corresponding optimal forward and backward TPB coefficients are denoted as m F ch,opt and m B ch,opt Shown as:

[0163] A forward 3D-LUT can be used to forward reshape an input HDR color (or codeword) in an input HDR domain or color space, such as the R.2020 domain or color space (which includes the a color space), to a forward reshaped SDR color (or codeword) in a forward reshaped SDR domain or color space.

[0164] A color space identified within an input HDR domain or color space, such as the R2020 (container) domain or color space, may be used to clip out HDR colors or codewords expressed outside the color space. Forward TPB coefficients may be applied to input HDR colors or codewords within a color space identified within the input HDR domain or color space, such as the R2020 (container) domain or color space, to generate predicted or forward-reshaped SDR colors or codewords for each lookup node or entry in the forward 3D-LUT. As a result, the forward 3D-LUT includes multiple lookup nodes or entries, each of which maps or forward-reshapes a respective input (cross-channel or three-channel) HDR color or codeword to a corresponding predicted or forward-reshaped (cross-channel or three-channel) SDR color or codeword.

[0165] In some operating scenarios, the forward 3D-LUT construction process includes a first step in which a 3D uniform sampling grid is prepared in the input HDR domain or color space, such as the R2020 (container) domain or color space.

[0166] By way of example and not limitation, the input HDR domain or color space is the R.2020 YCbCr color space HLG, which includes three dimensions or channels: a Y axis, a Cb axis, and a Cr axis. The input HDR values ​​(

number

number

[0167] Although the valid values ​​in the entire R.2020 YCbCr color space may be limited (so as not to cover the entire 3D cube of YCbCr values), the entire sampling value grid may be used to cover the entire 3D cube of YCbCr values ​​to reduce the possibility of codeword deviations caused by the compression operation.

[0168] For simplicity, (i,j,k) may be vectorized or denoted as p. Thus, the sampled values ​​in R.2020 YCbCr are

number

number

number

[0169] The forward 3D-LUT construction process includes a second step in which the sampled values ​​in the R.2020 YCbCr color space HLG can be converted to corresponding values ​​(or colors) in the R.2020 RGB color space HLG as follows:

number

[0170] The forward 3D-LUT construction process is performed by converting the transformed values ​​in the R.2020 RGB color space HLG to the optimized values ​​for the parameter a (a opt The third step involves converting the color values ​​(or colors) corresponding to the RGB color space HLG (denoted as a

number

[0171] The forward 3D-LUT construction process includes a fourth step in which (a) the converted values ​​in the RGB color space HLG may be clipped as follows:

number

[0172] The forward 3D-LUT construction process includes a fifth step in which the clipped values ​​in a RGB color space HLG can be converted to corresponding values ​​(or colors) in the R.2020 RGB color space HLG as follows:

number

[0173] The forward 3D-LUT construction process includes a sixth step in which values ​​in the R.2020 RGB color space HLG derived using equation (47) above can be converted to corresponding values ​​(or colors) in the R.2020 YCbCr color space HLG as follows:

number

[0174] The forward 3D-LUT construction process is as follows: for each lookup node / entry in the forward 3D-LUT, the optimized forward TPB coefficients are mapped or forward reshaped to the SDR color or codeword (or codeword value), as follows:

number

number

[0175] In the above equation (49), the forward generator matrix S F ch is the input HDR color or codeword (or codeword value) V, as follows: YCbCr (FL),(R2020) can be constructed or constructed using as input to the forward TPB basis functions.

number

[0176] Mapped or forward reshaped SDR color or codeword (or codeword value)

number

[0177] An inverse 3D-LUT such as that described above can be used to inversely reshape reshaped SDR colors (or codewords) in a forward-reshaped SDR domain or color space to inversely reshaped HDR colors (or codewords) in an inverse-reshaped HDR domain or color space, such as the R.2020 domain or color space (which includes the a color space).

[0178] In some operating scenarios, a reverse 3D-LUT construction process, which may be simpler than the forward 3D-LUT construction process described above, may be implemented or executed to construct or configure the reverse 3D-LUT. In some operating scenarios, to speed up or reduce computation, the reverse 3D-LUT may be constructed from pre-calculated TPB basis function values ​​from sampled values ​​and optimized reverse TPB coefficients generated for an optimized value for the parameter a. The reverse 3D-LUT or lookup nodes / entries therein may be pre-constructed and applied with a relatively simple lookup operation in the reverse path or a corresponding reverse reshaping operation performed therein on the input image at runtime before being deployed at runtime to process an input image, such as a forward-reshaped SDR image.

[0179] The pre-constructed backward 3D-LUT can be deployed at the decoder side, where a downstream device on the receiving side at the decoder side can receive and decode a video signal encoded with a forward-reshaped SDR image in a forward-reshaped SDR domain or color space, and apply backward reshaping to the forward-reshaped SDR image using the 3D-LUT to generate a backward-reshaped HDR image in a backward-reshaped HDR domain or color space, such as a color space included in the R.2020 domain or color space.

[0180] The shape formed or defined by the complete boundary of the forward reshaped SDR domain or color space mapped using the TPB-based forward reshaping operation may not be a simple 3D cube, but the forward reshaped SDR values ​​in the forward reshaped SDR domain or color space may be clipped to the narrowest or smallest 3D cube (e.g., a 3D rectangle, a rescaled 3D cube from a 3D rectangle, etc.) that contains or supports all the forward reshaped SDR values ​​represented in all lookup nodes / entries of the forward 3D-LUT, for example, without clipping these SDR values ​​represented in the forward 3D-LUT.

[0181] In some operating scenarios, the backward 3D-LUT construction process includes a first step in which a complete sampling value grid can be prepared in a forward reshaped SDR domain or color space, such as the entire R.709 SDR YCbCr color space.

[0182] The reverse 3D-LUT construction process includes a second step in which minimum and maximum values ​​for each dimension or (color) channel in the forward reshaped SDR domain or color space are determined among the forward reshaped SDR values ​​in the forward 3D-LUT and can be used to restrict the (input) data range of the forward reshaped SDR (e.g., YCbCr, etc.) colors or codewords that serve as input for the reverse path.

[0183] The inverse 3D-LUT construction process involves, for each lookup node / entry in the inverse 3D-LUT, optimizing the inverse TPB coefficients to obtain a mapped or inversely reshaped HDR color or codeword (or codeword value), using an inverse generator matrix S constructed with the forward reshaped SDR color or codeword (or codeword value) within the clipped or limited (input) data range derived in the second step of the inverse 3D-LUT construction process as input parameters to the inverse TPB basis functions.B ch (e.g., as shown in equation (29) above) is used.

[0184] The mapped or inversely reshaped HDR colors or codewords (or codeword values) may be used as lookup values ​​of lookup nodes / entries of the inverse 3D-LUT, while the input or forward reshaped SDR colors or codewords (or codeword values) used as input parameters to the inverse TPB basis functions may be used as lookup keys of lookup nodes / entries of the inverse 3D-LUT. At runtime, the mapped or inversely reshaped HDR colors or codewords can simply be looked up in the inverse 3D-LUT using the lookup keys as the input or forward reshaped SDR colors or codewords to be inversely reshaped into the mapped or inversely reshaped HDR colors or codewords.

[0185] [White-box Inverse Optimization (WB)] Forward TPB-based reshaping or a corresponding forward 3D-LUT may not be used in all video capture and / or editing devices due to the cost and / or computational overhead associated with implementing or working with forward TPB-based reshaping or a forward 3D-LUT in video capture and / or editing applications.

[0186] In some operating scenarios, a hardware-based solution, such as one implemented using an available ISP, may be used in a video capture and / or editing device to perform HDR-SDR conversion, such as HDR HLG to SDR image generation. An existing ISP pipeline deployed with the device can operate with one or more programmable parameters to generate, convert, and / or output an SDR image based in part or in whole on a corresponding HDR (e.g., HLG, etc.) image captured by the device.

[0187] In the forward path, the programmable parameters for the ISP pipeline can be specifically set or configured to cause the ISP pipeline to output an SDR image that approximates as closely as possible a reference SDR image generated using a white-box HDR-to-SDR conversion function.

[0188] In the reverse path, reverse TPB-based reshaping or a corresponding reverse 3D-LUT may be used to generate a reverse-reshaped HDR image from the SDR image output from the ISP pipeline. The reverse TPB coefficients used in the reverse path may be optimized to cover as much of the reverse-reshaped HDR domain or color space as possible, such as the R.2020 color space.

[0189] FIG. 3B illustrates an exemplary process flow for determining or generating optimized values ​​for one or more programmable parameters in an ISP pipeline used to generate or output an SDR image from an HDR image (ISP).

[0190] In some operating scenarios, the ISP pipeline may be a relatively limited programmable module implemented using ISP (hardware) to approximate a white-box HDR-to-SDR conversion function, such as a white-box (e.g., known, well-defined, etc.) conversion function, to convert an HDR HLG image to a corresponding SDR image in block 330.

[0191] The ISP pipeline may include or implement (1) a first set of three one-dimensional lookup tables (1D-LUTs) for HLG RGB to linear RGB conversion, followed by (2) a 3x3 matrix for conversion from an HDR image in an HDR domain or color space, such as the R.2020 color space, to an SDR domain or color space, such as the R.709 color space, followed by (3) a second set of three 1D-LUTs implementing a BT.1886 linear-to-nonlinear (gamma) conversion using a (e.g., standards-based) HLG optical-to-optical transfer function (OOTF).

[0192] For purposes of example only, the HDR image may be an HDR HLG image retrieved from an image dataset or database such as that shown in Figure 3B. One or more programmable parameters for the ISP pipeline may be a design parameter (γ BT1886 (shown as

[0193] The process flow of FIG. 3B involves determining the design parameter γ for the ISP SDR image to best approximate the reference SDR image generated in block 330. BT1886 The process flow may be used to search for an optimized value of the design parameter γ in an iterative order, which may be sequential or non-sequential. BT1886 The method may iterate through multiple candidate values ​​of .

[0194] Block 322 determines the design parameter γ BT1886 The current (e.g., to be iterated) value of the design parameter γ BT1886 The design parameter γ is selected as the next candidate value (e.g., initially as the first candidate value, etc.) among the plurality of candidate values ​​of γ. BT1886 Given the current value of , the ISP SDR image can be compared to the reference SDR image in one or more subsequent process flow blocks of FIG. 3B.

[0195] Block 324 involves applying the first set of 1D-LUTs to the HDR HLG(RGB) image retrieved from the database to generate a corresponding HDR linear RGB image, as follows:

number

[0196] Block 326 includes applying a 3x3 matrix to the corresponding HDR linear RGB image to generate a corresponding SDR linear RGB image. The 3x3 matrix may be given as:

number

[0197] The SDR linear RGB codeword in the corresponding SDR linear RGB image generated using the 3 × 3 matrix in equation (52) above is l ch (R709) It may be shown as:

[0198] Block 328 maps a second set of 1D-LUTs to the SDR linear RGB codeword l in the corresponding SDR linear RGB image. ch (R709) to obtain the corresponding ISP SDR codeword (s ch (R709) This includes generating a

[0199] In some operating scenarios, as described above, the second set of 1D-LUTs merges or combines the linear-to-nonlinear SDR transformation with the HLG OOTF given as follows:

number

number

number

[0200] The intermediate SDR codeword in equation (53) above

number

[0201]

number

[0202] where γ BT1886 represents the above design parameters.

[0203] The blocks 322-328 of FIG. 3B, or the combination of the first set of 1D-LUTs represented in equation (51) above, the 3×3 matrix represented in equation (52) above, and the second set of 1D-LUTs represented in equations (53)-(56) above, when implemented using an ISP pipeline, provide a design parameter γ BT1886, and calculate the HDR HLG-SDR conversion function (f HLG→ISPSDR (shown as

[0204] Block 332 calculates the HDR HLG-SDR conversion function f HLG→ISPSDR and determining the difference (e.g., quality, etc.) between the ISP SDR image generated from the white-box HDR-to-SDR conversion function in block 330. A quality assessment function such as MSE, RMSE, SAD, PSNR, SSIM, etc. may be used to calculate the difference.

[0205] Block 334 determines the design parameter γ BT1886 The current candidate value of the design parameter γ BT1886 If so, process flow proceeds to block 336. If not, process flow returns to block 322.

[0206] Block 336 determines the design parameter γ BT1886 The design parameter γ is selected to have an optimum or optimized value. BT1886 The selection of the optimized value of the design parameter γ is selected from multiple candidate values ​​such that the difference between the ISP SDR image and the reference SDR image calculated using a relatively large image data set or database is minimized, as follows: BT1886 For a particular value of (γ BT1886,opt This can be formulated as an optimization problem to find

number

[0207] gamma BT1886,opt An example value of may be, but is not limited to, 2.115.

[0208] [SDR-PQ TPB optimization] As described above, in a "WFB" use case or operating scenario, the BESA algorithm may be used under the SLBC framework to generate optimized reshaping operation parameters used by forward and backward reshaping operations to generate a forward-reshaped SDR image and a backward-reshaped HDR image. In comparison, in a "WB" use case or operating scenario, only backward reshaping may be performed on a non-forward-reshaped SDR image, such as an ISP SDR image generated using an ISP pipeline implemented with a video capture and / or editing device, encoded into an output video signal under the SLiDM framework to generate a corresponding backward-reshaped or reconstructed HDR image.

[0209] In "WB" use cases or operating scenarios, forward TPB reshaping from an input HDR domain or color space, such as the R.2020 HDR color space HLG, to a forward-reshaped SDR domain or color space, such as the R.709 SDR YCbCr color space, may not be used to help utilize the full range of codewords supported by the R.709 SDR YCbCr color space. ISP SDR codewords in ISP SDR images encoded into an output video signal are often severely restricted in the R.709 color space portion generated or supported by a video capture and / or editing device or an ISP pipeline implemented therein. It may be difficult for ISP SDR codewords in the severely restricted R.709 color space portion to be further extended or mapped back into the inversely reshaped or reconstructed HDR domain or color space by inverse reshaping, such as inverse TPB-based reshaping.

[0210] However, in "WB" use cases or operating scenarios, as in "WFB" use cases or operating scenarios, optimized reshaping operating parameters, such as TPB coefficients, can help increase the coverage of supported color spaces in which the inversely reshaped HDR images produced in the reshaping optimizations should be represented. In many "WB" operating scenarios, the inversely reshaped or reconstructed HDR domain or color space achieved using inverse TPB-based reshaping can be only slightly larger than the R.709 color space compared to the largest (a) color space achievable in "WFB" use cases or operating scenarios.

[0211] For purposes of example only, in an operating scenario such as that shown in FIG. 1B, the ISP SDR domain or color space may be that of R.709 or a severely restricted portion thereof in the ISP pipeline, while the (original or inversely reshaped) HDR domain or color space may be that of R.2020.

[0212] By way of example and not limitation, a subset or subspace in the R.2020 color space (or an HDR color space for representing a reshaped HDR image) may be defined or characterized by a particular white point and three particular primary colors (red, green, and blue). The particular white point may be selected or fixed to be the D65 white point.

[0213] The inversely reshaped HDR color space for representing the inversely reshaped or reconstructed HDR image is shown as (b) color space. Therefore, the CIExy coordinates of the primary colors and white point that define the (b) color space are (R x (b) ,R y (b) ), (G x (b) ,G y (ab) ), (B x (b) ,B y (b) ) and (W x(b) ,W y (b) (b) The white point of the color space (W x (b) ,W y (b) ) can be specified as the D65 white point.

[0214] (b) Each of the primary colors of the color space is a primary color (G x (P3 ),G y (P3)) and each primary color of the R.2020 color space (G x (R2020) ,G y (R2020) ) Any point along the line between two respective primaries in the P3 color space and the R.2020 color space can be expressed as a linear combination of these two respective primaries with a weighting factor b as follows: R x (b) =bR x (P3) +(1-b)R x (R2020) (58-1) R y (b) =bR y (P3) +(1-b)R y (R2020) (58-2) G x (b) =bG x (P3) +(1-b)G x (R2020) (59-1) G y (b) =bG y (P3) +(1-b)G y (R2020) (59-2) B x (b) =bB x (P3) +(1-b)B x (R2020) (60-1) B y (b) =bG y (P3) +(1-b)B y (R2020) (60-2)

[0215] Therefore, the optimization problem of finding maximum support from the (b) color space for the R.2020 color space can be simplified to the problem of selecting a weighting factor b. When b = 0, the (b) color space is the entire R.2020 color space. When b = 1, the (b) color space is the P3 color space, as shown in Figure 2K (where the (b) color space is denoted as "TPB" or "TPB Covered Colors (b Color Space)"). Figures 2H-2J show three exemplary (b) color spaces for b = 0.25, 0.50, and 0.75, respectively.

[0216] As shown in Figures 2H-2K, the TPB optimization problem reduces to finding or searching for the smallest possible value for the parameter b in order to cover as much of the R.2020 color space as possible with the inversely reshaped or reconstructed HDR domain or color space.

[0217] 3C shows an exemplary process flow for finding the smallest possible value for parameter b. The process flow of FIG. 3C may iterate through multiple candidate values ​​for parameter b in an iterative order, which may be sequential or non-sequential.

[0218] Block 342 includes selecting the current (e.g., iterated, etc.) value of parameter b as the next candidate value (e.g., initially the first candidate value, etc.) among multiple candidate values ​​of parameter b. Given the candidate inversely reshaped HDR color space with the current value of parameter b, optimized reshaping operation parameters, such as optimized TPB coefficients, can be obtained in one or more subsequent process flow blocks of FIG. 3C.

[0219] Block 344 includes constructing sample points or preparing two sampled data sets in a candidate inversely reshaped HDR color space. By way of example and not limitation, the candidate inversely reshaped HDR color space may be a Hybrid Log Gamma (HLG) RGB color space (referred to as "(b) RGB color space HLG").

[0220] Similar to block 304 of FIG. 3A, the first of the two sampled data sets is a uniformly sampled data set of color patches. Each color patch in the uniformly sampled data set of color patches includes three dimensions, denoted as R, G, and B axes for the R, G, and B component colors, respectively. (b) A uniformly sampled data set from the RGB color space HLG.

number

number

[0221] For simplicity, (i,j,k) can be vectorized or simply denoted as p. Correspondingly, the uniformly sampled data points or RGB colors

number

number

number

[0222] The second of the two sampled data sets prepared or constructed in block 344 is a neutral color data set. This second data set includes a plurality of neutral colors or neutral color patches (also called gray colors or gray color patches).

[0223] The second data set may be used to preserve input gray color patches in the input domain as output gray color patches in the output domain when the input gray color patches in the input domain are mapped or reshaped to output gray color patches in the reshaping operations described herein. The input gray color patches in the input domain (or input color space) may be given increased weighting in the optimization problem compared to other color patches to reduce the likelihood that these input gray color patches will be mapped to non-gray color patches in the output domain (or output color space) by the reshaping operations.

[0224] The second data set (gray color data set or gray color data set) may be prepared or constructed by uniformly sampling R, G, B values ​​along a line connecting between the first gray color (0,0,0) and the second gray color (1,1,1) in the RGB domain (e.g., (b) RGB color space HLG, etc.), as follows:n This results in nodes or gray color patches.

number

[0225] All N in the second data set n The nodes can be grouped or aggregated into a neutral color vector / matrix as follows:

number

[0226] The neutral color vector / matrix in equation (8) above is expressed as follows: t Repeat (a positive integer greater than or equal to 1) times to find N in the second data set (which is repeated here). n N t It can generate neutral color patches.

number

[0227] The repetition of neutral colors in the second data set increases the weighting of neutral or gray colors relative to other colors, and therefore neutral colors may be more preserved in the optimization problem than other colors.

[0228] The first data set of (all sampled) colors and the second (repeated) data set of neutral colors in equations (61) and (64) can be aggregated or placed together into a single combined vector / matrix as follows:

number

[0229] Combined Vector / Matrix

number

number

number

[0230] Block 346 calculates the combination vector / matrix in (b) RGB color space (or (b) RGB color space HLG) as follows:

number

number

number

[0231] Block 348 converts the vector / matrix in the R.2020 RGB color space HLG as follows:

number

number

number

number

[0232]

number

number

[0233] Block 350 converts the vector / matrix in the R.2020 RGB color space HLG as follows:

number

number

number

[0234]

number

number

[0235] Block 352 uses the current value of parameter b as input to the backward TPB optimization algorithm to generate optimized backward TPB coefficients for TPB-based reshaping for the (b) color space.

number

[0236] To do this, the inverse generator matrix is ​​composed of the TPB basis functions and

number

number

[0237] The backward prediction error can be determined by comparing the backward reshaped HDR color patches or codewords aggregated into a vector / matrix with the reference HDR color patches or codewords in the per-channel backward observation vector / matrix derived and minimized when solving the TPB optimization problem. The per-channel backward observation vector / matrix can be generated or pre-computed from equation (71) above, stored or cached in computer memory, and fixed at every iteration as follows:

number

[0238] Optimized values ​​of the backward TPB coefficients (e.g., per channel) (m B ch ) can be generated via a least-squares solution to an optimization problem that minimizes the difference between the inversely reshaped HDR color patch or codeword and the reference HDR color patch or codeword.

number

[0239] For each channel ch, the predicted (or inversely reshaped or reconstructed) HDR codeword per channel can be calculated as follows:

number

[0240] Block 354 involves determining whether the current candidate value for parameter b is the last candidate value of multiple candidate values ​​for parameter b. If so, process flow proceeds to block 356. If not, process flow returns to block 342.

[0241] Block 356 involves selecting an optimal or optimized value for parameter b, and calculating (or simply selecting those already calculated) optimized forward and reverse TPB coefficients in the (b) RGB color space that correspond to the optimized value of parameter b.

[0242] Similar to the "WFB" use case or operating scenario, in the "WB" use case or operating scenario, (b) color space is also part of the optimization. (b) color space and the backward TPB coefficients can be optimized together. Thus, the optimization problem becomes:

number

number

[0243] An example of an optimized value for the parameter b may be, but is not limited to, 1, which corresponds to the P3 color space.

[0244] [Black-box TPB backward optimization on a single device (BB1)] Some (dual-mode) video capture devices or mobile devices support both SDR and HDR capture modes and can therefore output either SDR or HDR images. Some (mono-mode) video capture devices or mobile devices support only SDR capture mode. Under the techniques described herein, SDR-HDR mappings can be modeled or generated using a dual-mode video capture device. Some or all of these SDR-HDR mappings can then be applied, for example, in a downloaded and / or installed video capture application, to SDR images captured by either the dual-mode video capture device or the mono-mode video capture device to "upconvert" the SDR images obtained in SDR capture mode to corresponding HDR images, regardless of whether the video capture device supports HDR capture mode.

[0245] In some operating scenarios, the computing environment or processing power within a mobile device may be limited, so a static (inverse) 3D-LUT derived at least in part from the TPB basis functions and optimized TPB coefficients may be used to represent a static SDR-HDR mapping that maps or inversely reshapes SDR images, such as all SDR images in an SDR video sequence, to generate corresponding HDR images, such as all HDR images in a corresponding HDR video sequence.

[0246] The static SDR-HDR mapping may be modeled at least in part based on multiple image pairs of HDR and SDR images captured by a particular camera of a dual-mode video capture device. Each image pair in the HDR and SDR image pairs includes an SDR image and an HDR image corresponding to the SDR image. The SDR and HDR images depict the same visual scene in the real world, subject to spatial alignment errors caused by spatial translation that may occur between a first point in time when the SDR image is captured by a particular camera operating in SDR capture mode and a second point in time when the HDR image is captured by the same camera operating in HDR capture mode. For example, for each visual scene in the physical world, the video capture device may operate in HDR mode to capture the HDR image in an image pair and the SDR image in the same image pair using SDR mode. This may be repeated to generate multiple image pairs of HDR and SDR images to generate a relatively large number of different color patches (or a relatively large number of different colors or different codewords) used to generate the static SDR-HDR mapping.

[0247] There are several challenges in modeling static SDR-HDR mapping using captured SDR and HDR images. First, although both SDR and HDR capture modes may use the same ISP within the same video capture device, the video capture device may apply different image capture settings (e.g., exposure settings for a particular camera) to the same scene to obtain the designed optimized picture quality in the captured SDR and HDR images. As a result, static SDR-HDR mapping can model or approximate to some extent the HDR image capture settings actually implemented by the video capture device. Second, the spatial alignment between the captured SDR and HDR images used to derive the image pair of the SDR and HDR images may not be accurate. For example, selecting a particular capture mode between the SDR and HDR capture modes may be performed via touching the screen of the video capture device, which may cause the video capture device to lose or move away from its previous spatial position and / or orientation. Furthermore, temporal alignment between a captured SDR video / image sequence and a captured HDR video / image sequence can easily occur when these SDR and HDR sequences are captured at different time instances or durations. Some depicted visual objects in these sequences may move in or out of the camera field of view. Some local regions within the camera field of view may be occluded or unoccluded over time.

[0248] In some operational scenarios, a registration operation can be performed to resolve spatial and / or temporal registration issues between a captured SDR image and a corresponding captured HDR image to generate a registered SDR image and a corresponding registered HDR image and include the registered SDR image and the corresponding registered HDR image in an image pair of the SDR image and the HDR image. The image pair can then be used to determine or establish SDR color patches or colors in the registered SDR image and corresponding HDR color patches or colors in the corresponding registered HDR image. These SDR and HDR color patches or colors can be used to derive at least some color patches in a set of corresponding SDR and HDR color patches or colors for the purpose of deriving or generating a static SDR-HDR mapping.

[0249] The SDR and HDR video capture process may be performed to provide comprehensive coverage of different scenes and / or exposures at different times of day, such as from early morning to late night, for both indoor and outdoor scenes. The same video capture device may be used in both SDR and HDR modes to capture SDR and HDR video results for each scene. In some operating scenarios, specific frames / images, such as the first frame / image of each of a video sequence, may be selected or extracted for the purpose of constructing or deriving multiple image pairs of SDR and HDR images to be included in a training image dataset or database.

[0250] FIG. 3D shows an exemplary process flow for generating optimized reshaping operational parameters, such as optimized TPB coefficients or optimal TPB coefficients, using corresponding (or matching) SDR and HDR color patches or colors in a captured SDR image and a captured HDR image (corresponding to the captured SDR image) captured by a video capture device operating in an SDR capture mode and an HDR capture mode, respectively.

[0251] Block 362 includes receiving a captured SDR image and locating or extracting a set of SDR image feature points within the captured SDR image. Block 364 includes receiving a captured HDR image and locating or extracting a set of HDR image feature points within the captured HDR image. The set of SDR image feature points and the set of HDR image feature points may be the same set of feature point types.

[0252] The set of feature point types may include a wide variety of different feature point types, including, but not limited to, any, some, or all of: Binary-Robust-Invariant-Scalable-Keypoints (BRISK) features detected using the BRISK algorithm, corners detected using Features from Accelerated-Segment-Test (FAST) algorithm, features detected from the KAZE algorithm, corners detected using the Minimum Eigenvalue algorithm, features generated using the Maximally-Stable-Extremal-Regions (MSER) ​​algorithm, features detected using the Oriented-FAST-and-Rotated (ORB) algorithm, features extracted from the Scale-Invariant-Feature-Transform (SIFT) algorithm, features extracted from the Speeded-Up-Robust-Features (SURF) algorithm, etc.

[0253] Some or all of the feature points in the set of SDR image feature points and the set of HDR image feature points may each be represented by a feature vector or descriptor, such as an array of feature values ​​(e.g., numerical values).

[0254] Block 366 includes matching some or all of the feature points in the set of SDR image feature points with some or all of the feature points in the set of HDR image feature points. For each feature point of a particular type in the SDR image in the set of SDR image feature points, a matching metric may be calculated between the feature point in the SDR image and each feature point of the same type in the HDR image in the set of HDR image feature points. In a non-limiting example, the matching metric may be calculated as the sum of absolute differences (SAD) between a first feature vector representing the feature point in the SDR image and a second feature vector representing the feature point in the HDR image, or another metric function used to measure the difference between the two feature points. A particular feature point with the lowest SAD or matching metric may be selected from among the feature points of the same type in the HDR image as a possible match for the feature point in the SDR image. In response to determining that the lowest SAD or matching metric is lower than the matching difference threshold, the particular feature point in the HDR image may be identified or determined as a match for the feature point in the SDR image, and the feature point in the SDR image and the particular feature point in the HDR image form a matched SDR and HDR feature point pair. Otherwise, the particular feature point may not be identified as such a match. This matching operation may be performed iteratively for all feature points in the set of SDR image feature points, thereby resulting in a set of matched SDR and HDR feature point pairs.

[0255] Block 368 includes calculating a geometric transformation, such as a 3x3 2D affine transformation, between the SDR image and the HDR image based in part or in whole on the set of matched SDR and HDR feature point pairs. The geometric transformation may be calculated or derived using coordinates (e.g., pixel rows and pixel columns) of matched feature points in the SDR and HDR images from the set of matched SDR and HDR feature point pairs.

[0256] For each (e.g., k-th) pair of matched SDR and HDR feature points, the 2D coordinates of the HDR feature point in the pair are (τ ix ,τ iy ) and τ i =[τ ix τ iy 1] in the first vector, while the 2D coordinates of the SDR feature points are (η ix ,η iy ) and the second vector η i =[η ix η iy 1] may be included.

[0257] The total number of matched SDR and HDR feature pairs in the set of matched SDR and HDR feature pairs is denoted as N. The vectors generated from the 2D coordinates of the feature points in all the pairs of matched SDR and HDR feature points in the set of matched SDR and HDR feature pairs can be aggregated together as follows:

number

[0258] As mentioned above, the geometric transformation may be expressed as a 3x3 matrix as follows:

number

[0259] Mathematically, a first vector and a second vector representing the SDR feature points and the HDR feature points, respectively, in each pair of matched SDR and HDR feature points in the set of matched SDR and HDR feature point pairs can be related to each other through a 3×3 matrix representing a geometric transformation as follows:

number

[0260] Alternatively, for all N pairs of matched features in the set of matched SDR and HDR feature pairs, the vectors representing the SDR and HDR feature points can be related through a 3x3 matrix representing a geometric transformation as follows:

number

[0261] The values ​​of the matrix elements in the 3×3 matrix representing the geometric transformation may be generated or obtained as a solution (e.g., a least-squares solution, etc.) of an optimization problem that minimizes the transformation or registration error. More specifically, the optimization problem may be formulated as follows:

number

[0262] The optimized or optimal values ​​for the 3x3 matrix representing the geometric transformation can be obtained via a least squares solution to the optimization problem in equation () above, as follows:

number

[0263] Block 370 includes applying a geometric transformation to one of the SDR image and the HDR image. In this example, the transformation is applied to the HDR image, whereby HDR pixel positions in the HDR image are shifted to match SDR pixel positions in the SDR image for each of the three channels, e.g., Y, Cb, and Cr, in which the HDR image is represented. As a result, most HDR pixels in the HDR image are spatially aligned with corresponding SDR pixels in the SDR image. The remaining HDR pixels and the remaining SDR pixels with which the HDR pixels are not spatially aligned may be excluded (e.g., assigned out-of-range pixel values) from being used as (valid) color patches or colors for purposes of generating static SDR-HDR mapping or TPB coefficients or static inverse 3D-LUTs.

[0264] Block 372 involves finding corresponding valid color patches or colors in the SDR and HDR images. The valid color patches or colors can be obtained from the codeword values ​​of spatially aligned pixels in the SDR and HDR images. As described above, the spatially aligned pixels can be generated by applying a geometric transformation generated from the matched SDR and HDR feature points.

[0265] In some operational scenarios, after (initial) matched SDR and HDR feature points having matching metrics lower than a minimum matching difference threshold are identified, a further matching threshold may be applied to select or distinguish a subset of final matched SDR and HDR feature points from or among the (initial) matched SDR and HDR feature points, in order to increase the spatial alignment accuracy or reliability of the geometric transformation.

[0266] For each spatially aligned pixel in the SDR image, an SDR codeword, such as SDR Y / Cb / Cr values, may be determined for the spatially aligned pixel in the SDR image. For co-located or spatially aligned pixels in the spatially transformed HDR image generated using a geometric transformation that correspond to spatially aligned pixels in the SDR image, a corresponding HDR codeword, such as corresponding HDR Y / Cb / Cr values, may be determined for the spatially aligned pixel in the HDR image. If a transformed HDR pixel is not available (e.g., has a value of 0), the HDR pixel may be discarded or may be prevented from being considered as part of the matching color patch or codeword.

[0267] For each pair (e.g., the i-th pair) of matched SDR and HDR color patches or colors generated from a pair of spatially aligned SDR and HDR pixels in the SDR and HDR images, the matched SDR and HDR color patches or colors (or codeword values), respectively, may be given as follows:

number

[0268] The matched SDR and HDR color patches (or codeword values) in all pairs of matched SDR and HDR color patches, as generated from all pairs of spatially aligned SDR and HDR pixels in the SDR and HDR images of all image pairs, can be aggregated together from all images and merged into two matrices as follows:

number

[0269] Block 374 calculates the matrices in equations (88) and (89) above.

number

[0270] The inverse generator matrix for each channel is given by the matrix

number

number

[0271] The per-channel observation matrix for SDR-HDR mapping is given by the matrix

number

number

[0272] The optimized or optimal backward TPB coefficients for a channel ch can be solved via a least squares solution as follows:

number

[0273] The predicted or inversely reshaped HDR value for channel ch can be calculated as follows:

number

[0274] [Black Box Reverse Reshaping Optimization in Two Devices (BB2)] In a "BB2" operating scenario, (reverse) reshaping mapping may be used to map an SDR image (ISP-captured) captured by a first video capture device in SDR capture mode to generate an (ISP-mapped) HDR image that simulates the HDR appearance of an (ISP-mapped) HDR image captured by a second, different video capture device operating in HDR capture mode. In some operating scenarios, the first video capture device may be a relatively low-end mobile phone capable of capturing only SDR images or pictures, while the second video capture device may be a relatively high-end phone capable of capturing HDR images or pictures. Although the first and second video capture devices may operate with different hardware configurations and capabilities, reshaping mapping is used to reshape the (ISP-captured) SDR image captured by the first device into an (ISP-mapped) HDR image that approximates the (ISP-mapped) HDR image captured by the second device.

[0275] The reshaping (SDR-HDR) mapping may be modeled at least in part based on a plurality of image pairs formed by (training) SDR images captured by a first device and (training) HDR images captured by a second device, where each image pair of an HDR image and an SDR image includes an SDR image and an HDR image corresponding to the SDR image.

[0276] In some "BB2" operating scenarios, some or all of the image pairs may include captured SDR and HDR images depicting visual scenes, such as natural indoor / outdoor scenes, in the real world that are subject to spatial and / or temporal alignment errors with respect to the first and second video capturing devices (e.g., initially, before the image alignment operation). Similar to the "BB1" use cases or operating scenarios, in these "BB2" use cases or operating scenarios, the image alignment operation between corresponding SDR and HDR images in the image pair may be performed using, for example, a geometric transformation constructed using selected feature points extracted from the SDR and HDR images. The SDR and HDR color patches or colors determined from the aligned SDR and HDR images in the image pair may then be used to generate the reshaping (SDR-HDR) mapping described herein.

[0277] In some "BB2" operating scenarios, instead of or in addition to capturing images from natural scenes, some or all of the image pairs in the plurality of image pairs may include captured SDR and HDR images (e.g., initially, before image alignment operations, etc.) showing a color chart displayed on one or more reference image displays of the same type, e.g., in a laboratory environment. For example, the color chart can be generated as a 16-bit full HD RGB (color chart) TIFF image. These TIFF images including the color chart can be displayed as a perceptually quantized (PQ) video signal on a reference image display, such as a PRM TV. The color charts displayed on the reference image displays can be captured by first and second video capture devices, respectively.

[0278] FIG. 2N illustrates an exemplary TIFF image including a color chart. As shown, the color chart may be a central square in a TIFF image captured between a first device and a second device, including multiple color blocks having distinct sets of colors to be matched. Each color in the distinct set of colors may be different from all other colors in the distinct set of colors. The distinct set of colors may be displayed by a reference image display using a distinct set of pixel / codeword values. Thus, each color in the distinct set of colors may correspond to a respective pixel or codeword value in the distinct set of pixel / codeword values.

[0279] The different TIFF images may include different color charts with different multiple colors or different sets of distinct colors, and each color chart in each of the different TIFF images may correspond to a respective multiple distinct color set in the different multiple colors or distinct color sets.

[0280] Each of the four corner rectangles in the TIFF image of Figure 2N includes a checkerboard pattern. Different TIFF images containing different color charts may include the same corner rectangle or the same checkerboard pattern. The checkerboard patterns at the four corners of the TIFF image can be used as a spatial key or reference mark for estimating a projective transformation, as described in more detail below. Additionally, optionally or alternatively, the TIFF image may include one or more numbers, binary coding, a QR code, etc., as a unique identifier (ID) assigned or used to identify the TIFF image, the color chart therein, or multiple or sets of colors therein, etc.

[0281] In order to find as many color correspondence / mapping relationships as possible between the first video capture device and the second video capture device, the color chart of the TIFF image described herein may include as many different colors (sets or multiple colors) as possible corresponding to as many different pixel / codeword values ​​as possible according to the display capability of the reference image display to distinguish different colors. Furthermore, the color chart may be displayed at different overall intensities or illuminances in order to allow the first capture device and the second capture device to have different exposure settings under different illumination conditions. Thus, the same color displayed in the color chart of the TIFF image can be displayed at different intensities or illuminances on the reference image display.

[0282] 3E shows an exemplary process flow for generating multiple different color charts contained in multiple different TIFF images. To generate each color chart, the mean and variance values ​​of the colors (e.g., RGB, etc.) in the color chart may be determined and used to randomly generate colors using a statistical distribution (of some type). In some operating scenarios, the statistical distribution represents a beta distribution or distribution type. Different color charts with different colors or distinct sets of colors can be generated with different combinations of mean and variance values ​​of the statistical distribution or distribution type.

[0283] Block 382 determines the set of possible mean values ​​(M PQ This includes defining or determining the

[0284] In an operating scenario where the reference image display used to display the TIFF image containing the color chart is a PRM TV, the maximum and minimum luminance values ​​supported by the reference image display may range between 1000 nit and 0 nit, or an even larger dynamic range and contrast ratio. By way of example and not limitation, the minimum and maximum PQ values ​​that can be displayed by the reference image display without clipping may be given as P0 = L2PQ(0) and P1 = L2PQ(1000), respectively, where L2PQ(·) denotes the mapping function from (linear or non-PQ) luminance to (non-linear or PQ) luma codewords in the PQ domain or color space.

[0285] In a non-limiting example, the set M PQ is L on a logarithmic scale lower From L upper The lower bound L can be defined as a set of different PQ luma codewords corresponding to (linear) luminance values ​​evenly distributed within a (linear) luminance value range up to L. Depending in part or in whole on the image capture capabilities of the first and second devices (e.g., to capture the darkest and brightest luminances), lower is 10 -4 while the upper limit L upper is 10 2.99 (e.g., a fractional exponent value selected for numerical stability). A logarithmic scale is used because the luma codewords in captured SDR and HDR images can be approximately linear with respect to the logarithm of luminance in the range of luminance values. The total number of possible mean values, or magnitude |M PQ | is |M PQ It may be set as |=2048, but is not limited to this.

[0286] Block 384 involves defining or determining a set of possible shape coefficients to be used to generate possible variance values ​​for a statistical distribution or distribution type (eg, a beta distribution, etc.).

[0287] In some operating scenarios, the Beta distribution used to randomly select colors or codeword values ​​is defined or has support over a scaled or closed value interval [0,1]. Given PQ values ​​over the PQ codeword value range of [P0,P1], the scaled mean μ (0<μ<1 or within the closed value interval of the Beta distribution) may be calculated as follows: μ=(μ PQ -P0) / ((P1-P0), where μ PQ ∈M PQ (94)

[0288] Scaled value interval or closed value interval [0, Scaled variance (σ 2 ) can be adaptively set to be proportional to μ(1−μ) to avoid overexposure / underexposure caused by large variance (e.g., luminance, etc.) of color blocks during SDR and HDR image capture by the first and second devices, as follows: σ 2 =μ(1-μ) / θ (95) where θ denotes a shape factor that influences or determines the shape of the Beta distribution (e.g., more compressed, more expanded, etc.). The value of the shape factor can be selected from a set of possible shape factors Θ, allowing the generated color chart to have a relatively high diversity in codeword values ​​and / or colors resulting therefrom. A non-limiting example of a set of possible shape factors may be Θ={3, 6, 9, 12}.

[0289] Block 386 involves generating a plurality of all unique combinations of scaled means and shape factors using the set of possible means and the set of shape factors for different instances of the statistical (e.g., beta) distribution. In this example, the total number of unique combinations is |M PQ4=8192, where |Θ} denotes the total number of elements or magnitudes in the set Θ. In some operating scenarios, a different TIFF image containing a different color chart may be generated for each unique combination of scaled mean and shape factor in a plurality of all unique combinations of scaled mean and shape factor, thereby increasing the total number of different color charts or corresponding different TIFF images to |M PQ It is given as |×|Θ|.

[0290] Block 388 includes selecting the current combination of scaled mean and shape factor values ​​from among all the multiple unique combinations of scaled mean and shape factor values ​​to be the next combination, where the multiple different pixel or codeword values ​​or sets thereof respectively specify or define multiple different colors or sets thereof for the current color chart corresponding to the current combination of scaled mean and shape factor values, and the current scaled mean μ for the current combination and the current scaled variance σ for the current combination. 2 can be generated from the beta distribution given

[0291] Block 390 involves calculating or defining the beta distribution as follows: f(x;α,β)=(x α-1 (1-x) β-1 ) / B(α,β), x∈[0,1] (96) where α and β are the beta distribution parameters and B(α,β) is a normalization constant. The beta distribution parameters α and β can be derived from the mean and variance of the beta distribution as follows: α=μν and β=(1-μ)ν, ν=μ(1-μ)σ 2 -1 (97)

[0292] Block 392 involves generating a set of different colors (or pixel / codeword values) (e.g., 144, etc.) for the current color chart or current color chart image based on a set of random numbers (e.g., 144, etc.) generated from a beta distribution. Each random number, denoted as x (0≦x≦1), can be scaled back or converted to a pixel or codeword value (referred to as a "PQ value") for each corresponding channel within the PQ value range [P0,P1] as follows: x PQ =x(P1-P0)+P0(98)

[0293] The set of PQ values ​​generated from scaling or converting the set of random numbers back to the PQ value range can then be generated from the set of random numbers and used as the per-channel codeword values ​​or set of codewords for the set of colors in the current color chart. In some operating scenarios, for each color chart, the same beta distribution (with the same mean and shape coefficient) is used with, for example, three different sets of random numbers to generate per-channel codewords for each of multiple channels (e.g., RGB, etc.) of the color space in which the display image is to be rendered by the reference image display. Thus, the average color generated from averaging the colors represented in all color blocks in the color chart approaches a neutral or gray color.

[0294] Block 394 involves generating a central block of the current color chart or corresponding TIFF image. The current color chart may include a set of color blocks (e.g., 2D squares), each of which is a single color given by a particular (cross-channel or composite) pixel or codeword value with three per-channel pixel or codeword values. Three sets of per-channel codewords generated from the same beta distribution but with three different sets of random numbers may be used by the reference image display to drive the rendering of the red, blue, and green channels, respectively, of the color blocks in the current color chart or set of color blocks in the current TIFF image.

[0295] Additionally, optionally or alternatively, the RGB values ​​used to specify or generate the background color of the background of the color chart image may be calculated using the mean PQ value μ from which the scaled mean of the beta distribution is derived. PQ may be set to

[0296] Block 396 involves determining whether the current combination of scaled mean values ​​and shape factor values ​​is the last combination among all the multiple unique combinations of scaled mean values ​​and shape factors. If so, the process flow for generating multiple different color charts or different color chart images ends. If not, the process flow returns to block 388.

[0297] By way of example and not limitation, different colors or color blocks such as 144, c The total number of colors or color blocks generated from multiple different color charts and color chart images may be |M PQ |×|Θ|×n c= 2048 × 4 × 144 = 1179648 colors or color blocks, which can be used to determine correspondence / matching relationships between SDR color patches or codewords captured in SDR capture mode using a first video capture device and corresponding HDR color patches or codewords captured in HDR capture mode using a second video capture device.

[0298] In some operating scenarios, multiple TIFF images each containing multiple color charts can be repeatedly displayed or rendered on a reference image display, such as a PRM TV, at a constant playback frame rate and captured into SDR and HDR images by a first video capture device and a second video capture device, respectively. Because the illumination of the reference image display can vary with viewing angle, the SDR and HDR images can be separately captured by a first video capture device and a second video capture device (e.g., two phones) positioned at the same position and orientation with respect to or relative to the reference image display. Thus, SDR and HDR video signals or bitstreams containing captured SDR and HDR images of the color charts or TIFF images as rendered on the reference image display can be generated by the first and second devices or their cameras, respectively. Because the playback frame rate of the reference image display can be kept constant, the frame number of the color chart can be relatively easily determined to establish an image pair in which the SDR image captured by the first video capture device and the HDR image captured by the second video capture device depict the same color chart among the multiple color charts. The captured SDR and HDR images may be represented in the SDR and HDR YCbCr (or YUV) color space, for example in a YUV image / video file.

[0299] 3F shows an exemplary process flow for matching SDR and HDR colors between image pairs of a captured SDR image and an HDR image for each color chart among a plurality of color charts. To extract the colors of each individual color block from the captured SDR and HDR images of a color chart image that includes a color chart, the captured SDR and HDR images can be transformed into the same layout as the (original) color chart image displayed on the reference image display.

[0300] Block 3002 includes receiving a captured (SDR or HDR) checkerboard image, which may be taken by a camera of one of the first and second video capture devices from a displayed checkerboard image displayed on a reference image display. Blocks 3002-3006 of the same process flow of FIG. 3F may be performed with respect to a captured checkerboard image taken by a camera of the other of the first and second video capture devices from the same displayed checkerboard image displayed on a reference image display.

[0301] Block 3004 includes detecting checkerboard corners from a checkerboard image captured by a camera of the device.

[0302] Block 3006 includes calculating or calibrating camera distortion coefficients of the camera using the checkerboard corners detected from the captured checkerboard image. The 3D (reference) coordinates of the checkerboard image displayed on the reference image display may first be determined in a 3D coordinate system that is stationary relative to the reference image display. Camera parameters, such as distortion coefficients and other intrinsic parameters of the camera used by the device to generate the captured checkerboard image, can be calculated as optimized values ​​that produce the best mapping (e.g., minimum error or minimum discrepancy) between the 3D reference coordinates of the checkerboard corners and the 2D (image) coordinates of the checkerboard corners in the captured checkerboard image. This calibration process can be performed in either the YUV color space or the RGB color space to which the displayed or captured checkerboard image can be converted or represented.

[0303] 2M and 2N show two exemplary checkerboard images detected from captured (HDR and SDR, respectively) images of the checkerboard images. A plurality (e.g., 100, etc.) of captured checkerboard images can be used to determine, calculate, or calibrate distortion coefficients or other intrinsic parameters of the respective cameras used in the first and second video capture devices to achieve relatively high accuracy (e.g., within half a pixel of camera reprojection error) in projecting the captured images by the respective cameras of the displayed image on the reference image display onto the displayed image on the reference image display.

[0304] The (camera-specific) distortion coefficients and intrinsic parameters of the cameras of the first and second video capture devices obtained in processing blocks 3002-3006 can then be used to analyze or correlate between captured SDR and HDR images from TIFF images containing the respective color charts as follows:

[0305] Block 3008 includes receiving a captured (SDR or HDR) image of the TIFF image displayed on the reference image display, including a color chart and checkerboard corners (or a checkerboard pattern within the corners). The captured (SDR or HDR) image may be captured by a camera whose distortion coefficients were previously obtained in block 3006 using a previously captured checkerboard image (which may not include a color chart).

[0306] Block 3010 involves correcting or de-distorting the captured image of the displayed TIFF image to compensate for camera lens distortion of the camera, for example, using distortion coefficients obtained in a camera calibration process such as block 3006.

[0307] Block 3012 includes detecting checkerboard corners in the captured image, which are captured from checkerboard corners (e.g., four corners) in the TIFF image displayed on a reference image display, which may be a PPM TV.

[0308] Block 3014 involves estimating a projective transformation between the image coordinates of the captured image and the image coordinates of the original TIFF image displayed on the reference image display and captured in the received captured image.

[0309] Block 3016 includes using the estimated projective transformation to rectify the captured image to have the same layout as the original TIFF image. Using the estimated projective transformation, the captured image can be rectified to have the same spatial layout as the original TIFF image.

[0310] Figure 2O shows (a) a captured image from the (displayed) TIFF image, (b) a modified captured image generated by correcting the captured image using camera distortion correction and a projective transformation, and (c) the original TIFF color chart image displayed on a reference image display. Because the modified captured image in Figure 2O(b) is generated from a distortion removal or correction operation performed on the captured image in Figure 2O(a), some pixels in the modified image in Figure 2O(b) may be undetermined.

[0311] Block 3018 involves locating and extracting (sets of) individual color blocks from the color chart within the modified captured image. As used herein, a color block is designated using a single corresponding (e.g., RGB, YCbCr, composite, etc.) pixel or codeword value. Corresponding individual codewords that designate each individual color block in the captured image can be determined based on the pixel values ​​or codeword values ​​of the pixels within these individual color blocks in the modified captured image. Furthermore, individual original codewords that designate each individual original color block in the original TIFF image that corresponds to each individual color block of the modified captured image can be determined from the RGB or YUV file of the original TIFF image.

[0312] The correspondence / mapping relationship between the SDR color blocks or codewords extracted from the (modified) captured SDR image of the original color chart image and the HDR color blocks or codewords extracted from the (modified) captured HDR image of the same original color chart image can be established based in part or in whole on the individual original color blocks or codewords in the original TIFF image from which both the (modified) SDR image and the HDR image are derived.

[0313] Using a plurality of TIFF images, a plurality of SDR images captured by a first video capture device in SDR capture mode from the displayed TIFF images, and a plurality of HDR images captured by a second video capture device in HDR capture mode from the same displayed TIFF images, a plurality of correspondences or mapping relationships may be established between a set of SDR colors and a set of HDR images.

[0314] Several techniques may be used to map an SDR image to an HDR image.

[0315] In a first approach, in some "BB2" operating scenarios, such as in some "BB1" operating scenarios, the mapped SDR and HDR color patches or colors can be used to generate (e.g., static) SDR-HDR reshaping operating parameters such as TPB coefficients. These TPB coefficients may be combined with TPB basis functions using the SDR codeword of the SDR image as input parameters to predict the corresponding HDR codeword of the reconstructed HDR image. A static 3D-LUT may be pre-constructed and deployed on a video capture device (e.g., the first video capture device) to map the SDR image captured by the first video capture device to a reconstructed HDR image that simulates the HDR appearance of the second video capture device. Because the HDR color blocks or codewords used to generate the reshaping operation parameters are extracted from a captured (training) HDR image of the second video capture device, the predicted HDR codewords generated using the reshaping operation parameters are likely to provide a mapped HDR appearance in the reconstructed HDR image that is similar to the actual HDR appearance of the actual HDR image captured by the second video capture device.

[0316] In a second approach, in some "BB2" operating scenarios, non-TPB optimization may be used to generate non-TPB reshaping operation parameters used in non-TPB reshaping operations to map a captured SDR image taken by a first video capture device to a reconstructed HDR image that simulates the HDR appearance of a second video capture device. The 3D-LUT for the non-TPB reshaping operation may be constructed directly without the relatively high continuity and smoothness supported by the TPB reshaping operation. This non-TPB second approach offers relatively high design freedom and can be relatively flexible, whether or not different colors represented in nearby 3D-LUT entries / nodes may have relatively high continuity and smoothness. In some operating scenarios, the relatively high design freedom may make the non-TPB second approach more suitable than the first approach for two-device or "BB2" operating scenarios in which an SDR image from one device is mapped to an HDR image that simulates the HDR appearance of another device.

[0317] A second non-TPB approach may utilize a 3D mapping table (3DMT) to construct a backward lookup table (BLUT) representing a backward reshaping mapping. The BLUT may be used to map SDR (e.g., cross-channel, 3-channel, etc.) codewords of an SDR image to HDR (e.g., chroma channel, per-chroma channel, etc.) codewords of a reconstructed HDR image. Exemplary operations for constructing a BLUT-like backward reshaping mapping from a 3DMT can be found in U.S. Patent Application No. 17 / 054,495, filed May 9, 2019, the contents of which are incorporated herein by reference in their entirety as if fully set forth herein.

[0318] For purposes of example only, an SDR codeword may be represented in three dimensions or channels Y, Cb, and Cr in the SDR YCbCr color space. The 3D mapping table may be generated from a 3D histogram having multiple 3D histogram bins. The multiple 3D histogram bins may be represented by a set of three positive integers, Q y , Q Cb , Q Cr , whereby (Q y ×Q Cb ×Q Cr ) 3D histogram bins.

[0319] Multiple 3D histogram bins (Ω Q,s Each 3D histogram bin (denoted as q) in the 3D histogram (denoted as q) corresponds to a respective bin index or three respective quantized channel values ​​q=(q y ,q Cb ,q Cr ) and store pixel counts of all SDR color patches or codewords within the SDR color space partition represented by the 3D histogram bins. All bin indices (or quantized channel values) of multiple 3D histogram bins can be aggregated into a set of bin index values ​​denoted as Q, where Q=[Q y ,Q C0 ,Q Cr ].

[0320] Additionally, the sum of the HDR codewords mapped to the SDR codewords in each 3D histogram in the multiple 3D histograms may be calculated and stored for the 3D histogram bins. Y Q,v , Ψ Cb Q,v and Ψ Cr Q,v Let denote the sum of the HDR codewords (also called "mapped HDR luma and chroma values").

[0321] An exemplary procedure for generating SDR pixel counts and mapped HDR luma and chroma values ​​for 3D histogram bins of the histogram is shown in Table 2 below. [Table 2]

[0322] (s q y,(B) ,s q Cb,(B) ),s q Cr,(B) Let q denote the (representative) SDR codeword at the center of the qth 3D histogram bin. The representative SDR codewords for all 3D histogram bins are fixed for all SDR images / frames and can be pre-computed with the exemplary procedure shown in Table 3 below. [Table 3]

[0323] Next, 3D histogram bins with non-zero (SDR) pixel counts may be identified and retained in the plurality of 3D histogram bins, and all other 3D histogram bins with zero (SDR) pixel counts are discarded or removed from the plurality of 3D histogram bins.

[0324] q0, q1,.,q k-1 , each of which is a non-zero (SDR) pixel count or Ω q Q,s Let ∑ k = 1, ∑ k = 2, ∑ k = 3, ∑ k = 4, ∑ k = 5, ∑ k = 6, ∑ k = 7, ∑ k = 8, ∑ k = 9, ∑ k = 10, ∑ k = 11, ∑ k = 12, ∑ k = 13, ∑ k

number

number

[0325] As noted above, exemplary operations for constructing a 3D-LUT from a 3-DMT that functions as a non-TPB backward reshaping mapping or BLUT (e.g., for predicting HDR chroma codewords per chroma channel from cross-channel SDR codewords) are described in the above-referenced U.S. patent application Ser. No. 17 / 054,495.

[0326] In some "BB2" operating scenarios, luma inverse reshaping mapping may be used to inversely reshape an SDR luma codeword of an SDR image into a predicted or inversely reshaped HDR luma codeword of a corresponding HDR image using a GPR-based model. Exemplary generation of luma reshaping mapping using a GPR-based model can be found in U.S. Provisional Application No. 62 / 887,123, "Efficient user-defined SDR-to-HDR conversion with model templates," filed August 15, 2019, by Guan-Ming Su and Harshad Kadu, and PCT Application No. PCT / US2020 / 046032, filed August 12, 2020, the contents of which are incorporated herein by reference as if fully set forth herein. For example, a CDF matching curve may be generated from each training SDR-HDR image pair of multiple training SDR-HDR image pairs. Using a set of SDR points (e.g., 15, etc.) uniformly sampled across the SDR codeword range (e.g., the entire SDR codeword range, etc.), a corresponding mapped HDR codeword can be found in each of the CDF matching curves. For each sample SDR point in the set of uniformly sampled SDR points, a GPR model can be constructed based on a histogram having a histogram of multiple (e.g., 128, etc.) luma bins and used to generate a corresponding luma reshaping mapping.

[0327] [TPB Optimization in Editing] Image / video editing is a common application in video capture devices, such as mobile devices, to allow users to adjust color, contrast, brightness, or other user preferences. Image / video editing can be done or performed on the encoder side before the captured images are compressed into a (compressed) video signal or bitstream. Additionally, optionally or alternatively, image / video editing can be done or performed on the decoder side after the video signal bitstream is decoded or decompressed.

[0328] The image / video editing operations described herein can be performed in either or both the HDR domain and the SDR domain. Image / video editing operations in the HDR domain can be performed relatively easily. For example, after HDR content or images are edited, the edited HDR content or images can be passed as input or reference HDR images to a video pipeline that generates corresponding reshaped content or images or ISP SDR content or images that are encoded into a video signal or bitstream. In comparison, image / video editing operations in the SDR domain can be relatively difficult. While HDR-to-SDR mapping and / or SDR-to-HDR mapping can be designed or generated to ensure or enhance restorability between the SDR domain and the HDR domain, image / video editing operations performed using SDR images encoded into a video signal or bitstream are likely to pose difficulties or challenges to restoring or inversely reshaping the edited SDR images to generate reconstructed HDR images that approximate the reference HDR images.

[0329] FIG. 3G illustrates exemplary image / video editing operations at the encoder side performed using an upstream device, such as a video capture device, or a video encoder. As shown, an input HDR image may be represented in an input HDR domain or color space (denoted as “HLG YCbCr”). A TPB-based forward reshaping operation (denoted as “Forward TPB”) may be performed to forward reshape the input HDR image into a forward-reshaped SDR image represented in a forward-reshaped SDR domain or color space (denoted as “SDR YCbCr”). An image / video editing operation (denoted as “RGB domain editing”) may be performed to edit the forward-reshaped SDR image to generate an edited SDR image represented in an edited forward-reshaped SDR domain or color space (denoted as “Edited SDR YCbCr”). The edited SDR image represented in the edited forward-reshaped SDR domain or color space may be compressed or coded into a video signal or bitstream using one or more video codecs of the upstream device.

[0330] 3H illustrates exemplary decoder-side image / video editing operations performed using a downstream receiving device or video decoder. As shown, a video signal or bitstream is decoded by one or more video codecs in the downstream device into an SDR image represented in the SDR domain or color space (denoted as "SDR YCbCr"). An image / video editing operation (denoted as "RGB domain editing") may be performed to edit the SDR image to generate an edited SDR image represented in the edited forward-reshaped SDR domain or color space (denoted as "edited SDR YCbCr"). A TPB-based backward reshaping operation (denoted as "backward TPB") may be performed on the edited SDR image represented in the edited forward-reshaped SDR domain or color space to generate a backward-reshaped or reconstructed HDR image in the reconstructed HDR domain or color space (denoted as "PQ YCbCr").

[0331] As shown in Figures 3G and 3H, an SDR image encoded into a video signal at the encoder side or decoded from a video signal may be represented in an SDR domain or color space (e.g., forward reshaped), such as an SDR YCbCr domain or color space, in which a video codec can be programmed or developed to perform image processing operations relatively efficiently. By way of example and not limitation, image / video editing operations can be performed in an SDR RGB domain or color space, as shown in Figures 3G and 3H. RGB-to-YCbCr or YCbCr-to-RGB conversion may be performed through a conversion (e.g., standard-specified, etc.), such as an SMPTE range conversion.

[0332] As more colors are pushed into an SDR domain or color space, codewords in a forward-reshaped domain or color space, such as the SDR YCbCr color space, may exceed a codeword value range, such as the SMPTE range, in the SDR RGB color space where image / video editing operations are often performed. Applying a YCbCr-to-RGB transform that converts YCbCr codewords within the YCbCr codeword value range (normalized to a value range of [0,1]) of the SDR YCbCr color space to RGB codewords in the SDR RGB color space may cause some of the RGB codewords to exceed or surpass the SMPTE range (normalized to a value range of [0,1]) of the SDR RGB color space. While codeword values ​​outside these ranges can be clipped, the clipped codeword values ​​may not be recoverable in a backward reshaping operation (e.g., TPB-based, etc.) on the original, unclipped HDR codewords. Further operations described herein may be used to improve recoverability and reduce or avoid visual artifacts in image / video editing applications.

[0333] 3I-3M show several example solutions for clipping out-of-range codewords in image / video editing applications.

[0334] In some operating scenarios, as shown in Figure 3I, out-of-range SDR RGB codewords resulting from a YCbCr-to-RGB conversion (denoted as "YCbCr-to-RGB clipping conversion") performed on an SDR YCbCr image at the encoder or decoder side can be clipped to a valid or specified codeword value range such as [0,1], e.g., everything greater than 1 is hard clipped to 1, everything less than 0 is hard clipped to 0, etc. Image / video editing operations ("editing in the RGB domain") can then be performed using SDR RGB codewords within the valid or specified codeword value range in the SDR RGB color space. The edited SDR RGB codeword in the SDR RGB color space can be converted and clipped by an RGB-to-YCbCr conversion (denoted as "RGB-to-YCbCr clipping conversion") to a converted edited SDR YCbCr codeword (denoted as "edited SDR YCbCr") within a valid or specified codeword value range, such as [0,1], of the converted edited SDR YCbCr color space.

[0335] In some operating scenarios, as shown in FIG. 3J , out-of-range SDR RGB codewords generated from a YCbCr-to-RGB conversion (denoted as a “YCbCr-to-RGB non-clipping conversion”) performed on an SDR YCbCr image at the encoder or decoder side are not clipped to a valid or specified codeword value range, such as [0, 1]. An image / video editing operation (“editing in the RGB domain”) can then be performed using the unclipped (input) SDR RGB codewords within a valid or specified codeword value range in the SDR RGB color space to generate edited SDR codewords. The YCbCr-to-RGB conversion with the non-clipping operation allows the original information in the (input) SDR YCbCr codewords to be preserved in the (input) SDR RGB codewords received by the image / video editing operation. The image / video editing operation captures all possible actual codeword values ​​in the input SDR RGB codewords, including those less than 0 or greater than 1, e.g., a limited value range of

[0001] . As a result, some information in the input SDR YCbCr codeword can be preserved both in the input SDR RGB codeword and in the edited SDR RGB codeword. An SDR RGB boundary operation (denoted as "RGB boundary

[0001] clipping") can be performed, for example, optionally based on a user preference of a user operating the video editing application, to scale down or push an edited SDR RGB codeword that is outside the valid or specified range of [0,1] into the valid or specified range of [0,1], thereby generating a scaled edited SDR RGB codeword that is within the valid or specified range of [0,1].The scaled edited SDR RGB codeword in the SDR RGB color space can be converted and clipped to a transformed scaled edited SDR YCbCr codeword (denoted as "edited SDR YCbCr") within a valid or specified codeword value range, such as [0,1], of the transformed scaled edited SDR YCbCr color space by an RGB-to-YCbCr conversion (denoted as "RGB-to-YCbCr clipping conversion").

[0336] In some operating scenarios, as shown in FIG. 3K, out-of-range SDR RGB codewords generated from a YCbCr-to-RGB conversion (denoted as a "YCbCr-to-RGB non-clipping conversion") performed on an SDR YCbCr image at the encoder or decoder side are not clipped to a valid or specified codeword value range, such as [0, 1]. An image / video editing operation ("editing in the RGB domain") can then be performed using the unclipped (input) SDR RGB codewords within a valid or specified codeword value range in the SDR RGB color space to generate edited SDR codewords. The YCbCr-to-RGB conversion with the non-clipping operation allows the original information in the (input) SDR YCbCr codewords to be preserved in the (input) SDR RGB codewords received by the image / video editing operation. The image / video editing operation captures all possible actual codeword values ​​in the input SDR RGB codewords, including those less than 0 or greater than 1, e.g., a limited value range of

[0001] . As a result, some information in the input SDR YCbCr codeword can be preserved both in the input SDR RGB codeword and in the edited SDR RGB codeword. Unlike that shown in Figure 3J, in the operating scenario shown in Figure 3K, no SDR RGB boundary operation may be performed. The edited SDR RGB codeword in the SDR RGB color space can be converted and clipped to a converted edited SDR YCbCr codeword (denoted as "edited SDR YCbCr") within a valid or specified codeword value range, such as [0,1], in the converted edited SDR YCbCr color space by an RGB-to-YCbCr conversion (denoted as "RGB-to-YCbCr clipping conversion").

[0337] In some operating scenarios, as shown in FIG. 3L, out-of-range SDR RGB codewords generated from a YCbCr-to-RGB conversion (denoted as a "YCbCr-to-RGB non-clipping conversion") performed on an SDR YCbCr image at the encoder or decoder side are not clipped to a valid or specified codeword value range, such as [0, 1]. An image / video editing operation ("editing in the RGB domain") can then be performed using the unclipped (input) SDR RGB codewords within a valid or specified codeword value range in the SDR RGB color space to generate edited SDR codewords. The YCbCr-to-RGB conversion with the non-clipping operation allows the original information in the (input) SDR YCbCr codewords to be preserved in the (input) SDR RGB codewords received by the image / video editing operation. The image / video editing operation captures all possible actual codeword values ​​in the input SDR RGB codewords, including those less than 0 or greater than 1, e.g., a limited value range of

[0001] . As a result, some information in the input SDR YCbCr codeword can be preserved both in the input SDR RGB codeword and in the edited SDR RGB codeword. An SDR RGB boundary operation (denoted as an "RGB boundary clipping 3D-LUT") can be performed, for example, optionally based on the user preferences of a user operating the video editing application, to scale down or push edited SDR RGB codewords outside the (3D) boundary supported by the TPB-based reshaping operation into the boundary, thereby generating scaled edited SDR RGB codewords supported by the TPB-based reshaping operation or within a boundary clearly defined in the TPB-based reshaping operation. The boundary supported by the TPB-based reshaping operation does not have to be a regular shape or a cube simply defined by a fixed value range of [0, 1].The scaled edited SDR RGB codeword in the SDR RGB color space can be converted and clipped to a transformed scaled edited SDR YCbCr codeword (denoted as "edited SDR YCbCr") within a valid or specified codeword value range, such as [0,1], of the transformed scaled edited SDR YCbCr color space by an RGB-to-YCbCr conversion (denoted as "RGB-to-YCbCr clipping conversion").

[0338] The (input) SDR YCbCr image may be obtained from the original HDR image through forward reshaping or ISP processing using a known (e.g., white-box, ISP, etc.) HDR-to-SDR forward transform or mapping, and the boundaries may be determined and represented as a (TPB boundary clipping) 3D-LUT based partially or fully on the HDR-to-SDR forward transform or mapping and / or any applicable color space transformation matrix. For example, boundary pixel or codeword values ​​in the forward reshaped color space or ISP SDR color space can be determined using full grid sampling data covering the entire HDR domain or color space in which the original HDR image is represented.

[0339] After image / video editing operations ("editing in the RGB domain"), the edited SDR RGB codewords may be scaled or squeezed into 3D shapes delineated or enclosed by boundaries, or irregularly clipped. In some operating scenarios, the scaled edited SDR RGB codewords in the SDR RGB color space can be converted to transformed scaled edited SDR YCbCr codewords by an RGB-to-YCbCr conversion ("RGB-to-YCbCr clipping conversion") without further clipping by an RGB-to-YCbCr conversion ("RGB-to-YCbCr clipping conversion"), because the transformed scaled edited SDR YCbCr codewords are already positioned within corresponding boundaries in the SDR YCbCr color space supported by the TPB-based inverse reshaping operation. In these operating scenarios shown in FIG. 3J, the maximum number of colors can be preserved while avoiding the generation of color artifacts in the converted scaled edited SDR YCbCr codeword and in the inverse reshaped HDR codeword generated from the converted scaled edited SDR YCbCr codeword by the TPB-based inverse reshaping operation.

[0340] In some operating scenarios, as shown in FIG. 3M, out-of-range SDR RGB codewords generated from a YCbCr-to-RGB conversion (denoted as a "YCbCr-to-RGB non-clipping conversion") performed on an SDR YCbCr image at the encoder or decoder side are not clipped to a valid or specified codeword value range, such as [0, 1]. An image / video editing operation ("editing in the RGB domain") can then be performed using the unclipped (input) SDR RGB codewords within a valid or specified codeword value range in the SDR RGB color space to generate edited SDR codewords. The YCbCr-to-RGB conversion with the non-clipping operation allows the original information in the (input) SDR YCbCr codewords to be preserved in the (input) SDR RGB codewords received by the image / video editing operation. The image / video editing operation captures all possible actual codeword values ​​in the input SDR RGB codewords, including those less than 0 or greater than 1, e.g., a limited value range of

[0001] . As a result, some information in the input SDR YCbCr codeword can be preserved both in the input SDR RGB codeword and in the edited SDR RGB codeword. The edited SDR RGB codeword in the SDR RGB color space can be converted and clipped to a converted edited SDR YCbCr codeword (denoted as "edited SDR YCbCr") within a valid or specified codeword value range, such as [0,1], in the converted edited SDR YCbCr color space by an RGB-to-YCbCr conversion (denoted as "RGB-to-YCbCr clipping conversion").The transformed edited SDR YCbCr codeword (“edited SDR YCbCr”) in the transformed edited SDR YCbCr color space can be clipped by an SDR YCbCr boundary operation (denoted as “YCbCr boundary clipping 3D-LUT”), which, for example, can be optionally performed based on a user preference of a user operating the video editing application to scale down or push a transformed edited SDR YCbCr codeword outside a (3D) boundary supported by the TPB-based reshaping operation inside the boundary, thereby generating a scaled transformed edited SDR YCbCr codeword (“edited SDR YCbCr”) in the scaled transformed edited SDR YCbCr color space that is supported by or within a boundary well-defined in the TPB-based reshaping operation. The boundaries supported by TPB-based reshaping operations do not have to be regular shapes such as cubes with sides defined simply as the valid or specified (normalized) range of [0,1].

[0341] The (input) SDR YCbCr image may be obtained from the original HDR image through forward reshaping or ISP processing using a known (e.g., white-box, ISP, etc.) HDR-to-SDR forward transform or mapping, and the boundaries may be determined and represented in a (TPB boundary clipping) 3D-LUT based on the HDR-to-SDR forward transform or mapping and / or any applicable color space transformation matrix. For example, boundary pixel or codeword values ​​in the forward reshaped color space or ISP SDR color space can be determined using full grid sampling data covering the entire HDR domain or color space in which the original HDR image is represented.

[0342] [Construction of boundary clipping 3D-LUT and clipping] HDR-to-SDR forward reshaping or mapping, such as TPB-based forward reshaping, may be a nonlinear function. While input HDR codewords in an input HDR domain or color space may be well organized within a simple 3D cube, the boundaries of mapped SDR codewords generated from (nonlinear or TPB) HDR-to-SDR mapping may be relatively irregular, unlike a 3D cube.

[0343] To perform (TPB) boundary clipping for relatively irregular boundaries, a (TPB) boundary clipping 3D-LUT may be constructed. The 3D-LUT can perform a lookup using an SDR codeword as an input (or lookup key), and in response to determining that the SDR codeword is within the boundary, return the SDR codeword as a value. Alternatively, in response to determining that the SDR codeword is outside the boundary, the 3D-LUT can return a clipped SDR codeword that is different from the original SDR codeword, where the former is within the boundary.

[0344] In some operating scenarios, boundary clipping can be implemented at two levels: at the first level, regular clipping is performed using a range defined by minimum and maximum values ​​(or lower and upper limits) that define the (3D) codeword range; at the second level, irregular clipping is performed using a (TPB) boundary clipping 3D-LUT.

[0345] 3N shows an exemplary process flow for constructing a boundary clipping 3D-LUT (e.g., TPB). The left side of the process flow shown in FIG. 3N can be implemented or performed to construct a 3D boundary using alphaShape technology. The right side of the process flow shown in FIG. 3N can be implemented or performed to generate a 3D-LUT for irregular clipping using the constructed 3D boundary.

[0346] Block 3022 calculates a 3D uniform sampling grid or set of sampled values ​​in an input HDR domain or color space, such as the R.2020 (container) domain or color space.

number

[0347] The sampled values ​​in the R.2020 YCbCr color space HLG may be converted to corresponding values ​​in the R.2020 RGB color space HLG as shown in equation (44) above. The converted values ​​in the R.2020 RGB color space HLG are converted to the optimized value for the parameter a (a opt ) may be further converted to a corresponding value in a RGB color space HLG, corresponding to (a). The converted value in (a) RGB color space HLG may be clipped as shown in equation (46) above and converted to a corresponding clipped value in R.2020 RGB color space HLG as shown in equation (47) above. The clipped value in R.2020 RGB color space HLG derived in equation (47) above may be converted to a corresponding clipped value V in R.2020 YCbCr color space HLG as shown in equation (48) above. YCbCr (FL),(R2020) may be converted to

[0348] Block 3024 converts the clipped or constrained (HDR YCbC HLG) values ​​V YCbCr (FL),(R2020)The method includes applying TPB-based forward reshaping (referred to as "forward TPB" in FIG. 3N) using the clipped HDR values ​​V in the R.2020 YCbCr color space HLG as input parameters to the forward TPB basis functions. YCbCr (FL),(R2020) The forward generator matrix S (in equation (50) above) constructed by F ch is used with or multiplied by, as follows:

number

number

[0349] Block 3046 is as follows:

number

number

[0350] R.709 RGB values ​​converted from R.709 YCbCr values ​​may contain values ​​outside a valid or specified range, such as the value range [0,1].

[0351] For each channel, the minimum and maximum values ​​of the R.709 RGB values ​​given in equation (100) may be measured or determined as follows:

number

[0352] These extreme values ​​can be used as lower and upper bounds in block 3030 to construct or prepare a uniformly sampled 3D grid or set of sampling values ​​in the unclipped RGB domain or color space.

[0353] FIG. 2P shows an example distribution of R.709 RGB values ​​(in an SDR RGB image) converted from R.709 YCbCr values. As shown, the distribution of R.709 RGB values ​​in an SDR RGB image is an irregular shape other than a 3D cube. The irregular shape with a 3D boundary represents the space of maximum supported SDR RGB colors that can be mapped back to reconstructed HDR colors or inversely reshaped in the reconstructed HDR domain or color space without losing information. Any SDR RGB colors outside this irregular shape or range will suffer information loss in the inverse reshaping operation.

[0354] After image / video editing operations are performed on an SDR RGB image, the resulting (3D) codeword range or distribution of edited codewords or colors may be wider (e.g., much wider) than the irregular shape in the SDR RGB color space, thereby resulting in many SDR RGB codewords that are not defined or supported by the inverse reshaping operation.

[0355] Furthermore, it can be difficult to characterize, represent, or approximate the actual 3D boundary of an irregular shape using analytical formulas or multiple 2D planes that act to cut a 3D cube in RGB color space into the irregular shape.

[0356] As mentioned above, a two-level clipping solution may be used to clip the edited codeword back to the maximum supported RGB color space as represented by the irregular shape. More specifically, at the first level:

number

number

[0357] Block 3028 of FIG. 3N represents a set of SDR RGB codewords.

number

number

[0358] The alphaShape object created to represent an alphaShape, e.g., in a non-convex region, is a representation of an SDR RGB codeword.

number

number

[0359] The function that constructs the alphaShape is f αS (Φ,r α ), where r α represents the radius parameter.

[0360] The first input parameter to this function is a set of SDR RGB codewords.

number

number

[0361] Figures 2Q to 2T show different r α A set of SDR RGB points or codewords using

number

[0362] alphaShape provides a bounding polyhedron as a clipping boundary for irregular shapes, making it possible to determine whether a 3D point represented by an SDR RGB codeword is inside or outside the irregular shape.

[0363] I αS Let (αS,x) denote the binary function used to determine whether a 3D point or SDR RGB codeword (denoted as x) is inside the irregular shape. Assume that a returned binary value of "1" indicates inside the irregular shape and a returned binary value of "0" is outside.

[0364] NN αS Let (αS,x) denote an index function that receives a given query 3D point or SDR RGB codeword, such as x, as a second input parameter and returns the nearest neighbor in αS for the query 3D point or SDR RGB codeword x.

[0365] As mentioned above, a boundary clipping 3D-LUT can be constructed in the SDR RGB domain or color space to clip any SDR RGB codewords outside of an irregular shape representing the space of maximally supported colors by (TPB-based forward and backward) reshaping operations.

[0366] Block 3030 includes constructing the boundary clipping 3D-LUT as a full-grid 3D-LUT with multiple nodes / entries that contain the query SDR codeword as a lookup key and the returned SDR codeword of the query SDR codeword as a value.

[0367] The extrema determined for each color channel (in block 3026 of FIG. 3N)

number

number

number

[0368] Therefore, the total number of query SDR codewords represented in the 3D-LUT is N u =N R N G N B It may be given as:

[0369] Block 3032 includes beginning to perform a node processing loop for each node / entry in the 3D-LUT by selecting a current node / entry from among multiple nodes / entries in the 3D-LUT (e.g., in a sequential or non-sequential loop / iteration order, etc.).

[0370] For simplicity, (i,j,k) in the above equation (103) can be vectorized as p. The current node / entry may be the pth node / entry among multiple nodes / entries in the 3D-LUT. The lookup key of the pth node / entry is u in the LHS of the above equation (103). p The pth represented query may be represented by an SDR RGB codeword, denoted as

number

[0371] Block 3034 calculates the current or pth represented query SDR RGB codeword u for the current node / entry. p is the alphaShape αS generated in block 3028 (709) (or a shape constructed using the alpha shape construction function as shown in equation (102) above). p alphaShapeαS (709) In response to determining that the value is within the range 0 to 1000, process flow proceeds to block 3036. Otherwise, process flow proceeds to block 3040.

[0372] The plurality of nodes / entries in the 3D-LUT includes a subset of nodes or entries, each of which is an alphaShape αR(709) Thus, a subset of nodes or entries are represented by alphaShapeαR, each with a lookup key specified by the query SDR RGB codeword represented in (709) Contains the nodes or entries that are considered to be within the

[0373] Block 3040 calculates the alphaShape αR for the current node / entry. (709) The closest node / entry is found by finding the nearest represented query SDR RGB codeword among the subset of nodes / entries in

number

number

[0374] In some operating scenarios, the nearest represented query SDR RGB codeword for the current or pth represented query SDR RGB codeword is

number

number

[0375] Nearest represented query SDR RGB codeword

number

number

[0376] Block 3036 calculates the current or pth represented query SDR RGB codeword u p as a return value for the current or pth node entry in the 3D-LUT (in addition to being the lookup key), and determining whether the current or pth node / entry is the last node or entry among the multiple nodes / entries in the 3D-LUT. In response to determining that the current or pth node / entry is the last node or entry among the multiple nodes / entries in the 3D-LUT, process flow proceeds to block 3038. If not, process flow returns to block 3032.

[0377] Block 3038 converts the 3D-LUT into a final boundary clipping 3D-LUT (f 3DLUT αS This includes outputting it as (denoted as ().

[0378] Final boundary clipping 3D-LUT f 3DLUT αS An exemplary procedure for generating () is shown in Table 5 below. [Table 5]

[0379] (Final) boundary clipping 3D-LUT f 3DLUT αS Given (), boundary clipping can be performed relatively efficiently on the (e.g., edited, etc.) SDR image using the two-level solution described above. As described above, regular clipping can be performed to ensure that all (e.g., edited, etc.) SDR RGB codewords in the SDR image are within the extreme values ​​or upper / lower limits for the SDR RGB codewords in the SDR RGB color space in which the image / video editing operations on the SDR image were performed. Then, after regular clipping, the (final) boundary clipping 3D-LUT f is calculated using the (regularly clipped, if applicable) SDR codewords as lookup keys. 3DLUT αS (), boundary clipping, and the (final) boundary clipping 3D-LUT f 3DLUT αS Irregular clipping is performed by outputting the return value from () as the output (further irregularly clipped, if applicable) SDR codeword in the output (e.g., clipped, edited, etc.) SDR image.

[0380] In various operating scenarios, including but not limited to those illustrated in Figures 3I-3M, the regular and / or irregular clipping operations described herein can be applied in video capture and / or editing applications to ensure maximum support for reconstructing HDR images from SDR images and to prevent / reduce visual artifacts in the reconstructed HDR images.

[0381] [Example Process Flow] 4A illustrates an exemplary process flow according to an embodiment of the present invention. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. In block 402, the system described herein constructs sampled high dynamic range (HDR) color space points distributed across an HDR color space. The HDR color space is parameterized by primary color scaling parameters having candidate values ​​selected from among a plurality of candidate values. The primary color scaling parameters are used to calculate the color space coordinates of at least one of a plurality of primary colors describing the HDR color space.

[0382] In block 404, the system generates from the sampled HDR color space points in the HDR color space: (a) reference standard dynamic range (SDR) color space points represented in the reference SDR color space, (b) input HDR color space points represented in the input HDR color space, and (c) reference HDR color space points represented in the reference HDR color space points.

[0383] In block 406, the system executes a reshaping operation optimization algorithm to generate a chain of optimized forward reshaping mappings and optimized reverse reshaping mappings. The reshaping operation optimization algorithm uses the reference SDR color space points, the input HDR color space points, and the reference HDR color space points as inputs.

[0384] In an embodiment, an optimized forward reshaping mapping is used to forward reshape an input HDR image in an input HDR color space into a forward reshaped SDR image in a forward reshaped SDR color space, and an optimized backward reshaping mapping is used to backward reshape the forward reshaped SDR image in the forward reshaped SDR color space into a backward reshaped HDR image.

[0385] In an embodiment, the sampled HDR color space points are constructed in HDR color space without using any images.

[0386] In an embodiment, multiple chains of optimized forward reshaping mappings and optimized inverse reshaping mappings are generated by a reshaping operation optimization algorithm for multiple candidate values ​​of the primary color scaling parameter, and each chain in the multiple chains of optimized forward reshaping mappings and optimized inverse reshaping mappings includes a respective optimized forward reshaping mapping and a respective optimized inverse reshaping mapping.

[0387] In an embodiment, the sampled HDR color space points are mapped to reference SDR color space points based at least in part on a predetermined HDR-to-SDR mapping.

[0388] In an embodiment, a set of prediction errors is calculated for a plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings, each set of prediction errors in the set of prediction errors is calculated for a respective chain in the plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings, and the set of prediction errors is used to select a particular candidate value from a plurality of candidate values ​​of the primary color scaling parameters.

[0389] In an embodiment, a particular candidate value of the primary color scaling parameters is used to generate a particular chain of a particular optimized forward reshaping mapping and a particular optimized inverse reshaping mapping.

[0390] In an embodiment, the particular optimized forward reshaping mapping is represented in a forward reshaping three-dimensional lookup table.

[0391] In an embodiment, the specific optimized inverse reshaping mapping is represented in an inverse reshaping three-dimensional lookup table.

[0392] In an embodiment, a video encoder applies an optimized forward reshaping mapping to a sequence of input HDR images to generate a sequence of forward reshaped SDR images, and encodes the sequence of forward reshaped SDR images into a video signal.

[0393] In an embodiment, a video decoder decodes a sequence of forward-reshaped SDR images from a video signal and applies an optimized backward reshaping mapping to the sequence of forward-reshaped SDR images to generate a sequence of backward-reshaped HDR images.

[0394] In an embodiment, a sequence of display images derived from the sequence of inversely reshaped HDR images is rendered on an image display operating in conjunction with a video decoder.

[0395] In an embodiment, the HDR color space and the input HDR color space share a common white point.

[0396] In an embodiment, the reshaping operation optimization algorithm represents a Backward-Error-Subtraction-for-Signal-Adjustment (BESA) algorithm.

[0397] 4B illustrates an exemplary process flow according to an embodiment of the present invention. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. In block 422, the systems described herein construct sampled HDR color space points distributed across the HDR color space. The HDR color space is parameterized by primary color scaling parameters having candidate values ​​selected from among a plurality of candidate values. The primary color scaling parameters are used to calculate the color space coordinates of at least one of a plurality of primary colors that describe the HDR color space.

[0398] In block 424, the system generates from the sampled HDR color space points in the HDR color space: (a) input SDR color space points represented in the input SDR color space, and (b) reference HDR color space points represented in the reference HDR color space points.

[0399] In block 426, the system executes a reshaping operation optimization algorithm to generate an optimized inverse reshaping mapping. The reshaping operation optimization algorithm receives as input the input SDR color space points and the reference HDR color space points.

[0400] In an embodiment, an inverse reshaping mapping is used to inversely reshape an SDR image in an input SDR color space into an inversely reshaped HDR image.

[0401] In an embodiment, the sampled HDR color space points are constructed in HDR color space without any images.

[0402] In an embodiment, a plurality of optimized inverse reshaping mappings are generated by a reshaping operation optimization algorithm for a plurality of candidate values ​​of the primary color scaling parameter, and each optimized inverse reshaping mapping in the plurality of optimized inverse reshaping mappings comprises a respective optimized inverse reshaping mapping.

[0403] In an embodiment, a plurality of sets of prediction errors are calculated for a plurality of optimized inverse reshaping mappings, each set of prediction errors in the plurality of sets of prediction errors is calculated for a respective optimized inverse reshaping mapping in the plurality of optimized inverse reshaping mappings, and the plurality of sets of prediction errors are used to select a particular candidate value from a plurality of candidate values ​​of the primary color scaling parameter.

[0404] In an embodiment, the sampled HDR color space points are processed by the programmable ISP pipeline into input SDR color space points based at least in part on optimized values ​​of programmable configuration parameters of the programmable ISP pipeline.

[0405] In an embodiment, optimized values ​​of the programmable configuration parameters of the programmable ISP pipeline are determined by minimizing the approximation error between an ISP SDR image generated by the programmable ISP pipeline from an HDR image and a reference SDR image generated by applying a predetermined HDR-SDR mapping to the same HDR image.

[0406] 4C illustrates an exemplary process flow according to an embodiment of the present invention. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. At block 442, the system described herein extracts a set of SDR image feature points from the training SDR image and a set of HDR image feature points from the training HDR image.

[0407] In block 444, the system matches a subset of one or more SDR image features in the set of SDR image features with a subset of one or more HDR image features in the set of HDR image features.

[0408] In block 446, the system generates a geometric transformation to spatially align a set of SDR pixels in the training SDR image with a set of HDR pixels in the training HDR image using a subset of one or more SDR image feature points and a subset of one or more HDR image feature points.

[0409] In block 448, the system determines a set of SDR color patch and HDR color patch pairs from the set of SDR pixels in the training SDR image and the set of HDR pixels in the training HDR image after the training SDR image and the training HDR image are spatially aligned by a geometric transformation.

[0410] At block 450, the system generates an optimized SDR-HDR mapping based at least in part on a set of SDR and HDR color patch pairs derived from the training SDR and HDR images.

[0411] In block 452, the system applies the optimized SDR-HDR mapping to one or more non-training SDR images to generate one or more corresponding non-training HDR images.

[0412] In an embodiment, training SDR images and training HDR images are captured from a three-dimensional (3D) visual scene by a capture device operating in an SDR capture mode and an HDR capture mode, respectively.

[0413] In an embodiment, the training SDR image and the training HDR image form a pair of training SDR image and training HDR image among a plurality of pairs of training SDR image and training HDR image, and the optimized SDR-HDR mapping is generated based at least in part on a set of a plurality of SDR color patch and HDR color patch pairs derived from the plurality of pairs of training SDR image and training HDR image.

[0414] In an embodiment, each SDR image feature point in the subset of one or more SDR image features is matched with a respective HDR image feature point in the subset of one or more HDR image features, and the SDR image features and HDR image features are extracted from the training SDR image and the training HDR image, respectively, using a common feature point extraction algorithm.

[0415] In an embodiment, the common feature point extraction algorithm represents one of a binary robust invariant scalable keypoint algorithm, an accelerated feature from segment test algorithm, a KAZE algorithm, a minimum eigenvalue algorithm, a maximum stable extremum region algorithm, a fast orientation rotation algorithm, a scale invariant feature transformation algorithm, an accelerated robust feature algorithm, etc.

[0416] 4D shows an exemplary process flow according to an embodiment of the present invention. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. At block 462, the system described herein performs a respective camera distortion correction operation on each training image in the training SDR image and training HDR image pair to generate a respective undistorted image in the undistorted training SDR image and undistorted training HDR image pair.

[0417] In block 464, the system generates each projective transform in the pair of SDR image projective transform and HDR image projective transform using the corner pattern marks detected from each undistorted image in the pair of undistorted training SDR image and undistorted training HDR image.

[0418] In block 466, the system applies each projective transformation in the pair of SDR image projective transformation and HDR image projective transformation to a respective undistorted image in the pair of undistorted training SDR image and undistorted training HDR image to generate a respective modified image in the pair of modified training SDR image and modified training HDR image.

[0419] In block 468, the system extracts a set of SDR color patches from the modified training SDR image and a set of HDR color patches from the modified training HDR image.

[0420] In block 470, the system generates an optimized SDR-HDR mapping based at least in part on the set of SDR and HDR color patches derived from the training SDR and HDR images.

[0421] In block 472, the system applies the optimized SDR-HDR mapping to one or more non-training SDR images to generate one or more corresponding non-training HDR images.

[0422] In an embodiment, training SDR images and training HDR images are captured from a common color target image by a first capture device operating in an SDR capture mode and a second capture device operating in an HDR capture mode, respectively.

[0423] In an embodiment, a common color chart image is selected from a plurality of color chart images, each color chart image comprising a separate distribution of color patches arranged on a two-dimensional color chart.

[0424] In an embodiment, the distinct distributions of color patches are generated using random colors randomly selected from a common statistical distribution having a particular combination of statistical mean and variance.

[0425] In an embodiment, a common color chart image is rendered on a first capture device and a second capture device and captured from the screen of a common reference image display.

[0426] In an embodiment, each camera distortion correction operation is based at least in part on camera-specific distortion coefficients generated from a camera calibration process performed with the camera used to acquire the training images.

[0427] In an embodiment, the set of SDR color patches and the set of HDR color patches are used to derive a three-dimensional mapping table (3DMT), and an optimized SDR-HDR mapping is generated based at least in part on the 3DMT.

[0428] In an embodiment, the optimized SDR-HDR mapping represents one of a TPB-based mapping or a non-TPB-based mapping.

[0429] In an embodiment, the optimized SDR-HDR mapping is one of a static mapping applied to all non-training SDR images represented in the video signal, or a dynamic mapping generated based at least in part on the distribution of particular values ​​of SDR codewords of non-training SDR images among the non-training SDR images represented in the video signal.

[0430] 4E illustrates an exemplary process flow according to an embodiment of the present invention. In some embodiments, one or more computing devices or components (e.g., an encoding device / module, a transcoding device / module, a decoding device / module, an inverse tone mapping device / module, a tone mapping device / module, a media device / module, an inverse mapping generation and application system, etc.) may perform this process flow. In block 482, the system described herein constructs sampled HDR color space points distributed across the HDR color space used to represent the reconstructed HDR image.

[0431] In block 484, the system converts the sampled HDR color space points to SDR color space points in a first SDR color space in which the SDR image to be edited by the editing device is represented.

[0432] In block 486, the system determines a bounding SDR color space rectangle based on extreme SDR codeword values ​​of SDR color space points in the first SDR color space and determines an irregular three-dimensional shape from the distribution of SDR color space points.

[0433] In block 488, the system constructs sampled SDR color space points distributed over a bounded SDR color space rectangle in the first SDR color space.

[0434] In block 490, the system uses the sampled SDR color space points and the irregular shape to generate a boundary clipping 3D-LUT, which uses the sampled SDR color space points as lookup keys.

[0435] In block 492, the system performs a clipping operation on the edited SDR image in the first SDR color space based at least in part on the boundary clipping 3D-LUT to generate a boundary-clipped edited SDR image in the first SDR color space.

[0436] In an embodiment, the clipping operation includes first performing regular clipping on the edited SDR image using a bounded SDR color space rectangle to generate a regularly clipped edited SDR image, and then performing irregular clipping on the regularly clipped edited SDR image using a 3D-LUT to generate a bounded clipped edited SDR image.

[0437] In an embodiment, a set of one or more SDR pixels in an SDR image to be edited are edited from one or more first luminance values ​​to one or more second luminance values ​​in the edited image, the one or more second luminance values ​​being different from the one or more first luminance values.

[0438] In an embodiment, a set of one or more SDR pixels in an SDR image to be edited are edited in the edited image from one or more first color difference values ​​to one or more second color difference values, the one or more second color difference values ​​being different from the one or more first color difference values.

[0439] In an embodiment, image details that are shown in the SDR image to be edited are removed in the edited SDR image.

[0440] In an embodiment, image details not depicted in the SDR image to be edited are added to the edited SDR image.

[0441] In an embodiment, the 3D-LUT includes one or more nodes, each node including a lookup key and a lookup value, the lookup key being equal to the lookup value, and the lookup key being inside an irregular shape.

[0442] In an embodiment, the 3D-LUT includes one or more nodes, each node including a lookup key and a lookup value, the lookup key being outside the irregular shape and the lookup value being inside the irregular shape.

[0443] In an embodiment, the lookup value is determined based on an exponential function that takes the irregular shape and the lookup key as input and returns the closest neighbor to the lookup key as output.

[0444] In embodiments, a computing device such as a display device, a mobile device, a set-top box, a multimedia device, etc. is configured to perform any of the above methods. In embodiments, an apparatus includes a processor and is configured to perform any of the above methods. In embodiments, a non-transitory computer-readable storage medium is provided that stores software instructions, which when executed by one or more processors, cause any of the above methods to be performed.

[0445] In one embodiment, a computing device includes one or more processors and one or more storage media that store a set of instructions that, when executed by the one or more processors, cause any of the methods described above to be performed.

[0446] Although separate embodiments are discussed herein, it is noted that any combination of the embodiments and / or sub-embodiments discussed herein may be combined to form further embodiments.

[0447] Exemplary Computer System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprised of electronic circuits and components, integrated circuit (IC) devices such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or apparatuses including one or more of such systems, devices, or components. The computers and / or ICs may execute, control, or perform instructions related to adaptive perceptual quantization of images with extended dynamic range, such as those described herein. The computers and / or ICs may calculate any of the various parameters or values ​​associated with the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0448] Certain implementations of the present invention include computer processors that execute software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. may implement the methods for adaptive perceptual quantization of HDR images, as described above, by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention may be in any of a variety of formats. The program product may include physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, electronic data storage media including flash RAM, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0449] Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including reference to "means") should be interpreted as including any component that performs the function of the described component (e.g., is functionally equivalent) as an equivalent of that component, including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the present invention.

[0450] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hardwired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) permanently programmed to perform the techniques, or may include one or more general-purpose hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to achieve the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices incorporating hardwired and / or program logic to implement the techniques.

[0451] 5 is a block diagram illustrating a computer system 500 upon which embodiments of the present invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled to bus 502 for processing information. Hardware processor 504 may be, for example, a general-purpose microprocessor.

[0452] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions executed by processor 504. Main memory 506 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored on a non-transitory storage medium accessible to processor 504, render computer system 500 a special-purpose machine customized to perform the operations specified in the instructions.

[0453] Computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions.

[0454] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.

[0455] Computer system 500 may implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, makes computer system 500 a special-purpose machine or programs it to be a special-purpose machine. According to one embodiment, the techniques described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0456] The term "storage medium" as used herein refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific fashion. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, or any other memory chip or cartridge.

[0457] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media involves transferring information to and from storage media. For example, transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0458] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.

[0459] Computer system 500 also includes a communication interface 518 coupled to bus 502. The communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0460] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 provides data communication services through the worldwide packet data communication network now commonly referred to as the "Internet" 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals through communication interface 518 on network link 520, which carry the digital data to and from computer system 500, are exemplary forms of transmission media.

[0461] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.

[0462] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution.

[0463] [Equivalents, Extensions, Substitutes and Others] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what the claimed embodiments of the invention are and are intended by the applicant to be the claimed embodiments of the invention is the set of claims issuing from this application, in the particular form in which such claims are issued, including any subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall control the meaning of such terms as used in the claims. Accordingly, no limitations, elements, properties, features, advantages, or attributes not expressly recited in a claim should in any way limit the scope of such claims. Accordingly, the specification and drawings are to be regarded in an illustrative sense, and not in a restrictive sense.

[0464] Enumerated exemplary embodiments The present invention may be embodied in any of the forms described herein, including, but not limited to, the following enumerated exemplary embodiments (EEE), which describe the structure, features, and functions of some portions of the embodiments of the present invention.

[0465] EEE1. constructing sampled high dynamic range (HDR) color space points distributed across an HDR color space, the HDR color space being parameterized by primary color scaling parameters having candidate values ​​selected from a plurality of candidate values, the primary color scaling parameters being used to calculate color space coordinates of at least one of a plurality of primary colors describing the HDR color space; generating, from the sampled HDR color space points in the HDR color space, (a) reference standard dynamic range (SDR) color space points represented in a reference SDR color space, (b) input HDR color space points represented in an input HDR color space, and (c) reference HDR color space points represented in a reference HDR color space point; running a reshaping operation optimization algorithm to generate a chain of optimized forward reshaping mappings and optimized reverse reshaping mappings, the reshaping operation optimization algorithm using the reference SDR color space points, the input HDR color space points, and the reference HDR color space points as inputs; Including, the optimized forward reshaping mapping is used to forward reshape an input HDR image in the input HDR color space into a forward reshaped SDR image in a forward reshaped SDR color space, and the optimized backward reshaping mapping is used to backward reshape the forward reshaped SDR image in the forward reshaped SDR color space into a backward reshaped HDR image.

[0466] EEE2. The method of EEE1, wherein the sampled HDR color space points are constructed in the HDR color space without using any images.

[0467] EEE3. The method of any one of EEE1 and EEE2, wherein a plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings are generated by the reshaping operation optimization algorithm for the plurality of candidate values ​​of the primary color scaling parameter, each chain in the plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings including a respective optimized forward reshaping mapping and a respective optimized inverse reshaping mapping.

[0468] EEE4. 8. The method of any one of EEE1 to EEE3, wherein the sampled HDR color space points are mapped to the reference SDR color space points based at least in part on a pre-defined HDR-to-SDR mapping.

[0469] EEE5. 8. The method of any one of EEE1 to EEE4, wherein a plurality of sets of prediction errors are calculated for a plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings, each set of prediction errors being calculated for a respective chain of the plurality of chains of optimized forward reshaping mappings and optimized inverse reshaping mappings, and the plurality of sets of prediction errors are used to select a particular candidate value from the plurality of candidate values ​​of the primary color scaling parameter.

[0470] EEE6. The method according to EEE5, wherein the particular candidate values ​​of the primary color scaling parameters are used to generate a particular chain of a particular optimized forward reshaping mapping and a particular optimized inverse reshaping mapping.

[0471] EEE7. The method of EEE6, wherein the specific optimized forward reshaping mapping is represented in a forward reshaping three-dimensional lookup table.

[0472] EEE8. The method according to EEE6, wherein the specific optimized inverse reshaping mapping is represented in an inverse reshaping three-dimensional lookup table.

[0473] EEE9. 9. The method of any one of EEE6 to EEE8, wherein a video encoder applies the optimized forward reshaping mapping to a sequence of input HDR images to generate a sequence of forward reshaped SDR images, and encodes the sequence of forward reshaped SDR images into a video signal.

[0474] EEE10. 10. The method of any one of EEE6 to EEE9, wherein a video decoder decodes a sequence of forward reshaped SDR images from a video signal and applies the optimized inverse reshaping mapping to the sequence of forward reshaped SDR images to generate a sequence of inverse reshaped HDR images.

[0475] EEE11. The method of EEE10, wherein a sequence of display images derived from the sequence of inversely reshaped HDR images is rendered on an image display operating in conjunction with the video decoder.

[0476] EEE12. The method of any one of EEE1 to EEE10, wherein the HDR color space and the input HDR color space share a common white point.

[0477] EEE13. The method of any one of EEE1 to EEE10, wherein the reshaping performance optimization algorithm represents a Backward-Error-Subtraction-for-signal-Adjustment (BESA) algorithm with neutral color preservation.

[0478] EEE14. constructing sampled high dynamic range (HDR) color space points distributed across an HDR color space, the HDR color space being parameterized by primary color scaling parameters having candidate values ​​selected from a plurality of candidate values, the primary color scaling parameters being used to calculate color space coordinates of at least one of a plurality of primary colors describing the HDR color space; generating, from the sampled HDR color space points in the HDR color space, (a) input standard dynamic range (SDR) color space points represented in an input SDR color space, and (b) reference HDR color space points represented in a reference HDR color space point; executing a reshaping operation optimization algorithm to generate an optimized inverse reshaping mapping, the reshaping operation optimization algorithm receiving as input the input SDR color space points and the reference HDR color space points; Including, The method, wherein the inverse reshaping mapping is used to inversely reshape an SDR image in the input SDR color space into an inversely reshaped HDR image.

[0479] EEE16. The method of any one of EEE14 and EEE15, wherein a plurality of optimized inverse reshaping mappings are generated by the reshaping operation optimization algorithm for the plurality of candidate values ​​of the primary color scaling parameter, each optimized inverse reshaping mapping in the plurality of optimized inverse reshaping mappings comprising a respective optimized inverse reshaping mapping.

[0480] EEE17. The method of any one of EEE14 to EEE16, wherein a plurality of sets of prediction errors are calculated for the plurality of optimized inverse reshaping mappings, each set of prediction errors in the plurality of sets of prediction errors is calculated for a respective optimized inverse reshaping mapping in the plurality of optimized inverse reshaping mappings, and the plurality of sets of prediction errors are used to select a particular candidate value from the plurality of candidate values ​​for the primary color scaling parameter.

[0481] EEE18. 8. The method of any one of EEE14 to EEE17, wherein the sampled HDR color space points are processed into the input SDR color space points by a programmable image signal processor (ISP) pipeline based at least in part on optimized values ​​of programmable configuration parameters of the ISP pipeline.

[0482] EEE19. 9. The method of any one of EEE14 to EEE18, wherein the optimized values ​​of the programmable configuration parameters of the programmable ISP pipeline are determined by minimizing an approximation error between an ISP SDR image generated by the programmable ISP pipeline from an HDR image and a reference SDR image generated by applying a predefined HDR-SDR mapping to the same HDR image.

[0483] EEE20. Extracting a set of standard dynamic range (SDR) image feature points from training SDR images and a set of high dynamic range (HDR) image feature points from training HDR images; matching a subset of one or more SDR image features in the set of SDR image features with a subset of one or more HDR image features in the set of HDR image features; generating a geometric transformation using the subset of one or more SDR image feature points and the subset of one or more HDR image feature points to spatially align a set of SDR pixels in the training SDR image with a set of HDR pixels in the training HDR image; determining a set of pairs of SDR color patches and HDR color patches from the set of SDR pixels in the training SDR image and the set of HDR pixels in the training HDR image after the training SDR image and the training HDR image are spatially aligned by the geometric transformation; generating an optimized SDR-HDR mapping based at least in part on the set of SDR and HDR color patch pairs derived from the training SDR image and the training HDR image; applying the optimized SDR-HDR mapping to one or more non-training SDR images to generate one or more corresponding non-training HDR images; A method comprising:

[0484] EEE21. The method of EEE20, wherein the training SDR images and the training HDR images are captured from a three-dimensional (3D) visual scene by a capture device operating in an SDR capture mode and an HDR capture mode, respectively.

[0485] EEE22. The method of any one of EEE20 to EEE21, wherein the training SDR image and the training HDR image form a pair of training SDR and HDR images among a plurality of pairs of training SDR and training HDR images, and the optimized SDR-HDR mapping is generated based at least in part on a set of a plurality of SDR and HDR color patch pairs derived from the plurality of pairs of training SDR and training HDR images.

[0486] EEE23. 3. The method of any one of EEE20 to EEE22, wherein each SDR image feature point in the subset of one or more SDR image features is matched with a respective HDR image feature point in the subset of one or more HDR image features, and the SDR image features and the HDR image features are extracted from the training SDR image and the training HDR image, respectively, using a common feature point extraction algorithm.

[0487] EEE24. The method according to EEE23, wherein the common feature point extraction algorithm represents one of a binary robust invariant scalable keypoint algorithm, an accelerated feature from segment test algorithm, a KAZE algorithm, a minimum eigenvalue algorithm, a maximum stable extremum region algorithm, a fast orientation rotation algorithm, a scale invariant feature transformation algorithm, or an accelerated robust feature algorithm.

[0488] EEE25. performing a respective camera distortion correction operation on each training image in the pair of training SDR and training HDR images to generate a respective undistorted image in the pair of undistorted training SDR and training HDR images; generating a projective transformation in a pair of SDR image projective transformation and HDR image projective transformation using the corner pattern marks detected from each undistorted image in the pair of undistorted training SDR image and undistorted training HDR image; applying each projective transformation in the pair of SDR image projective transformation and HDR image projective transformation to a respective undistorted image in the pair of undistorted training SDR image and undistorted training HDR image to generate a respective modified image in the pair of modified training SDR image and modified training HDR image; extracting a set of SDR color patches from the modified training SDR image and a set of HDR color patches from the modified training HDR image; generating an optimized SDR-HDR mapping based at least in part on the set of SDR color patches and the set of HDR color patches derived from the training SDR image and the training HDR image; applying the optimized SDR-HDR mapping to one or more non-training SDR images to generate one or more corresponding non-training HDR images; A method comprising:

[0489] EEE26. 8. The method of claim 8, wherein the training SDR images and the training HDR images are captured from a common color target image by a first capture device operating in an SDR capture mode and a second capture device operating in an HDR capture mode, respectively.

[0490] EEE27. The method of EEE26, wherein the common color chart image is selected from a plurality of color chart images, each color chart image comprising a separate distribution of color patches arranged on a two-dimensional color chart.

[0491] EEE28. The method according to EEE27, wherein the distinct distributions of color patches are generated using random colors randomly selected from a common statistical distribution having a particular combination of statistical mean and variance.

[0492] EEE29. The method of any one of EEE25 to EEE28, wherein the common color chart image is rendered on the first capture device and the second capture device and captured from a screen of a common reference image display.

[0493] EEE30. The method of any one of EEE25 to EEE29, wherein the each camera distortion correction operation is based at least in part on camera-specific distortion coefficients generated from a camera calibration process performed with the camera used to acquire the training images.

[0494] EEE31. 10. The method of any one of EEE25 to EEE30, wherein the set of SDR color patches and the set of HDR color patches are used to derive a three-dimensional mapping table (3DMT), and the optimized SDR-HDR mapping is generated based at least in part on the 3DMT.

[0495] EEE32. The method of any one of EEE25 to EEE31, wherein the optimized SDR-HDR mapping represents one of a Tensor Product B-Spline (TPB) based mapping or a non-TPB based mapping.

[0496] EEE33. 10. The method of any one of EEE25 to EEE32, wherein the optimized SDR-HDR mapping is one of a static mapping applied to all non-training SDR images represented in the video signal, or a dynamic mapping generated based at least in part on a distribution of particular values ​​of SDR codewords of non-training SDR images among the non-training SDR images represented in the video signal.

[0497] EEE34. constructing sampled high dynamic range (HDR) color space points distributed across an HDR color space used to represent a reconstructed HDR image; converting the sampled HDR color space points to SDR color space points in a first standard dynamic range (SDR) color space representing an SDR image to be edited by an editing device; determining a bounded SDR color space rectangle based on extreme SDR codeword values ​​of the SDR color space points in the first SDR color space and determining an irregular three-dimensional shape from the distribution of the SDR color space points; constructing sampled SDR color space points distributed across the bounded SDR color space rectangle in the first SDR color space; generating a boundary clipping 3D-LUT using the sampled SDR color space points and the irregular shape, the boundary clipping 3D-LUT using the sampled SDR color space points as lookup keys; performing a clipping operation on the edited SDR image in the first SDR color space based at least in part on the boundary clipping 3D-LUT to generate a boundary-clipped edited SDR image in the first SDR color space; A method comprising:

[0498] EEE35. 8. The method of claim 6, wherein the clipping operation comprises first performing regular clipping on the edited SDR image using the bounded SDR color space rectangle to generate a regular clipped edited SDR image, and then performing irregular clipping on the regular clipped edited SDR image using the 3D-LUT to generate the bounded clipped edited SDR image.

[0499] EEE36. The method of any one of EEE34 and EEE35, wherein a set of one or more SDR pixels in the SDR image to be edited are edited in the edited image from one or more first luminance values ​​to one or more second luminance values, the one or more second luminance values ​​being different from the one or more first luminance values.

[0500] EEE37. 10. The method of any one of EEE34 to EEE36, wherein a set of one or more SDR pixels in the SDR image to be edited are edited in the edited image from one or more first chrominance values ​​to one or more second chrominance values, the one or more second chrominance values ​​being different from the one or more first chrominance values.

[0501] EEE38. The method of any one of EEE34 to EEE37, wherein image details that are shown in the SDR image to be edited are removed in the edited SDR image.

[0502] EEE39. The method of any one of EEE34 to EEE38, wherein image details not depicted in the SDR image to be edited are added to the edited SDR image.

[0503] EEE40. The method of any one of EEE34 to EEE40, wherein the 3D-LUT comprises one or more nodes, each node comprising a lookup key and a lookup value, the lookup key being equal to the lookup value, and the lookup key being inside the irregular shape.

[0504] EEE41. 3. The method of any one of EEE34 to EEE40, wherein the 3D-LUT comprises one or more nodes, each node comprising a lookup key and a lookup value, the lookup key being outside the irregular shape and the lookup value being inside the irregular shape.

[0505] EEE42. The lookup value is determined based on an exponential function that takes the irregular shape and the lookup key as input and returns the nearest neighbor of the lookup key as output.

[0506] EEE43. 10. An apparatus including a processor and configured to perform the method of any one of EEE1 to EEE42.

[0507] EEE44. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing, with one or more processors, the method of any one of EEE1 to EEE42.

[0508] EEE45. A computer system configured to carry out the method of any one of EEE1 to EEE42.

Claims

1. Extracting a set of standard dynamic range (SDR) image feature points from training SDR images and a set of high dynamic range (HDR) image feature points from training HDR images; matching a subset of one or more SDR image features in the set of SDR image features with a subset of one or more HDR image features in the set of HDR image features; generating a geometric transformation using the subset of one or more SDR image feature points and the subset of one or more HDR image feature points to spatially align a set of SDR pixels in the training SDR image with a set of HDR pixels in the training HDR image; determining a set of pairs of SDR color patches and HDR color patches from the set of SDR pixels in the training SDR image and the set of HDR pixels in the training HDR image after the training SDR image and the training HDR image are spatially aligned by the geometric transformation; generating an optimized SDR-HDR mapping based at least in part on the set of SDR and HDR color patch pairs derived from the training SDR image and the training HDR image; applying the optimized SDR-HDR mapping to one or more non-training SDR images to generate one or more corresponding non-training HDR images; A method comprising:

2. The method of claim 1 , wherein the training SDR images and the training HDR images are captured from a three-dimensional (3D) visual scene by a capture device operating in an SDR capture mode and an HDR capture mode, respectively.

3. 2. The method of claim 1 , wherein the training SDR image and the training HDR image form a pair of training SDR image and training HDR image among a plurality of pairs of training SDR image and training HDR image, and the optimized SDR-HDR mapping is generated based at least in part on a set of a plurality of pairs of SDR color patches and HDR color patches derived from the plurality of pairs of training SDR image and training HDR image.

4. 2. The method of claim 1 , wherein each SDR image feature point in the subset of one or more SDR image features is matched with a respective HDR image feature point in the subset of one or more HDR image features, and the SDR image feature points and the HDR image feature points are extracted from the training SDR image and the training HDR image, respectively, using a common feature point extraction algorithm.

5. The training SDR image and the training HDR image are obtained by performing a respective camera distortion correction operation on each distorted training image in a pair of distorted training standard dynamic range (SDR) images and distorted training high dynamic range (HDR) images to generate a respective training image in the pair of training SDR images and training HDR images; the set of SDR image feature points and the set of HDR image feature points correspond to corner pattern marks; 2. The method of claim 1 , wherein using the subset of one or more SDR image feature points and the subset of one or more HDR image feature points includes generating each projective transformation in a pair of SDR image projective transformation and HDR image projective transformation using the corner pattern marks detected from each training image in the pair of training SDR image and training HDR image.

6. 6. The method of claim 5, wherein the training SDR images and the training HDR images are captured from a common color target image by a first capture device operating in an SDR capture mode and a second capture device operating in an HDR capture mode, respectively.

7. The method of claim 5 , wherein the common color chart image is rendered on the first capture device and the second capture device and is captured from a screen of a common reference image display.

8. The method of claim 5 , wherein the respective camera distortion correction operations are based at least in part on camera-specific distortion coefficients generated from a camera calibration process performed with the camera used to acquire the training images.

9. 6. The method of claim 5, wherein the set of SDR color patches and the set of HDR color patches are used to derive a three-dimensional mapping table (3DMT), and the optimized SDR-HDR mapping is generated based at least in part on the 3DMT.

10. The method of claim 5 , wherein the optimized SDR-HDR mapping represents one of a tensor product B-spline (TPB) based mapping or a non-TPB based mapping.

11. constructing sampled high dynamic range (HDR) color space points distributed across an HDR color space used to represent a reconstructed HDR image; converting the sampled HDR color space points to SDR color space points in a first standard dynamic range (SDR) color space representing an SDR image to be edited by an editing device; determining a bounded SDR color space rectangle based on extreme SDR codeword values ​​of the SDR color space points in the first SDR color space and determining an irregular three-dimensional (3D) shape from the distribution of the SDR color space points; constructing sampled SDR color space points distributed across the bounded SDR color space rectangle in the first SDR color space; generating a boundary clipping 3D lookup table (3D-LUT) using the sampled SDR color space point and the irregular shape, the boundary clipping 3D-LUT including lookup keys and corresponding lookup values, the boundary clipping 3D-LUT using the sampled SDR color space point as a lookup key, such that when the lookup key is inside the irregular shape, the lookup key is equal to the lookup value, and when the lookup key is outside the irregular shape, the lookup value is determined based on an exponential function that takes the irregular shape and the lookup key as input and returns as lookup value the closest neighbor inside the irregular shape to the lookup key; performing a clipping operation on the edited SDR image in the first SDR color space based at least in part on the boundary clipping 3D-LUT to generate a boundary-clipped edited SDR image in the first SDR color space; A method comprising:

12. 12. The method of claim 11, wherein the clipping operation includes first performing regular clipping on the edited SDR image using the bounded SDR color space rectangle to generate a regular clipped edited SDR image, and then performing irregular clipping on the regular clipped edited SDR image using the 3D-LUT to generate the bounded clipped edited SDR image.

13. Apparatus including a processor and configured to perform the method of any one of claims 1 to 12.

14. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing, with one or more processors, the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Guided color grading for extended dynamic range

    JP2016536873A

  • Display management for high dynamic range video

    JP2018510574A

  • Iterative optimization of reshaping functions in single-layer HDR image codec

    WO2021216767A1