A method, apparatus, and storage medium for processing entanglement and shaping images with neighborhood consistency.

By employing the wrapping and shaping technique and utilizing the forward and backward shaping functions modeled by TPB, the problem of detail loss and color distortion when mapping high dynamic range video content to a low bit depth codec is solved, achieving efficient video content reconstruction and image quality preservation.

CN116391353BActive Publication Date: 2026-05-05DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2021-11-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing techniques tend to cause loss of detail and color distortion when mapping high dynamic range (HDR) video content to low bit depth codecs, especially when using 3D lookup table (3D-LUT) mapping, resulting in banding artifacts and visual artifacts.

Method used

By employing the winding shaping technique, the forward and backward shaping functions of the sensing-sensitive channel are modeled using tensor product B-splines (TPB). Combined with the out-of-loop training and in-loop prediction stages, the function parameters used in the shaping process are optimized to ensure neighborhood consistency and reversibility in the low-level depth domain, thereby reducing artifacts and distortion.

Benefits of technology

It effectively maps high-bit-depth video content to low-bit-depth codecs, maintaining image quality, avoiding banding artifacts and color distortion, and improving the reversibility and image quality of reconstructed video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116391353B_ABST
    Figure CN116391353B_ABST
Patent Text Reader

Abstract

Receive the input image at the first bit depth in the input domain. Perform a forward shaping operation on the input image to generate a forward-shaped image at the second bit depth in the shaping domain. Encode the image container containing the image data obtained from the forward-shaped image into an output video signal at the second bit depth.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 112,336, filed November 11, 2020, and European Patent Application No. 20206922.5, filed November 11, 2020, both of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure generally relates to image processing operations. More specifically, embodiments of this disclosure relate to video codecs. Background Technology

[0004] As used herein, the term "dynamic range (DR)" can refer to the ability of the human visual system (HVS) to perceive a range of intensity (e.g., luminance, brightness) in an image, such as from the darkest black (darkness) to the brightest white (highlight). In this sense, DR relates to the intensity "scene-referred". DR can also refer to the ability of a display device to fully or approximately render a specific breadth of intensity range. In this sense, DR relates to the intensity "display-referred". Unless a particular meaning is explicitly specified to have a specific implication at any point in the description herein, it should be inferred that the terms can be used interchangeably in either sense, for example.

[0005] As used herein, the term "high dynamic range (HDR)" refers to a DR width spanning approximately 14 to 15 or more orders of magnitude across the human visual system (HVS). In practice, the DR, which represents a broad range of intensity that humans can simultaneously perceive relative to HDR, may be slightly truncated. As used herein, the terms "enhanced dynamic range (EDR)" or "visual dynamic range (VDR)" can be associated, individually or interchangeably, with this type of DR: DR that can be perceived within a scene or image by the human visual system (HVS), including eye movements, allowing for some changes in light adaptability on the scene or image. As used herein, EDR can refer to a DR spanning 5 to 6 orders of magnitude. While this may be slightly narrower than HDR relative to a reference real-world scene, EDR indicates a wide DR width and can also be referred to as HDR.

[0006] In fact, an image comprises one or more color components in a color space (e.g., luminance Y and chrominance Cb and Cr), where each color component is represented by a per pixel. n Bit precision representation (e.g., n = 8). Using non-linear light intensity coding (e.g., gamma coding), where... nImages with a dynamic range ≤ 8 (e.g., a color 24-bit JPEG image) are considered to have a standard dynamic range, where... n Images with a dynamic range greater than 8 can be considered images with enhanced dynamic range.

[0007] A reference electro-optical transfer function (EOTF) for a given display characterizes the relationship between the color values ​​(e.g., luminance) of the input video signal and the color values ​​(e.g., screen luminance) of the output screen generated by the display. For example, ITU Rec.ITU-R BT.1886, “Reference electro-optical transfer function for flat panel displays used in HDTV studio production” (March 2011) defines a reference EOTF for flat panel displays, the contents of which are incorporated herein by reference in their entirety. Given a video stream, information about its EOTF can be embedded in the bitstream as (image) metadata. The term “metadata” in this document refers to any auxiliary information transmitted as part of the encoded bitstream and used to assist the decoder in rendering the decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as described herein.

[0008] As used herein, the term "PQ" refers to Perceived Luminance Amplitude Quantization. The human visual system responds to increasing light levels in a highly nonlinear manner. The human ability to perceive stimuli is influenced by factors such as the luminance of the stimulus, the size of the stimulus, the spatial frequency constituting the stimulus, and the luminance level to which the eye adapts at a particular moment of viewing the stimulus. In some embodiments, the perceived quantizer function maps linear input gray levels to output gray levels that better match the contrast sensitivity threshold in the human visual system. An example PQ mapping function is described in SMPTE ST 2084:2014 "High Dynamic Range EOTF of Mastering Reference Displays" (hereinafter referred to as "SMPTE"), which is incorporated herein by reference in its entirety, wherein, given a fixed stimulus size, for each luminance level (e.g., stimulus level, etc.), the minimum visible contrast step size at said luminance level is selected based on the most sensitive adaptation level and the most sensitive spatial frequency (according to the HVS model).

[0009] Supports 200 to 1,000 cd / m³ 2A display with a brightness of nits or nits represents a lower dynamic range (LDR) associated with EDR (or HDR), also known as standard dynamic range (SDR). EDR content can be displayed on an EDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or higher). Such a display can be defined using an alternative EOTF that supports high brightness capabilities (e.g., 0 to 10,000 nits or higher). In SMPTE 2084 and Rec. ITU-R BT. 2100, “ Image parameter values ​​for high dynamic range television for use in production and international program exchange An example of such an EOTF is defined in "[Image Parameter Values ​​for High Dynamic Range Television Used in Production and International Program Exchange]" (06 / 2017). As the inventors understand it herein, an improved technique is desired for encoding high-quality video content data that can be rendered on a variety of display devices.

[0010] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Attached Figure Description

[0011] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:

[0012] Figure 1 The illustration shows an example image processing pipeline for performing wrap shaping operations;

[0013] Figure 2A The illustration shows an example mapping from the input domain to the reshaping domain of the reference shape; Figure 2B The diagram illustrates an example flow for generating specific operation parameter values ​​in primary integer shaping; Figure 2C The illustration shows example procedures for primary and secondary plastic surgery; Figures 2D to 2F The diagram illustrates an example process for determining the optimal scaling factor;

[0014] Figure 3A The illustration shows the prediction error in forward reshaping; Figure 3BFigure 3C illustrates the convergence of prediction error in backward shaping; Figure 3D and Figure 3E illustrate example visualizations of a complete donut in a color cube; Figure 3F illustrates an example visualization of a complete cylinder in a color cube; Figure 3G illustrates an example visualization of a partial cylinder in a color cube. Figure 3H The illustration shows example pixels of the input image in a grid representation; Figure 3I The figure shows the average values ​​of non-empty clusters in the input image; Figure 3J The illustration shows example centers of non-empty clusters in the input image; Figure 3K The illustration shows the example centers of non-empty clusters and neighboring empty clusters in the input image; Figure 3L The illustration shows an example distribution of pixels in an input image visualized in a color cube; Figure 3M The illustration shows example points generated from the pixel distribution in the input image; Figure 3N The illustration shows example points transformed forward by wrapped reshaping;

[0015] Figure 4A and Figure 4B The example process flow is illustrated; and

[0016] Figure 5 A simplified block diagram of an example hardware platform is shown, on which a computer or computing device as described herein can be implemented. Detailed Implementation

[0017] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent that this disclosure may be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing this disclosure. Summary of the Invention

[0018] The techniques described herein can be implemented to wind perceptually sensitive channels among multiple color space channels used to represent video data into geometric or topological shapes, such as near-circular or toroidal shapes, to increase the total effective number of codewords for the perceptually sensitive channels. In some operational scenarios, the winding of perceptually sensitive (or dominant) channels can be modeled using tensor product B-splines (TPBs), which offer relatively high flexibility in representing or approximating mappings / functions used for forward and backward shaping.

[0019] Example reshaping operations are described in U.S. Provisional Patent Application Serial No. 62 / 136,402, filed March 20, 2015 (also published January 18, 2018, as U.S. Patent Application Publication Serial No. 2018 / 0020224) and PCT Application Serial No. PCT / US2019 / 031620, filed May 9, 2019, the entire contents of which are incorporated herein by reference as if fully set forth herein. Example constructions of the forward and backward shaping functions are described in GM. Su’s U.S. Provisional Patent Application Serial No. 63 / 013,063, filed April 21, 2020, “Reshaping functions for HDR imaging with continuity and reversibility constraints”, GM. Su and H. Kadu’s U.S. Provisional Patent Application Serial No. 63 / 013,807, filed April 22, 2020, “Iterative optimization of reshaping functions in single-layer HDR image codec”, and PCT / US2021 / 028475, filed April 21, 2021, the contents of which are incorporated herein by reference in their entirety as if fully set forth herein.

[0020] In some approaches, high dynamic range (HDR) or wide color gamut (WCG) video signals typically use a relatively high number of bits per channel (e.g., no less than 12 bits per channel) to support encoding the rich colors and wide range of brightness levels represented in HDR or WCG video content.

[0021] However, many end-user video devices may be equipped with popular video compression codecs that are only capable of encoding / decoding relatively low bit-per-channel video signals, such as 8- or 10-bit video signals. As used herein, the “bits per channel” of a color channel may be referred to as the “bit depth” of the color channel.

[0022] To achieve widespread or ubiquitous distribution of HDR video content on such a limited-bit-depth codec infrastructure, HDR video content can be mapped to the actual bit depth supported by these codecs through forward shaping operations. For example, a video encoder can map (or forward-shape) the input or source video content from an input 16-bit HDR video signal into an output or shaped 8-bit video signal for transmission to a recipient device, such as an end-user video device with an 8-bit codec. The recipient device with the shaped 8-bit video signal can then transform the backward-shaped (or back-shaped) video content from the received shaped 8-bit video signal to generate or reconstruct 16-bit HDR video content for display. The reconstructed 16-bit HDR video content generated by the recipient device can be rendered on a display device to approximate the input or source 16-bit HDR video content that has been forward-shaped to an 8-bit video signal by the video encoder.

[0023] In some methods, simple bit-depth truncation is performed to forward-shape the input or source 16-bit HDR video content. After conversion or forward-shaping to a shaped 8-bit video signal, details in the source HDR video content may be irreversibly lost. Therefore, under these methods, the reconstructed HDR video content generated from such truncated shaped 8-bit video signals may tend to display banding or contour artifacts.

[0024] In some other methods, a three-dimensional lookup table (3D-LUT) mapping can be used to transform or forward-shape the input or source codeword represented in a higher bit-depth video signal into a shaped codeword in a lower bit-depth video signal. Any valid color value from a higher bit-depth domain or color space to a lower bit-depth domain or color space can be supported as long as there is still space fill throughout the entire lower bit-depth domain or color space (e.g., color CRT).

[0025] By using codeword interleaving or 3D-LUT mapping with sparse codeword arrangements, and simultaneously packing or squeezing codewords more efficiently in the lower bit depth domain or color space, it may result in assigning neighboring colors in the higher bit depth domain or color space to completely different locations in the lower bit depth domain or color space. As a result, color distortion and visual artifacts may occur, especially in scenarios where image processing, such as subsampling or filtering, is applied to the lower bit depth shaping domain. This is due to the nonlocality of image processing, which creates new codeword values ​​in the lower bit depth shaping domain that do not correspond to neighboring codewords in the original higher bit depth domain or color space.

[0026] In stark contrast, the techniques described in this paper can be implemented to increase the availability of codewords in the lower bit depth domain (or color space) while maintaining neighborhood consistency between the higher and lower bit depth domains (or color spaces), thereby avoiding or reducing banding artifacts and color distortion in shaped video content that would occur with other methods. Additionally, alternatively, or even more alternatively, the mapping process can be implemented using closed-form equations or solutions, such as the forward and / or backward shaping with winding operations described herein, thereby enhancing reversibility and image quality in the reconstructed video.

[0027] Some or all of the operations described herein can be implemented using an out-of-loop training phase, an in-loop prediction phase, or a combination of both. In some operational scenarios, during the training phase, target (shaping) objects / forms can be designed, obtained, or generated using training data with different operational parameters. Later, during the deployment or prediction phase, a specific target object / form can be selected from these target objects / forms for the input or source higher bit-depth video signal or the input or source image therein. The specifically selected target object / form can be applied during the winding process to wind or shape the input or source higher bit-depth video signal or the input or source image into a lower bit-depth video signal or the shaped image therein, represented in the shaping domain or color space.

[0028] In some operational scenarios, the winding process can be modeled as applying image processing operations using TPB-based functions / mappings to forward and backward shaping functions / mappings.

[0029] Optimized values ​​of some or all TPB coefficients of a TPB-based function / mapping used in the winding process can be obtained or pre-configured via an iterative algorithm / method / process to minimize prediction error and / or enhance prediction quality using training data. As a result, a set of forward and backward shaping functions can be obtained or stored in a shaping function data store during the training phase and can be retrieved later in the deployment or prediction phase.

[0030] During the deployment or prediction phase, given an input image, (1) optimized pre-reshaping operations (such as applying scaling / offset to predefined or stored shaping functions) can be performed to enhance the accuracy of shaping operations, and (2) optimized selection operations (such as selecting a specific shaping function for the input image from a shaping function data store) can be performed to minimize prediction errors.

[0031] During the deployment or prediction phase, image processing operations can be safely performed on shaped video data within the shaping domain of the shaped signal. When performing image processing operations in the shaping domain, new codewords created during image processing can be processed (e.g., specifically, appropriately).

[0032] The techniques described in this article are not limited to the scenario of shaping video signals using AVC 8-bit compression. Some or all of these techniques can be used to compress high-bit-depth video signals of various relative high bit depths (such as 12-bit, 16-bit, etc.) into lower-bit-depth (such as 8-bit, 10-bit, 12-bit, etc.) video signals by winding perceptibly important or dominant channels / axis into a 3D space represented in a low-bit-depth shaping domain or color space. In other methods, the available codewords for encoding or representing perceptibly important or dominant channels / axis in the shaped signal can be increased to 2 or 4 times the available codewords, and even more through adaptive prescaling.

[0033] In some operational scenarios, the shaped video content generated using the techniques described herein (which may be 8, 10, or 12 bits per channel) can be stored in a baseband signal container for storage or transmitted in a baseband signal used with an SDI or HDMI cable.

[0034] The example embodiments described herein relate to encoding video images. An input image of first depth in the input domain is received from an input video signal of first depth. In the shaping domain, the first depth is higher than the second depth. A forward shaping operation is performed on the input image to generate a forward-shaped image of second depth in the shaping domain. The forward shaping operation includes winding input codewords in the input image along the non-wrapped axis of the input domain into shaped codewords in the forward-shaped image along the wrapped axis of the shaping domain. An image container containing image data obtained from the forward-shaped image is encoded into an output video signal of second depth. The image data in the image container causes a receiving device of the output video signal to construct a backward-shaped image of third depth for rendering on a display device. The third depth is higher than the second depth.

[0035] The example embodiments described herein relate to decoding video images. An image container containing image data derived from a forward-shaped image in a shaping domain is decoded from a video signal. A source image of a first depth in the source domain has been forward-shaped to generate a forward-shaped image of a second depth lower than the first depth. A backward-shaped operation is applied to the image data decoded from the video signal to generate a backward-shaped image in a target domain. The backward-shaped image has a third depth higher than the second depth, and the backward-shaped operation includes unwinding codewords in the image data along the winding axis of the shaping domain into backward-shaped codewords in the backward-shaped image on the non-winding axis of the target domain. A display image generated from the backward-shaped image is rendered on a display device.

[0036] Example image processing pipeline

[0037] Figure 1 The illustration depicts an example image processing pipeline for performing wrapping and shaping operations. Some or all of the encoding and / or decoding operations in the image processing pipeline can be performed by one or more computing devices with audio and video codecs, such as those defined by ATSC, DVB, DVD, Blu-ray, and other standards / specifications. The image processing pipeline can be implemented on the decoder side only, on the encoder side only, or a combination of the encoder and decoder sides.

[0038] An input or source HDR image 102 with a high bit depth (e.g., 12-bit, 16-bit, etc.) per channel in the input or source domain or color space can be received via an input or source high bit depth video signal. A reconstructed HDR image 110 with a high bit depth per channel in the output or target domain or color space can be generated at the end of the image processing pipeline. The reconstructed HDR image (110) may, but is not necessarily limited to, having the same high bit depth as the input or source HDR image (102).

[0039] In some operational scenarios, the input or source HDR image (102) can be any image in a sequence of consecutive input or source HDR images received and processed by an image processing pipeline. Some or all of the consecutive input or source HDR images can be generated from images captured in analog or digital formats or computer-generated images through video editing or transformation operations (e.g., automatic without human input, manual, automatic with human input, etc.), color grading operations, etc. Consecutive input or source HDR images can be images associated with one or more of the following: film releases, archived media programs, media program libraries, video recording / editing, media programs, television programs, user-generated video content, etc.

[0040] Figure 1 The output or target HDR image (110) can be a corresponding reconstructed image from a sequence of consecutive output or target HDR images generated from an image processing pipeline. The consecutive output or target HDR images depict the same visual semantic content as depicted in the consecutive input or source HDR images. The display image can be obtained directly or indirectly (e.g., through additional display management operations, etc.) from the consecutive output or target HDR images and rendered on an HDR display. Example HDR displays may include, but are not limited to, image displays that operate in conjunction with televisions, mobile devices, home theaters, etc.

[0041] like Figure 1As shown, the input or source HDR image (102) can be converted into a forward-shaped EDR image 106 with a low-bit depth per channel in the integer domain or color space (e.g., 8-bit, 10-bit, 12-bit, etc.) by a wrap-around forward-shaping operation 104. In some operating scenarios, the forward-shaped EDR image (106) in the integer domain or color space can be contained or encapsulated in an 8-bit image container, which is supported, for example, by various video codecs.

[0042] The wrapping forward shaping operation (104) can be specially selected or designed to generate a forward-shaped EDR image (106) that introduces minimal or no distortion in the reconstructed or backward-shaped HDR image (110) in the output or target domain or color space.

[0043] The reconstructed or backward-shaped HDR image (110) can be generated by wrapping backward-shaped 108 from the decoded EDR image 106' at low bit depth per channel in the shaping domain or color space. The distortion introduced by the wrapping forward and backward-shaped operations (104 and 108) in the reconstructed or backward-shaped HDR image (110) can be estimated or measured with reference to the input or source HDR image (102).

[0044] In some operational scenarios, the decoded EDR image (106') can be the same as the forward-shaped EDR image (106), which is subject to quantization and / or encoding errors during the process of encoding the forward-shaped EDR image (106) into a low-bit depth shaped video signal and decoding the decoded EDR image (106') from the low-bit depth shaped video signal.

[0045] In some operational scenarios, the forward-shaped EDR image (106) is transformed by a truncated field transform 112 into a truncated forward-shaped EDR image 114 at a low bit depth per channel in the truncated integer domain or a new transform domain (e.g., with floating-point precision, etc.). One or more image processing operations 116 are applied to the truncated forward-shaped EDR image (114) in the truncated integer domain to generate a processed truncated forward-shaped EDR image 114' at a low bit depth per channel in the truncated integer domain. Example image processing operations may include, but are not limited to, one or more of the following: combination of neighborhood information, linear or nonlinear combination of neighborhood information, upsampling, downsampling, deblocking, low-pass filtering, etc. The processed truncated forward-shaped EDR image (114') is inversely transformed by the inverse truncated field transform 112' into a processed forward-shaped EDR image with a low bit depth per channel in the truncated or new transform domain (in floating-point precision).

[0046] The truncated field transform (112), the inverse truncated field transform (112'), and the image processing operation (114) can be performed on the encoder side or the decoder side. By way of example and not limitation, these operations can be performed on the encoder side.

[0047] The processed forward-shaped EDR image at low bit depth for each channel in the shaping domain can be encoded into a low bit-depth shaped video signal. The decoded EDR image (106') can be identical to the processed forward-shaped EDR image at low bit depth for each channel in the shaping domain, subject to quantization and / or encoding errors during the encoding of the processed forward-shaped EDR image at low bit depth for each channel in the shaping domain into a low bit-depth shaped video signal and the decoding of the decoded EDR image (106') from the low bit-depth shaped video signal.

[0048] To pack or squeeze codewords in a high-bit-depth HDR image (102) into an 8-bit image container, a wrapping forward shaping operation (104) can be performed in many different ways using various forward shaping maps or functions.

[0049] As an example and not a limitation, the wrapping forward shaping operation (104) performs two-stage shaping. The first stage (which may be referred to as secondary shaping) involves per-channel pre-shaping or scaling, which linearly scales a limited range of input data to the full data range of each color channel of the input video data in the high bit-depth HDR image (102). The second stage (which may be referred to as primary shaping) involves cross-color-channel nonlinear forward shaping using a TPB-based forward shaping function.

[0050] Similarly, the wrapping backward shaping operation (108) includes (1) applying cross-color channel nonlinear backward shaping using a TPB-based backward shaping function, and (2) applying per-channel backward shaping or inverse scaling, which can inversely scale the entire data range to a limited input data range for each color channel of the reconstructed video data in the reconstructed HDR image (110).

[0051] The truncated field transform (112) and the inverse truncated field transform (112') are applied to ensure that the codewords in the processed forward-shaped EDR image generated from the image processing operation (114) and the transforms (112 and 112') remain within the specified codeword value range, such that no codewords are outside the design form or shape in the 8-bit wrapped integer domain or color space. Without these transforms (112 and 112'), the codewords generated by the image processing operation (114) may become out of range and outside the design form or shape in the 8-bit wrapped integer domain or color space, thereby ultimately generating visible color artifacts in the reconstructed HDR image (110).

[0052] Basic plastic surgery during the training phase

[0053] Shaping operations can be applied in various ways to different operational scenarios. In some scenarios, fully dynamic shaping can be applied. In these scenarios, shaping functions can be constructed to perform end-to-end shaping between the original (e.g., input, source, etc.) domain and the shaped (e.g., output, target, etc.) domain. Given that video content can have different specific color distributions, theoretically, this function can be selected for specific video content. This fully dynamic shaping method may require relatively high computational complexity and cost because the end-to-end optimization process used to obtain the optimal operational parameters of the shaping function requires computationally complex and cost-intensive algorithms. Furthermore, numerical instability may occur under this method, leading to slow or non-convergent approaches when reaching the optimal or convergent solution.

[0054] In some operational scenarios, two-stage shaping can be applied to mitigate issues related to computational complexity, cost, and numerical stability. The end-to-end shaping process under this method includes primary shaping and secondary shaping.

[0055] Primary shaping, or the second stage, involves relatively heavy computation and can be performed offline or as an out-of-loop operation for training. Secondary shaping, or the first stage, is relatively lightweight and can be performed online or as an in-loop operation.

[0056] As used in this paper, an out-of-loop operation can refer to an operation performed during offline training that is not part of the real-time image encoding / decoding operation. An in-loop operation can refer to an operation performed as part of the real-time image encoding / decoding operation or the runtime online deployment process.

[0057] To help provide the lightest or least complex computation during the (online) deployment phase, some or all of the computationally intensive tasks can be performed offline in the primary shaping or second phase, i.e., obtaining trained operational coefficients for shaping and storing the trained operational coefficients in a shaping coefficient data store.

[0058] It can be done during the deployment phase in the secondary shaping or the first phase. Figure 1 Such as 102, it receives the input image and performs a relatively light thinning operation.

[0059] During the training phase, before processing the input image, the training data or training images can be used to train a primary shaping function, such as the forward and backward shaping functions based on TPB. In some operational scenarios, the training phase may be divided or separated into two parts: (1) reference map creation and (2) TPB training.

[0060] Reference mapping creation involves two sets of control variables. The first set of control variables involves selecting a reference mapping object, shape, or form among different reference mapping objects, shapes, or forms supported in the winding shaping domain.

[0061] The second of the two sets of control variables mentioned above relates to the geometric parameters of a given reference mapping object, shape, or form. These geometric parameters determine the width, height, and other properties of the given reference mapping object, shape, or form.

[0062] The reference mapping shape or form described herein refers to a geometric or topological object, shape, or form within a winding shaping domain. Examples of geometric or topological objects, shapes, or forms that can serve as reference mapping objects, shapes, or forms as described herein include, but are not limited to, any of the following: toroidal objects / shapes / forms, toroidal objects / shapes / forms, cylindrical shapes / forms, spiral objects / shapes / forms, etc.

[0063] like Figure 2A As illustrated, the codewords in the HDR image (which can be the training HDR image during the training phase, or the input or source HDR image during the deployment prediction phase, such as...) Figure 1 The codewords in color cube (202) can be represented or distributed in color cube (202). A mapping function 204 based on two sets of control variables can be used to map or convert codewords in color cube (202) into corresponding codewords from multiple available reference mapping shapes / forms with specific geometric parameters in a shaped object list (206). Multiple available reference mapping shapes / forms in the shaped object list (206) can be created, defined, or designed during the training phase.

[0064] Two sets of control variables influence or at least partially control the quantization and dequantization during the end-to-end forward and backward shaping processes, and thus influence or at least partially control the prediction accuracy or error during the forward and backward shaping processes. The operational parameters in the two sets of control variables, representing different reference mapping objects, shapes, or forms supported in the winding shaping domain or color space, can be trained during the training phase and stored as templates in the operational parameter dataset. Each template can include specific combinations (e.g., different, etc.) of the operational parameters from the two sets of control variables (as generated during training). The combinations of specific values ​​in the templates can be specifically selected during the training phase, or selected from multiple candidate value combinations or sets to obtain the best prediction results during the training phase and produce the minimum error in the training / test data.

[0065] During the training phase, TPB training can be used to generate (e.g., general, trained, TPB-based, etc.) shaping functions for different templates corresponding to different combinations of values ​​from the two sets of control variables. Gradient-based BESA (Backward Error Subtraction Signal Adjustment) algorithms / methods can be used or executed as part of TPB training to generate or obtain specific (e.g., optimized, trained, etc.) forward and backward shaping coefficients that define or specify the TPB shaping function for each template corresponding to the respective combinations of values ​​from the two sets of control variables. Figure 2B An example of TPB training using the gradient-based BESA algorithm / method is illustrated, which will be discussed in more detail later. The example BESA algorithm / method can be found in the aforementioned U.S. Provisional Patent Application Serial No. 63 / 013,807.

[0066] Secondary and primary cosmetic procedures

[0067] Figure 2C The illustration shows an example flow for secondary and primary shaping, which can be performed partially or entirely during the deployment or prediction phase. As shown, during the deployment or prediction phase, an input image or source image 222 (represented as a "normalized high dynamic range image") in the high bit depth domain or color space is received (e.g., ...). Figure 1 102), used to wrap and shape it into a shaped object / shape / form or a shaped image 230 (represented as "normalized forward-shaped object image") in a low-bit depth domain or color space (e.g. Figure 1 (106).

[0068] In some operational scenarios, the input or source codewords in the input or source image (222) are normalized to a specific codeword range such as [0, 1]. Similarly, the shaped codewords in the shaped image (230) are normalized to a specific codeword range such as [0, 1].

[0069] During the deployment or prediction phase, the input or source HDR image may or may not have a color or codeword distribution that completely occupies the entire codeword space (such as a 3D color cube in the input or source HDR domain). A pre-shaping (or secondary shaping) phase may be implemented or performed to help make full use of the available capacity in the codeword space, where the codewords of the input or source HDR image (222) can be represented or hosted.

[0070] like Figure 2CAs illustrated, the pre-shaping (or secondary shaping) stage includes box 224, in which one or more optimal scaling factors / values ​​are determined in response to determining the finite range occupied by the input image or source image in each of one or more channels. The pre-shaping (or secondary shaping) stage further includes box 226, in which codewords in one or more finite ranges can be pre-scaled or re-scaled from one or more finite ranges (e.g., linearly, non-linearly, etc.) to one or more complete codeword ranges in one or more channels using one or more optimal scaling factors / values, respectively.

[0071] In box 228, the intermediate image, which includes scaled codewords generated in box 226 by scaling input or source codewords in the input or source image (222), can be shaped by primary shaping based at least in part on a predefined or trained TPB shaping function specifically selected for the input or source HDR image, thereby generating a shaped image (230) in the shaping domain or color space.

[0072] In some operational scenarios, assuming complete occupation of the codeword space in the input or source HDR domain or color space, a predefined or trained integer function can be generated or obtained during the training phase for a corresponding template with specific values ​​of two sets of control variables. The integer coefficients of the defined or specified predefined or trained integer function can be trained to perform integer or full mapping on codewords throughout the entire codeword space.

[0073] The shaped image (230) (or a processed version of the shaped image (230) generated by truncated field transform, image processing operations, and / or inverse truncated field transform) can be encoded into a (shaped) video signal. The receiving device of the video signal can decode the shaped image (230) or its processed version and perform backward shaping and inverse / inverse scaling on the shaped image (230) or its processed version, thereby generating a high-bit-depth reconstructed image in the high-bit-depth output or target domain or color space.

[0074] Figure 2D The illustration shows the deployment or prediction phase with a given input or source image (e.g., Figure 2CIn the case of 222, etc., an example process for determining the optimal scaling factor or value 238 is provided. Specific (e.g., optimized, etc.) values ​​for the operating parameters used for primary and secondary shaping can be generated or obtained via a search 236 based on a cost function such as mean squared error (MSE). These operating parameters may include operating parameters for primary shaping represented in TPB list 234. These operating parameters may further include operating parameters for secondary shaping, such as scaling factors / values ​​represented in scaling list 232. Specific (e.g., optimized, etc.) values ​​or settings (238) for the operating parameters in TPB list (234) and scaling list (232) can be generated or obtained via search (236).

[0075] Primary Integer Functions

[0076] Forward and backward shaping operations in primary shaping can be performed, at least in part, through mathematical functions and / or parametric equations. The target shaping object / form / shape, or its associated transformation, can be represented by mathematical equations. In some cases, hard-coded, fixed-form equations may be acceptable, but in others they are not. To help support multiple shaping objects / forms / shapes, the shaping functions described in this paper can be designed with relatively large or maximum flexibility and degrees of freedom.

[0077] In some operational scenarios, the parameters, coefficients, and / or values ​​of (multiple) shaping functions associated with a specific target object / form / shape can be included in the encoded bitstream of the video signal. The receiving device of the video signal on the decoder side can decode and use these parameters, coefficients, and / or values ​​to establish or reconstruct (multiple) shaping functions, such as inverse shaping functions, without using hard-coded fixed-form equations.

[0078] As an example and not a limitation, the shaping functions described herein can be specified using flexible mappings or function constructions, such as geometric functions / transformations, multivariate regression (MMR), and transformation functions / mappings based on tensor product B-splines (TPB). Specific (e.g., optimized, etc.) values ​​or settings of the coefficients or operating parameters of shaping functions or mappings based on MMR and TPB can be generated or obtained through iterative algorithms, which will be explained in further detail later.

[0079] Examples of MMR operations are described in U.S. Patent No. 8,811,490, the entire contents of which are incorporated herein by reference as fully set forth herein. Examples of TPB operations are described in U.S. Provisional Application Serial No. 62 / 908,770 (Attorney General's File No. 60175-0417), filed October 1, 2019, entitled "TENSOR-PRODUCT B-SPLINE PREDICTOR," the entire contents of which are incorporated herein by reference as fully set forth herein.

[0080] Geometric transformations (e.g., transformations of a shape or form (such as a color cube) into one of a reference integer shape or form (such as a cylindrical shape or form, a toroidal shape or form, a torus shape or form, etc.)) can be used to transform input codewords in an input domain or color space (such as those codewords that are best scaled and represented in a color cube) into shaped codewords in a reference integer shape or form within an integer domain. Geometric transformations can be represented by functional expressions such as 3D geometric transformation functions.

[0081] In some operational scenarios, the geometric transformations described in this paper can include transformations for each dimension or channel that can be performed individually or separately via cross-channel transformation functions. The per-channel transformation is represented as... f x , f y and f z The per-channel transformation in geometric transformation can be given as follows:

[0082]

[0083] (1-1)

[0084]

[0085] (1-2)

[0086]

[0087] (1-3)

[0088] In some operational scenarios, the transformation or transformation function in the above expression (1) can be represented using a 3 × 3 matrix. In these operational scenarios, the geometric transformation represented by the transformation or transformation function in the above expression (1) can be rewritten via matrix multiplication, and the elements of the matrix (e.g., a single-column matrix) can be transformed or changed from a given point in the input domain or color space to the corresponding point in the output domain or color space, as shown below:

[0089]

[0090] (2-1)

[0091] Or use different symbols as shown below:

[0092]

[0093] (2-2)

[0094] In expression (2-1), the 1 × 3 column on the left-hand side (LHS) can represent the corresponding point in the output domain or color space, and can be represented as in expression (2-2). In expression (2-1), the 1 × 3 column on the right-hand side (RHS) can represent a given point in the input domain or color space, and can be represented as in expression (2-2). The 3 × 3 matrix on the right-hand side (RHS) of expression (2-1) can include matrix elements representing the per-channel transformation, and can be represented as in expression (2-2). .

[0095] matrix Each element of a matrix is ​​not necessarily limited to any particular functional form. For example, an element may or may not be a polynomial. In some operational scenarios, the matrix... Each element can be given in closed form or as a formula.

[0096] MMR transformation (e.g., used to shape an input codeword into a shaped codeword) can be represented as: The MMR mapping or function is used to represent this. The MMR mapping or function can accept three input parameters. And generate or output a single value. As shown below:

[0097] (3)

[0099] MMR mappings or functions can have a predefined format that involves higher powers of the input variables / independent variables and preselected cross terms specified using some or all of the input variables / independent variables. A non-restrictive example form of an MMR mapping or function can be second-order, as shown below:

[0100] (4)

[0102] in, This represents the MMR coefficients to be trained or learned from training data, which includes triples. and corresponding target The training dataset.

[0103] For a three-channel input image represented in the input or source domain or color space, three predictors can be specified or learned to shape the codewords in the three channels of the input or source domain or color space into shaped codewords in the three target channels of the output or target domain or color space. Each of the three predictors can be used to shape the codewords in the three channels of the input or source domain or color space into shaped codewords in the corresponding target channels of the output or target domain or color space.

[0104] The TPB transformation (e.g., used to shape an input codeword into a shaped codeword) can be represented as: The TPB mapping or function is used to represent this. The TPB mapping or function (which represents a relatively advanced and powerful cross-channel prediction function) can use B-splines as basis functions and accepts three input parameters. and generate or output a single value. As shown below:

[0105] (5)

[0107] The internal structure of the TPB function can be derived along its three input channels or dimensions by having and The basis functions are represented by the following equations, as shown below:

[0108] (6)

[0110] in, and Indicates the node index; This represents the TPB coefficients to be generated or learned from the training data.

[0111] The composite basis functions in the above expression (6) It can be given as a product of individual basis functions in the three channels or dimensions, as shown below:

[0112] (7)

[0114] The TPB function can have much larger coefficients than the MMR, thus providing additional degrees of freedom and capability to model complex mappings in image reshaping operations.

[0115] Optimized values ​​of operating parameters in primary plastic shaping

[0116] Figure 2BThe diagram illustrates an example flow for generating specific (e.g., optimized, etc.) values ​​of operational parameters in primary integer shaping. In some operational scenarios, iterative algorithms, such as the BESA algorithm, can be implemented in TPB training as part of the training phase to minimize prediction errors arising from cascaded forward and backward integer shaping or prediction operations performed in primary integer shaping. Cascaded forward and backward integer shaping or prediction operations can be performed using TPB or MMR forward and backward integer shaping functions.

[0117] As an example, not a limitation, such as Figure 2B As illustrated, in TPB training, given the input values ​​of the training input or source image in the input / original domain or color space represented by color cube 202. The TPB forward-shaping function 212 can be applied to obtain or generate the corresponding forward-shaping value of the forward-shaping image (represented as "forward-shaping object") in the integer domain or color space. Additionally, given a forward-shaped image ("forward-shaped object"), the corresponding forward-shaped value... The TPB backward shaping function 214 can be applied to obtain or generate the corresponding reconstructed or backward-shaped values ​​of the reconstructed or backward-shaped image in the reconstructed or output / target domain or color space. The constructed or output / target domain or color space can be the same as the input / original domain or color space.

[0118] The optimized values ​​of the operation parameters in the forward and backward shaping functions of TPB can be obtained or generated through the least squares solution of an optimization problem that minimizes the difference or prediction error between the reconstructed image generated by the forward and backward shaping operations and the training input or source image.

[0119] In some operational scenarios, the prediction error can be calculated as In some other operational scenarios, the prediction error can be calculated as... ,in, This indicates that the input values ​​include those from the input / raw domain or color space. ; This represents the reconstructed values, including those in the input / original domain or color space. ; This indicates that the forward integer function is in the form of... This represents the spatial gradient at a point (e.g., a 3 × 3 matrix, etc.), which will be discussed in more detail later. Adding spatial gradients This allows the iterative algorithm to converge faster than without a spatial gradient, because adding a spatial gradient to the prediction error helps update or accelerate the changes in the shaped value during successive iterations.

[0120] like Figure 2BAs illustrated, for each iteration, the iterative algorithm may include propagating the prediction error calculated in box 218 (represented as "Error Propagation + Gradient Determination") to an integer domain or color space with an integer value. For example, the prediction error may be subtracted from the integer value as... Then, during the TPB training phase, the updated shaped values... It is used as input for the next iteration, or for the next execution of box 218 (“Error Propagation + Gradient Determination”).

[0121] More specifically, in each iteration, the updated shaped value It can remain unchanged and serve as the target value for the operating parameters used to optimize the TPB forward shaping function (212). Meanwhile, the input / original value... It can continue to serve as the target value for the total operating parameters used to optimize the forward and backward shaping functions (212 and 214) of TPB.

[0122] The aforementioned operations can be iterated until the optimized values ​​of the operation parameters (or TPB coefficients) of the forward and backward integer functions (212 and 214) of TPB converge, such as until the total value change measure between two consecutive iterations is less than the minimum value change threshold.

[0123] In some operational scenarios, MMR shaping functions can be used for forward and backward shaping operations as described in this paper. In these scenarios, optimized values ​​of the operational parameters (or MMR coefficients) of the MMR forward and backward shaping mapping / functions can be generated or obtained, similar to... Figure 2B The diagram illustrates the method for generating or obtaining the optimized values ​​of the operation parameters (or TPB coefficients) of the TPB forward and backward shaping mapping / functions.

[0124] During the training phase, one or more reference maps corresponding to one or more corresponding forward shaping functions can be specified or defined using optimized values ​​of operating parameters such as the set of control variables used by these forward shaping functions and / or TPB parameters / coefficients and / or MMR parameters / coefficients. For example, through methods such as Figure 2B The iterative algorithm illustrated can generate or obtain some or all optimized values ​​of the operating parameters.

[0125] Each of one or more reference maps maps the codewords of the input image distributed in the input domain or color space (such as a color cube) to the codewords distributed in the shaped object or form (such as a cylindrically shaped object).

[0126] Represented as The unnormalized spatial gradient can be obtained by taking the reference map (or forward shaping function) in the input values ​​in the input domain or color space (e.g., three, etc.) channels / dimensions. The derivative is used to calculate the value at each point, as shown below:

[0127] (8)

[0129] These spatial gradients can be normalized in each channel / dimension, as shown below:

[0130] For each dimension (9)

[0132] The normalized spatial gradient constitutes the Jacobian matrix of the reference mapping (or forward shaping function) (denoted as...). or ), as shown below:

[0133] (10)

[0135] in, f The transformation function form representing the reference map (or forward shaping function) is as shown in the per-channel transformation map in expression (1) (where... f x , f y and f z As f (vector / matrix components), MMR mapping / function in expression (3) or TPB mapping / function in expression (5); , , Indicates relative to the color channel x , y , z The partial reciprocal or difference.

[0136] The normalized spatial gradient can be rewritten in functional form using a vector and a Jacobian matrix, as shown below:

[0137] (11)

[0139] In some operational scenarios, the spatial gradient of the reference map (or forward shaping function) can be numerically computed or generated. For each point in the shaping domain or color space (as represented by the forward-shaped value generated from the corresponding point represented by the corresponding input value in the input / original domain or color space), neighboring or adjacent points can be identified (e.g., the nearest, etc.) and used to compute one or more distances or distance vectors in each dimension / channel between said point and its neighborhood. The distance vectors described herein can be normalized to one (1).

[0140] In operational scenarios where spatial gradients are not used to compute the prediction error between the reconstructed image and the input image that leads to reconstruction using forward and backward shaping, the prediction errors in both forward and backward shaping may diverge or converge relatively slowly.

[0141] Conversely, in operational scenarios where spatial gradients are used to compute the prediction error between the reconstructed image and the input image that leads to reconstruction using forward and backward shaping, such as Figure 3A The prediction error in the forward shaping illustrated diverges because the prediction error propagates relatively efficiently to the shaped value. Meanwhile, as... Figure 3B As illustrated, the prediction error in backward shaping converges as more and more iterations are performed.

[0142] Target / reference shape or form in the shaping domain

[0143] For example, forward and backward shaping functions based on MMR or TPB can model the forward and backward mapping process between (1) the input / raw high-bit depth domain or color space and (2) the low-bit depth shaping domain or color space. MMR or TPB shaping functions can be used with any function in various shapes or forms with codewords distributed in the shaping domain.

[0144] Under the techniques described herein, forward and backward shaping functions can be constructed on the following two inputs: (1) an input shape or form, such as a color cube representing the input / original domain or color space, and (2) a target / reference (shaped) shape or form, such as a cylinder, torus, torus, twisted torus, etc. representing the shaping domain.

[0145] Given these two inputs, you can perform the following: Figure 2B The illustrated gradient-based BESA algorithm / method and other iterative algorithms / methods obtain optimized values ​​of the operating parameters (e.g., MMR coefficients, TPB coefficients, etc.) of the shaping function (e.g., MMR, TPB, etc.).

[0146] A non-wound axis can refer to a single channel or dimension or a combination of multiple channels or dimensions (e.g., linear, etc.) in the input video within the input domain, while a wound axis can refer to an axis mapped from a non-wound axis in the wound shaping domain through (multiple) wound transformations, having more available codewords than the non-wound axis. In some operational scenarios, the wound axis in the shaping domain has a different geometry than the non-wound axis; wherein the total number of available codewords on the wound axis is greater than the total number of available codewords on the non-wound axis. In various operational scenarios, many different target / reference (shaped) shapes or forms can be used for wound shaping. Wound shaping can be used to transform perceptibly significant channels or visually dominant color channels (e.g., dimensions, axes, luminance or brightness channels, specific color channels in the color space, linear or nonlinear combinations of two or more channels, etc.) in the input / original high-bit depth domain or color space into the same visually dominant color channel, which is not necessarily represented as a linear axis, but rather as a nonlinear (or wound) axis of a reference shape or form in the low-bit depth shaping domain. Perceptually significant or visually dominant color channels can be statically specified or dynamically identified at runtime. Dominant color channels, which are initially represented as axes with fewer codewords in the input / original domain or color space (such as linear axes in a color cube), are “banded” or “wound” into nonlinear or wrapped axes of reference shapes or forms (such as circular axes, the longest axis of a toroidal shape, the longest axis of a toroidal shape, etc.).

[0147] As a result, the dominant color channel represented in a reference shape or form can obtain more codewords in the shape domain with the reference shape or form compared to color channels represented by linear axes under other methods. The relatively large number of codewords obtained by binding or wrapping the linear axes of the color cube representing the dominant color channel into nonlinear axes of the reference shape or form (possibly at the cost of reducing the available codewords for non-dominant color axes) can be used to avoid observer-perceptible binding or false contour artifacts in the reconstructed image.

[0148] Examples of shapes or forms in the shaping domain described herein (e.g., reference, shaped, etc.) that can be used to provide a nonlinear axis with a relatively large number of available codewords for the perceptually dominant channel in an input image may include, but are not limited to, any of the following: complete torus, partial torus, complete cylinder, partial cylinder, etc.

[0149] For illustrative purposes only, the codewords encoded in each of the input / raw and shaped domains, and in each of the input and shaped video signals, can be normalized to a value range such as [0, 1]. Additionally, alternatively, or alternatively, the image container can be used to store (e.g., final, intermediate, input, shaped, etc.) image content, as if using codewords of a non-negative integer type in each of the input / raw and shaped domains.

[0150] Transformation functions can be used to map input codewords in an input shape or form (such as a color cube) to a reference shape or form with a nonlinear axis that is perceptually dominant. While the reference shape or form can be visualized within the color cube, it should be noted that the shaped codewords are restricted to residing within the reference shape or form and are not permitted to reside outside the reference shape or form in other parts of the color cube used for visualization purposes.

[0151] For illustrative purposes only, normalized linear color coordinates (e.g., R, G, and B coordinates, respectively, denoted as x, z, and y) can be used in the coordinate system of a color cube as described herein. It should be noted that some or all of the techniques described herein can be similarly applied to nonlinear RGB domains or color spaces, as well as nonlinear RGB or non-RGB coordinates in linear or nonlinear non-RGB domains or color spaces.

[0152] The circular shape or form in the shaping domain

[0153] In some operational scenarios, transformation functions can be used to bind or wrap one or more input axes in the input color cube into a reference shape or form (or shape domain) represented by the toroidal shape in the visual color cube.

[0154] For example, a transformation function can be used to transform coordinates ( x , y , z The three input axes of the input color cube are bound or wrapped together by coordinates ( , , The reference shape or form (or integer domain) represented by the complete torus in the visual color cube is shown below:

[0155]

[0156] (12-1)

[0157]

[0158] (12-2)

[0159]

[0160] (12-3)

[0161]

[0162] (12-4)

[0163] in, z Indicates three input channels ( x , y , z The dominant channel or axis that the viewer can perceive in the image; Indicates the winding axis that transitions from the dominant channel or axis perceived by the viewer; , and Represents three axes in a visualization space or cube, in which a winding shape or form with winding axes can be visualized.

[0164] In the above expression (12), the parameter denoted as w is used to control the width of the complete ring. The example data range for w is between [0, 0.5]. If w is 0.5 (corresponding to the widest width), the inner ellipse of the complete ring shrinks to or becomes a point. On the other hand, if w is 0 (corresponding to the narrowest width), the complete ring shrinks to or becomes a one-dimensional ellipse, such as the outer ellipse of the complete ring.

[0165] Figure 3C illustrates an example view of the complete torus visualized in the visualization color cube. As can be seen in the above expression (12), the dominant channel z Binding or wrapping by periodic functions such as sine or cosine functions. When the width of the ring controlled by parameter w is very thin and very close to the boundary, given the same 8-bit image container, by controlling the dominant channel z The available codewords provided by binding or wrapping are approximately 4.44 times the available codewords in the unbound or wrapped axis of the integer domain, as shown below:

[0166] (13)

[0168] A complete toroidal shape can be used to completely occupy the integer domain along the longest direction (e.g., projected diagonally) provided by the visual color cube corresponding to the image container. However, a drawback is that the complete toroidal shape is end-to-end and therefore has no error tolerance. Since the starting point of the complete toroidal shape connects back to its ending point, any small value change or error (e.g., caused by image processing performed within the integer domain) can significantly alter the values ​​in the reconstructed image. A black pixel in the input image might become a white pixel in the reconstructed image due to a small value change or error introduced (e.g., by image processing operations) into the entangled and connected shape within the integer domain. This is analogous to the "overflow" and "underflow" problems in computation.

[0169] In some operational scenarios, gaps can be inserted or used in the integer domain represented by a toroidal shape to help avoid or improve such overflow or underflow problems, and to help tolerate or better handle small value variations or errors that may be introduced by various image processing operations involved in the end-to-end image transfer between the received input image and the reconstructed image to approximate the input image.

[0170] The parameter g can be introduced into the transformation function to control the size or range of the gaps inserted into the toroidal shape. The parameter g can have values ​​in the range [0, 1]. If g is 1, the toroidal shape shrinks or becomes flat. If g is 0, the toroidal shape becomes a complete toroidal shape.

[0171] Indicates the input color cube ( x , y , z ) to a ring-shaped shaping domain with gaps ( , , The transformation function of the mapping of ) can be given as follows:

[0172]

[0173] (14-1)

[0174]

[0175] (14-2)

[0176]

[0177] (14-3)

[0178]

[0179] (14-4)

[0180] in, zIndicates three input channels ( x , y , z The dominant channel or axis that the viewer can perceive in the image; Indicates the winding axis that transitions from the dominant channel or axis perceived by the viewer; , and Represents three axes in a visualization space or cube, in which a winding shape or form with winding axes can be visualized.

[0181] Figure 3D illustrates an example view of a partial annulus visualized in a color cube. As can be seen in expression (14) above, the dominant channel... z It is bound or wrapped by periodic functions such as sine or cosine functions. Additionally, gaps are inserted to prevent "overflow / underflow". When the width of the ring, controlled by parameter w, is very thin and very close to the boundary, given the same 8-bit image container subtracted by the gap portion, by adjusting the dominant channel... z The available codewords provided by binding or wrapping are approximately 4.44 times the available codewords in the unbound or wrapped axis of the shaping domain. Times, as shown below:

[0182] (15)

[0184] Another solution to the overflow / underflow problem is to use multiple small gaps interleaved within the complete ring, instead of inserting a single, relatively large gap into the complete ring.

[0185] Similar to a toroidal shape with a single gap, a parameter `w` can be introduced to control the width of a toroidal shape with multiple gaps or gap sections. The value of `w` can range between [0, 0.5]. If `w` is 0.5, the inner ellipse will become a point. If `w` is 0, the toroidal shape will become an ellipse, which is the outer ellipse of the toroidal shape. Additionally, a parameter `g` can be introduced to control the total gap size or range of the gap sections, such as four gap sections placed in four different locations or positions on the toroidal shape. If the total gap size or range `g` is 0.2, then four gaps or gap sections of 0.05 are placed in four different locations or positions on the toroidal shape.

[0186] Indicates the input color cube ( x , y , z ) to the shaping domain with four gaps in a ring shape ( , , The transformation function of the mapping of ) can be given as follows:

[0187] if

[0188]

[0189] (16-1)

[0190] if

[0191]

[0192] (16-2)

[0193] if

[0194]

[0195] (16-3)

[0196] if

[0197]

[0198] (16-4)

[0199]

[0200] (16-5)

[0201]

[0202] (16-6)

[0203]

[0204] (16-7)

[0205] in, z Indicates three input channels ( x , y , z The dominant channel or axis that the viewer can perceive in the image; Indicates the winding axis that transitions from the dominant channel or axis perceived by the viewer; , and Represents three axes in a visualization space or cube, in which a winding shape or form with winding axes can be visualized.

[0206] Figure 3E illustrates an example view of a partial annulus visualized in a color cube. As can be seen in expression (16) above, the dominant channel... zIt is bound or wrapped by periodic functions such as sine or cosine functions. Additionally, four gap sections are inserted to prevent "overflow / underflow". When the width of the ring, controlled by parameter w, is very thin and very close to the boundary, given the same 8-bit image container subtracted by the gap sections, the dominant channel... z The available codewords provided by binding or wrapping are approximately 4.44 times the available codewords in the unbound or wrapped axis of the shaping domain. Times, as shown below:

[0207] (17)

[0208] While toroidal shapes may have a relatively high theoretical maximum available codewords, they can have drawbacks. For example, for codewords that deviate further from the winding axis (e.g., projected onto the diagonal of a visualized color cube), the quantization error can become increasingly large in the integer domain represented by a toroidal shape. Since the total mapping error can become relatively large, the average performance of all color values ​​can be significantly affected by the toroidal shape. Furthermore, toroidal shapes may encounter numerical problems such as numerical instability when using TPB or MMR integer functions to model forward and backward integer mappings or operations.

[0209] cylindrical shape or form in the shaping domain

[0210] In some operational scenarios, to help solve problems in toroidal shapes or forms, transformation functions can be used to bind or wrap one or more input axes in the input color cube into a reference shape or form (or shape domain) represented by a cylindrical shape in the visualized color cube.

[0211] For example, a transformation function can be used to transform coordinates ( x , y , z The three input axes of the input color cube are bound or wrapped together by coordinates ( , , The reference shape or form (or integer domain) represented by the complete cylinder in the visualization color cube is shown below:

[0212]

[0213] (18-1)

[0214]

[0215] (18-2)

[0216]

[0217] (18-3)

[0218]

[0219] (18-4)

[0220] in, z Indicates three input channels ( x , y , z The dominant channel or axis that the viewer can perceive in the image; Indicates the winding axis that transitions from the dominant channel or axis perceived by the viewer; , and Represents three axes in a visualization space or cube, in which a winding shape or form with winding axes can be visualized.

[0221] In the above expression (18), the parameter denoted as w is used to control the width of the complete cylinder. The example data range for w is between [0, 0.5]. If w is 0.5 (corresponding to the widest width), the inner circle of the complete cylinder shrinks to or becomes a point. On the other hand, if w is 0 (corresponding to the narrowest width), the complete cylinder shrinks to or becomes a one-dimensional circle, such as the outer circle of the complete cylinder.

[0222] Figure 3F illustrates an example view of the complete cylinder visualized in the visualization color cube. As can be seen in the above expression (18), the dominant channel z Binding or winding by periodic functions such as sine or cosine functions. When the width of the cylinder, controlled by parameter w, is very thin and very close to the boundary, given the same 8-bit image container, by controlling the dominant channel... z The available codewords provided by binding or wrapping are close to the available codewords in the unbound or wrapped axis of the shaping domain. Times, as shown below:

[0223] (19)

[0225] While a full cylindrical shape can be used to completely occupy the full-length domain provided by the visual color cube corresponding to the image container (e.g., projected as a diagonal direction), the same problems encountered in a full torus shape (such as overflow / underflow issues) may also exist in a full cylindrical shape.

[0226] In some operational scenarios, gaps can be inserted or used in the cylindrical shape of the integer domain to help avoid or improve such overflow or underflow problems, and to help tolerate or better handle small value variations or errors that may be introduced by various image processing operations involved in the end-to-end image transfer between the received input image and the reconstructed image to approximate the input image.

[0227] The parameter g can be introduced into the transformation function to control the size or range of the gap inserted into the cylindrical shape. The parameter g can have values ​​in the range [0, 1]. If g is 1, the cylindrical shape shrinks or becomes flat. If g is 0, the cylindrical shape becomes a complete cylinder.

[0228] Indicates the input color cube ( x , y , z ) to a cylindrical shaping domain with gaps ( , , The transformation function of the mapping of ) can be given as follows:

[0229]

[0230] (20-1)

[0231]

[0232] (20-2)

[0233]

[0234] (20-3)

[0235]

[0236] (20-4)

[0237] in, z Indicates three input channels ( x , y , z The dominant channel or axis that the viewer can perceive in the image; Indicates the winding axis that transitions from the dominant channel or axis perceived by the viewer; , and Represents three axes in a visualization space or cube, in which a winding shape or form with winding axes can be visualized.

[0238] Figure 3G illustrates an example view of a portion of a cylinder visualized in a color cube. As can be seen in the above expression (20), the dominant channel zIt is bound or wrapped by periodic functions such as sine or cosine functions. Additionally, gaps are inserted to prevent "overflow / underflow". When the width of the ring, controlled by parameter w, is very thin and very close to the boundary, given the same 8-bit image container subtracted by the gap portion, by adjusting the dominant channel... z The available codewords provided by binding or wrapping are close to the available codewords in the unbound or wrapped axis of the shaping domain. Times, as shown below:

[0239] (twenty one)

[0241] The theoretical codeword increments for the different example shapes discussed in this article are listed in Table 1 below. ).

[0242] Table 1

[0243]

[0244] As can be seen above, a complete torus has the best codeword increase, while a cylinder with a gap has the worst. On the other hand, for codewords located at or near the boundaries of the color cube, torus shapes suffer more color distortion because these codewords are mapped to ellipses smaller than those located elsewhere. In contrast, cylindrical shapes provide relatively uniform distortion across codewords at different locations. Additionally, as noted, shapes without gaps may suffer from overflow / underflow, thus leading to color artifacts. Utilizing these observations, in some operational scenarios, torus or cylindrical shapes with gaps can be used as target (e.g., reference, shaped, etc.) shapes or forms in the shaping domain to increase the availability of codewords in the dominant axis. Specifically, in some operational scenarios, cylindrical shapes with gaps can be used to increase the available codewords with relatively uniform distortion.

[0245] It should be noted that, in various embodiments, in order to increase the availability of codewords in the dominant axis, shapes or forms other than cylindrical or toroidal shapes or forms can be used as target (e.g., reference, shaped, etc.) shapes or forms in the shaping domain.

[0246] Construct content-related optimized integer functions

[0247] A simple approach to constructing an integer function is to use a static mapping that maps a fully input 3D color cube representing the input / original domain to a predefined reference or shaped form within the integer domain, regardless of or depending on the content-dependent color distribution in the input / original domain. This is likely the fastest solution for performing integer shaping.

[0248] On the other hand, if actual content information, such as codewords or color distribution in the input image, is considered when selecting a specific reference or shaped shape or form, and when selecting specific operational parameters related to that reference or shaped shape or form, the prediction error can be further improved or reduced. The shaping function generated under this method can be called a dynamic shaping function.

[0249] Different approaches, such as (1) fully dynamic approach / solution and (2) two-stage approach / solution, can be implemented or applied to construct dynamic integer functions.

[0250] To fully capture or account for specific codeword or color distributions in the input image of the input video signal, a fully dynamic solution can be executed on-the-fly during the deployment or prediction phase. Primary and secondary shaping can be combined as a single-trigger or processing unit to generate specific values ​​for operational parameters used in the shaping function, such as scaling factors / values, MMR coefficients, and / or TPB coefficients. While content-relevant and capable of fully utilizing or accounting for the actual codeword or color distribution in the input image, a fully dynamic solution may not use pre-built shaping functions obtained or generated during the training phase, and thus may result in relatively high computational costs and complexity.

[0251] In a fully dynamic solution, four different methods / approaches can be used to group the input codewords or colors before passing information about the codewords or colors to the shaping function optimization process.

[0252] The first of the four methods / approaches for implementing a fully dynamic solution can be called the pixel-based approach, in which all pixels in the input image or their codewords (e.g., ...) are used. Figure 3H The image shown is considered as input to TPB training (at runtime, not during the training phase). Because this method involves using all pixels or codewords during the optimization process, it can result in relatively large storage space and relatively high computational cost or complexity in the optimization process used to generate or obtain optimized values ​​for operating parameters (such as MMR or TPB coefficients in the forward and backward shaping functions).

[0253] The second of the four methods / approaches for implementing a fully dynamic solution can be termed the gridded color clustering averaging method, in which pixels or their codewords in the input image are first divided into multiple (e.g., uniform, non-uniform, etc.) non-overlapping (e.g., relatively small, etc.) cubes or clusters. An average value can be calculated for each cube / cluster. Figure 3IAs illustrated, the individually computed average of a cube / cluster can be considered as input to a shaping function optimization process that acts on the optimized values ​​of operating parameters (such as MMR or TPB coefficients in forward and backward shaping functions). The cube / cluster can be non-empty, comprised of multiple pixel values ​​or codewords exceeding a minimum number threshold, etc. Exemplary advantages of this method may include reduced memory usage and lower computational cost and complexity.

[0254] The third of the four methods / approaches for implementing a fully dynamic solution can be called the gridded color clustering center method. For example... Figure 3J The illustration does not show the average value of each cube or cluster, but rather considers the center of each cube or cluster as the input to a shaping function optimization process that acts on the optimized values ​​of the operating parameters (such as the MMR or TPB coefficients in the forward and backward shaping functions). As in the second method, the cubes / clusters in the third method can be non-empty, occupied by multiple pixel values ​​or codewords exceeding a minimum threshold, etc.

[0255] The fourth of the four methods / approaches for implementing a fully dynamic solution can be called the gridded clustering augmented center method. To mitigate or improve numerical problems associated with relatively small non-empty cubes or the total number of clusters, empty clusters near non-empty clusters can be selected. The centers of non-empty clusters and the empty clusters near them (such as...) can be considered. Figure 3K The diagram shows the input to the shaping function optimization process that acts on the generated or obtained optimized values ​​of the operating parameters (such as the MMR or TPB coefficients in the forward and backward shaping functions).

[0256] Figure 3L The illustration shows an example distribution of pixels or their codewords in an input image within a color cube representing the input / original domain or color space.

[0257] Figure 3M The illustration shows the process from the input image. Figure 3L Example points are generated from the distribution of pixels or codewords. In some operational scenarios, in addition to the data points corresponding to non-empty cubes, additional (data) points can be generated for empty neighbor cubes using a gridded clustering augmentation center method. These points generated from the input image can be used as part of the input in the optimization process for generating or obtaining optimized values ​​for operational parameters (such as the MMR or TPB coefficients in the forward and backward shaping functions of the input image) (e.g., in TPB training).

[0258] Figure 3N The illustration shows the shaping process from... Figure 3MExamples of point transformations are forward-transformed points. During optimization, these forward-transformed points, mapped from points generated from the input image through wrapping shaping, can be used as part of the input along with points generated from the input image (e.g., in TPB training, etc.).

[0259] Two-stage plastic surgery

[0260] To avoid numerical stability issues that may be encountered in fully dynamic methods, a two-stage shaping method can be used. As mentioned earlier, a two-stage shaping method may include two stages / processes: (1) an offline training stage / process, and (2) an online deployment or prediction stage / process.

[0261] Primary shaping (or the second stage) results in relatively high computational complexity / cost. To facilitate or induce the lightest computation during the deployment or prediction phase / process, specific values ​​for some or all operational parameters, such as MMR or TPB shaping coefficients, can be pre-trained offline during the training phase / process and stored as an operational parameter dataset in an operational parameter data store accessible to the relatively lighter secondary shaping (or the first stage) during the deployment or prediction phase / process. Secondary shaping (the first stage) can be performed in response to receiving the actual input image during the deployment or prediction phase.

[0262] As mentioned earlier, during the training phase, primary integer functions such as forward and backward integer functions based on TPB can be trained in two parts: (1) reference map creation and (2) TPB training.

[0263] Reference mapping creation involves two sets of control variables. The first set of control variables involves selecting the reference mapping shape or form from the different reference mapping shapes or forms supported in the winding shaping domain (such as toroidal shapes, cylindrical shapes, etc.). The first set of control variables used for shape selection can be represented as s.

[0264] The second set of control variables pertains to the geometric parameters of a given reference mapping shape or form. These geometric parameters determine the width, height, and other properties of the given reference mapping shape or form, such as width w, gap size, or range g.

[0265] These two sets of control variables affect quantization and / or dequantization during forward and backward shaping or mapping operations, thereby affecting the prediction accuracy of the reconstructed output image.

[0266] A list of shaped objects can be created, defined, or designed during the training phase (e.g., Figure 2AMultiple available reference mapping shapes / forms in (e.g., 206). Operational parameters represented in two sets of control variables (e.g., (s, w, g) for different reference mapping shapes or forms supported in the shaped object list (206)) can be trained during the training phase and stored as templates (or operational parameter datasets) in the operational parameter data store.

[0267] Each template can include specific combinations (e.g., different values) of operational parameters (e.g., generated during training) from two sets of control variables (e.g., (s, w, g)). The specific combinations of values ​​in the template can be specifically chosen during the training phase, or selected from multiple candidate value combinations or sets to achieve the best prediction results during training and produce minimal error on the training / test data, for example, using gradient-based BESA algorithms / methods.

[0268] During the deployment or prediction phase, given an input / original image, a specific template can be determined or selected from multiple templates to provide specific values ​​for the operational parameters in primary shaping. Additionally, during the deployment or prediction phase, specific values ​​for the operational parameters in secondary shaping can be determined or generated at runtime to minimize prediction errors in the (final) reconstructed image that depicts the same visual semantic content as the input / original image.

[0269] The secondary shaping described in this article may include applying the following linear scaling steps to input codewords in the input / original image:

[0270] Step 1: Given a minimum value (represented as...) v L ) and maximum value (represented as v H The channel, within the value range [ v L , v H Scaling codewords from [0, 1].

[0271] Step 2: Specify the scaling factor x Scale the value range [0, 1] to [ ].

[0272] It should be noted that the scaling factor described in this article x Each channel can be scaled differently because each such channel can have a different range of values. v L , v HUsing the scaling factors described herein, an input / raw image in an input video signal can be scaled to generate a scaled image in a scaled signal. The scaled image in the scaled signal can then be fed or provided to primary shaping (e.g., MMR, TPB, etc.).

[0273] The prediction error that measures the difference between the reconstructed image and the input / original image can come from the following two offsetting sources.

[0274] First, shaping or mapping operations (such as those implemented using MMR or TPB algorithms / methods) can introduce relatively high reconstruction errors at the boundaries of a specific reference or shaped (forward) shape in the shaping domain. The smaller the scaling factor / value, the fewer boundary points are near the boundaries of a specific reference or shaped (forward) shape. Therefore, reducing the scaling (factor) value can reduce this type of reconstruction error.

[0275] Second, quantization error increases as the scaling factor value decreases because multiple input codewords / values ​​from the input / raw domain are more likely to be mapped to a single codeword / value in the integer domain. Therefore, increasing the scaling factor value can reduce this type of error.

[0276] An optimized solution (e.g., an optimized value for the scaling factor) can depend on how the cost function used to measure the prediction error is defined or specified, and on the assumptions made using the cost function. Additionally, alternatively, or otherwise, different search algorithms can be developed or used to support or achieve different optimization objectives.

[0277] In some operational scenarios, a separate cost function can be defined for each channel in one or more channels of the input / source domain or color space.

[0278] Figure 2E The diagram illustrates an example flow for optimizing scaling factor values, for example, for the dominant channel in the input / source domain or color space. For illustrative purposes only, cost functions, such as convex-like functions of the scaling factor value, can be used in scaling factor value optimization. It should be noted that the cost function does not have to be a purely convex function or a globally convex function, but can monotonically increase or decrease to an inflection point and then move in the opposite direction. Relatively fast search algorithms can be used in conjunction with the cost function to search for the optimal solution or the optimal scaling factor value.

[0279] The optimization problem is to find an optimal scaling factor value that will be used with a template selected or specified in the primary and secondary shaping processes to achieve the minimum prediction error in the reconstructed image compared to the input / original image.

[0280] Let the scaling factor x function f ( x) is a cost function that measures the prediction error generated throughout the process of scaling, forward shaping, backward shaping, reverse / inverse scaling, etc.

[0281] Optimized scaling factor value It can be generated or obtained as a solution to an optimization / minimization problem, which is minimized by a cost function. f ( x The prediction error of the measurement is as follows:

[0282] (twenty two)

[0284] In some operational scenarios, the golden section search algorithm can be applied to the above expression (22) to determine the minimum value of the cost or error function.

[0285] like Figure 2E As illustrated, box 252 includes performing initialization to define the left boundary of the search. = 0 and search right boundary = 1, and calculate or generate its corresponding cost / error value. and .

[0286] Two interior points can be selected or defined { }and{ }. Also calculate or generate the interior points { }and{ Corresponding cost / error value and .

[0287] It is possible and Select or define interior points between { } satisfies the following relationship:

[0288] (twenty three)

[0290] It is possible and Select or define interior points between { } satisfies the following relationship:

[0291] (twenty four)

[0293] and The cost / error values ​​at these four points can be used to determine whether the convergence / exit criteria are met and / or where the scaling factor value for the optimization that meets the convergence / exit criteria can be found.

[0294] In some operational scenarios, the convergence / exit criteria are met if the following conditions are true:

[0295] Where B is the positive constant / threshold. (25)

[0297] In response to determining the convergence / exit criteria satisfying expression (25), from four points { , Identify or select the minimum cost / error value from the cost / error values. Generate or obtain the point corresponding to the minimum cost / error value as the optimal solution. The process ends or exits.

[0298] In response to determining that the convergence / exit criteria in expression (25) are not satisfied, at four points { , In the context of}, the minimum cost / error (function) value among its corresponding cost / error values ​​is determined to belong to the point. still (or belongs to a point) still .

[0299] In response to determining the minimum cost / error (function) value among its corresponding cost / error values, the point belongs to... or The process flow proceeds to frame 254 to begin the next iteration. In the next iteration, the three points { , } was reused; Become a new point ; Become a new point Furthermore, box 258 includes generating new points based on the above expression (23). The process flow then returns to box 252.

[0300] Otherwise, in response to determining the minimum cost / error (function) value among its corresponding cost / error values, it belongs to the point. or The process flow proceeds to frame 256 to begin the next iteration. In the next iteration, the three points { , } was reused; Become a new point X 1; X 3. Become a new point X2. Further, box 260 includes generating new points based on the above expression (24). The process flow then returns to box 252.

[0301] In some operational scenarios, a cross-color channel cost function can be defined for one or more channels of the input / source domain or color space.

[0302] Figure 2F The diagram illustrates an example flow for optimizing scaling factor values, for example, for three color channels in the input / source domain or color space. A cost function (which may or may not be a convex function) can be used for scaling factor value optimization.

[0303] The optimization problem is to find an optimal scaling factor value that will be used with a template selected or specified in the primary and secondary shaping processes to achieve the minimum prediction error in the reconstructed image compared to the input / original image.

[0304] Let three scaling factors function The cost function measures the prediction error generated throughout the process of scaling, forward shaping, backward shaping, reverse / inverse scaling, etc.

[0305] Optimized scaling factor value It can be generated or obtained as a solution to an optimization / minimization problem, which is minimized by a cost function. The prediction error of the measurement is as follows:

[0306] (26)

[0308] In some operational scenarios, the Nelder-Mead algorithm can be applied to the above expression (26) to determine the minimum value of the cost or error function.

[0309] Figure 2F Box 262 includes four initial points for initializing the scaling factors, each represented by one of the following four vectors: = (0.3, 0.3, 0.3) = (0.3, 0.3, 1) = (0.3,1,1) and = (1,1,1).

[0310] Box 264 includes the four points based on their cost / error (function) values. , , , } Classify or sort the points, and rename them according to their cost / error (function) values. , , , }, so that:

[0311] (27)

[0313] Box 266 includes determining whether the convergence / exit criteria are met. In some operational scenarios, the convergence / exit criteria are met if the following conditions are true:

[0314] (28)

[0316] in, Th This indicates a preset or dynamically configurable threshold, such as 100, 50, etc.

[0317] In response to determining the convergence / exit criteria satisfying expression (28), from four points { , , , Identify or select the minimum cost / error value from the cost / error values. Generate or obtain the point corresponding to the minimum cost / error value as the optimal solution. The process ends or exits.

[0318] Box 268 includes, in response to determining that the convergence / exit criteria in expression (28) are not satisfied, calculating , and The centroid of the point. This centroid is represented as... As shown below:

[0319] (29)

[0321] Box 270 includes The reflection point is calculated as follows As shown below:

[0322] (30)

[0324] Box 272 includes determining the reflection point. Is it better (in terms of prediction error) than the second worst in this example? But not the optimal point. ),or

[0325] (31)

[0327] In response to determining the reflection point Better than the second worst ( The process flow proceeds to frame 264. = Otherwise, the process proceeds to frame 274.

[0328] Box 274 includes determining whether the reflection point error value is better than the optimal point ( ), as shown below:

[0329] (32)

[0331] Box 282 includes a response to determining the reflection point error value better than the optimal point ( The expansion point x is calculated as follows. e :

[0332] (33)

[0334] Box 284 includes determining whether the extension point is superior to the reflection point, as shown below:

[0335] (34)

[0337] In response to the determination that the extension point is superior to the reflection point, the process flow proceeds to frame 264. = Otherwise, in response to the determination that the extension point is not superior to the reflection point, the process flow proceeds to box 264. = .

[0338] Box 276 includes a response to determining that the error value of the reflection point is not better than the optimal point ( The contraction point is calculated as follows. :

[0339] (35)

[0341] Box 278 includes determining whether the contraction point is better than the worst point, as shown below:

[0342] (36)

[0344] In response to the determination that the contraction point is better than the worst point, the process flow proceeds to frame 264. = .

[0345] Box 280 includes a method to shrink (candidate) points by replacing all points except the best point (x1) in response to determining that the shrinking point is not better than the worst point, as follows:

[0346] ,for i = 2, 3, 4 (37)

[0348] Then, the process flow uses these replaced points to proceed to box 264.

[0349] Cut-off field transformation

[0350] The winding shaping described in this paper can map or shape an input image from (e.g., linear, perceptual, etc.) a high-bit depth domain or color space to (e.g., linear, perceptual, etc.) a lower-bit depth winding nonlinear shaping domain or color space. However, a winding domain with nonlinearity can make it difficult or even impossible to perform (e.g., commonly used, etc.) image processing algorithms or operations. To address this issue, truncated field transforms can be designed and used to convert the winding domain into a truncated transform domain, such as the linear domain in which these image processing algorithms or operations can be performed. In some operational scenarios, truncated field transforms allow neighboring codeword values ​​in the input domain to remain neighboring codeword values ​​in the truncated transform domain before winding.

[0351] For different target shapes or forms (e.g., reference, shaped, etc.), different truncation field transformations can be designed or used to ensure that image processing algorithms or operations that generate codeword values ​​outside the selected target shape or form can remain within that range.

[0352] For illustrative purposes only, the cutoff field transformation can be applied to a cylindrical shape, which serves as an example target shape or form for winding shaping. It should be noted that in various other embodiments, some or all of the techniques associated with the cutoff field transformation can be applied to target shapes or forms other than cylindrical shapes.

[0353] make b This represents the bit depth of the nonlinear shaping domain (e.g., 8 bits). Let ( , , This represents the coordinates of a point in an integer domain (e.g., 8 bits, etc.). A forward truncated field transformation (or simply truncated field transformation) can convert this point (e.g., with 8-bit integer precision, etc.) from the integer domain to a coordinate value representing the coordinates of a point in an integer domain. , , ) truncated field domain (e.g., in 32 or 64 floating point or double precision, etc.).

[0354] As used herein, a truncated field is transformed from a reference or shaped shape or form with a winding axis by a truncated field transformation, and transformed back to the reference or shaped shape or form by an inverse truncated field transformation. Image processing operations (such as those designed or implemented for linear or nonlinear domains without a winding axis) can be applied as intended to generate processed image data in the truncated field. Additionally, optionally, or alternatively, a cropping operation can be performed on the processed image data in the truncated field to avoid or prevent new image data values ​​generated according to the image processing operation from overflowing the boundaries of the reference or shaped shape or the shape transformed back by the inverse truncated field transformation.

[0355] We can first calculate a common variable, which can represent the distance to the center of the cylinder, as shown below:

[0356] (38)

[0358] according to and The specific location or value can be considered in the following eight (8) cases:

[0359] if and and ,but

[0360]

[0361] (39-1)

[0362]

[0363] (39-2)

[0364]

[0365] (39-3)

[0366] if and and ,but

[0367]

[0368] (40-1)

[0369]

[0370] (40-2)

[0371]

[0372] (40-3)

[0373] if and and ,but

[0374]

[0375] (41-1)

[0376]

[0377] (41-2)

[0378]

[0379] (41-3)

[0380] if and and ,but

[0381]

[0382] (42-1)

[0383]

[0384] (42-2)

[0385]

[0386] (42-3)

[0387] if and and ,but

[0388]

[0389] (43-1)

[0390]

[0391] (43-2)

[0392]

[0393] (43-3)

[0394] if and and ,but

[0395]

[0396] (44-1)

[0397]

[0398] (44-2)

[0399]

[0400] (44-3)

[0401] if and and ,but

[0402]

[0403] (45-1)

[0404]

[0405] (45-2)

[0406]

[0407] (45-3)

[0408] if and and ,but

[0409]

[0410] (46-1)

[0411]

[0412] (46-2)

[0413]

[0414] (46-3)

[0415] Where acos() represents the inverse cosine function; asin() represents the arcsine function. The truncated field transform in the above expressions (38) to (46) allows the neighboring codeword values ​​in the input domain to remain the neighboring codeword values ​​in the truncated transform domain transformed from the cylindrical shape by the truncated field transform before being wound into a cylindrical shape.

[0416] After applying a forward truncation field transform to convert points in the wrapped-shape domain to corresponding points in the truncation field domain, image processing operations such as filtering (e.g., averaging filtering) can be performed on the corresponding points in the truncation field domain. Figure 1(e.g., 116) to generate processed points in the truncated field. In some operational scenarios, points in the truncated field can be represented by floating-point or double-precision codeword values ​​to avoid or reduce any information loss.

[0417] After completing the image processing operations, the processed points can be converted back to the wrapped-shape domain using a backward or inverse truncation field transform, as shown below:

[0418]

[0419] (47-1)

[0420]

[0421] (47-2)

[0422]

[0423] (47-3)

[0424] In some operational scenarios, image processing operations may generate processed points in a truncated field that correspond to points located outside the target shape or form representing the nonlinear winding domain (within the visual color cube used to host the nonlinear winding domain or target shape or form). These points outside the target shape or form (in the visual color space) can be truncated to ensure that these points are constrained or located within the target shape or form.

[0425] For illustrative purposes only, the target shape or form is the same as the cylindrical shape in the preceding example. Given w as the cylinder width of the cylindrical shape, and b The distance from the center of the cylindrical shape to the bit depth of the integer domain can be calculated as follows:

[0426] (48)

[0428] For the cropping operation, three different scenarios can be considered, as shown below:

[0429] for ( ),in :

[0430]

[0431] (49-1)

[0432]

[0433] (49-2)

[0434]

[0435] (49-3)

[0436] for ( ),in :

[0437]

[0438] (50-1)

[0439]

[0440] (50-2)

[0441]

[0442] (50-3)

[0443] otherwise

[0444] , , (51)

[0446] Example process flow

[0447] Figure 4A An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform this process flow. In block 402, the image processing system receives an input image of a first depth in the input domain from an input video signal of a first depth, the first depth being higher than a second depth in the shaping domain.

[0448] In block 404, the image processing system performs a forward shaping operation on the input image to generate a forward-shaped image with a second bit depth in the shaping domain. The forward shaping operation includes winding input codewords in the input image along the non-winding axis of the input domain into a shaped codeword in the forward-shaped image on the winding axis of the shaping domain.

[0449] In block 406, the image processing system encodes an image container containing image data obtained from a forward-shaped image into an output video signal of a second bit depth. The image data in the image container causes the receiving device of the output video signal to construct a backward-shaped image of a third bit depth for rendering on a display device. The third bit depth is greater than the second bit depth.

[0450] In an embodiment, the image processing system is further configured to perform the following operations: apply a forward truncation field transform to the forward-shaped image to generate an intermediate image in a truncation field; perform one or more image processing operations on the intermediate image to generate a processed intermediate image; and apply an inverse truncation field transform to the processed intermediate image to generate the image data in the image container.

[0451] In an embodiment, the processed intermediate image includes one or more cropped codeword values ​​generated according to a cropping operation to ensure that all codeword values ​​in the processed intermediate image are within the target spatial shape representing the shaping domain.

[0452] In an embodiment, the second bit depth represents one of the following: 8 bits, 10 bits, 12 bits, or another bit depth lower than the first bit depth.

[0453] In an embodiment, the forward shaping operation is based on a set of forward shaping maps; at least in part based on the input codewords in the input image, a set of operation parameter values ​​used in the set of forward shaping maps is selected from multiple sets of operation parameter values; each of the multiple sets of operation parameter values ​​is optimized to minimize the prediction error of the corresponding training image cluster in multiple training image clusters.

[0454] In an embodiment, the set of forward shaping maps includes one or more of the following: multivariate regression (MMR) mapping, tensor product B-spline (TPB) mapping, forward shaping lookup table (FLUT) or other types of forward shaping mapping.

[0455] In this embodiment, each set of operating parameters is generated based on the Backward Error Subtraction Signal Adjustment (BESA) algorithm.

[0456] In an embodiment, the input codewords in the input image are scaled using one or more scaling factors; the one or more scaling factors are optimized at runtime using a search algorithm based on one or more of the following: the golden section algorithm, the Nelder-Mead algorithm, or another search algorithm.

[0457] In an embodiment, a set of mapping functions is used to map the input shape representing the input domain to a target reference shape representing the shape domain; the target reference shape is selected from multiple different reference shapes corresponding to multiple different candidate shape domains, based at least in part on the distribution of input codewords in the input image.

[0458] In embodiments, the plurality of different reference shapes include at least one of the following: a complete geometry, a complete geometry with a single gap portion inserted, a complete geometry with multiple gap portions inserted, a toroidal shape, a cylindrical shape, a torus shape, a homomorphic shape to a cube, or a non-homomorphic shape to a cube.

[0459] In the embodiments, the input domain represents one of the following: RGB color space, YCbCr color space, perceptual quantization color space, linear color space, or another color space.

[0460] In an embodiment, the first depth represents one of the following: 12 bits, 16 bits or more bits, or another number of bits higher than the second depth.

[0461] Figure 4B An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., encoding devices / modules, transcoding devices / modules, decoding devices / modules, inverse tone mapping devices / modules, tone mapping devices / modules, media devices / modules, inverse mapping generation and application systems, etc.) may perform this process flow. In block 452, the image processing system decodes from a video signal an image container containing image data obtained from a forward-shaped image in the shaping domain, wherein a source image of a first depth in the source domain has been forward-shaped to generate the forward-shaped image of a second depth lower than the first bit depth.

[0462] In block 454, the image processing system applies a backward shaping operation to the image data decoded from the video signal to generate a backward-shaped image in a target domain, the backward-shaped image having a third bit depth higher than the second bit depth, the backward shaping operation including unwinding codewords in the image data along the winding axis of the shaping domain into backward-shaped codewords in the backward-shaped image on the non-winding axis of the target domain.

[0463] In box 456, the image processing system renders a display image generated from the backward-shaped image on a display device.

[0464] In embodiments, computing devices such as display devices, mobile devices, set-top boxes, and multimedia devices are configured to perform any of the methods described above. In embodiments, an apparatus includes a processor and is configured to perform any of the methods described above. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause any of the methods described above to be performed.

[0465] In one embodiment, a computing device includes one or more processors and one or more storage media storing an instruction set that, when executed by the one or more processors, causes any of the methods described above to be performed.

[0466] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.

[0467] Example computer system implementation

[0468] Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions relating to adaptive perceptual quantization of images with enhanced dynamic range, as described herein. The computer and / or IC may calculate any of the various parameters or values ​​relating to the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.

[0469] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention can take any of a variety of forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0470] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to said components (including references to “devices”) should be interpreted as including any component that performs the function of the described component as an equivalent of said component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions in the illustrated exemplary embodiments of the invention.

[0471] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that incorporates hardwired and / or program logic to implement the techniques.

[0472] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0473] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in non-transitory storage media accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.

[0474] Computer system 500 further includes read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions of processor 504. Storage device 510 (such as a magnetic disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.

[0475] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.

[0476] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the instruction sequence contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used instead of or in combination with software instructions.

[0477] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.

[0478] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0479] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 fetches and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0480] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0481] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0482] Computer system 500 can send messages and receive data, including program code, through multiple networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0483] The received code may be executed by processor 504 and / or stored in storage device 510 or other non-volatile storage device for later execution upon receipt.

[0484] Equivalents, extensions, alternatives and miscellaneous

[0485] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the claimed embodiments of the invention, and of the applicant's view, is the set of claims published in specific form according to this application, wherein such claim publication includes any subsequent corrections. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and the drawings should be viewed in an illustrative rather than restrictive sense.

[0486] Exemplary examples of enumeration

[0487] This invention may be practiced in any of the forms described herein, including but not limited to the enumerated example embodiments (EEE) that describe some parts of the structure, features and functions of embodiments of the invention.

[0488] EEE 1. A method comprising:

[0489] The input image at the first bit depth in the input domain is received from the input video signal at the first bit depth, wherein the first bit depth is higher than the second bit depth in the shaping domain;

[0490] A forward shaping operation is performed on the input image to generate a forward-shaped image of the second bit depth in the shaping domain. The forward shaping operation includes winding input codewords in the input image along the non-winding axis of the input domain into shaped codewords in the forward-shaped image on the winding axis of the shaping domain.

[0491] An image container containing image data obtained from the forward-shaped image is encoded into an output video signal of the second bit depth. The image data in the image container causes a receiving device of the output video signal to construct a backward-shaped image of the third bit depth for rendering on a display device. The third bit depth is higher than the second bit depth.

[0492] EEE 2. The method as described in EEE 1, wherein the winding axis of the shaping domain has a geometry different from that of the non-winding axis; wherein the total number of available codewords on the winding axis is greater than the total number of available codewords on the non-winding axis.

[0493] EEE 3. The method as described in EEE 1 or 2, further comprising:

[0494] A forward truncation field transform is applied to the forward-shaped image to generate an intermediate image in the truncation field.

[0495] Perform one or more image processing operations on the intermediate image to generate a processed intermediate image;

[0496] An inverse truncated field transform is applied to the processed intermediate image to generate the image data in the image container.

[0497] EEE 4. The method as described in EEE 1 to 3, wherein the processed intermediate image includes one or more cropped codeword values ​​generated according to a cropping operation to ensure that all codeword values ​​in the processed intermediate image are within the target spatial shape representing the integer domain.

[0498] EEE 5. The method of any one of EEE 1 to 4, wherein the second bit depth represents one of the following: 8 bits, 10 bits, 12 bits or another bit depth less than the first bit depth.

[0499] EEE 6. The method of any one of EEE 1 to 5, wherein the forward shaping operation is based on a set of forward shaping maps; wherein a set of operation parameter values ​​used in the set of forward shaping maps is selected from a plurality of sets of operation parameter values ​​based at least in part on input codewords in the input image; wherein each set of operation parameter values ​​in the plurality of sets of operation parameter values ​​is optimized to minimize the prediction error of the corresponding training image cluster in a plurality of training image clusters.

[0500] EEE 7. The method as described in EEE 6, wherein the set of forward shaping maps includes one or more of the following: multivariate regression (MMR) mapping, tensor product B-spline (TPB) mapping, forward shaping lookup table (FLUT) or other types of forward shaping mapping.

[0501] EEE 8. The method as described in EEE 6 or 7, wherein each set of operating parameters in the plurality of sets of operating parameter values ​​is generated based on the Backward Error Subtraction Signal Adjustment (BESA) algorithm.

[0502] EEE 9. The method as described in EEE 8, wherein the prediction error propagating in the BESA algorithm is calculated at least in part based on (a) the difference between the input value and the reconstructed value and (b) the spatial gradient obtained as the partial reciprocal of the cross-channel forward shaping function used to generate the reconstructed value.

[0503] EEE 10. The method of any one of EEE 1 to 9, wherein the input codewords in the input image are scaled using one or more scaling factors; wherein the one or more scaling factors are optimized at runtime using a search algorithm based on one or more of the following: the golden section algorithm, the Nelder-Mead algorithm, or another search algorithm.

[0504] EEE 11. The method of any one of EEE 1 to 10, wherein an input shape representing the input domain is mapped to a target reference shape representing the shape domain using a set of mapping functions; wherein the target reference shape is selected from a plurality of different reference shapes corresponding to a plurality of different candidate shape domains based at least in part on the distribution of input codewords in the input image.

[0505] EEE 12. The method as described in EEE 11, wherein the plurality of different reference shapes includes at least one of the following: a complete geometry, a complete geometry with a single gap portion inserted, a complete geometry with multiple gap portions inserted, a toroidal shape, a cylindrical shape, a toroidal shape, another shape that is the same shape as a cube or another shape that is different from a cube.

[0506] EEE 13. The method of any one of EEE 1 to 12, wherein the input domain represents one of the following: RGB color space, YCbCr color space, perceptual quantization color space, linear color space or another color space.

[0507] EEE 14. The method of any one of EEE 1 to 13, wherein the first bit depth represents one of the following: 12 bits, 16 bits or more bits or another bit depth higher than the second bit depth.

[0508] EEE 15. The method of any one of EEE 1 to 14, wherein the forward shaping operation represents a primary shaping operation; the method further comprises: performing a secondary shaping operation on the input image to linearly scale a finite range of input data to the full range of data for each color channel, wherein the secondary shaping operation includes one or more of the following: per-channel pre-shaping or scaling.

[0509] EEE 16. The method as described in EEE 15, wherein image metadata is generated based on the operation parameters used in the primary shaping operation and the secondary shaping operation; wherein the image metadata is provided to the receiving device in the output video signal at the second bit depth.

[0510] EEE 17. The method of any one of EEE 1 to 16, wherein the secondary shaping operation is performed in a first stage prior to the second stage of performing the primary shaping operation.

[0511] EEE 18. The method of any one of EEE 1 to 17, wherein the combination of the secondary shaping operation and the primary shaping operation is performed in a single combination phase during runtime.

[0512] EEE 19. The method of any one of EEE 1 to 18, wherein the shaping domain is represented by a toroidal shape.

[0513] EEE 20. The method of any one of EEE 1 to 19, wherein the shaping domain is represented by a cylindrical shape.

[0514] EEE 21. A method comprising:

[0515] Decode an image container from a video signal containing image data obtained from a forward-shaped image in the shaping domain, wherein the source image at a first depth in the source domain has been forward-shaped to generate the forward-shaped image at a second depth lower than the first depth.

[0516] A backward shaping operation is applied to the image data decoded from the video signal to generate a backward-shaped image in a target domain, the backward-shaped image having a third bit depth higher than the second bit depth, the backward shaping operation comprising unwinding codewords in the image data along the winding axis of the shaping domain into backward-shaped codewords in the backward-shaped image on the non-winding axis of the target domain;

[0517] The display image generated from the backward-shaped image is rendered on the display device.

[0518] EEE 22. An apparatus comprising a processor and configured to perform any of the methods described in EEE 1 to 21.

[0519] EEE 23. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for performing a method according to any one of the methods described in EEE 1 to 21 using one or more processors.

Claims

1. A method for image processing, comprising: The input image at the first bit depth in the input domain is received from the input video signal at the first bit depth, wherein the first bit depth is higher than the second bit depth in the shaping domain; A forward shaping operation is performed on the input image to generate a forward-shaped image of the second bit depth in the shaping domain. The forward shaping operation includes winding an input codeword in the input image along a non-winding axis of the input domain into a shaped codeword in the forward-shaped image on a winding axis of the shaping domain, wherein the non-winding axis is represented as a linear axis in the input domain, and wherein the winding axis is represented by a non-linear spatial shape in the shaping domain. An image container containing image data obtained from the forward-shaped image is encoded into an output video signal of the second bit depth, wherein the image data in the image container causes the receiving device of the output video signal to construct a backward-shaped image of the third bit depth for rendering on a display device, the third bit depth being greater than the second bit depth. In order to generate the image data in the image container, the method further includes: A forward truncation field transform is applied to the forward-shaped image in the shaping domain to transform the forward-shaped image in the shaping domain into an intermediate image in the truncation field domain, wherein the forward truncation field transform ensures that the codewords remain within a specified codeword value range; Perform one or more image processing operations on the intermediate image to generate a processed intermediate image; and An inverse truncated field transform is applied to the processed intermediate image to generate the image data in the image container.

2. The method as described in claim 1, wherein, The shaping domain is a winding shaping domain, and wherein the winding axis refers to an axis in the winding shaping domain that is mapped from the non-winding axis through one or more winding transformations, wherein the one or more winding transformations are one or more geometric transformations.

3. The method as described in claim 1 or 2, wherein, The non-wound axis refers to a single channel or dimension or a combination of multiple channels or dimensions in the input image within the input domain.

4. The method as described in claim 1 or 2, wherein, The winding axis of the shaping domain has a different geometry than the non-winding axis; wherein the total number of available codewords on the winding axis is greater than the total number of available codewords on the non-winding axis.

5. The method of claim 1, wherein, The processed intermediate image includes one or more cropped codeword values ​​generated according to the cropping operation, to ensure that all codeword values ​​in the processed intermediate image are within the target space shape representing the integer domain.

6. The method as described in claim 1 or 2, wherein, The second bit depth represents one of the following: 8 bits, 10 bits, 12 bits, or another bit depth lower than the first bit depth.

7. The method as described in claim 1 or 2, wherein, The forward shaping operation is based on a set of forward shaping maps; wherein a set of operation parameter values ​​used in the set of forward shaping maps are selected from multiple sets of operation parameter values ​​based at least in part on input codewords in the input image; wherein each set of operation parameter values ​​is optimized to minimize the prediction error of the corresponding training image cluster in multiple training image clusters.

8. The method of claim 7, wherein, The set of forward shaping mappings includes one or more of the following: multivariate regression (MMR) mapping, tensor product B-spline (TPB) mapping, forward shaping lookup table (FLUT), or other types of forward shaping mappings.

9. The method of claim 7, wherein, Each of the multiple sets of operating parameter values ​​is generated based on the BESA algorithm adjusted by backward error subtraction signal.

10. The method of claim 9, wherein, The prediction error propagated in the BESA algorithm is calculated at least in part based on a) the difference between the input value and the reconstructed value and b) the spatial gradient obtained as the partial derivative of the cross-channel forward shaping function used to generate the reconstructed value.

11. The method as claimed in claim 1 or 2, wherein, The input codewords in the input image are scaled using one or more scaling factors; wherein the one or more scaling factors are optimized at runtime using a search algorithm based on one or more of the following: the golden section algorithm, the Nelder-Mead algorithm, or another search algorithm.

12. The method as claimed in claim 1 or 2, wherein, The input shaping representing the input domain is mapped to a target reference shape representing the shaping domain using a set of mapping functions; wherein the target reference shape is selected from multiple different reference shapes corresponding to multiple different candidate shaping domains, at least in part based on the distribution of input codewords in the input image.

13. The method of claim 12, wherein, The plurality of different reference shapes include at least one of the following: a complete geometry, a complete geometry with a single gap portion inserted, a complete geometry with multiple gap portions inserted, a toroidal shape, a cylindrical shape, or a toroidal shape.

14. The method as claimed in claim 1 or 2, wherein, The input domain represents one of the following: RGB color space, YCbCr color space, perceptual quantization color space, linear color space, or another color space.

15. The method as claimed in claim 1 or 2, wherein, The first bit depth represents one of the following: 12 bits, 16 bits or more, or another number of bits higher than the second bit depth.

16. The method as claimed in claim 1 or 2, wherein, The forward shaping operation represents a primary shaping operation; the method further includes: performing a secondary shaping operation on the input image to linearly scale a finite range of input data to the full range of data for each color channel, wherein the secondary shaping operation includes one or more of the following: per-channel pre-shaping or scaling.

17. The method of claim 16, wherein, Image metadata is generated based on the operating parameters used in the primary shaping operation and the secondary shaping operation; wherein, the image metadata is provided to the receiving device in the output video signal at the second bit depth.

18. The method of claim 16, wherein, The secondary shaping operation is performed in the first stage, prior to the second stage of the primary shaping operation.

19. The method of claim 16, wherein, The combination of the secondary shaping operation and the primary shaping operation is executed in a single combination phase during runtime.

20. The method of claim 1 or 2, wherein, The shaping domain is represented by a toroidal shape.

21. The method as claimed in claim 1 or 2, wherein, The shaping domain is represented by a cylindrical shape.

22. A method for image processing, comprising: Decode an image container from a video signal containing image data obtained from a forward-shaped image in the shaping domain, wherein the source image at a first depth in the source domain has been forward-shaped to generate the forward-shaped image at a second depth lower than the first depth. A backward shaping operation is applied to the image data decoded from the video signal to generate a backward-shaped image in a target domain, the backward-shaped image having a third bit depth higher than the second bit depth, the backward shaping operation comprising unwinding codewords in the image data along a winding axis of the shaping domain into backward-shaped codewords in the backward-shaped image along a non-winding axis in the target domain, wherein the non-winding axis is represented as a straight line in the target domain, and wherein the winding axis is represented in the shaping domain with a spatial shape different from the straight line; Render a display image generated from the backward-shaped image on a display device; The image data in the image container is generated through operations including the following: A forward truncation field transform is applied to the forward-shaped image in the shaping domain to transform the forward-shaped image in the shaping domain into an intermediate image in the truncation field domain, wherein the forward truncation field transform ensures that the codewords remain within a specified codeword value range; Perform one or more image processing operations on the intermediate image to generate a processed intermediate image; and An inverse truncated field transform is applied to the processed intermediate image to generate the image data in the image container.

23. An apparatus for image processing, comprising a processor configured to perform the method as claimed in any one of claims 1 to 22.

24. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon for performing the method according to any one of claims 1 to 22 using one or more processors.

Citation Information

Patent Citations

  • Multiple color channel multiple regression predictor

    US8811490B2

  • Signal reshaping approximation

    US20180020224A1