Video distribution system capable of changing dynamic range
The video delivery system addresses color shift issues in SDR-to-HDR conversion using a 5D grid-based LUT for chroma offset correction, ensuring efficient and high-quality image conversion.
Patent Information
- Application Number
- JP2024573308
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-14
- Filing Date
- 2023-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Existing video delivery systems face challenges in efficiently converting Standard Dynamic Range (SDR) images to High Dynamic Range (HDR) images, particularly due to color shift issues that occur during the upconversion process, which are not adequately addressed by current methods.
A video delivery system utilizing a 5-dimensional grid-based look-up table (LUT) for chroma offset values, combined with a reshaping function index map, to perform color shift correction by applying chroma offsets to each pixel, ensuring compatibility with existing display-management functions without requiring modifications.
Effectively mitigates color shift between SDR and HDR images, enabling seamless conversion while maintaining image quality and compatibility with various display devices.
Smart Images

Figure 2025522416000036 
Figure 2025522416000037 
Figure 2025522416000038
Abstract
Description
Technical Field
[0001] 1. Cross - reference to related applications This application claims priority from European Patent Application No. 22178928.2 and US Provisional Patent Application No. 63 / 351,855 (both filed on June 14, 2022), each of which is incorporated herein by reference in its entirety.
[0002] 2. Field of disclosure Various exemplary embodiments relate to image processing operations, and more specifically, but not limited to, video codecs.
Background Art
[0003] 3. Background This section introduces aspects that may be useful for facilitating a better understanding of the present disclosure. Thus, the description of this section should be read from this perspective and should not be understood as an admission as to what is prior art or what is not prior art.
[0004] As used herein, the term "dynamic range" (DR) may relate to the ability of the human visual system (HVS) to perceive the range of intensities (e.g., luminance, luma) within an image, e.g., from the darkest black (shadow) to the brightest white (highlight). In this sense, DR relates to "scene - referenced" intensities. DR may also relate to the ability of a display device to adequately or approximately render a particular width of intensity range. In this sense, DR relates to "display - referenced" intensities. Unless explicitly specified to have a particular meaning in any point of the description herein, this term should be presumed to be used interchangeably in either sense.
[0005] As used herein, the term "high dynamic range" (HDR) relates to a DR span of 14 to 15 digits or more of the human visual system (HVS). Practically, the DR that a human can simultaneously perceive over a broad span in the intensity range may be somewhat truncated compared to HDR. As used herein, the term "enhanced dynamic range" (EDR) or "visual dynamic range" (VDR) may relate, individually or interchangeably, to the DR perceivable within a scene or image by the human visual system, including eye movements, taking into account some light adaptation changes across the scene or image. In this specification, EDR may relate to a DR spanning 5 to 6 digits. Although perhaps somewhat narrower in comparison to true scene-based HDR, EDR nonetheless represents a wide DR span and may also be referred to as HDR.
[0006] Practically, an image includes one or more color components of a color space (e.g., luma Y, chroma Cb, and Cr), and each color component is represented with an accuracy of n bits per pixel (e.g., n = 8). Using non-linear luminance encoding (e.g., gamma encoding), an image with n ≤ 8 (e.g., a 24-bit color JPEG image) is considered a standard dynamic range (SDR) image, while an image with n > 8 is considered an enhanced dynamic range image.
[0007] The reference electro-optical transfer function (EOTF) of a given display characterizes the relationship between the color values (e.g., luminance) of an input video signal and the output screen color values (e.g., screen luminance) generated by the display. For example, ITU Rec. ITU-R BT.1886, "Reference electro-optical transfer function for flat panel displays used in HDTV studio production" (March 2011), defines a reference EOTF for flat panel displays. When a video stream is provided, information regarding its EOTF may be embedded in the bitstream as (image) metadata. The term "metadata" as used herein relates to any auxiliary information transmitted as part of an encoded bitstream. Such metadata may be used to assist a decoder in rendering the decoded image and may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as further described herein.
[0008] As used herein, the term "PQ" refers to perceptual luminance amplitude quantization. The human visual system responds in a highly non-linear manner to increases in the level of light. A human's ability to perceive a stimulus is affected by the luminance of the stimulus, the size of the stimulus, the spatial frequency that makes up the stimulus, and the luminance level to which the eye has adapted at the particular instant the stimulus is being viewed. In some cases, the PQ function maps linear input gray levels to output gray levels that better match the contrast sensitivity thresholds in the human visual system. An exemplary PQ mapping function is described in SMPTE ST 2084:2014, "High dynamic range EOTF for mastering reference displays" (hereinafter, "SMPTE"), which is hereby incorporated by reference in its entirety.
[0009] 200~1000cd / m 2Alternatively, a display that supports nit luminance is typical of a low dynamic range (LDR), also known as standard dynamic range (SDR), as compared to EDR (or HDR). EDR content can be displayed on an EDR display that supports a higher dynamic range (e.g., from 1,000 nits to 5,000 nits or more). Such a display may be defined using an alternative EOTF that supports high luminance capabilities (e.g., 0 to 10,000 nits or more). Examples of such EOTFs are defined in SMPTE 2084 and Rec. ITU-R BT.2100, "Image parameter values for high dynamic range television for production and international programme exchange" (06 / 2017). These are incorporated herein by reference in their entirety.
[0010] Patent Document 1 discloses an adaptive local reshaping method for SDR-HDR upconversion. Using the luma codeword in the input image, a global index value is generated for selecting a global reshaping function for the input image with a relatively low dynamic range. Image filtering is applied to the input image to generate a filtered image. The filtered values of the filtered image provide a measure of the local luma level in the input image. Using the global index value and the filtered values of the filtered image, a local index value is generated for selecting a specific local reshaping function for the input image. By reshaping the input image with the specific local reshaping function selected using the local index value, a reshaped image with a relatively high dynamic range is generated.
Patent Document 1
[0011] Patent Document 2 discloses a method for color correction in high dynamic range (HDR) video using a 2D look-up table (LUT). The color correction can be applied in a decoder after decoding an HDR video signal. For example, the color correction can be applied before, during, or after chroma upsampling of the HDR video signal. The 2D LUT can include a representation of the color space of the HDR video signal. The color correction can include applying triangular interpolation to sample values of color components in the color space. The 2D LUT can be estimated by an encoder and signaled to a decoder. The encoder can determine to reuse a previously signaled 2D LUT or use a new 2D LUT. [Patent Document 2] WO2017 / 059415A1 [Summary of the Invention] [Means for Solving the Problems]
[0012] This specification discloses various embodiments of a video delivery system that can perform color shift correction in an HDR image generated from an SDR image. In an exemplary embodiment, the color shift correction is performed using a pre-computed look-up table (LUT) representing a 5-dimensional (5D) grid, and the LUT is addressable using a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value. Linear interpolation can be used to obtain chroma offset values for any point in the corresponding 5D parameter space that is not a grid point. Also disclosed herein is an exemplary sequential iterative minimization method using a properly constructed cost function that can be used to populate the LUT. Advantageously, the disclosed exemplary embodiments of color shift correction are compatible with existing display-management functions and do not require modification thereof.
[0013] According to an exemplary embodiment, a video distribution system is provided that can change the dynamic range of an input image. The distribution system includes a memory that stores a plurality of chroma offset values corresponding to lattice points of a 5-dimensional grid; and a processor that converts an input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the processor generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes for identifying a corresponding reshaping function to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; and generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by addressing the lattice points using the respective indexes, the display management metadata values, and three respective pixel values of corresponding pixels of the input image.
[0014] According to another exemplary embodiment, a method for changing the dynamic range of an input image is provided. The method includes converting an input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the conversion being performed using a plurality of pre-calculated chroma offset values corresponding to the grid points of a five-dimensional grid; the conversion comprising: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes identifying a corresponding reshaping function to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; and generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by addressing the grid points using the respective indexes, the display management metadata values, and the three respective pixel values of the corresponding pixel of the input image.
[0015] According to yet another exemplary embodiment, there is provided a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including a method of changing a dynamic range of an input image. The method includes converting an input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the conversion being performed using a plurality of pre-computed chroma offset values corresponding to grid points of a five-dimensional grid; the conversion comprising: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indices identifying a corresponding reshaping function to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; and generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by addressing the grid points using the respective indices, the display management metadata values, and three respective pixel values of a corresponding pixel of the input image.
[0016] According to yet another exemplary embodiment, a method is provided for generating a plurality of chroma offset values for performing color shift correction in an output video generated by changing the dynamic range of an input image. The method defines a five-dimensional grid having first, second, third, fourth, and fifth dimensions respectively representing a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value; defines a cost function for quantifying at least one cost for a color difference between the input image and the output image; for each grid point of the five-dimensional grid, performs sequential iterative minimization of the cost function to determine each set of chroma offset values; and arrays each set of chroma offset values in an electronically addressable look-up table addressable using a set of discrete values corresponding to the first, second, third, fourth, and fifth dimensions.
Brief Description of the Drawings
[0017] Other aspects, features, and advantages of the various disclosed embodiments will become more fully apparent from the following detailed description and the accompanying drawings, by way of example.
[0018]
Figure 1
[0019]
Figure 2
[0020]
Figure 3A
Figure 3B
Figure 3C
[0021]
Figure 4A
Figure 4B
Figure 4C
[0022]
Figure 5A
Figure 5B
[0023]
Figure 6A
Figure 6B
[0024]
Figure 7
[0025]
Figure 8
Embodiments for Carrying Out the Invention
[0026] The present disclosure and aspects thereof can be embodied in various forms including hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces, as well as application programming interfaces, and hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc. The above is only intended to give a general concept of various aspects of the present disclosure and is in no way intended to limit the scope of the present disclosure.
[0027] In the following description, numerous details such as optical device configurations, timings, operations, etc. are described in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to those skilled in the art that these specific details are merely illustrative and are not intended to limit the scope of the present application.
[0028] Furthermore, although the present disclosure mainly focuses on examples where various circuits are used in digital projection systems, it will be understood that these are merely examples. It should be further understood that the disclosed systems and methods can be used in any device that needs to project light, such as in movie theaters, consumer, and other commercial projection systems, head-up displays, virtual reality displays, etc. The disclosed systems and methods can be implemented in additional display devices having OLED displays, LCD displays, quantum dot displays, etc.
[0029] Many consumer desktop displays can support a luminance of 200 - 300 cd / m 2 or nits. Many consumer HDTVs are in the range of 300 - 500 nits, and new models are 1000 nits (cd / m 2) is reached. As the availability of HDR content increases with the advancement of both image capture devices (e.g., cameras) and HDR displays (e.g., the PRM-4200 Professional Reference Monitor from Dolby Laboratories), HDR content may be color graded and displayed on HDR displays that support a higher dynamic range (e.g., from 1000 nits to 5000 nits or more).
[0030] Some embodiments may benefit from at least some of the features disclosed in International Patent Application "ADAPTIVE LOCAL RESHAPING FOR SDR-TO-HDR UP-CONVERSION" by T-W. Huang et al., PCT / US2021 / 053241, filed on October 1, 2021, which is incorporated herein by reference in its entirety.
[0031] In this specification, the term "metadata" relates to any auxiliary information that is transmitted as part of an encoded bitstream and aids a decoder in rendering the corresponding image. For television broadcasts and video streaming, video metadata can be used to provide side information about a particular video and audio stream or file. The metadata can be either directly embedded in the video or included as a separate file within a container such as MP4 or MKV. The metadata can include information about the entire video stream or file, or about a particular video frame. The metadata is created by cameras, encoders, and other video processing elements (see, e.g., 115, 120 of FIG. 1), and can include, but is not limited to, time stamps, video resolution, digital film grain parameters, color space or color gamut information, reference display parameters, auxiliary signal parameters, file size, closed captions, audio language, advertisement insertion points, color space, error messages, etc. Additional examples of metadata related to the disclosed embodiments are described herein below.
[0032] In some embodiments disclosed below in this specification, the image metadata includes L1 metadata. As used herein, the term "L1 metadata" represents one or more of the minimum (L1-min), intermediate (L1-mid), and maximum (L1-max) luminance values associated with a particular portion of video content, such as an input frame or image. The L1 metadata is associated with a video signal. To generate the L1 metadata, a per-frame analysis at the pixel level of the video content is preferably performed on the encoding side. Alternatively, the analysis may also be performed on the decoding side. The analysis describes the distribution of luminance values over a defined portion of the video content covered by an analysis pass, such as a single frame or a series of frames like a scene. The L1 metadata can be calculated in an analysis pass covering a series of frames such as a single video frame and / or scene. The L1 metadata has various values derived between analysis passes, is associated with each portion of the video content from which the L1 metadata was calculated, and together forms the L1 metadata associated with the video signal. Such L1 metadata can include at least one of (i) an L1-min value representing the lowest black level in each portion of the video content, (ii) an L1-mid value representing the average luminance level across each portion of the video content, and (iii) an L1-max value representing the highest luminance level in each portion of the video content. The L1 metadata can be generated for and attached to each video frame and / or each scene encoded in the video signal. The L1 metadata may also be generated for regions of an image, and such L1 metadata may be referred to as local L1 values. The L1 metadata can be calculated by converting the RGB data to a luma-chroma format (e.g., YCbCr) and then calculating one or more of the minimum, intermediate (average), and maximum values in the Y plane, or they can be calculated directly in the RGB space.
[0033] In some embodiments, the L1-min value may represent the minimum value of the PQ-encoded min(RGB) values of each part of the video content (e.g., a video frame or an image), considering only the active area (e.g., by excluding gray bars or black bars, letterbox bars, etc.), where min(RGB) represents the minimum value of the pixel color component values {R, G, B}. The L1-mid value and the L1-max value can be calculated similarly. In particular, in an exemplary embodiment, L1-mid may represent the average of the PQ-encoded max(RGB) values of the image, and L1-max may represent the maximum value of the PQ-encoded max(RGB) values of the image, where max(RGB) represents the maximum value of the pixel color component values {R, G, B}. In some embodiments, the L1 metadata may be normalized to be within the range [0, 1].
[0034] Video coding according to an exemplary embodiment FIG. 1 shows an exemplary process of a video delivery pipeline (100) showing various stages from video capture to video content display according to an embodiment. A sequence of video frames (102) can be captured or generated using an image generation block (105). The video frame (102) may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frame (102) may be captured on film by a film camera. The film can then be converted to a digital format to provide video data (107).
[0035] In the production phase (110), video data (107) can be edited to provide a video production stream (112). The data of the video production stream (112) can then be provided to a processor (or one or more processors such as a central processing unit, CPU) in a post-production block (115) for post-production editing. Post-production editing of the block (115) can include, for example, adjusting or modifying the color or brightness in a specific area of an image according to the creative intention of the video producer to improve the image quality or achieve a specific look for the image. Other edits (such as scene selection and sequencing, image cropping, addition of computer-generated visual special effects, etc.) are performed in the block (115) to generate a “final” version (117) of the production for distribution. During post-production editing (115), the video image can be viewed on a reference display (125).
[0036] Following post-production (115), the final version (117) of the video data can be delivered to an encoding block (120) for downstream distribution to decoding and playback devices such as televisions, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) can include audio and video encoders such as those defined by ATSC, DVB, DVD, Blu-Ray, and other delivery formats to generate an encoded bitstream (122). Some of the methods described herein below can be executed by a corresponding processor in the encoding block (120). For example, the encoding block (120) may be configured to perform SDR-HDR local reshaping and color shift correction, as described in more detail below. At the receiver, the encoded bitstream (122) can be decoded by a decoding unit (130) to generate a corresponding decoded signal (132) representing a copy or a close approximation of the signal (117). The receiver may be attached to a target display (140) that may have somewhat or completely different characteristics from the reference display (125). In such a case, a display management (DM) block (135) can be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Some of the methods described herein below can be executed by the decoding unit (130) and / or the display management block (135). Depending on the embodiment, the decoding unit (130) and the display management block (135) may include individual processors or may be based on a single integrated processing unit.
[0037] In one exemplary embodiment, SDR-HDR local reshaping can take a 3-channel (Y, Cb, Cr) input SDR image and predict a 3-channel (Y, Cb, Cr) output HDR image using a pre-trained reshaping function. To view the HDR image on the target display (140), the DM block (135) can further process the HDR image based on the corresponding metadata to generate a display-mapped signal (137) representing the DM image according to the target display luminance.
[0038] In most cases, the chromaticity of the HDR image and the DM image is equal to or close to the chromaticity of the SDR image, even if the luminance differs due to upconversion and enhancement. However, in some cases, the difference in chromaticity may become significant, and such a difference is called an SDR-HDR color shift. The color shift from SDR to HDR can occur for several reasons. For example, since the pre-trained reshaping function and the DM are typically trained on natural images, such training may cause larger errors and / or color shifts for colors that are not very prevalent in natural images. Also, the DM block (135) may perform clipping on pixels near the color space boundary, thereby amplifying an existing color shift or introducing a new color shift.
[0039] In one exemplary embodiment, the DM block (135) is not modified and can remain the same as, for example, a legacy video delivery pipeline. Rather, the exemplary embodiments of the proposed color shift correction framework are designed to correct the end-to-end color shift between the SDR image and the final DM image by adding a chroma offset to the Cb and Cr channels of the HDR image. Since the end-to-end process from the SDR image to the DM image is typically highly non-linear, the required chroma offset is determined using a 5D look-up table (LUT) on a 5D grid. Here, the five dimensions are the SDR pixel values (3D), the reshaping function index, and the DM metadata L1-mid. The resulting 5D space is referred to herein as the HSWLM space, where H represents hue, S represents saturation, W represents the scaled intensity value, L represents the local reshaping function index, and M represents the metadata. The 5D LUT can be populated with values using an appropriate cost function by successive iterative minimization, as will be described in further detail below with reference to FIGS. 7-8, for example.
[0040] Color shift correction FIG. 2 shows an exemplary process (200) of a video delivery pipeline (100) according to an embodiment. The process (200) can typically receive an input SDR image (117) and generate a corresponding output DM image (137) (see also FIG. 1). During the process (200), a chroma offset can be added to the initial HDR image (216) so that any color shift that may occur for the reasons described above can be mitigated.
[0041] The process (200) includes a reshaping block (210) and a chroma offset processing block (220). The reshaping block (210) is configured to generate an initial HDR image (216) based on an input SDR image (117). The reshaping block (210) includes generating a reshaping function index map (212) and performing processing directed to applying a pre-trained reshaping function (214) to the input SDR image (117). In some embodiments, the reshaping block (210) may be implemented as disclosed in the above-cited international patent application PCT / US2021 / 053241.
[0042] In one exemplary embodiment, the reshaping block (210) includes SDR-HDR local reshaping, and the 3-channel (Y, Cb, Cr) output HDR image (216) is predicted using the 3-channel (Y, Cb, Cr) input SDR image (117) and a set of pre-trained reshaping functions (214). The reshaping function index map (212) is created to indicate which reshaping function (214) is used for different pixels.
[0043] Let the Y, Cb, and Cr channels of the i-th pixel in the input SDR image S (117) be s i YCbCr =(s i Y , s i Cb , s i Cr ), and let the reshaping function index of the i-th pixel in the reshaping function index map L (212) be l i . The initial HDR image V (init) (216) generated by the SDR-HDR local reshaping of the reshaping block (210), the i-th pixel v i (init),YCbCr =(v i (init),Y , v i (init),Cb , v i (init),Cr ) can be expressed as follows.
Number
[0044] As a non-limiting example, the processing of a single image / frame is described below. Based on the provided description, one of ordinary skill in the art will be able to implement the corresponding processing of multiple frames without any undue experimentation, since, for example, the disclosed processing typically or explicitly does not rely on time information. For the sake of simplicity of description and without any implication of limitation, the description is given for normalized pixel values, i.e., values belonging to the range [0,1]. Also, the conversion between color spaces is described as being able to be performed as needed. For example, if there is a value s i YCbCr in the YCbCr color space, it is assumed that there is also a corresponding value s i RGB in the RGB color space. One of ordinary skill in the art will readily understand how to perform such a conversion between color spaces.
[0045] To view an HDR image, such as an initial HDR image (216) or an output HDR image (240), on a target display (140), the processing of the DM block (135) is typically applied to the HDR image, and the result of such DM processing is the output DM image (137). Such DM processing can be controlled by the above-described metadata L1-min, L1-mid, and L1-max, which represent the minimum value, average value, and maximum value of the RGB channels of the HDR image, respectively. However, in an exemplary SDR-HDR upconversion scenario, L1-min and L1-max may be set to constant values in at least some cases. Thus, some embodiments may rely only on the L1-mid value.
[0046] If the chroma offset processing block (220) does not exist, the initial HDR image (216) is directly applied to the DM block (135). The R, G, and B channels of the i-th pixel in the initial HDR image V (init) (216) are represented as v i (init),RGB =(v i (init),R ,v i (init),G ,v i (init),B ). The DM metadata L1-mid of the initial HDR image can be calculated as follows.
Number
Number
Number
[0047] In some cases, ^s i(init),YCbCr and s i YCbCr may have similar chromaticities. However, in other cases, the chromaticity difference between ^s i(init),YCbCr and s i YCbCr i.e., the color shift, may become noticeable to an observer. Therefore, the processing implemented in the chroma offset processing block (220) is directed to significantly reducing or completely eliminating such differences. In one exemplary embodiment, the chroma offset processing block (220) uses a 5D grid and a corresponding 5D LUT (228).
[0048] A grid of dimension D (e.g., D = 5) can be constructed by sampling along each dimension in the corresponding multi-dimensional space R D For computational efficiency, a uniformly spaced grid can be used, and the coordinate values in the same dimension are uniformly sampled. For the d-th dimension, N d values are sampled starting from the initial value p d with an interval b d Assuming that the i d -th sample is then the p d + i d b d where i d = 0, 1, …, N d - 1, and d = 0, 1, …, D - 1. This sampling defines the grid X. The lattice point X D-1 with index i = (i0, i1, …, i i ) is represented as follows.
Equation
[0049] In general, a LUT such as a 5D LUT (228) can be constructed to model any function φ on a grid. More specifically, the LUT function Φ on grid X returns the output value Φ i corresponding to the LUT for each input grid point X i = φ(X i ). For example, in one embodiment, the function Φ can be a LUT representing a function φ that takes as input SDR pixel values, reshaping function indices, and DM metadata L1-mid, and then provides as output the chroma offset. In various embodiments, the outputs of the functions φ and Φ can be in scalar or vector form. In the non-limiting examples described below, the output is a 2D vector for the chroma offsets in the Cb and Cr channels.
[0050] In one exemplary embodiment, linear interpolation can be used to handle input values that are not on the grid. For example, such linear interpolation may follow the same definition as that used in ordinary bilinear or trilinear interpolation, and the output values are based on the corresponding linear interpolation in each dimension. To facilitate the calculation, one exemplary embodiment may rely on normalized grid coordinates. This approach provides a normalized grid that starts from 0 and progresses at unit intervals. Given an input x = (x0, x1, …, x D-1 ), the corresponding normalized grid coordinates can be obtained by shifting and scaling as
Equation
Equation
Equation
[0051] Furthermore, the normalized grid
Number
Number
Number
Number
[0052] Figures 3A - 3C pictorially show applying the above indexing method to an exemplary 2D grid of size 4×4 (tilde - marked X) according to an embodiment. More specifically, Figure 3A shows two grid dimensions, represented as dimension 0 and dimension 1 respectively. Figure 3B shows the indexing, where the lattice points (shown as nodes) are indexed as described above. Figure 3C shows the indexing, where the hypercubes are indexed as described above.
[0053] Since linear interpolation is linear in each dimension, such linear interpolation can be performed using the normalized coordinates [tilde-x] and the normalized grid [tilde-X] instead of the corresponding non-normalized entities. In one exemplary embodiment, the processing steps for performing linear interpolation may include finding the unit hypercube that contains tilde-x and then performing the interpolation using the distance between tilde-x and the vertices of the unit hypercube. The index of the unit hypercube that contains tilde-x is [Number] represented as. The clipping function clip3 is defined as clip3(x,a,b)=min(max(x,a),b). When tilde-x is on the boundary between hypercubes, it can be seen that that particular tilde-x is assigned to the hypercube that is in the direction away from the origin.
[0054] Let the interpolation result of the input x be denoted as ^φ(x). Based on the above definition of linear interpolation, the interpolation result ^φ(x) can be expressed as follows. [Number] Note that since the normalized grid has an interval of 1, the normalization factor in this linear interpolation has already been processed. In one exemplary embodiment, when the grid is relatively dense, the interpolation result ^φ(x) may typically be very close to the actual function output φ(x).
[0055] The weights for the above linear interpolation depend only on the distance between tilde-x and tilde-X i' within the same unit hypercube. For computational efficiency, the weights can be pre-computed for a plurality of possible distances to enable a lookup at runtime. For example, the unit hypercube can be quantized and the corresponding interpolation weights can be stored in a LUT. When the quantization is relatively dense, the output of the LUT is typically relatively close to the corresponding non-quantized interpolation result.
[0056] Consider a unit hypercube located at the origin. Such a unit hypercube has 2 D vertices V = {0, 1} D In one exemplary embodiment, the vertices can be indexed by their coordinates, i.e., V k = vector k = (k0, k1, …, k D-1 )). If the unit hypercube is uniformly quantized to M d points within the closed interval [0, 1] for each dimension d = 0, 1, … D-1, then the quantization grid Q can be defined such that the quantization points with index j = (j0, j1, …, j D-1 ) are represented as follows.
Equation
[0057] For node Q j the linear interpolation weights for vertex k can be calculated as follows.
Equation
[0058] Figures 4A - 4C show an example of calculating interpolation weights in a 2D unit hypercube using a quantization grid Q of size 4 × 5 according to an embodiment. More specifically, Figure 4A shows the two dimensions of the unit hypercube, shown as dimension 0 and dimension 1 respectively. Figure 4B shows the quantization grid Q for the two dimensions of the unit hypercube, with the indices V k = k = (k0, k1, …, k D-1 ) of the corresponding 4 vertices explicitly shown. Figure 4C shows, for the 4 vertices of the unit hypercube shown in Figure 4B, node Q 2,1shows the coordinates and the corresponding calculated weights.
[0059] To use the LUT W, for an input q = (q0, q1, …, q D-1 ) within the unit hypercube, the input is mapped to the nearest node Q j in the direction towards the origin. Here:
Number
Number
Number
Number
[0060] In one embodiment, the 5D grid X can be defined within the HSWLM space described above. To obtain better numerical stability and interpolation quality, the 5D grid X can be aligned with the boundaries of the effective input parameter space. Such alignment can typically help to properly perform interpolation for input points located near the boundaries. For the reshaping function index and the DM metadata L1-mid, since their original values are independent (separated) from the other dimensions, the values can be used for the grid X. For example, changing the reshaping function index and the DM metadata L1-mid does not invalidate valid inputs at other points. On the other hand, for the SDR pixel values, when using the YCbCr color space, 0 ≦ s i Y , s i Cb , s i CrNot all combinations of ≤ 1 are valid. One way to prevent invalid inputs in the YCbCr color space is to cross-reference the YCbCr input with other color spaces. For example, in the RGB or HSV color spaces, all combinations of [0,1] values are valid (see also Figure 6B). Here, RGB represents red, green, and blue, and HSV represents hue, saturation, and value.
[0061] In one exemplary embodiment, the scaled HSV color space for grid creation can be designed such that the density of the grid is proportional to the perceived color difference. For example, the V component can be scaled in a non-linear manner such that the density of the grid remains approximately the same for different luminance values. Let the H, S, and V channels of the i-th pixel in the input SDR image S(117) be s i HSV =(s i H ,s i S ,s i V ) be denoted. Then, the HSW color space can be defined as follows.
Equation
[0062] Figures 5A - 5B graphically show the effect of V - W scaling on the probability density function (PDF) of luminance Y according to an embodiment. The scaling can be performed, for example, in the HSW channel block (224) of process (200). In this particular example, FIG. 5A graphs the PDF as a function of luminance Y for a grid created in the HSV color space and then converted to the YCbCr color space. FIG. 5B graphs the PDF as a function of luminance Y for a grid created in the HSW color space and then converted to the YCbCr color space. Comparing the two PDFs, it is apparent that the PDF of FIG. 5B is advantageously more uniform than the PDF of FIG. 5A. Here, the luminance Y is within the typical SMPTE range (SMPTE represents the Society of Motion Picture and Television Engineers).
[0063] Figures 6A - 6B are diagrams showing an exemplary grid X created in the HSW color space and converted to the YCbCr color space and the RGB color space, respectively, according to an embodiment. In this particular example, the grid size is 13×5×9, and the range is [0,1] for each of the HSW dimensions. From FIG. 6A, it can be seen that the grid Q in the YCbCr color space only occupies a part of the [0,1][0,1][0,1] cube. In contrast, the grid Q in the RGB color space occupies the entire [0,1][0,1][0,1] cube. In both cases, the lattice points are properly aligned with the respective valid color space boundaries. The edges connecting R = G = B = 0 or R = G = B = 1 on the RGB color space boundary have hue values {0,1 / 6,2 / 6,3 / 6,4 / 6,5 / 6}. Therefore, in order to place lattice points exactly on these edges, the grid size in the H dimension can be selected according to Equation 6n + 1, where n is a positive integer. The other edges on the RGB color space boundary have a saturation value of 1, according to which, regardless of any particular grid size, they have lattice points on them.
[0064] Typically, the DM metadata L1-mid of the HDR image represents the average of the RGB channels of the image. However, in process (200), the RGB channels of the output HDR image (240) depend on the chroma offset (230) (see FIG. 2). As a result, it is necessary to estimate the DM metadata L1-mid. According to an exemplary embodiment, such an estimate (222) is obtained using the initial HDR image (216). More specifically, the estimated value (222) of the DM metadata L1-mid, denoted as ^m〔m with ^〕, can be calculated as the average of the Y channel of the initial HDR image (216) as follows.
Equation
[0065] Referring back to FIG. 2, in one exemplary embodiment, the 5D LUT Φ (228) used in the chroma offset processing block (220) of process (200) can be defined on a 5D grid in the HSWLM space. In operation, the 5D LUT Φ (228) outputs a chroma offset (230) in response to an input vector (226) defined in the HSWLM space. The input vector (226) is composed of the HSW channel block (224), the reshaping function index map (212), and the estimated value (222) of the DM metadata L1-mid. An exemplary training process that can be used to load values into the 5D LUT (228) will be described in more detail below (see, for example, FIGS. 7-8). The chroma offset (230) obtained using the 5D LUT (228) is added to the initial HDR image (216) (232), thereby generating the output HDR image (240). The DM block (135) then processes the output HDR image (240) to generate the output DM image (137).
[0066] In one exemplary embodiment, the above linear interpolation may be used to determine the chroma offset (230) for different pixels, for example, on a pixel-by-pixel basis. Here, the linear interpolation operation is denoted as ^φ. For the i-th pixel, the chroma offset r i CbCr =(r i Cb ,r i Cr ) can be calculated as follows.
Equation
Equation
Equation
Equation
[0067] Parameter learning FIG. 7 is a flowchart showing a method (700) for loading values into a 5D LUT (228) according to an embodiment. In one exemplary embodiment, the method (700) depends on a cost function (704), which may typically include a color shift term and one or more regularization terms. For each selected grid point (706) on a suitable (e.g., pre-defined as described above) 5D grid (702), the method (700) is configured to find a chroma offset (710) corresponding to an approximate minimum value of the cost function (704) through a sequential iterative minimization process (708), and to update the emerging 5D LUT (228) using the found chroma offset (712). The method (700) further includes repeating the set of processing operations (708), (710), (712) for a plurality of different selected grid points (706). The end of this iterative cycle occurs when an exit condition (714) is met. After termination, the method (700) includes outputting the loaded 5D LUT and storing it as the 5D LUT (228) (see also FIG. 2).
[0068] In one exemplary embodiment, the cost function (704) can be constructed to drive the sequential iterative minimization process (708) to find a nearly optimal chroma offset (710) that corrects the aforementioned color shift on the 5D grid X (702) for a plurality of grid points. For computational efficiency, the processing operations (708), (710), (712) can be configured to process one grid point at a time. Hereinafter, for the selected grid point X i (706) at index i, the H, S, and W channels at that point are s i HSW =(s i H ,s i S ,s i W )=(X i,0 ,X i,1 ,X i,2 ) and represented as; the reshaping function index is l i =X i,3is represented as; the DM metadata L1-mid(222) is m i =X i,4 is represented as. The corresponding HDR value of the initial HDR image (216) is v i (init),YCbCr =f l (B) (s i YCbCr ).
[0069] When the chroma offset (230) is not applied, the HDR value v i (init),YCbCr is passed to the DM block (135) to obtain the initial DM value ^s i (init),YCbCr =f (DM) (v i (init),YCbCr ,m i ). On the other hand, when the chroma offset r CbCr =(r Cb ,r Cr ) is applied to the Cb channel and Cr channel of the HDR value v i (init),YCbCr , a new HDR value v i YCbCr =v i (init),YCbCr +(0,r Cb ,r Cr ) is generated for the HDR image (240), and the final output DM value for the DM image (137) is ^s i YCbCr =f (DM) (v i YCbCr ,m i ).
[0070] In one exemplary embodiment, it is desirable for the chroma offset (710) to significantly reduce (e.g., to an imperceptible level) or completely eliminate the SDR-HDR [from SDR to HDR] color shift. However, it is also desirable for the chroma offset (710) not to generate artifacts or change the "look" of the image. Thus, the cost function (704) can be constructed to include a color shift term and one or more regularization terms to ensure stability. Since the cost function (704) considers only one lattice point (706) at a time, the cost function (704) can be constructed such that it is not affected by the position of the lattice points (706) on the 5D grid (702), and further, is not affected by any particular topological features of the 5D grid (702).
[0071] For illustrative purposes and without any implied limitation, the example of the cost function (704) described below includes the following terms: the color difference cost E hue , the offset cost E off , the luminance change cost E lum , the chroma change cost E sat , and the valid range cost E valid . The total cost function E total (704) is defined as follows. E total =E hue +λ off E off +λ lum E lum +λ sat E sat +E valid (18) where λ off , λ lum , and λ sat are weighting constants. For the term E valid , since the output value of this particular term is either 0 or infinity, the weighting constant is 1. Exemplary values for the other weighting constants are λ off = 0.01, λ lum = 0.04, and λ satIt may be 0.0025. In various other embodiments, the cost function (704) may have more or fewer cost terms. Some of the terms of such other cost functions (704) may be different from the exemplary cost terms listed above.
[0072] Color difference cost E hue is a color shift term. In one exemplary embodiment, the color shift may be measured by the difference in hue in the HSV color space. Note that when the SDR image has neutral colors, its hue may be indeterminate and its saturation may be 0. Measuring the color shift by the change in saturation in the HSV color space can typically help to appropriately handle such situations. The H, S, and V channels of the input SDR value and the final output DM value are respectively [Number] are represented as. Then, the color difference cost E hue can be defined as follows. [Number] Here, θ hue , θ sat is the threshold of an acceptable color shift. The function diff hue (a, b) = min(|a - b|, 1 - |a - b|) can be used to measure the difference in hue. This is because, for example, the maximum difference in hue is 0.5. From Equation (19), for non-neutral SDR values, i.e., s i S > 0, if diff hue (^s i H , s i H ) 2 ≤ θ hue then the color difference cost E hue is 0. On the other hand, for neutral SDR values, i.e., s i S = 0, when (^s i S ) 2 ≤ θ sat the color difference cost Ehue is still 0. In one exemplary embodiment, the parameter values for Equation (19) are θ hue = 0.008 2 and θ sat = 0.008 2 may be.
[0073] Offset cost E off is a regularization term configured to regularize the chroma offset within a reasonable range and avoid overfitting. Such an offset cost E off can be defined as follows.
Equation
[0074] Luminance change cost E lum is a regularization term configured to regularize the change in luminance caused by the chroma offset. Such a luminance change cost E lum can be defined as follows.
Equation
[0075] The chroma change cost E sat is a regularization term configured to regularize the change in chroma caused by the chroma offset. Such a chroma change cost E sat can be defined as follows.
Equation
[0076] The valid range cost E valid is a regularization term configured to limit the new HDR value to the valid range. Such a valid range cost E valid can be defined as follows.
Equation
[0077] FIG. 8 is a flowchart showing a sequential iterative minimization process (708) according to an embodiment. The sequential iterative minimization process (708) uses a cost function (704), e.g., the total cost function E of Equation (18). total The input to the sequential iterative minimization process (708) includes a lattice point X i (804) and an initial value (802) of a chroma offset and a step size. The output of the sequential iterative minimization process (708) includes a chroma offset (710) corresponding to the minimum value of the cost function (704). The chroma offset (710) thus obtained may typically be stored in a 5D LUT (228). For computational efficiency, the sequential iterative minimization process (708) is configured to find the chroma offset (710) within a relatively small local range specified for a computational block (806). A corresponding processing loop including a block (808) for calculating the value of the cost function (704) is executed until convergence (814) or a maximum number of iterations t max (812) is reached. The step size can be changed (typically reduced) in a change block (816) to better correspond the chroma offset (710) to the actual minimum value of the cost function (704) within the local range used. For computational efficiency, the step size is not allowed to be smaller than a specified fixed minimum step size checked in a step size check block (818).
[0078] In an exemplary embodiment, the initial value (802) of the chroma offset may be set to r0 CbCr =(0, 0). In the t-th iteration step for t≧1, the local range R t CbCr for the processing block (806) can be set as follows.
Equation
Equation
[0079] When the sequential iterative minimization process (708) reaches the maximum number of iterations t max in the counter block (812), the chroma offset (710) is set to r CbCr* = r t CbCr . Otherwise, if it is determined that the estimated chroma offset is at a local minimum (814), that is, r t CbCr = r t-1 CbCr , the step size for the next iteration step may be reduced to Δr t+1 = αΔr t . Here, α < 1 is a constant. The value of α can typically be selected to achieve the desired convergence rate. If r t CbCr ≠ r t-1 CbCr , the process may proceed with Δr t+1 = Δr t . The accuracy of the chroma offset (710) can typically be controlled by the constant Δr min . If Δr t+1 < Δr min , it is considered that the convergence criterion is satisfied, and the output chroma offset (710) is r CbCr* = r t CbCrIt is set to. In one exemplary embodiment, the following parameter values may be used. Δr1 = 10 -3 , Δr min = 10 -6 , α = 0.5, t max = 100. In some embodiments, Δr min may be in the range between about 10 -3 and 10 -6 , and the visual quality of the corresponding output HDR image (137) may still be acceptable for certain applications even if it is at the top of this range.
[0080] For example, according to the exemplary embodiments disclosed above, with reference to the summary section and / or any one or part or all of FIGS. 1 to 8 and any combination thereof, an apparatus including a video distribution system capable of changing the dynamic range of an input image is provided. The distribution system includes a memory (e.g., 228 in FIG. 2) that stores a plurality of chroma offset values corresponding to the grid points of a 5D grid, and a processor (e.g., 120 in FIG. 1) that converts an input image (e.g., 117 in FIG. 2) having a first dynamic range (e.g., SDR in FIG. 2) into a corresponding output image (e.g., 240 in FIG. 2) having a larger second dynamic range (e.g., HDR in FIG. 2). The processor is configured to generate an intermediate image (e.g., 216 in FIG. 2) having a second dynamic range by reshaping the input image (e.g., 210 in FIG. 2) (the reshaping is performed using a reshaping function index map (e.g., 212 in FIG. 2) having respective indexes for identifying the corresponding reshaping function applied to each pixel of the intermediate image); estimate display management metadata values corresponding to the intermediate image (e.g., 222 in FIG. 2); and generate an output image by applying respective chroma offsets (e.g., 230 in FIG. 2) to each pixel of the intermediate image (e.g., 232 in FIG. 2), where each chroma offset is determined from the plurality of chroma offset values by addressing the grid points using three respective pixel values of the respective index, the display management metadata value, and the corresponding pixel of the input image.
[0081] In some embodiments of the above apparatus, the first dynamic range is a standard dynamic range (e.g., SDR, FIG. 2), and the second dynamic range is a high dynamic range (e.g., HDR, FIG. 2).
[0082] In some embodiments of any of the above apparatus, the three respective pixel values are the hue value, saturation value, and intensity value of the corresponding pixel of the input image.
[0083] In some embodiments of any of the above devices, the processor is further configured to non-linearly rescale the intensity values of the input image (e.g., V-to-W, 224 in FIG. 2), and the three respective pixel values are the hue value, saturation value, and rescaled intensity value of the corresponding pixel of the input image.
[0084] In some embodiments of any of the above devices, the video distribution system is configured to generate a display-adapted image (e.g., 137 in FIG. 2) by applying display management processing to the output image.
[0085] In some embodiments of any of the above devices, the video distribution system includes a video encoder (e.g., 120, FIG. 1) that includes at least a portion of the processor.
[0086] In some embodiments of any of the above devices, the plurality of chroma offset values are arranged in memory within a lookup table that can be addressed using a reshaping function index value, metadata value, hue value, saturation value, and intensity value.
[0087] In some embodiments of any of the above devices, the display management metadata value corresponding to the intermediate image is the L1-mid luminance value.
[0088] In some embodiments of any of the above devices, the processor is further configured to perform linear interpolation of chroma offset values (e.g., equations (6)-(11)) to determine each chroma offset.
[0089] For example, according to another exemplary embodiment disclosed above, with reference to the summary section and / or any one or part or all of FIGS. 1 to 8, a method for changing the dynamic range of an input image is provided. This method includes converting an input image (e.g., 117 in FIG. 2) having a first dynamic range (e.g., SDR in FIG. 2) into a corresponding output image (e.g., 240 in FIG. 2) having a larger second dynamic range (e.g., HDR in FIG. 2). The conversion is performed using a plurality of pre-computed chroma offset values corresponding to the lattice points of a 5D grid. The conversion includes generating an intermediate image (e.g., 216 in FIG. 2) having a second dynamic range by reshaping the input image (e.g., 210 in FIG. 2) (the reshaping is performed using a reshaping function index map (e.g., 212 in FIG. 2) having respective indices for identifying the corresponding reshaping function applied to each pixel of the intermediate image); estimating display management metadata values corresponding to the intermediate image (e.g., 222 in FIG. 2); and generating the output image by applying respective chroma offsets (e.g., 230 in FIG. 2) to each pixel of the intermediate image (e.g., 232 in FIG. 2). Each chroma offset is determined from the plurality of chroma offset values by addressing the lattice points using three respective pixel values of the respective index, the display management metadata value, and the corresponding pixel of the input image.
[0090] In some embodiments of the above method, the method further includes non-linearly rescaling the intensity values of the input image (e.g., V-to-W, 224, FIG. 2), and the three respective pixel values are the hue value, saturation value, and rescaled intensity value of the corresponding pixel of the input image.
[0091] In some embodiments of any of the above methods, the method further includes generating a display-adapted image (e.g., 137 in FIG. 2) by applying display management processing to the output image.
[0092] In some embodiments of any of the above methods, the display management metadata value corresponding to the intermediate image is the L1-mid luminance value.
[0093] In some embodiments of any of the above methods, the converting further includes performing linear interpolation of chroma offset values (e.g., equations (6)-(11)) to determine respective chroma offsets.
[0094] In some embodiments of any of the above methods, the plurality of chroma offset values are arranged in a lookup table addressable using a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value.
[0095] For example, according to yet another exemplary embodiment disclosed above, in the summary section and / or with respect to any one or a part or all of FIGS. 1-8, a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including a method of changing a dynamic range of an input image is provided. The method includes converting an input image (e.g., 117 in FIG. 2) having a first dynamic range (e.g., SDR in FIG. 2) to a corresponding output image (e.g., 240 in FIG. 2) having a larger second dynamic range (e.g., HDR in FIG. 2). The conversion is performed using a plurality of pre-computed chroma offset values corresponding to grid points of a 5D grid. The conversion includes generating an intermediate image (e.g., 216 in FIG. 2) having a second dynamic range by reshaping the input image (e.g., 210 in FIG. 2) (reshaping is performed using a reshaping function index map (e.g., 212 in FIG. 2) having respective indices identifying corresponding reshaping functions for each pixel of the intermediate image); estimating display management metadata values corresponding to the intermediate image (e.g., 222 in FIG. 2); and generating the output image by applying respective chroma offsets (e.g., 230 in FIG. 2) to each pixel of the intermediate image (e.g., 232 in FIG. 2), where each chroma offset is determined from the plurality of chroma offset values by addressing the grid points using three respective pixel values of the respective index, the display management metadata value, and the corresponding pixel of the input image.
[0096] For example, according to yet another exemplary embodiment disclosed above, with reference to, in the summary section and / or any one or part or all of any combination of FIGS. 1-8, a method for generating a plurality of chroma offset values for performing color shift correction in an output image generated by changing the dynamic range of an input image is provided. This method includes, respectively, defining a five-dimensional grid (e.g., 702 in FIG. 7) having first, second, third, fourth, and fifth dimensions representing a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value; defining a cost function (e.g., 704 in FIG. 7) for quantifying at least a cost (e.g., Equation (19)) for a color difference between the input image and the output image; for each grid point of the five-dimensional grid (e.g., 804 in FIG. 8), performing sequential iterative minimization of the cost function to determine each set of chroma offset values; and arranging each set of chroma offset values in an addressing electronic look-up table (e.g., 228 in FIG. 2) using a set of discrete values corresponding to the first, second, third, fourth, and fifth dimensions.
[0097] In some embodiments of the above method, defining the cost function includes including one or more regularization terms configured to keep sequential iterative minimization within a valid range in the cost function.
[0098] In some embodiments of any of the above methods, sequential iterative minimization is performed within a local range of parameters (e.g., 806 in FIG. 8) that is narrower than the full range of parameters.
[0099] In some embodiments of any of the above methods, sequential iterative minimization is performed using a variable step size (e.g., 818 in FIG. 8).
[0100] Regarding the processes, systems, methods, heuristics, etc. described in this specification, while the steps of such processes, etc. are described as occurring according to a certain ordered sequence, it should be understood that such processes may be implemented using the steps described herein in an order other than the order described. Further, certain steps may be executed simultaneously, other steps may be added, or certain steps described herein may be omitted. In other words, the description of the processes in this specification is provided for the purpose of exemplifying certain embodiments and should in no way be construed as limiting the scope of the claims.
[0101] Therefore, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided will be apparent upon reading the above description. The scope should not be determined with reference to the above description, but rather with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. Future developments are expected and intended in the technologies described herein, and the disclosed systems and methods to be incorporated into such future embodiments. In short, it should be understood that this application is capable of modification and variation.
[0102] All terms used in the claims are intended to be given their broadest reasonable interpretation and their ordinary meaning as would be understood by one of ordinary skill in the art of the technology described herein, unless explicitly indicated to the contrary herein. In particular, the use of singular articles such as "a", "the", "said", etc. should be read as reciting one or more of the indicated elements, unless the claim states a clear limitation to the contrary.
[0103] The abstract of the present disclosure is provided to enable a reader to quickly ascertain the nature of the technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, for the purpose of better presenting the flow of the disclosure, it can be seen that various features are grouped together in various embodiments. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are explicitly recited in each claim. Rather, as reflected by the following claims, the subject matter of the invention lies in less features than all of the features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the detailed description herein, with each claim standing on its own as a separately claimed subject matter.
[0104] This disclosure includes references to exemplary embodiments, but this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which would be apparent to those skilled in the art to which this disclosure pertains, are considered to be within the principles and scope of the disclosure, for example, as set forth in the following claims.
[0105] Some embodiments may be implemented as a circuit-based process, including possible implementation on a single integrated circuit.
[0106] Some embodiments may be embodied in the form of methods and apparatuses for carrying out those methods. Some embodiments may also be embodied in the form of program code recorded on a tangible medium such as a magnetic recording medium, an optical recording medium, a solid state memory, a floppy disk, a CD-ROM, a hard drive, or any other non-transitory machine-readable storage medium, where the program code, when loaded and executed by a machine such as a computer, causes the machine to become an apparatus for practicing the patented invention. Some embodiments may also be embodied in the form of program code, including being loaded and / or executed by a machine, such as program code stored on a non-transitory machine-readable storage medium, where the program code, when loaded and executed by a machine such as a computer or a processor, causes the machine to become an apparatus for practicing the patented invention. When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates in a manner similar to specific logic circuitry.
[0107] Unless explicitly stated otherwise, each numerical value and range should be interpreted as approximate as if the word “about” or “approximately” preceded the value or range.
[0108] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the subject matter recited in the claims to facilitate interpretation of the claims. Such use should not be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0109] The elements in the following method claims, if any, are recited in a particular order using corresponding labeling, but unless the claim recitation otherwise implies a particular order for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular order.
[0110] References to "one embodiment" or "an embodiment" in this specification mean that the particular features, structures, or characteristics described in connection with that embodiment can be included in at least one embodiment of the present disclosure. The appearances of the phrase "in an embodiment" in various places in this specification do not necessarily all refer to the same embodiment, and distinct or alternative embodiments do not necessarily exclude each other. The same applies to the term "implementation".
[0111] Unless otherwise specified in this specification, the use of ordinal adjectives such as "first", "second", "third", etc. to refer to a particular one of a plurality of similar objects only indicates that different instances of such similar objects are being referred to, and does not imply that the similar objects so referred to must be in a corresponding order or sequence in time, in space, in ranking, or in any other way.
[0112] Unless otherwise specified in this specification, in addition to its plain meaning, the connective "when" may also be interpreted to mean, or alternatively to mean, "at the time of" or "upon" or "in response to determining" or "in response to detecting", and that interpretation may depend on the corresponding specific context. For example, the phrases "when determined" or "when [stated condition] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "at the time of detecting [stated condition or event]" or "in response to detecting [stated condition or event]".
[0113] Also, for the purposes of this description, the terms "coupled", "coupling", "coupled to", "connected", "connecting", or "connected to" refer to any manner known in the art or later developed in which energy is permitted to be transferred between two or more elements, and the intervention of one or more additional elements, while not necessary, is contemplated. Conversely, terms such as "directly coupled to", "directly connected to", etc. mean that no such additional elements are present.
[0114] As used herein with respect to elements and standards, the term compatible means that an element communicates with other elements in a manner fully or partially specified by the standard and is recognized by the other elements as being fully capable of communicating with the other elements in the manner specified by the standard. Compatible elements need not operate internally in the manner specified by the standard.
[0115] The functionality of the various elements shown in the figures, including any functional blocks labeled "processor" and / or "controller", can be provided through dedicated hardware, as well as through the use of hardware capable of executing software in association with appropriate software. When provided by a processor, the functionality can be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors some of which can be shared. Further, the explicit use of the term "processor" or "controller" should not be construed to refer only to hardware capable of executing software, and can implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage. Conventional and / or custom other hardware may also be included. Similarly, any switches shown in the figures are conceptual only. Their functionality may be implemented through the operation of program logic, through dedicated logic, through an interaction of program control and dedicated logic, or even manually, and the specific technique can be selectable by the implementer as more specifically understood from the context.
[0116] As used in this application, the terms "circuit" and "circuits" refer to (a) a circuit implementation of only hardware (such as an implementation with only analog and / or digital circuits), (b) (where applicable) (i) a combination of analog and / or digital hardware circuits and software / firmware, and (ii) any portion of a hardware processor (including a digital signal processor), software, and memory that operate together to perform various functions in a device such as a mobile phone or a server, and (c) a hardware circuit and / or processor such as a microprocessor or a portion of a microprocessor that requires software (such as firmware) for operation, but the software may not be present when not required for operation. The definition of a circuit may apply to all uses of this term in this application, including any claims. As a further example, as used in this application, the term "circuit" also covers a hardware circuit or processor (or multiple processors), or a portion of a hardware circuit or processor, and its (or their) accompanying software and / or firmware implementation. The term "circuit" covers, for example, a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit in a server, a cellular network device, or other computing or network device, when applicable to an element of a particular claim.
[0117] It should be understood by those skilled in the art that any block diagram in this specification represents a conceptual diagram of an exemplary circuit embodying the principles of the present disclosure. Similarly, any flowchart, flow diagram, state transition diagram, pseudocode, etc. is substantially represented in a computer-readable medium and thus represents various processes that can be executed by such a computer or processor, whether or not a computer or processor is explicitly shown.
[0118] The "Summary of the Invention" in this specification is intended to introduce some exemplary embodiments, and additional embodiments are described in the "Detailed Description of the Invention" and / or with reference to one or more drawings. The "Summary of the Invention" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0119] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). 〔EEE1〕 A video distribution system capable of changing the dynamic range of an input image, the distribution system comprising: a memory storing a plurality of chroma offset values corresponding to grid points of a 5D grid; a processor for converting the input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the processor: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes identifying a corresponding reshaping function to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by specifying the grid points using the respective indexes, the display management metadata values, and three respective pixel values of corresponding pixels of the input image; Video distribution system. 〔EEE2〕 The first dynamic range is a standard dynamic range; The second dynamic range is a high dynamic range. The video distribution system described in EEE1. [EEE3] In the video distribution system according to EEE1 or 2, each of the three pixel values is a hue value, a saturation value, and a luminance value of a corresponding pixel of the input image. [EEE4] The processor is further configured to non-linearly rescale the intensity value of the input image; In the video distribution system according to any one of EEE1 to 3, each of the three pixel values is a hue value, a saturation value, and a rescaled intensity value of the corresponding pixel of the input image. The video distribution system according to any one of EEE1 to 3. [EEE5] The video distribution system according to any one of EEE1 to 4, wherein the video distribution system is configured to generate a display-adapted image by applying a display management process to the output image. [EEE6] The video distribution system according to any one of EEE1 to 5, comprising a video encoder including at least a part of the processor. [EEE7] In the video distribution system according to any one of EEE1 to 6, the plurality of chroma offset values are arranged in the memory in a lookup table addressable using a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value. [EEE8] In the video distribution system according to any one of EEE1 to 7, the display management metadata value corresponding to the intermediate image is an average luminance value. [EEE9] The video distribution system according to any one of EEE1 to 8, wherein the processor is further configured to perform linear interpolation of the chroma offset values to determine the respective chroma offsets. 〔EEE10〕 A method for changing the dynamic range of an input image, the method comprising: converting the input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the conversion being performed using a plurality of pre-calculated chroma offset values corresponding to grid points of a five-dimensional grid; the conversion comprising: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes for identifying corresponding reshaping functions to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by specifying the grid points using the respective indexes, the display management metadata values, and three respective pixel values of the corresponding pixel of the input image; Method. 〔EEE11〕 further comprising non-linearly rescaling the intensity values of the input image, wherein the three respective pixel values are a hue value, a saturation value, and a rescaled intensity value of the corresponding pixel of the input image. The method according to EEE10. 〔EEE12〕 The method according to EEE10 or 11, further comprising the step of generating a display-adapted image by applying display management processing to the output image. [EEE13] The method according to any one of EEE10 to 12, wherein the display management metadata value corresponding to the intermediate image is an average luminance value. [EEE14] The method according to any one of EEE10 to 13, wherein the converting further comprises performing linear interpolation of the chroma offset values to determine the respective chroma offsets. [EEE15] The method according to any one of EEE10 to 14, wherein the plurality of chroma offset values are arranged in a lookup table addressable using a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value. [EEE16] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations including the method according to any one of EEE10 to 15. [EEE17] A method of generating a plurality of chroma offset values for performing color shift correction in an output video generated by changing the dynamic range of an input image, the method comprising: defining a 5D grid having first, second, third, fourth, and fifth dimensions respectively representing a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value; defining a cost function for quantifying at least one cost for a color difference between the input image and the output image; for each grid point of the 5D grid, performing sequential iterative minimization of the cost function to determine each set of the chroma offset values; In an electronic look-up table addressable using a set of discrete values corresponding to the first, second, third, fourth, and fifth dimensions, arranging each set of the chroma offset values, Method. 〔EEE18〕 Defining the cost function includes including one or more regularization terms configured to keep the successive iterative minimization within a valid range in the cost function, the method according to EEE17. 〔EEE19〕 The method according to EEE17 or 18, wherein the successive iterative minimization is performed within a local range of parameters narrower than the entire range of parameters. 〔EEE20〕 The method according to any one of EEE17 to 19, wherein the successive iterative minimization is performed using a variable step size. 〔EEE21〕 For the t-th iteration step of the successive iterative minimization, each best chroma offset r t CbCr is
Number
Claims
1. A video delivery system capable of changing the dynamic range of an input image, the delivery system comprising: a memory storing a plurality of chroma offset values corresponding to grid points of a 5D grid; a processor for converting the input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the processor: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes for identifying a corresponding reshaping function to be applied to each pixel of the intermediate image; estimating display management metadata values corresponding to the intermediate image; generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by specifying the grid points using the respective indexes, the display management metadata values, and three respective pixel values of the corresponding pixel of the input image; A video delivery system.
2. The first dynamic range is a standard dynamic range; The second dynamic range is a high dynamic range. The video delivery system according to claim 1.
3. The three respective pixel values are the hue value, saturation value, and luminance value of the corresponding pixel of the input image. The video delivery system according to claim 1.
4. The processor is further configured to non-linearly rescale the intensity value of the input image; The three respective pixel values are the hue value, saturation value, and rescaled intensity value of the corresponding pixel of the input image. The video delivery system according to claim 1.
5. The video delivery system is configured to generate a display-adapted image by applying a display management process to the output image. The video delivery system according to claim 1.
6. The video distribution system according to claim 1, comprising a video encoder including at least a part of the processor.
7. The video distribution system according to any one of claims 1 to 6, wherein the plurality of chroma offset values are arranged in the memory in a look-up table addressable using a reshaping function index value, a metadata value, a hue value, a saturation value, and an intensity value.
8. The video distribution system according to claim 1, wherein the display management metadata value corresponding to the intermediate image is an average luminance value.
9. The video distribution system according to claim 1, wherein the processor is further configured to perform linear interpolation of the chroma offset values to determine the respective chroma offsets.
10. A method for changing the dynamic range of an input image, the method comprising: converting the input image having a first dynamic range into a corresponding output image having a larger second dynamic range, the conversion being performed using a plurality of pre-calculated chroma offset values corresponding to grid points of a five-dimensional grid; the conversion comprising: generating an intermediate image having the second dynamic range by reshaping the input image, the reshaping being performed using a reshaping function index map having respective indexes for identifying a corresponding reshaping function applied to each pixel of the intermediate image; estimating a display management metadata value corresponding to the intermediate image; generating the output image by applying respective chroma offsets to each pixel of the intermediate image, the respective chroma offsets being determined from the plurality of chroma offset values by specifying the grid points using the respective indexes, the display management metadata value, and three respective pixel values of the corresponding pixel of the input image; method.
11. further comprising non-linearly rescaling the intensity value of the input image, wherein the three respective pixel values are a hue value, a saturation value, and a rescaled intensity value of the corresponding pixel of the input image. The method according to claim 10.
12. The method according to claim 10, further comprising the step of generating a display-adapted image by applying display management processing to the output image.
13. The method according to claim 10, wherein the display management metadata value corresponding to the intermediate image is an average luminance value.
14. The method according to claim 10, wherein the converting further comprises performing linear interpolation of the chroma offset values to determine the respective chroma offsets.
15. The method according to any one of claims 10 to 14, wherein the plurality of chroma offset values are arranged in a look-up table addressable using reshaping function index values, metadata values, hue values, saturation values, and intensity values.
Citation Information
Patent Citations
Color appearance preservation in video codecs
US20200351524A1
Adjustable trade-off between quality and computation complexity in video codecs
WO2021076822A1