Image reshaping in video coding using rate distortion optimization
Adaptive codeword allocation using RDO-based signal shaping addresses inefficiencies in HDR video coding, enhancing compression and visual quality by optimizing codeword distribution based on luminance variance.
Patent Information
- Application Number
- JP2025111199
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-14
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2039-02-13
AI Technical Summary
Existing video coding technologies struggle with inefficient compression and image quality at higher bit depths, particularly in high dynamic range (HDR) content, due to traditional nonlinear opto-electronic functions and electro-optical transfer functions that do not optimize rate-distortion performance.
Implementing a rate-distortion optimization (RDO) based signal shaping technique that allocates codewords adaptively to bins based on luminance variance, using a non-optimal but practical analytical solution to improve compression efficiency and subjective/objective quality metrics.
Enhances compression efficiency and visual quality by optimizing codeword allocation, improving both PSNR and subjective visual quality, especially for HDR content.
Smart Images

Figure 2025134977000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 792,122, filed January 14, 2019, U.S. Provisional Patent Application No. 62 / 782,659, filed December 20, 2018, U.S. Provisional Patent Application No. 62 / 772,228, filed November 28, 2018, U.S. Provisional Patent Application No. 62 / 739,402, filed October 1, 2018, U.S. Provisional Patent Application No. 62 / 726,608, filed September 4, 2018, U.S. Provisional Patent Application No. 62 / 691,366, filed June 28, 2018, and U.S. Provisional Patent Application No. 62 / 630,385, filed February 14, 2018, each of which is incorporated herein by reference in its entirety.
[0002] technology The present invention relates generally to image and video coding, and more particularly, to image shaping in video coding. [Background technology]
[0003] In 2013, the MPEG group of the International Organization for Standardization (ISO), in collaboration with the International Telecommunication Union (ITU), published the first draft of the HEVC (also known as H.265) video coding standard (Non-Patent Document 4). More recently, the group announced a call for evidence to support the development of a next-generation coding standard that will provide improved coding performance over existing video coding technologies.
[0004] As used herein, the term "bit depth" refers to the number of pixels used to represent one of the color components of an image. Traditionally, images were coded with 8 bits per pixel per color component (e.g., 24 bits per pixel), but current architectures may now support higher bit depths, such as 10 bits, 12 bits, or more.
[0005] In traditional image pipelines, captured images are quantized using a nonlinear opto-electronic function (OETF), which converts linear scene light into a nonlinear video signal (e.g., gamma-coded RGB or YCbCr). On the receiver, the signal is then processed by an electro-optical transfer function (EOTF), which converts the video signal values into output screen color values before being displayed on a display. Such nonlinear functions include the traditional "gamma" curve described in ITU-R Rec. BT.709 and BT.2020, the "PQ" (perceptual quantization) curve described in SMPTE ST2084, and the "Hybrid Log-gamma" or "HLG" curve described in Rec. ITU-R BT.2100. Summary of the Invention
[0006] As used herein, the term “forward reshaping” refers to the process of sample-to-sample or codeword-to-codeword mapping of a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma or PQ or HLG, etc.) to an image of the same or a different bit depth and different codeword distribution or representation. The shaping enables improved compressibility or improved image quality at a fixed bit rate. For example, but not by way of limitation, shaping may be applied to 10-bit or 12-bit PQ-encoded HDR video to improve coding efficiency in 10-bit video coding architectures. At the receiver, after decompressing the shaped signal, the receiver can apply an “inverse shaping function” to restore the signal to the original codeword distribution. As recognized herein by the inventors, improved techniques for integrated shaping and coding of images are desirable as development for the next generation of video coding standards begins. The methods of the present invention may be applicable to a variety of video content, including but not limited to standard dynamic range (SDR) and / or high dynamic range (HDR) content.
[0007] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any approach described in this section qualifies as prior art merely by virtue of its inclusion in this section. Likewise, it should not be assumed that problems identified with one or more approaches have been recognized by virtue of this section in any prior art unless specifically noted. [Brief explanation of the drawings]
[0008] Embodiments of the present invention are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference numerals refer to similar elements and in which:
[0009] [Figure 1A]Shows an exemplary process for a video delivery pipeline;
[0010] [Figure 1B] 1 shows an exemplary process for data compression using signal shaping according to the prior art;
[0011] [Figure 2A] 1 illustrates an exemplary architecture for an encoder using hybrid in-loop shaping in accordance with an embodiment of the present invention.
[0012] [Figure 2B] 1 illustrates an exemplary architecture for a decoder using hybrid in-loop shaping in accordance with one embodiment of the present invention.
[0013] [Figure 2C] 1 illustrates an example architecture for intra CU decoding using shaping according to an embodiment.
[0014] [Figure 2D] 1 illustrates an example architecture for inter-CU decoding using shaping according to an embodiment.
[0015] [Figure 2E] 1 illustrates an example architecture for intra-CU decoding within an inter-coded slice, according to an embodiment for luma or chroma processing.
[0016] [Figure 2F] 1 illustrates an exemplary architecture for intra-CU decoding within an inter-coded slice, according to an embodiment for chroma processing.
[0017] [Figure 3A] 1 illustrates an exemplary process for encoding video using a shaping architecture, according to one embodiment of the present invention.
[0018] [Figure 3B] 1 illustrates an exemplary process for decoding video using a shaping architecture, according to one embodiment of the present invention.
[0019] [Figure 4] 1 illustrates an exemplary process for reallocating codewords within a shaped domain, in accordance with an embodiment of the present invention.
[0020] [Figure 5] 1 illustrates an exemplary process for deriving a shaping threshold according to an embodiment of the present invention.
[0021] [Figure 6A] Figure 5 shows an exemplary data plot for deriving a shaping threshold according to the process illustrated and one embodiment of the present invention. [Figure 6B] Figure 5 shows an exemplary data plot for deriving a shaping threshold according to the process illustrated and one embodiment of the present invention. [Figure 6C] Figure 5 shows an exemplary data plot for deriving a shaping threshold according to the process illustrated and one embodiment of the present invention. [Figure 6D] Figure 5 shows an exemplary data plot for deriving a shaping threshold according to the process illustrated and one embodiment of the present invention.
[0022] [Figure 6E] 10 illustrates an example of codeword allocation according to bin distribution, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] Signal shaping and encoding techniques for compressing images using rate-distortion optimization (RDO) are described herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail in order to avoid unnecessarily obscuring, burying, or obscuring the present invention.
[0024] overview SUMMARY OF THE INVENTION
[0003] In an encoder, a processor receives an input image of a first codeword representation to be shaped into a second codeword representation, the second codeword representation allowing for more efficient compression than the first codeword representation, and generates a forward shaping function that maps pixels of the input image to the second codeword representation. To generate the forward shaping function, the encoder divides the input image into a plurality of pixel regions, assigns each of the pixel regions to one of a plurality of codeword bins according to a first luminance characteristic of the pixel region, calculates a bin metric for each of the plurality of codeword bins according to a second luminance characteristic of each of the pixel regions assigned to the respective codeword bins, assigns several codewords in the second codeword representation to each codeword bin according to the bin metric for each codeword bin and a rate-distortion optimization criterion, and generates the forward shaping function in response to the assignment of codewords in the second codeword representation to each of the plurality of codeword bins.
[0025] In another embodiment, in a decoder, a processor receives coded bitstream syntax elements that characterize a shaping model, the syntax elements including one or more of a flag indicating a minimum codeword bin index value to be used in the shaping construction process, a flag indicating a maximum codeword bin index value to be used in the shaping construction process, a flag indicating a shaping model profile type, the model profile type associated with default bin-related parameters including bin importance values, or a flag indicating one or more delta bin importance values to be used to adjust the default bin importance values defined in the shaping model profile. The processor determines, based on the shaping model profile, a default bin importance value for each bin and an assignment list of default numbers of codewords to be assigned to each bin according to the bin importance values. Then, for each codeword bin, the processor: Determine the bin importance value by adding the default bin importance value to the delta bin importance value; determining a number of codewords to be assigned to a codeword bin based on the bin importance value of the bin and the assignment list; A forward shaping function is generated based on the number of codewords assigned to each codeword bin.
[0026] In another embodiment, in a decoder, a processor receives an encoded bitstream including one or more encoded shaped images in a first codeword representation and metadata related to shaping information for the encoded shaped images. The processor generates an inverse shaping function and a forward shaping function based on metadata related to the shaping information, where the inverse shaping function maps pixels of the shaped image from a first codeword representation to a second codeword representation and the forward shaping function maps pixels of the image from the second codeword representation to the first codeword representation. The processor extracts from the coded bitstream a coded shaped image comprising one or more coded units, where for one or more coded units in the coded shaped image: For a shaped intra-coded coding unit (CU) in the coded shaped image, the processor: generating a first shaped reconstruction sample for the CU based on the shaped residual and the shaped prediction sample for the CU; generating a shaped loop filter output based on the first shaped reconstruction samples and loop filter parameters; applying an inverse shaping function to the shaped loop filter output to generate decoded samples of the coding unit in a second codeword representation; storing the decoded samples of the coding unit in a second codeword representation in a reference buffer; For a shaped inter-coded coding unit in a coded shaped image, the processor: applying a forward shaping function to the predicted samples stored in the reference buffer in a second codeword representation to generate second shaped predicted samples; generating a second shaped reconstructed sample of the coding unit based on the shaped residual in the coded CU and the second shaped predicted sample; generating a shaped loop filter output based on the second shaped reconstruction samples and the loop filter parameters; applying an inverse shaping function to the shaped loop filter output to generate samples of the coding unit in a second codeword representation; The samples of the coding unit in the second codeword representation are stored in a reference buffer. Finally, the processor generates a decoded image based on the stored samples in the reference buffer.
[0027] In another embodiment, in a decoder, a processor receives an encoded bitstream including one or more encoded shaped images in an input codeword representation and shaping metadata (207) for the one or more encoded shaped images in the encoded bitstream. The processor generates a forward shaping function (282) based on the shaping metadata, which maps pixels of the image from the first codeword representation to the input codeword representation. The processor generates an inverse shaping function (265-3) based on the shaping metadata or the forward shaping function, where the inverse shaping function maps pixels of the shaped image from the input codeword representation to the first codeword representation. The processor extracts an encoded shaped image including one or more coded units from the encoded bitstream, where: For an intra-coded coding unit (intra CU) in the coded shaped image, the processor: generating shaped reconstructed samples (285) of the intra CU based on the shaped residuals in the intra CU and the intra predicted shaped predicted samples; applying an inverse shaping function (265-3) to the shaped reconstructed samples of the intra CU to generate decoded samples of the intra CU in a first codeword representation; applying a loop filter (270) to the decoded samples of the intra CU to generate output samples of the intra CU; Store the output samples of the intra CU in a reference buffer; For an inter-coded CU (inter-CU) in the coded shaped image, the processor: applying a forward shaping function (282) to the inter predicted samples stored in the reference buffer in the first codeword representation to generate shaped predicted samples for the inter CU in the input codeword representation; generating shaped reconstructed samples for the inter-CU based on the shaped residuals for the inter-CU and the shaped predicted samples for the inter-CU; applying an inverse shaping function (265-3) to the shaped reconstructed samples of the inter-CU to generate decoded samples of the inter-CU in a first codeword representation; applying a loop filter (270) to the decoded samples of the inter-CU to generate output samples of the inter-CU; Store the inter-CU output samples in the reference buffer; A decoded image in the first codeword representation is generated based on the output samples in the reference buffer.
[0028] Video Delivery Processing Pipeline Example FIG. 1A illustrates an exemplary process of a conventional video distribution pipeline (100), showing various stages from video capture to video content display. A sequence of video frames (102) is captured or generated using an image generation block (105). The video frames (102) may be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide the video data (107). In a production phase (110), the video data (107) is edited to provide a video production stream (112).
[0029] The video data in the production stream (112) is then provided to a processor in block (115) for post-production editing. Post-production editing in block (115) may involve adjusting or modifying the color or brightness in individual areas of the image to improve image quality or achieve a particular look of the image according to the video creator's creative intent. This is sometimes referred to as "color timing" or "color grading." Other editing (e.g., scene selection and sequencing, image cropping, adding computer-generated visual special effects, etc.) may be performed in block (115) to provide a final version of the production (117) for distribution. During post-production editing (115), the video image is displayed on a reference display (125).
[0030] Following post-production (115), the video data of the final production (117) may be delivered to an encoding block (120) for downstream decoding and delivery to playback devices such as television sets, set-top boxes, and movie theaters. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-Ray, and other distribution formats, to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) that represents an identical or close approximation of the signal (117). The receiver may be attached to a target display (140), which may have completely different characteristics from the reference display (125). In that case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137).
[0031] Signal Reshaping FIG. 1B illustrates an exemplary process for signal shaping according to U.S. Patent No. 6,277,999. Given an input frame (117), a forward shaping block (150) analyzes the input and coding constraints and generates a codeword mapping function that maps the input frame (117) to a requantized output frame (152). For example, the input (117) may be encoded according to some electro-optical transfer function (EOTF) (e.g., gamma). In some embodiments, metadata may be used to convey information about the shaping process to a downstream device (e.g., a decoder). As used herein, the term "metadata" refers to any auxiliary information transmitted as part of an encoded bitstream that assists a decoder in rendering a decoded image. Such metadata includes, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters, as described herein.
[0032] Following encoding (120) and decoding (130), the decoded frames (132) may be processed by a backward (or inverse) shaping function (160), which converts the requantized frames (132) back to the original EOTF domain (e.g., gamma) for further downstream processing, such as the aforementioned display management process (135). In some embodiments, the backward shaping function (160) may be integrated with a dequantizer within the decoder (130), for example, as part of the dequantizer in an AVC or HEVC video decoder.
[0033] As used herein, the term "reshaper" may refer to a forward or inverse shaping function used when encoding and / or decoding digital images. Examples of shaping functions are discussed in U.S. Patent No. 6,275,999, which proposed an in-loop block-based image shaping method for high dynamic range video coding. The design allows for block-based shaping within the coding loop, but at the cost of increased complexity. Specifically, the design requires maintaining two sets of decoded image buffers: one set for inversely shaped (or unshaped) decoded pictures that can be used for both prediction without shaping and output to a display, and another set for forward-shaped decoded pictures that are used only for prediction with shaping. While forward-shaped decoded pictures can be computed on the fly, the complexity cost is very high, especially for inter-prediction (motion compensation with subpixel interpolation). In general, display-picture-buffer (DPB) management is complex and requires very careful attention, so as the inventors realize, a simplified method for encoding video is desirable.
[0034] In U.S. Patent No. 6,449,493, additional shaping-based codec architectures were presented, including architectures with an outer out-of-loop shaper, an in-loop intra-only shaper, an architecture with an in-loop shaper for the prediction residual, and a hybrid architecture that combines both intra-loop shaping and inter-residual shaping. The main goal of these proposed shaping architectures is to improve subjective visual quality. Thus, many of these approaches result in worse objective metrics, especially the well-known Peak Signal to Noise Ratio (PSNR) metric.
[0035] In this invention, a new shaper based on Rate-Distortion Optimization (RDO) is proposed. In particular, when the targeted distortion metric is MSE (Mean Square Error), the proposed shaper improves both the subjective visual quality and a commonly used objective metric based on PSNR, Bjontegaard PSNR (BD-PSNR), or Bjontegaard Rate (BD-Rate). It is noted that, without loss of generality, any of the proposed reconstruction architectures can be applied to one or more of the luma component, the chroma component, or a combination of the luma and chroma components.
[0036] Shaping based on Rate-Distortion Optimization Consider a shaped video signal represented by a bit depth of B bits in a given color component (e.g., B=10 for Y, Cb, and / or Cr), for a total of 2 B There are available codewords. The desired codeword range is [0 2 B ] into N segments or bins. k is the picture distortion D between the source picture and the decoded or reconstructed picture given a target bit rate R. After shaping mapping , denote the number of codewords in the kth segment or bin. Without loss of generality, D may be expressed as a measure of the sum of squared errors (SSE) between the corresponding pixel values of the source input (Source(i,j)) and the reconstructed picture (Recon(i,j)). D=SSE=Σ i,j Diff(i,j) 2 (1) where: Diff(i,j)=Source(i,j)-Recon(i,j) is.
[0037] The optimization shaping problem can be rewritten as follows: Given a bit rate R, we minimize D by k(k=0,1,…,N-1) where
number
[0038] Various optimization methods can be used to find the solution, but the optimal solution can be very complex for real-time encoding. In this invention, a non-optimal but more practical analytical solution is proposed.
[0039] Without loss of generality, consider an input signal represented by a bit depth of B bits (e.g., B=10), where the codeword is uniformly divided into N bins (e.g., N=32). By default, each bin has M a =2 B / N codewords (e.g., for N=32 and B=10, M a =32). Next, a more efficient codeword allocation based on RDO is demonstrated through an example.
[0040] As used herein, the term "narrow range" [CW1, CW2] refers to a continuous range of codewords between codewords CW1 and CW2, which does not include the full dynamic range [0 2 B -1]. For example, in one embodiment, the narrow range is a subset of [16*2 (B-8) ,235*2 (B-8) ] (e.g., for B=10, the narrow range includes the value [64 940]). o Assuming that the input signal has a narrow dynamic range, what is referred to as the "default" shaping will filter the signal over the full range [0 2 Bo -1]. Then, each bin is f =CEIL((2 Bo / (CW2-CW1))*M a ) codewords, or, for our example, B o If =B=10, M f=CEIL((1024 / (940-64))*32)=38 codewords. Here, CEIL(x) represents the ceiling function that maps x to the smallest integer greater than or equal to x. Without loss of generality, in the following examples, for simplicity, B o Assume =B.
[0041] For the same quantization parameter (QP), the effect of increasing the number of codewords in a bin is equivalent to allocating more bits to code the signal in the bin, which is therefore equivalent to lowering the SSE or improving the PSNR. However, a uniform increase in codeword allocation in each bin may not yield better results than coding without shaping, because the PSNR gain may not outweigh the increase in bitrate. That is, this is not a good tradeoff in terms of RDO. Ideally, we would like to allocate more codewords only to the bins that provide the best tradeoff in terms of RDO, i.e., a significant SSE reduction (PSNR increase) at the expense of a small bitrate increase.
[0042] In one embodiment, RDO performance is improved through adaptive piecewise shaping mapping. This method is applicable to any type of signal, including standard dynamic range (SDR) and high dynamic range (HDR) signals. Using the simple case above as an example, the goal of the present invention is to generate M a Pieces or M f The idea is to assign one of the code words.
[0043] At the encoder, given N codeword bins for the input signal, the average luminance variance of each bin can be approximated as: For each bin, the sum of the block variances (var bin (k)) and counter (c bin (k)) to 0. For example, for k=0,1,…,N-1, var bin (k)=0, cbin Let (k)=0. Divide the picture into L*L non-overlapping blocks (e.g., L=16). For each picture block, calculate the block luma mean and luma variance of block i (e.g., Luma_mean(i) and Luma_var(i)). Assign a block to one of N bins based on the block's mean luma. In one embodiment, if Luma_mean(i) is within the kth segment in the input dynamic range, then the total bin luma variance for the kth bin is incremented by the luma variance of the newly assigned block, and the counter for that bin is increased by 1. That is, if the ith pixel region belongs to the kth bin: var bin (k)=var bin (k)+Luma_var(i); (2) c bin (k)=c bin (k)+1 For each bin, calculate the average luminance variance for that bin by dividing the sum of the block variances within that bin by the counter, assuming that the counter is not equal to 0. Alternatively, bin If (k) is not 0, var bin (k)=var bin (k) / c bin (k) (3) Let's say.
[0044] Those skilled in the art will appreciate that alternative metrics other than luminance variance may be applied to characterize the sub-blocks, for example, standard deviation of luminance values, weighted luminance variance or luminance values, peak luminance, etc.
[0045] In one embodiment, the following pseudocode shows an example of how the encoder may adjust the bin assignments using the calculated metrics for each bin. For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k) <TH U (4) M k =M f ; else M k =M a ; / / (Note: This requires that each bin has at least M a To ensure that we have codewords. / / Alternatively, M a +1 codeword may be assigned.) end where TH U represents a predetermined upper threshold.
[0046] In another embodiment, the allocation may be performed as follows: For the kth bin, if there are no pixels in the bin M k =0; else if TH0 bin (k) <TH1(5) M k =M f ; else M k =M a ; end where TH0 and TH1 represent predetermined lower and upper thresholds.
[0047] In another embodiment For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k)>TH L (6) M k =M f ; else M k =M a ; end where TH L represents a predetermined lower threshold.
[0048] The above example uses two preselected numbers M f and M a It shows how to select the number of codewords for each bin from a threshold (e.g., TH U or TH L ) can be determined based on rate-distortion optimization, for example through exhaustive search. The threshold may be adjusted based on the quantization parameter value (QP). In one embodiment, for B=10, the threshold may range from 1,000 to 10,000.
[0049] In one embodiment, to speed up processing, the threshold value may be determined from a fixed set of values, e.g., {2,000, 3,000, 4,000, 5,000, 6,000, 7,000}, using a Lagrangian optimization method. For example, for each TH(i) value in the set, a compression test can be run with a fixed QP using a predefined training clip, and the value of the objective function J can be calculated. J is J(i)=D+λR (7) The optimal threshold can then be defined as the TH(i) value in said set for which J(i) is minimum.
[0050] In a more general case, a lookup table (LUT) can be predefined. For example, in Table 1, the first row lists the possible bin metrics (e.g., var binThe first row defines a set of thresholds that divide the entire range of (k) values into segments, and the second row defines the corresponding number of codewords (CWs) to be assigned in each segment. In one embodiment, one rule for constructing such a LUT is: if the bin variance is too large, we may need to spend more bits to reduce the SSE, so M a If the bin variance is very small, then M a A larger CW value can be assigned.
[0051] Table 1: Example LUT for codeword assignment based on bin variance threshold [Table 1]
[0052] Using Table 1, the mapping of thresholds to codewords can be generated as follows: For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k) <TH0 M k =CW0; else if TH0 bin (k) <TH1(8) M k =CW1; … else if TH p-1 bin (k) <TH p Mk=CW p ; … else if var bin (k)>TH q-1 Mk=CW q ; end
[0053] For example, given two thresholds and three codeword allocations, for B=10, in one embodiment, TH0=3,000, CW0=38, TH1=10,000, CW1=32, and CW2=28.
[0054] In another embodiment, the two thresholds TH0 and TH1 may be selected as follows: a) Consider TH1 as a very large number (even infinite) and select TH0 from a predetermined set of values, e.g., using the RDO optimization of equation (7). Given TH0, now define a second set of possible values for TH1, e.g., the set {10,000, 15,000, 20,000, 25,000, 30,000}, and apply equation (7) to identify the optimal value. This approach can be performed iteratively with a limited number of thresholds or until convergence.
[0055] After assigning codewords to bins according to one of the previously defined schemes, M k The sum of these values is the maximum (2 B ), or there are unused codewords. If there are unused codewords, one may decide to simply do nothing, or assign the unused codewords to a particular bin. On the other hand, if the algorithm allocates more codewords than are available, it may decide to reduce M by, for example, renormalizing the CW values. k You might want to readjust the values. Alternatively, you can generate a forward shaping function using the existing Mk values, but then k M k ) / 2 B The output values of the shaping function may be readjusted by scaling by: An example of a codeword reallocation technique is also described in US Pat. No. 6,239,999.
[0056] 4 shows an exemplary process for assigning codewords to a shaped domain in accordance with the RDO technique described above. In step 405, the desired shaped dynamic range is divided into N bins. After the input image is divided into non-overlapping blocks (step 410), for each block: Step 415 calculates its luminance characteristics (eg, mean and variance). Step 420 assigns each image block to one of N bins. Step 425 calculates the average luminance variance in each bin Given the values calculated in step 425, in step 430 each bin is assigned a number of code words according to one or more thresholds, for example using any of the code word assignment algorithms shown in equations (4) through (8). Finally, in step (435), the final code word assignments may be used to generate forward and / or inverse shaping functions.
[0057] In one embodiment, by way of example and not limitation, a forward LUT (FLUT) can be constructed using the following C code: tot_cw=2 B ; hist_lens=tot_cw / N; for (i=0;i <N;i++) { double temp=(double) M[i] / (double)hist_lens; / / M[i] is M k Compatible with for (j=0;j <hist_lens;j++) { CW_bins_LUT_all[i*hist_lens + j]=temp; } Y_LUT_all[0]=CW_bins_LUT_all[0]; for (i=1;i <tot_cw;i++) { Y_LUT_all[i]=Y_LUT_all[i-1]+CW_bins_LUT_all[i]; } for (i=0;i <tot_cw;i++) { FLUT[i]=Clip3(0,tot_cw-1,(Int)(Y_LUT_all[i]+0.5)); }
[0058] In one embodiment, the inverse LUT can be constructed as follows: low=FLUT[0]; high=FLUT[tot_cw-1]; first=0; last=tot_cw-1; for (i=1;i <tot_cw;i++) if(FLUT[0] <FLUT[i]) { first=i-1; break; } for (i=tot_cw-2;i>=0;i--) if(FLUT[tot_cw-1]>FLUT[i]) { last=i+1; break; } for (i=0;i <tot_cw;i++) if(i<=low) { ILUT[i]=first; } else if(i>=high) { ILUT[i]=last; } else { for(j=0;j <tot_cw-1;j++) if(FLUT[j]>=i) { ILUT[i]=j; break; }} }
[0059] Syntactically, we can reuse syntax proposed in previous applications, such as the piecewise polynomial modes or parametric models in US Pat. Nos. 5,629,999 and 5,749,102. Table 2 shows such an example for N=32 for equation (4). Table 2: Formatting syntax using the first parametric model [Table 2] where: reshaper_model_profile_type specifies the profile type to be used in the shaper construction process. A given profile specifies the number of bins, default bin importance or priority values, and default codeword assignments (e.g., M a and / or M f It may provide information about the default values used, such as the reshaper_model_scale_idx Specifies the index value of a scale factor (denoted ScaleFactor) used in the shaper construction process. The value of ScaleFactor allows for improved control of the shaper function to improve overall coding efficiency. reshaper_model_min_bin_idx Specifies the minimum bin index to be used in the shaper construction process. The value of reshaper_model_min_bin_idx must be in the range 0 to 31, inclusive. reshaper_model_max_bin_idx Specifies the maximum bin index to be used in the shaper construction process. The value of reshaper_model_max_bin_idx must be in the range 0 to 31, inclusive. reshaper_model_bin_profile_delta[i] specifies the delta value used to adjust the profile of the ith bin in the shaper construction process. The value of reshaper_model_bin_profile_delta[i] should range from 0 to 1 inclusive.
[0060] Table 3 shows another embodiment with an alternative, more efficient syntax representation. Table 3: Reconstruction syntax using the second parametric model [Table 3] where: resharper_model_delta_max_bin_idx is set equal to the maximum allowed bin index (e.g., 31) minus the maximum bin index used in the shaper construction process. reshaper_model_num_cw_minus1 plus one specifies the number of codewords to be signaled. reshaper_model_delta_abs_CW[i] specifies the i-th absolute delta codeword value. reshaper_model_delta_sign_CW[i] specifies the sign for the ith delta codeword, and: reshaper_model_delta_CW[i]=(1-2*reshaper_model_delta_sign_CW[i])* reshaper_model_delta_abs_CW [i]; reshaper_model_CW[i]=32+reshaper_model_delta_CW[i]. reshaper_model_bin_profile_delta[i]specifies the delta value used to adjust the profile of the ith bin in the shaper construction process. The value of reshaper_model_bin_profile_delta[i] ranges from 0 to 1 when reshaper_model_num_cw_minus1 is equal to 0. The value of reshaper_model_bin_profile_delta[i] ranges from 0 to 2 when reshaper_model_num_cw_minus1 is equal to 1. When reshaper_model_bin_profile_delta[i] is set equal to 0, CW=32; when reshaper_model_bin_profile_delta[i] is set equal to 1, CW=reshaper_model_CW[0]; when reshaper_model_bin_profile_delta[i] is set equal to 2, CW=reshaper_model_CW[1]. In one embodiment, reshaper_model_num_cw_minus1 is allowed to be greater than 1, so that reshaper_model_num_cw_minus1 and reshaper_model_bin_profile_delta[i] This allows for being signaled in ue(v) for more efficient coding.
[0061] In another embodiment, the number of codewords per bin may be explicitly defined, as set forth in Table 4. Table 4: Formatting syntax using the third model [Table 4] reshaper_model_number_bins_minus1 + 1 specifies the number of bins used for the luma component. In some embodiments, it may be more efficient for the number of bins to be a power of 2. The total number of bins is then given by its log2 representation, e.g. log2_reshaper_model_number_bins_minus_minus1 For example, in the case of 32 bins, log2_reshaper_model_number_bins_minus1 =4. reshaper_model_bin_delta_abs_cw_prec_minus1Adding 1 to it specifies the number of bits used for the representation of the syntax reshaper_model_bin_delta_abs_CW[i]. reshaper_model_bin_delta_abs_CW[i] Specifies the absolute delta codeword value for the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] Specifies the sign of reshaper_model_bin_delta_abs_CW[i] as follows: · When reshaper_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] has a positive value. · Otherwise (when reshaper_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] has a negative value. When reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. The variable RspDeltaCW[i] = (1 - 2 * reshaper_model_bin_delta_sign_CW [i]) * reshaper_model_bin_delta_abs_CW [i]; The variable OrgCW is set to (1 << BitDepthY) / (resharper_model_number_bins_minus_1 + 1); The variable RspCW[i] is derived as follows: if reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx then RspCW[i] = OrgCW + RspDeltaCW[i] else, RspCW[i] = 0
[0062] In one embodiment, assuming one of the above examples, for example, the sign word assignment according to Equation (4), an example of how to define the parameters in Table 2 includes the following: First, assign "bin importance" as follows: For the kth bin, if M k =0; bin_importance=0; else if M k ==M f bin_importance=2; (9) else bin_importance=1; end
[0063] As used herein, the term "bin importance" is a value assigned to each of the N codeword bins that indicates the importance of all codewords within that bin relative to the other bins in the shaping process.
[0064] In one embodiment, the default_bin_importance for reshaper_model_min_bin_idx through reshaper_model_max_bin_idx may be set to 1. The value of reshaper_model_min_bin_idx must be greater than or equal to 0. k The value of reshaper_model_max_bin_idx is set to the smallest bin index with M k reshaper_model_bin_profile_delta for each bin in [reshaper_model_min_bin_idx reshaper_model_max_bin_idx] is the difference between bin_importance and default_bin_importance.
[0065] Below we show an example of how to use the proposed parametric model to construct a Forward Reshaping LUT (FLUT) and an Inverse Reshaping LUT (ILUT). 1) Divide the luminance range into N bins (e.g., N=32) 2) Derive the bin-importance index for each bin from the syntax, e.g., For the kth bin, if reshaper_model_min_bin_idx<=k<=reshaper_model_max_bin_idx bin_importance[k]=default_bin_importance[k]+reshaper_model_bin_profile_delta[k]; else bin_importance[k]=0; 3) Automatically pre-assign codewords based on bin importance: For the kth bin, if bin_importance[k]==0 M k =0; else if bin_importance[k]==2 M k =M f ; else M k =M a ; end 4) Build a forward shaping LUT based on the codeword assignments for each bin by accumulating the codewords assigned for each bin. The sum should be less than or equal to the total codeword budget (e.g., 1024 for 10-bit full range). (See the first C code for an example.) 5) Build the inverse shaping LUT (see the first C code for an example).
[0066] From a syntax point of view, alternative methods are also applicable: the key is either explicitly or implicitly determined by the number of codewords in each bin (e.g., M for k=0, 1, 2, ..., N-1). k) in a bin. In one embodiment, the number of codewords in each bin can be explicitly specified. In another embodiment, the codewords can be specified differently. For example, the number of codewords in a bin can be determined using the difference between the number of codewords in the current bin and the previous bin (e.g., M_Delta(k)=M(k)-M(k-1)). In another embodiment, the number of most commonly used codewords (e.g., M M ), and the number of codewords in each bin can be specified by the difference of the number of codewords in each bin from this number (e.g., M_Delta(k)=M(k)-M M ) can be expressed as
[0067] In one embodiment, two formatting methods are supported: one designated "default formatter" and one designated M f is assigned to all bins. The other, denoted "adaptive shaper", applies the adaptive shaper described above. These two methods can be signaled to the decoder using a special flag, e.g., sps_reshaper_adaptive_flag, as in Patent Document 6 (e.g., use sps_reshaper_adaptive_flag=0 for the default shaper and sps_reshaper_adaptive_flag=1 for the adaptive shaper).
[0068] The present invention is applicable to any shaper proposed in U.S. Patent No. 6,049,499, such as an external shaper, an in-loop intra-only shaper, an in-loop residual shaper, or an in-loop hybrid shaper. As an example, FIGS. 2A and 2B show an exemplary architecture for hybrid in-loop shaping according to an embodiment of the present invention. In FIG. 2A, the architecture combines elements from both the in-loop intra-only shaping architecture (top of the figure) and the in-loop residual architecture (bottom of the figure). Under this architecture, for intra slices, shaping is applied to picture pixels, while for inter slices, shaping is applied to the prediction residual. In the encoder (200_E), two new blocks are added to a conventional block-based encoder (e.g., HEVC): a block (205) that estimates the forward shaping function (e.g., according to FIG. 4), a forward picture shaping block (210-1), and a forward residual shaping block (210-2). It applies forward shaping to one or more color components of the input video (117) or the prediction residual. In some embodiments, these two operations may be performed as part of a single image shaping block. Parameters (207) related to the determination of the inverse shaping function in the decoder may be passed to a lossless encoder block (e.g., CABAC 220) of the video encoder so that they can be embedded in the coded bitstream (122). In intra mode, the inverse shaping functions are: intra prediction (225-1), transform and quantization (T&Q), and inverse transform and inverse quantization (Q&Q). -1 &T -1 ) all use shaped pictures. In both modes, the picture stored in the DPB (215) is always in inverse shaping mode, which requires an inverse picture shaping block (e.g., 265-1) or an inverse residual shaping block (e.g., 265-2) before the loop filters (270-1, 270-2). As shown in FIG. 2A, an intra / inter-slice switch allows switching between the two architectures, depending on the type of slice being encoded. In another embodiment, in-loop filtering for intra slices may be performed before inverse shaping.
[0069] In the decoder (200_D), the following new canonical blocks are added to a conventional block-based decoder: a block (250) (shaper decode) that reconstructs the forward and backward shaping functions based on the encoded shaping function parameters (207); a block (265-1) that applies the inverse shaping function to the decoded data; and a block (265-2) that applies both the forward and inverse shaping functions to generate the decoded video signal (162). For example, in (265-2), the reconstructed value is given by Rec=ILUT(FLUT(Pred)+Res), where FLUT represents the forward shaping LUT and ILUT represents the inverse shaping LUT.
[0070] In some embodiments, the operations associated with blocks 250 and 265 may be combined into a single processing block. As shown in Figure 2B, an intra / inter slice switch allows switching between two modes depending on the type of slice in the encoded video picture.
[0071] FIG. 3A shows an exemplary process (300_E) for encoding video using a shaping architecture (e.g., 200_E) according to one embodiment of the present invention. If no shaping is enabled (path 305), encoding (335) proceeds as known in prior art encoders (e.g., HEVC). If shaping is enabled (path 310), the encoder may have the option to apply a predetermined (default) shaping function (315) or adaptively determine (325) a new shaping function based on picture analysis (320) (e.g., as described in FIG. 4). After encoding (330) the image using the shaping architecture, the remainder of the encoding follows the same steps as a conventional encoding pipeline (335). If adaptive shaping is used (312), metadata related to the shaping function is generated as part of the “encode shaping function” step (327).
[0072] FIG. 3B shows an exemplary process (300_D) for decoding video using a shaping architecture (e.g., 200_D) according to one embodiment of the present invention. If no shaping is enabled (path 340), the decoder decodes (350) the picture as in a conventional decoding pipeline, and then generates (390) an output frame. If shaping is enabled (path 360), the decoder determines whether to apply (375) a predetermined (default) shaping function or adaptively determine (380) a shaping function based on received parameters (e.g., 207). Following decoding (385) using the shaping architecture, the remainder of the decoding follows a conventional decoding pipeline.
[0073] As described in Patent Document 6 and herein above, the forward shaping LUT (FwdLUT) may be constructed by integration, while the inverse shaping LUT may be constructed based on backward mapping using a forward shaping LUT (FwdLUT). In one embodiment, the forward LUT may be constructed using piecewise linear interpolation. At the decoder, inverse shaping can be performed directly using the backward LUT or also by linear interpolation. The piecewise linear LUT is constructed based on the input pivot point and the output pivot point.
[0074] Let (X1,Y1), (X2,Y2) be two input pivot points and the corresponding output values for each bin. Input values X between X1 and X2 can be interpolated by: Y=((Y2-Y1) / (X2-X1))*(X-X1)+Y1 In a fixed-point implementation, the above equation can be rewritten as: Y=((m*X+2 FP_PREC-1 )>>FP_PREC)+c where m and c represent the scalar and offset for linear interpolation, and FP_PREC is a constant related to the fixed-point precision.
[0075] As an example, the FwdLUT can be constructed as follows: lutSize=(1< <BitDepth Y ) Variables: binNum=reshaper_model_number_bins_minus1+1 and binLen=lutSize / binNum Let's say. For the i-th bin, its two extreme pivots (e.g., X1 and X2) may be derived as X1=i*binLen and X2=(i+1)*binLen, and: binsLUT[0]=0; for(i=0; i <reshaper_model_number_bins_minus1+1; i++){ binsLUT[(i+1)*binLen]=binsLUT[i*binLen]+RspCW[i]; Y1=binsLUT[i*binLen]; Y2=binsLUT[(i+1)*binLen]; scale=((Y2-Y1)*(1< <FP_PREC)+(1<<(log2(binLen)-1)))> >(log2(binLen)); for(j=1;j <binLen;j++){ binsLUT[i*binLen+j]=Y1+((scale*j+(1<<(FP_PREC-1)))>>FP_PREC); } }
[0076] FP_PREC defines the fixed-point precision of the fractional part of the variable (e.g., FP_PREC=14). In some embodiments, binsLUT[] may be calculated with a higher precision than that of FwdLUT. For example, binsLUT[] values may be calculated as 32-bit integers, while FwdLUT may be 16-bit clipped binsLUT values.
[0077] Adaptive Threshold Derivation As mentioned above, during shaping, the codeword assignment is determined based on one or more thresholds (e.g., TH, TH U , T.H. L (e.g., using a threshold). In one embodiment, such a threshold may be adaptively generated based on content characteristics. Figure 5 shows an example process for deriving such a threshold, according to one embodiment. 1) In step 505, the luminance range of the input image is divided into N bins (e.g., N=32), where N is also denoted as PIC_ANALYZE_CW_BINS. 2) In step 510, image analysis is performed to calculate luminance characteristics for each bin. For example, the percentage of pixels in each bin (denoted as BinHist[b], b=1, 2, ..., N) may be calculated, where: BinHist[b] = 100 * (total number of pixels in bin b) / (total number of pixels in the picture) (10) As discussed earlier, another good metric of image quality is the average variance (or standard deviation) of the pixels in each bin, denoted as BinVar[b]. BinVar[b] may be calculated in "block mode," as in the steps described in the section leading to equations (2) and (3). Alternatively, the block-based calculation can be refined with a pixel-based calculation. For example, let vf(i) denote the variance associated with the group of pixels surrounding the i-th pixel within an m-by-m neighborhood window (e.g., m=5) centered on the i-th pixel. For example,
number
number
number
[0078] 3) In step 515, the average bin variances (and their corresponding indices) are sorted, for example, but not limited to, in descending order. For example, the sorted BinVar values may be stored in BinVarSortDsd[b], and the sorted bin indices may be stored in BinIdxSortDsd[b]. As an example, using C code, this process may be described as follows: for(int b=0; b <PIC_ANALYZE_CW_BINS; b++ / / Initialize (unsorted) { BinVarSortDsd[b]=BinVar[b]; BinIdxSortDsd[b]=b; } / / Sort (see code example in Appendix 1) bubbleSortDsd(BinVarSortDsd,BinIdxSortDsd,PIC_ANALYZE_CW_BINS); An exemplary plot of the sorted mean bin variance factors is shown in Figure 6A.
[0079] 4) Given the bin histogram values calculated in step 510, calculate and store the cumulative density function (CDF) in step 520 according to the order of the sorted mean bin variances. For example, if the CDF is stored in the array BinVarSortDsdCDF[b], then in one embodiment: BinVarSortDsdCDF[0]=BinHist[BinIdxSortDsd[0]]; for(int b=1; b <PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b]=BinVarSortDsdCDF[b-1]+BinHist[BinIdxSortDsd[b]]; } An exemplary plot (605) of the calculated CDF based on the data in Figure 6A is shown in Figure 6B. The CDF value versus sorted average bin variance pair {x = BinVarSortDsd[b], y = BinVarSortDsdCDF[b]} can be interpreted as "there are y% of pixels in the picture that have a variance greater than or equal to x" or "there are (100-y)% of pixels in the picture that have a variance less than x." 5) Finally, in step 525, given the CDF as a function of the sorted average bin variance value, BinVarSortDsdCDF[BinVarSortDsd[b]], thresholds can be defined based on bin variances and cumulative percentages.
[0080] Examples for determining a single threshold or two thresholds are shown in Figures 6C and 6D, respectively. H ) is used, as an example, T H means "k% of pixels are T H It may be defined as "average variance such that T H can be calculated by finding the intersection of the CDF plot (605) at k% (e.g., 610) (e.g., the BinVarSortDsd[b] value where BinVarSortDsdCDF=k%). For example, as shown in FIG. 6C, for k=50, T H = 2.5. Then, BinVar[b] <THをもつビンのためにM f codewords and assign M aAs a rule of thumb, it is preferable to assign a larger number of codewords to bins with smaller variance (e.g., for a 10-bit video signal with 32 bins, M f >32>M a ).
[0081] When using two thresholds, TH L and T.H. U An example of selecting TH is shown in Figure 6D. For example, without loss of generality, L 80% of the pixels are vf≧TH L may be defined as the variance with L =2.3), TH U 10% of all pixels are vf≧TH U may be defined as the variance with U =3.5). Given these thresholds, BinVar[b] <TH L For bins with M f The code words are expressed as BinVar[b]≧TH U For bins with M a TH codewords can be assigned. L and T.H. U For bins with BinVar between ≡ 0 and ≡ 10, the original number of codewords per bin (eg, 32 for B=10) may be used.
[0082] The above technique can be easily extended to the case of having more than two thresholds. This relationship is given by f , M a As a rule of thumb, in low-variance bins, more codewords should be allocated to boost PSNR (and lower MSE); for high-variance bins, fewer codewords should be allocated to save bits.
[0083] In one embodiment, a set of parameters (e.g., TH L , T.H. U , M a, M f If, for example, parameters such as those of f are manually obtained through, for example, exhaustive manual parameter tuning, this automatic method can be applied to design a decision tree for categorizing each content in order to automatically set optimal manual parameters. For example, content categories include: movies, TV, SDR, HDR, comics, nature, action, etc.
[0084] To reduce complexity, in-loop shaping can be restricted using various methods. When in-loop shaping is adopted in a video coding standard, these restrictions should be normative to ensure decoder simplification. For example, in one embodiment, loop filtering may be disabled for certain block coding sizes. For example, when nTbW * nTbH < TH, the intra and inter filter modes in the inter-slice can be disabled. Here, the variable nTbW specifies the transform block width, and the variable nTbH specifies the transform block height. For example, for TH = 64, blocks of sizes 4×4, 4×8, and 8×4 are disabled for both intra-mode and inter-mode filtering in inter-coded slices (or tiles).
[0085] Similarly, in another embodiment, chroma residual scaling based on loop filtering in the intra-mode in inter-coded slices (or tiles) may be disabled. Or it may also be disabled when it is effective to have separate loop filtering and chroma split trees.
[0086] Interaction with other coding tools Loop Filtering In Patent Document 6, it is described that the loop filter can operate either in the original pixel domain or in the shaped pixel domain. In one embodiment, it is proposed that the loop filtering is performed in the original pixel domain (after picture shaping). For example, in the hybrid in-loop shaping architectures (200_E and 200_D), for intra pictures, the inverse shaping (265-1) must be applied before the loop filter (270-1).
[0087] 2C and 2D show alternative decoder architectures (200B_D and 200C_D) in which inverse shaping (265) is performed after loop filtering (270) and immediately before storing the decoded data in the decoded picture buffer (DPB) (260). In the proposed embodiment, compared to the architecture of 200_D, the inverse residual shaping formula for inter-slices is modified, and inverse shaping is performed after loop filtering (270) (e.g., via an InvLUT() function or a lookup table). In this way, inverse shaping is performed after loop filtering for both intra- and inter-slices, and the reconstructed pixels before loop filtering are in the shaping domain for both intra- and inter-coded CUs. After inverse shaping (265), all output samples stored in the reference DPB are in the original domain. Such architectures allow both slice-based and CTU-based adaptation for in-loop shaping.
[0088] As shown in Figures 2C and 2D, in one embodiment, loop filtering (270) is performed in the shaping domain for both intra-coded and inter-coded CUs, and the inverse picture shaping (265) occurs only once, thus presenting a unified, simpler architecture for both intra-coded and inter-coded CUs.
[0089] To decode an intra-coded CU (200B_D), intra prediction (225) is performed on the shaped neighboring pixels. Given the residual Res and the predicted sample PredSample, the reconstructed sample (227) is derived as follows: RecSample=Res+PredSample (14) Given the reconstructed samples (227), loop filtering (270) and inverse picture shaping (265) are applied to derive RecSampleInDPB samples that are stored in the DPB (260), where: RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(Res+PredSample))) (15) where InvLUT() denotes the inverse shaping function or inverse shaping lookup table, and LPF() denotes the loop filtering operation.
[0090] In conventional coding, the inter / intra mode decision is based on calculating a distortion function (dfunc()) between the original samples and the predicted samples. Examples of such functions include sum of squared errors (SSE), sum of absolute differences (SAD), and others. When using shaping, on the encoder side (not shown), CU prediction and mode decision are performed in the shaping domain. That is, for mode decision, distortion=dfunc(FwdLUT(SrcSample)-RecSample) (16) where FwdLUT() represents the forward shaping function (or LUT) and SrcSample represents the original image sample.
[0091] For inter-coded CUs, at the decoder side (e.g., 200C_D), inter prediction is performed using the non-shaped domain reference pictures in the DPB. Then, in the reconstruction block 275, the reconstructed pixels (267) are derived as follows: RecSample=(Res+FwdLUT(PredSample)) (17) Given the reconstructed samples (267), loop filtering (270) and inverse picture shaping (265) are applied to derive RecSampleInDPB samples that are stored in the DPB, where: RecSampleInDPB=InvLUT(LPF(RecSample)))=InvLUT(LPF(Res+FwdLUT(PredSample)))) (18)
[0092] At the encoder side (not shown), intra prediction is performed in the shaped domain as follows: Res=FwdLUT(SrcSample)-PredSample (19a) Here, it is assumed that all neighboring samples (PredSample) used for prediction are already in the shaped domain. Inter prediction (e.g., using motion compensation) is performed in the non-shaped domain (i.e., directly using reference pictures from the DPB). That is, PredSample=MC(RecSampleinDPB) (19b) where MC() represents the motion compensation function. For fast mode decision where motion estimation and residuals are not generated, the distortion can be calculated using the following formula: distortion=dfunc(SrcSample-PredSample) However, for full modal decision, where the residual is generated, the mode decision is performed in the shaped domain. distortion=dfunc(FwdLUT(SrcSample)-RecSample) (20)
[0093] Block-level adaptation As explained earlier, the proposed in-loop shaper allows adapting shaping at the CU level, e.g., by setting the variable CU_reshaper on or off as needed. Under the same architecture, for an inter-coded CU, when CU_reshaper=off, the reconstructed pixels must be in the shaping domain even if the CU_reshaper flag is set to off for this inter-coded CU. RecSample=FwdLUT(Res+PredSample) (21) Thus, intra prediction always has neighboring pixels in the shaped domain. The DPB pixel can be derived as follows: RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(FwdLUT(Res+PredSample))) (22)
[0094] For intra-coded CUs, two alternative methods are proposed, depending on the encoding process: 1) All intra-coded CUs are coded with CU_reshaper=on. In this case, no additional processing is required since all pixels are already in the shaped domain. 2) Some intra-coded CUs can be coded using CU_reshaper=off. In this case, for CU_reshaper=off, when applying intra prediction, inverse shaping needs to be applied to neighboring pixels, so that intra prediction is performed in the original domain and the final reconstructed pixel needs to be in the shaped domain. That is, RecSample=FwdLUT(Res+InvLUT(PredSample)) (23) and, RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(FwdLUT(Res+InvLUT(PredSample))))) (24)
[0095] In general, the proposed architectures can be used in various combinations, such as in-loop intra-only shaping, in-loop shaping for prediction residuals only, or a hybrid architecture that combines both intra in-loop shaping and inter residual shaping. For example, to reduce latency in the hardware decode pipeline, for inter-slice decoding, intra prediction can be performed before inverse shaping (i.e., decode intra CUs within inter slices). An exemplary architecture (200D_D) of such an embodiment is shown in FIG. 2E. In the reconstruction module (285), for inter CUs from equation (17) (e.g., Mux enables outputs from 280 and 282), RecSample=(Res+FwdLUT(PredSample)) Here, FwdLUT(PredSample) denotes the output of the inter predictor (280) followed by forward shaping (282). Instead, for intra CUs (e.g., Mux enables the output from 284), the output of the reconstruction module (285) is RecSample=(Res+IPredSample) Here, IPredSample indicates the output of the intra prediction block (284). The inverse shaping block (265-3) is Y CU =InvLUT[RecSample] Execute.
[0096] Applying intra prediction in the shaped domain for inter slices is also applicable to other embodiments, including those shown in Figure 2C (where inverse shaping is performed after loop filtering) and Figure 2D. Special care must be taken in combined inter / intra prediction modes (i.e., where during reconstruction, some samples are from inter-coded blocks and some are from intra-coded blocks). This is because in all such embodiments, inter prediction is in the original domain, while intra prediction is in the shaped domain. When combining data from both inter- and intra-predicted coding units, prediction can be performed in either of the two domains. For example, when combined inter / intra prediction modes are performed in the shaped domain, PredSampleCombined=PredSampeIntra+FwdLUT(PredSampleInter) RecSample=Res+PredSampleCombined That is, the inter-coded samples in the original domain are shaped before addition. Otherwise, when combined inter / intra prediction modes are performed in the original domain: PredSampleCombined=InvLUT(PredSampeIntra)+PredSampleInter RecSample=Res+FwdLUT(PredSampleCombined) That is, the intra-predicted samples are inversely shaped to be in the original domain.
[0097] Similar considerations are applicable to the corresponding encoding embodiment, since the encoder (e.g., 200_E) includes a decoder loop that matches the corresponding decoder. As discussed above, equation (20) describes an embodiment in which the mode decision is performed in the shaping domain. In another embodiment, the mode decision may be performed in the original domain, i.e.: distortion=dfunc(SrcSample-InvLUT(RecSample))
[0098] For luma-based chroma QP offset or chroma residual scaling, the average CU luma value
number
[0099] Chroma QP derivation Similar to US Patent No. 6,119,239, the same proposed chroma DQP derivation process may be applied to balance the luma-chroma relationship caused by the shaping curve. In one embodiment, piecewise chroma DQP values may be derived based on the codeword assignment for each bin. For example: For the kth bin scale k =(M k / M a ); (twenty five) chromaDQP=6*log2(scale k ); end
[0100] Encoder Optimization As described in Patent Document 6, when lumaDQP is enabled, it is recommended to use pixel-based weighted distortion. When shaping is used, in one example, the required weights are adjusted based on a shaping function (f(x)). For example: W rsp =f'(x) 2 (26) Here, f'(x) represents the gradient of the shaping function f(x).
[0101] In another embodiment, the piecewise weights can be derived directly based on the codeword assignment for each bin. For example: For the kth bin, Wrsp (k)=(M k / M a ) 2 (27)
[0102] For the chroma components, the weights can be set to 1 or some scaling factor sf. To reduce chroma distortion, sf can be set to greater than 1. To increase chroma distortion, sf can be set to greater than 1. In some embodiments, sf can be used to compensate for equation (25). Since chromaDQP can only be set to integers, sf can be used to accommodate the fractional part of chromaDQP. Thus: sf=2 ((chromaDQP-INT(chromaDQP)) / 3)
[0103] In another embodiment, to control chroma distortion, the chromaQPOffset value can be explicitly set in the Picture Parameter Set (PPS) or slice header.
[0104] The shaper curve or mapping function need not be fixed for the entire video sequence. For example, it can be adapted based on the quantization parameter (QP) or target bitrate. In one embodiment, a more aggressive shaper curve can be used when the bitrate is low, and less aggressive shaping can be used when the bitrate is relatively high. For example, given 32 bins in a 10-bit sequence, each bin initially has 32 codewords. When the bitrate is relatively low, a codeword between
[2840] can be used to select a codeword for each bin. When the bitrate is high, a codeword between
[3133] can be selected for each bin, or simply an identity shaper curve can be used.
[0105] Given a slice (or tile), shaping at the slice (tile) level can be performed in a variety of ways that can trade off coding efficiency for complexity, including: 1) disabling shaping only on intra slices; 2) disabling shaping on specific inter slices, such as inter slices at a particular temporal level(s), or on inter slices that are not used for reference pictures or are considered to be less important reference pictures. Such slice adaptation can also be QP / rate dependent, so that different adaptation rules can be applied for different QPs or bitrates.
[0106] At the encoder, under the proposed algorithm, the variance is calculated for each bin (e.g., BinVar(b) in Equation (13)). Based on that information, codewords can be assigned based on each bin variance. In one embodiment, BinVar(b) may be inversely linearly mapped to the number of codewords in each bin b. In another embodiment, to inversely map the number of codewords in bin b, we use (BinVar(b)) 2 , sqrt(BinVar(b)) may also be used. Essentially, this approach allows the encoder to apply arbitrary codewords to each bin, beyond the simpler mappings used previously. Here, the encoder applies two upper-range values M f and M a (See, for example, Figure 6C) or three upper range values M f , 32 or M a (see, for example, FIG. 6D) is used to assign codewords in each bin.
[0107] 6E illustrates two codeword assignment schemes based on BinVar(b) values, where plot 610 illustrates codeword assignment using two thresholds, while plot 620 illustrates codeword assignment using an inverse linear mapping, where the codeword assignment for a bin is inversely proportional to its BinVar(b) value. For example, in one embodiment, the following code may be applied to derive the number of codewords in a particular bin (bin_cw): alpha=(minCW-maxCW) / (maxVar-minVar); beta=(maxCW*maxVar-minCW*minVar) / (maxVar-minVar); bin_cw=round(alpha*bin_var+beta);, where minVar represents the minimum variance across all bins, maxVar represents the maximum variance across all bins, and minCW, maxCW represent the minimum and maximum number of codewords per bin as determined by the shaping model.
[0108] Refine chroma QP offset based on luma In Patent Document 6, to compensate for the interaction between luma and chroma, an additional chroma QP offset (denoted as chromaDQP or cQPO) and a luma-based chroma residual scaler (cScale) are defined, for example: chromaQP=QP_luma+chromaQPOffset+cQPO (28) where chromaQPOffset represents the chroma QP offset and QP_luma represents the luma QP for the coding unit.
number
number
[0109] The luma-derived QP value (denoted as qPi) and the final chroma QP value (Qp C Given the nonlinear relationship between the Qp and qPi (e.g., Table 8-10 in Non-Patent Document 4, "Qp as a function of qPi for ChromaArrayType equal to 1"), C In one embodiment, cScale may be further adjusted as follows:
[0110] For example, as shown in Tables 8-10 of Non-Patent Document 4, the mapping between adjusted luma and chroma QP values is denoted as f_QPi2QPc(). chromaQP_actual=f_QPi2QPc[chromaQP]= =f_QPi2QPc[QP_luma+chromaQPOffset+cQPO] (31) To scale the chroma residual, we need to calculate a scale based on the actual difference between the actual chroma coding QP both before applying the cQPO and after applying the cQPO: QPcBase=f_QPi2QPc[QP_luma+chromaQPOffset]; QPcFinal=f_QPi2QPc[QP_luma+chromaQPOffset+cQPO]; (32) cQPO_refine=QPcFinal-QpcBase; cScale=pow(2,-cQPO_refine / 6)
[0111] In another embodiment, the chromaQPOffset can also be absorbed into cScale. For example, QPcBase=f_QPi2QPc[QP_luma]; QPcFinal=f_QPi2QPc[QP_luma+chromaQPOffset+cQPO]; (33) cTotalQPO_refine=QPcFinal-QpcBase; cScale=pow(2,-cTotalQPO_refine / 6)
[0112] As an example, as described in U.S. Patent No. 6,273,999, in one embodiment: Let CSCALE_FP_PREC=16 represent the precision parameter. Forward scaling: after the chroma residual is generated, but before transformation and quantization: C_Res=C_orig-C_pred C_Res_scaled=C_Res*cScale+(1<<(CSCALE_FP_PREC-1)))>>CSCALE_FP_PREC Inverse scaling: after chroma dequantization and inverse transform, but before reconstruction: C_Res_inv=(C_Res_scaled< <CSCALE_FP_PREC) / cScale C_Reco=C_Pred+C_Res_inv;
[0113] In an alternative embodiment, the operation for in-loop chroma shaping may be expressed as follows: On the encoder side, for the residual (CxRes=CxOrg−CxPred) of a chroma component Cx (e.g., Cb or Cr) of each CU or TU,
number
number
number
number
[0114] The use of cScale is not limited to chroma residual scaling for in-loop shaping. The same method can be applied to out-of-loop shaping. In out-of-loop shaping, cScale may be used to scale chroma samples. The operation is the same as for the in-loop approach.
[0115] On the encoder side, when calculating the chroma RDOQ (either when using QP offset or chroma residual scaling), the lambda modifier for the chroma adjustment also needs to be calculated based on the refined offset: Modifier=pow(2,-cQPO_refine / 3); New_lambda=Old_lambda / Modifier (38)
[0116] As noted in equation (35), using cScale may require a division operation in the decoder. To simplify the decoder implementation, the same functionality may be implemented using a division operation in the encoder and applying a simpler multiplication operation in the decoder. For example, cScaleInv=(1 / cScale) For example, in the encoder cResScale=CxRes*cScale=CxRes / (1 / cScale)=CxRes / cScaleInv Then, in the decoder CxRes=cResScale / cScale=CxRes*(1 / cScale)=CxRes*cScaleInv Let's say.
[0117] In an embodiment, each luma-dependent chroma scaling factor may be calculated for the corresponding luma range in the piece-wise linear (PWL) representation, rather than for each luma codeword value. Thus, instead of a 1024-entry LUT (for 10-bit luma codewords) (e.g., cScale[Y]), the chroma scaling factors may be stored in a smaller LUT (e.g., 16 or 32 entries), e.g., cScaleInv[binIdx]. The encoder-side and decoder-side scaling operations may be implemented in fixed-point integer arithmetic as follows: c'=sign(c)*((abs(c)*s+2 CSCALE_FP_PREC-1 )>>CSCALE_FP_PREC) where c is the chroma residual, s is the chroma residual scaling factor from cScaleInv[binIdx], where binIdx is determined by the corresponding mean luma value, and CSCALE_FP_PREC is a constant value for precision.
[0118] In one embodiment, the forward shaping function may be represented using N equal segments (e.g., N=8, 16, 32, etc.), but the inverse representation will include non-linear segments. From an implementation perspective, it is desirable to have a representation of the inverse shaping function using equal segments, but forcing such a representation can cause a loss of coding efficiency. As a compromise, in one embodiment, the inverse shaping function may be constructed using a "mixed" PWL representation that combines both equal and unequal segments. For example, when using eight segments, the full range may first be divided into two equal segments, and then each of these may be subdivided into four unequal segments. Alternatively, the full range may be divided into four equal segments, and then each of these may be subdivided into two unequal segments. Alternatively, the full range may be divided into several unequal segments, and then each unequal segment may be divided into multiple equal segments. Alternatively, the full range may be divided into two equal segments, and then each equal segment may be divided into equal subsegments, with the segment lengths in each group of subsegments being unequal.
[0119] For example, without limitation, with 1024 codewords, one could have: a) four segments with 150 codewords each and two segments with 212 codewords each, or b) eight segments with 64 codewords each and four segments with 128 codewords each. The general purpose of such combinations of segments is to reduce the number of comparisons required to identify a PWL partition index given a code value, thereby simplifying hardware and software implementation.
[0120] In an embodiment, for more efficient implementation related to chroma residual scaling, the following variations may be allowed: Disable chroma residual scaling when separate luma / chroma trees are used. Disable chroma residual scaling for 2x2 chroma. · For intra- and inter-coded units, use the predicted signal instead of the reconstructed signal.
[0121] As an example, given the decoder (200D_D) shown in Figure 2E for processing the luma component, Figure 2F shows an exemplary architecture (200D_DC) for processing the corresponding chroma samples.
[0122] As shown in Figure 2F, compared to Figure 2E, the following changes are made when processing chroma: · Forward and reverse shaping blocks (282 and 265-3) are not used. There is a new chroma residual scaling block (288), which essentially replaces the inverse shaping block (265-3) for luma. The reconstruction block (285-C) is modified to handle the color residuals in the original domain as described in equation (36): CxRec=CxPred+CxRes.
[0123] From equation (34), at the decoder side, CxResScaled represents the extracted scaled chroma residual signal after inverse quantization and transformation (before block 288), CxRes=CxResScaled*C ScaleInv Let CxRec denote the rescaled chroma residual generated by the chroma residual scaling block (288) used by the reconstruction unit (285-C) to calculate CxRec = CxPred + CxRes, where CxPred is generated by the intra (284) or inter (280) prediction block.
[0124] The value used for the transform unit (TU) may be shared by the Cb and Cr components and can be calculated as follows: If in intra mode, calculate the average of the intra predicted luma values; If in inter mode, calculate the average of the forward shaped inter predicted luma values, i.e. the average luma value is calculated in the shaped domain. If combined merge and intra prediction, calculate the average of the combined predicted luma value. For example, the combined predicted luma value may be calculated according to section 8.4.6.6 of Appendix 2. In one embodiment, avgY' TU Based on C ScaleInv Alternatively, a piecewise linear (PWL) representation of the shaping function can be given to calculate the value avgY' TU We can find the index idx that belongs to the inverse mapping PWL. Next, C ScaleInv =cScaleInv[idx] An exemplary implementation applicable to the Versatile Video Coding codec (Non-Patent Document 8) currently under development by ITU and ISO can be found in Appendix 2 (see, eg, Section 8.5.5.1.2).
[0125] Disabling luma-based chroma residual scaling for intra slices using bipartite trees may result in some loss in coding efficiency. To improve the effect of chroma shaping, the following methods may be used: 1. The chroma scaling factor may be kept the same for the entire frame, depending on the mean or median of the luma sample values. This removes the TU-level dependency on luma for chroma residual scaling. 2. The chroma scaling factor can be derived using the reconstructed luma values from neighboring CTUs. 3. The encoder can derive the chroma scaling factor based on the source luma pixels and send it in the bitstream at the CU / CTU level (e.g., as an index into a piecewise representation of the shaping function). The decoder can then extract the chroma scaling factor from the shaping function without relying on the luma data. Scale factors for CTUs can be derived and transmitted only for intra-slices, but can also be used for inter-slices. The additional signaling cost is incurred only for intra-slices and therefore does not affect coding efficiency in random access. 4. Chroma can be shaped at the frame level as luma, and the luma shaping curve is derived from the luma shaping curve based on correlation analysis between luma and chroma, which completely eliminates chroma residual scaling.
[0126] apply delta_qp In AVC and HEVC, a parameter delta_qp is allowed to modify the QP value for a coded block. In an embodiment, a luma curve in a shaper can be used to derive the delta_qp value. A piecewise luma DQP value can be derived based on the codeword allocation for each bin. For example: For the kth bin, scale k =(M k / M a ); (39) lumaDQP k =INT(6*log2(scale k )) where INT() can be CEIL(), ROUND(), or FLOOR(). The encoder can use a function of luma, such as average(luma), min(luma), max(luma), etc., to find the luma value for the block, and then use the corresponding lumaDQP value for the block. From equation (27), to obtain rate-distortion benefits, we use weighted distortion in the mode decision: W rsp (k)=scale k 2 It can be set as follows.
[0127] Shaping and bin count considerations In typical 10-bit video coding, it is preferable to use at least 32 bins for the shaping mapping, but to simplify decoder implementation, some embodiments may use fewer bins, for example 16 or even 8 bins. Given that the encoder may already have used 32 bins to analyze the sequence and derive the distributional codewords, it is possible to reuse the original 32-bin codeword distribution and derive a 16-bin codeword by adding, within each 32-bin, two of the corresponding 16-bins. That is, For i=0 to 15 CWIn16Bin[i]=CWIn32Bin[2i]+CWIn32Bin[2i+1]
[0128] For the chroma residual scaling factor, we can simply divide the codeword by 2 and point it to a 32-bin chromaScalingFactorLUT. For example, In32Bin
[32] ={0 0 33 38 38 38 38 38 38 38 38 38 38 38 38 38 38 33 33 33 33 33 33 33 33 33 33 33 33 33 0 0} Given, the corresponding 16-bin CW assignment is CWIn16Bin
[16] ={ 0 71 76 76 76 76 76 76 71 66 66 66 66 66 66 0} This approach can be extended to handle even fewer bins, for example, 8. For i=0 to 7 CWIn8Bin[i]=CWIn16Bin[2i]+CWIn16Bin[2i+1]
[0129] When using a narrow range of valid codewords (e.g., [64,940] for a 10-bit signal, [64,235] for an 8-bit signal), care should be taken not to consider mapping to codewords where the first and last bins are reserved. For example, for a 10-bit signal, with 8 bins, each bin has 1024 / 8=128 codewords, and the first bin is [0,127], but since the standard codeword range is [64,940], the first bin should only consider codeword [64,127]. If the input video does not have the full range [0,2 bitdepth A special flag (e.g., video_full_range_flag=0) may be used to inform the decoder that the first and last bins have a narrower range than [[(1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55,
[0130] By way of example, and not limitation, Appendix 2 provides an exemplary syntax structure and associated syntax elements for supporting shaping in the ISO / ITU Video Versatile Codec (VVC) (Non-Patent Document 8) according to one embodiment using the architecture shown in Figures 2C, 2E, and 2F, where the forward shaping function includes 16 segments. References Each of the references cited herein is hereby incorporated by reference in its entirety. [Non-Patent Document 1] "Exploratory Test Model for HDR extension of HEVC", K. Minoo et al., MPEG output document, JCTVC-W0092 (m37732), 2016, San Diego, USA [Patent Document 2] PCT application PCT / US2016 / 025082, "In-Loop Block-Based Image Reshaping in High Dynamic Range Video Coding," G.M. Su, filed March 30, 2016, also published as WO2016 / 164235 [Patent Document 3] U.S. Patent Application No. 15 / 410,563, Content-Adaptive Reshaping for High Codeword Representation Images, T. Lu et al., filed January 19, 2017 [Non-patent document 4] ITU-T H.265, "High efficiency video coding," ITU, Dec. 2016 [Patent Document 5] PCT application PCT / US2016 / 042229, Signal Reshaping and Coding for HDR and Wide Color Gamut Signals, P. Yin et al., filed July 14, 2016, also published as WO2017 / 011636 [Patent Document 6] PCT Patent Application PCT / US2018 / 040287, Integrated Image Reshaping and Video Coding, T. Lu et al., filed June 29, 2018 [Patent Document 7] J. Froehlich et al., "Content-Adaptive Perceptual Quantizer for High Dynamic Range Images," U.S. Patent Application Publication No. 2018 / 0041759, February 8, 2018 [Non-patent document 8] B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 3)," JVET output document, JVET-L1001, v9, uploaded, Jan. 8, 2019
[0131] Exemplary Computer System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprised of electronic circuits and components, integrated circuits (ICs) such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLCs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or apparatuses including one or more of such systems, devices, or components. The computers and / or ICs may execute, control, or execute instructions related to image signal shaping and encoding as described herein. The computers and / or ICs may calculate any of a variety of parameters or values associated with the signal shaping and encoding processes described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0132] Certain implementations of the present invention include a computer processor executing software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., can implement methods related to image signal shaping and encoding, such as those described above, by executing software instructions in a program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any non-transitory, tangible medium bearing a set of computer-readable signals that, when executed by a data processor, cause the data processor to perform the methods of the present invention. A program product according to the present invention may be in any of a wide variety of non-transitory, tangible forms. Program products may include, for example, magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, electronic data storage media including flash RAM, and the like. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0133] Where a component (e.g., a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including reference to "means") should be interpreted to include any component that performs the function of the described component (e.g., is functionally equivalent) as an equivalent of that component, including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the invention.
[0134] Equivalents, Extensions, Substitutes and Others Thus, exemplary embodiments relating to efficient signal shaping and coding of images are described. In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indication of what is, and is intended by the applicant to be, the invention is the set of claims issued on this application, including any subsequent amendments thereto, in the specific form in which such claims are permitted. The definitions expressly set forth herein for terms contained in such claims govern the meaning of such terms as used in those claims. Accordingly, no limitations, elements, properties, features, advantages, or attributes not expressly recited in a claim should in any way limit the scope of such claims. Accordingly, the specification and drawings are to be regarded in an illustrative and not a restrictive sense.
[0135] Itemized Exemplary Embodiments The present invention may be embodied in any of the forms described herein, including, but not limited to, the following enumerated example embodiments (EEE) that describe the structure, features, and functions of some portions of the present invention. [EEE1] 1. A method for adaptively shaping a video sequence using a processor, the method comprising: accessing, by a processor, the input image in a first codeword representation; generating a forward shaping function that maps pixels of the input image to a second codeword representation, the second codeword representation allowing for more efficient compression than the first codeword representation, wherein generating the forward shaping function comprises: Dividing the input image into a plurality of pixel regions; assigning each pixel region to one of a plurality of codeword bins according to a first luminance characteristic of each pixel region; calculating a bin metric for each of the plurality of codeword bins according to a second luminance characteristic of each of pixel regions assigned to each codeword bin; assigning a number of codewords in the second codeword representation to each codeword bin according to a bin metric for each codeword bin and a rate-distortion optimization criterion; generating the forward shaping function in response to an allocation of codewords in a second codeword representation to each of the plurality of codeword bins. method. [EEE2] The method of EEE1, wherein the first luminance characteristic of the pixel region comprises an average luminance pixel value within the pixel region. [EEE3] The method of EEE1, wherein the second luminance characteristic of the pixel region comprises a variance of luminance pixel values of the pixel region. [EEE4] The method of EEE3, wherein calculating a bin metric for a codeword bin includes calculating an average of the variance of luminance pixel values for all pixel regions assigned to that codeword bin. [EEE5] Assigning numbers of codewords in said second codeword representation to codeword bins according to their bin metrics: If no pixel region is assigned to that codeword bin, then do not assign a codeword to that codeword bin; assign the first number of codewords if the bin metric of that codeword bin is lower than the upper threshold; otherwise, assigning a second number of codewords to the codeword bin. Method described in EEE1. [EEE6] For a first codeword representation having a depth of B bits and a second codeword representation having a depth of Bo bits, and N codeword bins, the first number of codewords is M f =CEIL((2 Bo / (CW2-CW1))*M a ), and said second number of codewords comprises M a =2 B / N, where CW1 <CW2は、[0 2 B IEEE 802.11b, which represents two codewords in [IEEE 802.11b]. [EEE7] CW1=16×2 (B-8) and CW2 = 235 × 2 (B-8) The method of EEE6, wherein [EEE8] Determining the upper threshold comprises: Define a set of potential thresholds; For each threshold in the set of thresholds: Generate a forward shaping function based on the threshold; encoding and decoding a set of input test frames according to the shaping function and a bit rate R to generate an output set of decoded test frames; calculating an overall rate-distortion optimization (RDO) metric based on the input test frame and the decoded test frame; selecting as the upper threshold a threshold in the set of potential thresholds at which the RDO metric is minimum. Method described in EEE5. [EEE9] Calculating the RDO metric includes: calculating J=D+λR, where D represents a measure of distortion between pixel values of the input test frame and corresponding pixel values in the decoded test frame, and λ represents a Lagrange multiplier; Method described in EEE8. [EEE10] The method of EEE9, wherein D is a measure of the sum of squared differences between corresponding pixel values of the input test frame and the decoded test frame. [EEE11] The method of EEE1, wherein assigning a number of codewords in the second codeword representation to codeword bins according to their bin metrics is based on a codeword assignment lookup table, the codeword assignment lookup table defining two or more thresholds that divide a range of bin metric values into segments, and providing a number of codewords assigned to bins having a bin metric within each segment. [EEE12] The method according to EEE11, wherein, given a default codeword assignment to bins, bins with large bin metrics are assigned fewer codewords than the default codeword assignment, and bins with small bin metrics are assigned more codewords than the default codeword assignment. [EEE13] For a first codeword representation using B bits and N bins, the default codeword allocation per bin is M a =2 B The method according to EEE12 is given by / N. [EEE14] further comprising generating formatting information in response to the forward formatting function, the formatting information comprising: A flag indicating the minimum codeword bin index value used in the shaped reconstruction process; a flag indicating the maximum codeword bin index value used in the shaping construction process; Flags indicating the formatting model profile types, each of which is associated with default bin-related parameters; or One or more delta values used to adjust the default bin-related parameters including one or more of Method described in EEE1. [EEE15] further comprising assigning a bin importance value to each codeword bin, the bin importance value being: 0 if no codeword is assigned to that codeword bin; 2 if the codeword is assigned the first value of the codeword; otherwise it is 1. Method described in EEE5. [EEE16] Determining the upper threshold comprises: dividing a luminance range of pixel values in an input image into bins; determining, for each bin, a bin histogram value and an average bin variance value, wherein for a bin, the bin histogram value comprises the number of pixels in that bin out of the total number of pixels in the image, and the average bin variance value gives a metric of the average pixel variance of the pixels in that bin; sorting the average bin variance values to generate a sorted list of average bin variance values and a sorted list of average bin variance-value indices; calculating a cumulative density function as a function of the sorted average bin variance values based on the bin histogram values and the sorted list of average bin variance-value indices; determining an upper threshold based on a criterion satisfied by the value of the cumulative density function. Method described in EEE5. [EEE17] The cumulative density function may be calculated by: BinVarSortDsdCDF[0]=BinHist[BinIdxSortDsd[0]]; for(int b=1; b <PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b]=BinVarSortDsdCDF[b-1]+BinHist[BinIdxSortDsd[b]];} where b is the bin number, PIC_ANALYZE_CW_BINS is the total number of bins, BinVarSortDsdCDF[b] is the output of the CDF function for bin b, BinHist[i] is the bin histogram value for bin i, and BinIdxSortDsd[] represents the sorted list of average bin variance value indices. The method described in EEE16. [EEE18] The method of EEE16, wherein the upper threshold is determined as the average bin variance value for which k% of the CDF output is greater than or equal to the upper threshold, under the criterion that the average bin variance for k% of the pixels in the input image is greater than or equal to the upper threshold. [EEE19] The method according to EEE18, wherein k=50. [EEE20] 1. A method of reconstructing a shaping function in a decoder, the method comprising: receiving an encoded bitstream syntax element that characterizes a formatting model, the syntax element comprising: A flag indicating the minimum codeword bin index value used in the shaping construction process, A flag indicating the maximum codeword bin index value used in the shaping construction process, a flag indicating a well-formed model profile type, the model profile type being associated with default bin-related parameters, including bin importance values; or A flag indicating one or more delta bin importance values to be used to adjust the default bin importance values defined in the fairness model profile. and a step including one or more of: determining a default bin importance value for each bin based on the shaping model profile and an assignment list of default numbers of codewords to be assigned to each bin according to the bin importance value; For each codeword bin: Determine the bin importance value by adding the default bin importance value to the delta bin importance value; determining a number of codewords to be assigned to the codeword bin based on the bin importance value of the bin and the assignment list; generating a forward shaping function based on the number of codewords assigned to each codeword bin; method. [EEE21] The number of codewords M assigned to the kth codeword bin using the assignment list. k You can further determine: For the kth bin: If bin_importance[k]==0, M k =0; Otherwise, if bin_importance[k]==2, M k =M f year, If not, M k =M a and where M a and M f is an element of the assignment list, and bin_importance[k] represents the bin importance value of the kth bin. The method described in EEE20. [EEE522] 1. A method of reconstructing encoded data in a decoder having one or more processors, the method comprising: receiving an encoded bitstream (122) including one or more encoded shaped images in a first codeword representation and metadata (207) relating to shaping information for the encoded shaped images; generating (250) an inverse shaping function based on metadata related to the shaping information, the inverse shaping function mapping pixels of the shaped image from the first code word representation to a second code word representation; (250) generating a forward shaping function based on metadata related to the shaping information, the forward shaping function mapping pixels of an image from the second code word representation to the first code word representation; extracting from the coded bitstream a coded shaped image comprising one or more coded units, the step comprising, for one or more coded units in the coded shaped image: For each intra-coded coding unit (CU) in the coded shaped image: generating a first shaped reconstructed sample (227) for the CU based on the shaped residual and the first shaped predicted sample for the CU; generating (270) a shaped loop filter output based on the first shaped reconstructed samples and loop filter parameters; applying the inverse shaping function to the shaped loop filter output to generate decoded samples of the coding unit in the second codeword representation (265); storing decoded samples of the coding unit in the second codeword representation in a reference buffer; For inter-coded coding units within the coded formatted image: applying the forward shaping function to the predicted samples stored in the reference buffer in the second codeword representation to generate second shaped predicted samples; generating second shaped reconstructed samples of the coding unit based on the shaped residuals in the coded CU and the second shaped predicted samples; generating a shaped loop filter output based on the second shaped reconstructed samples and loop filter parameters; applying the inverse shaping function to the shaped loop filter output to generate samples of the coding unit in a second codeword representation; storing samples of the coding unit in the second codeword representation in a reference buffer; generating a decoded image based on the stored samples in the reference buffer; method. [EEE23] 23. An apparatus comprising a processor and configured to perform a method according to any one of EEE1 to 22. [EEE24] A non-transitory computer-readable storage medium having stored thereon computer-executable instructions for performing a method on one or more processors according to any of EEE1 to EEE22.
[0136] Appendix 1 Example implementation of bubble sort void bubbleSortDsd(double* array, int*idx, int n) { int i,j; bool swapped; for(i=0; i < n-1; i++) { swapped=false; for(j=0; j <n-i-1; j++) { if(array[j] <array[j+1]) { swap(&array[j],&array[j+1]); swap(&idx[j],&idx[j+1]); swapped=true; } } if(swapped==false) break; } }
[0137] Appendix 2 As an example, this appendix provides an exemplary syntax structure and associated syntax elements according to an embodiment that supports formatting in the Versatile Video Codec (VVC) (non-patent document 8), currently under joint development by ISO and ITU. Syntax elements that are new in existing draft versions are highlighted or explicitly noted. Equation numbers such as (8-xxx) indicate placeholders that will be updated as necessary in the final specification.
[0138] 7.3.2.1 Sequence Parameter Set RBSP Syntax [Table 5-1] [Table 5-2] [Table 5-3]
[0139] In 7.3.3.1 General Tile Group Header Syntax [Table 6-1] [Table 6-2] [Table 6-3]
[0140] Added new syntax tables and tile formatter models. [Table 7]
[0141] In General Sequence Parameter Set RBSP Semantics, the following semantics have been added: sps_reshaper_enabled_flag equal to 1 specifies that the reshaper is used in the coded video sequence (CVS). sps_reshaper_enabled_flag equal to 0 specifies that the reshaper is not used in the CVS.
[0142] Added the following semantics to the tile group header syntax: tile_group_reshaper_model_present_flag equal to 1 specifies that tile_group_reshaper_model() is present in the tile group header. tile_group_reshaper_model_present_flag equal to 0 specifies that tile_group_reshaper_model() is not present in the tile group header. If tile_group_reshaper_model_present_flag is not present, it is inferred to be equal to 0. tile_group_reshaper_enabled_flag equal to 1 specifies that the shaper is enabled for the current tile group. tile_group_reshaper_enabled_flag equal to 0 specifies that the shaper is not enabled for the current tile group. If tile_group_resharper_enable_flag is not present, it is inferred to be equal to 0. tile_group_reshaper_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current tile group. tile_group_reshaper_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current tile group. If tile_group_reshaper_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0143] Added tile_group_reshaper_model() syntax reshaper_model_min_bin_idx specifies the minimum bin (or piece) index used in the shaper construction process. The value of reshaper_model_min_bin_idx shall be in the range 0 to MaxBinIdx, inclusive. The value of MaxBinIdx shall be equal to 15. reshaper_model_delta_max_bin_idx specifies the value obtained by subtracting the maximum bin (or piece) index MaxBinIdx from the maximum bin index used in the reshaper configuration process. The value of reshaper_model_max_bin_idx is set equal to MaxBinIdx - reshaper_model_delta_max_bin_idx. One plus reshaper_model_bin_delta_abs_cw_prec_minus1 specifies the number of bits used for the representation of the syntax reshaper_model_bin_delta_abs_CW[i]. reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshaper_model_bin_delta_abs_CW[ i ] as follows. · When reshaper_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. · Otherwise (when reshaper_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is a negative value. If reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. The variable RspDeltaCW[i] = (1 - 2 * reshaper_model_bin_delta_sign_CW[i]) * reshaper_model_bin_delta_abs_CW[i]; The variable RspCW[i] is derived in the following steps: The variable OrgCW is set equal to (1 << BitDepthY) / (MaxBinIdx + 1). If reshaper_model_min_bin_idx<=i<=reshaper_model_max_bin_idx RspCW[i]=OrgCW+RspDeltaCW[i] Otherwise, RspCW[i]=0 The value of RspCW[i] ranges from 32 to 2*OrgCW-1 when the value of BitDepthYis is equal to 10. The variable InputPivot[i], for i ranging from 0 to MaxBinIdx+1 inclusive, is derived as follows: InputPivot[i]=i*OrgCW The variables ReshapePivot[i], for i in the range 0 to MaxBinIdx+1, inclusive, and the variables ScaleCoef[i] and InvScaleCoeff[i], for i in the range 0 to MaxBinIdx, inclusive, are derived as follows: shiftY=14 ReshapePivot[0]=0; for(i=0; i<=MaxBinIdx; i++){ ReshapePivot[i+1]=ReshapePivot[i]+RspCW[i] ScaleCoef[i]=(RspCW[i]*(1< <shiftY)+(1<<(Log2(OrgCW)-1)))> >(Log2(OrgCW)) if(RspCW[i]==0) InvScaleCoeff[i]=0 else InvScaleCoeff[i]=OrgCW*(1< <shiftY) / RspCW[i] } The variable ChromaScaleCoef[i], for i ranging from 0 to MaxBinIdx inclusive, is derived as follows: ChromaResidualScaleLut
[64] ={16384,16384,16384,16384,16384,16384,16384,8192,8192,8192,8192 ,5461,5461,5461,5461,4096,4096,4096,4096,3277,3277,3277,3277,2731,2731,2731,2731,2341,234 1,2341,2048,2048,2048,1820,1820,1820,1638,1638,1638,1638,1489,1489,1489,1489,1365,1365,1365,1365,1260,1260,1260,1260,1170,1170,1170,1170,1092,1092,1092,1024,1024,1024,1024}; shiftC=11 If (RspCW[i]==0) ChromaScaleCoef[i]=(1< <shiftC) Otherwise (RspCW[i]!=0) ChromaScaleCoef[i]=ChromaResidualScaleLut[Clip3(1,64,RspCW[i]>>1)-1] Note: An alternative implementation may unify luma and chroma scaling, making ChromaResidualScaleLut[] unnecessary. In that case, chroma scaling may be implemented as follows: shiftC=11 If (RspCW[i]==0) ChromaScaleCoef[i]=(1< <shiftC) Otherwise (RspCW[i] != 0), the following applies: BinCW=BitDepth Y >10?(RspCW[i]>>(BitDepth Y -10)):BitDepth Y <10?(RspCW[i]<<(10 BitDepth Y )):RspCW[i]; ChromaScaleCoef[i]=OrgCW*(1< <shiftC) / BinCW[i]
[0144] In the weighted sample prediction process for combined merge and intra prediction, add the following. Additions are highlighted. 8.4.6.6 Weighted Sample Prediction Process for Combined Merge and Intra Prediction The inputs to this process are: Current coding block width cbWidth Current coding block height cbHeight Two (cbWidth) x (cbHeight) arrays, preSamplesInter and preSamplesIntra Intra prediction mode predModeIntra · Variable cIdx that specifies the color component index The output of this process is a (cbWidth) x (cbHeight) array of predicted sample values, predSamplesComb. The variable bitDepth is derived as follows: If cIdx is equal to 0, bitDepth is BitDepth Y is set equal to Otherwise, bitDepth is BitDepth C is set equal to The predicted samples predSamplesComb[x][y] for x=0..cbWidth-1 and y=0..cbHeight-1 are derived as follows: The weight w is derived as follows: If predModeIntra is INTRA_ANGULAR50, w is specified in Table 8-10 with nPos equal to y and nSize equal to cbHeight. Otherwise, if predModeIntra is INTRA_ANGULAR18, w is specified in Table 8-10 with nPos equal to x and nSize equal to cbWidth. Otherwise, w is set to 4. [Table 8] The predicted samples predSamplesComb[x][y] are derived as follows: predSamplesComb[x][y]=(w*predSamplesIntra[x][y]+ (8-w)*predSamplesInter[x][y])>>3) (8-740) Table 8-10. Specification of w as a function of position nP and size nS [Table 9]
[0145] Added the following to the picture reconstruction process: 8.5.5 Picture Composition Process The inputs to this process are: Position (xCurr, yCurr) that specifies the top-left sample of the current block relative to the top-left sample of the current picture component The variables nCurrSw and nCurrSh specify the width and height of the current block, respectively. · Variable cIdx that specifies the color component of the current block A (nCurrSw) x (nCurrSh) array predSamples that specifies the predicted samples of the current block. A (nCurrSw) x (nCurrSh) array resamples that specifies the residual samples of the current block Depending on the value of the color component cIdx, the following assignments are made: If cIdx is equal to 0, recSamples is the reconstructed picture sample array S L The function clipCidx1 corresponds to Clip1 YCorresponds to. Otherwise, if cIdx is equal to 1, recSamples is the reconstructed chroma sample array S Cb The function clipCidx1 corresponds to Clip1 C Corresponds to. Otherwise (cIdx is equal to 2), recSamples is the reconstructed chroma sample array S Cr The function clipCidx1 corresponds to Clip1 C Corresponds to. [Table 10] Otherwise, the (nCurrSw) × (nCurrrSh) block of the reconstructed sample array recSamples at position (xCurr, yCurr) is derived as follows: For i=0..nCurrSw-1, j=0..nCurrSh-1, recSamples[xCurr+i][yCurr+j]=clipCidx1(predSamples[i][j]+resSamples[i][j]) [Outside 1] TIFF2025134977000027.tif7115
[0146] [Outside 2] TIFF2025134977000028.tif10115This section specifies picture reconstruction using the mapping process. Picture reconstruction using the mapping process for luma sample values is specified in 8.5.5.1.1. Picture reconstruction using the mapping process for chroma sample values is specified in 8.5.5.1.2. 8.5.5.1 Picture Reconstruction Using a Mapping Process for Luma Sample Values The inputs to this process are: A (nCurrSw) × (nCurrSh) array predSamples that specifies the luma prediction samples for the current block. A (nCurrSw) × (nCurrSh) array resamples that specifies the luma residual samples of the current block The output for this process is: (nCurrSw) × (nCurrSh) mapped luma prediction sample array predMapSamples (nCurrSw) × (nCurrSh) reconstructed luma sample array recSamples predMapSamples is derived as follows: If(CuPredMode[xCurr][yCurr]==MODE_INTRA)||(CuPredMode[xCurr][yCurr]==MODE_INTER && mh_intra_flag[xCurr][yCurr]) predMapSamples[xCurr+i][yCurr+j]=predSamples[i][j] where i=0..nCurrSw-1, j=0..nCurrSh-1 [Outside 3] TIFF2025134977000029.tif9115Otherwise((CuPredMode[xCurr][yCurr]==MODE_INTER && !mh_intra_flag[xCurr][yCurr])), the following applies: shiftY=14 idxY=predSamples[i][j]>>Log2(OrgCW) predMapSamples[xCurr+i][yCurr+j]=ReshapePivot[idxY] +(ScaleCoeff[idxY]*(predSamples[i][j]-InputPivot[idxY]) +(1<<(shiftY-1)))>>shiftY where i=0..nCurrSw-1, j=0..nCurrSh-1 [Outside 4] TIFF2025134977000030.tif9115recSamples is derived as follows: recSamples[xCurr+i][yCurr+j]=Clip1 Y (predMapSamples[xCurr+i][yCurr+j]+resSamples[i][j]]) where i=0..nCurrSw-1, j=0..nCurrSh-1 [Outside 5] TIFF2025134977000031.tif9115 8.5.5.1.2 Picture reconstruction using a mapping process for chroma sample values The inputs to this process are: A (nCurrSwx2) x (nCurrShx2) mapped array predMapSamples that specifies the mapped luma prediction samples of the current block. PredSamples, an array of (nCurrSw) x (nCurrSh) that specifies the chroma prediction samples for the current block resamples, an (nCurrSw) x (nCurrSh) array specifying the chroma residual samples of the current block The output of this process is the reconstructed chroma sample array recSamples. recSamples is derived as follows: ·If(!tile_group_reshaper_chroma_residual_scale_flag||((nCurrSw)x(nCurrSh)<=4)) recSamples[xCurr+i][yCurr+j]=Clip1 C (predSamples[i][j]+resSamples[i][j]) where i=0..nCurrSw-1, j=0..nCurrSh-1 [Outside 6] TIFF2025134977000032.tif9115·else (tile_group_reshaper_chroma_residual_scale_flag &&((nCurrSw)x(nCurrSh)>4)), then the following applies: The variable varScale is derived as follows: 1.invAvgLuma=Clip1 Y ((Σ i Σ j predMapSamples[(xCurr<<1)+i][(yCurr<<1)+j] +nCurrSw*nCurrSh*2) / (nCurrSw*nCurrSh*4)) 2. The variable idxYInv is calculated using the sample values invAvgLuma. [Outside 7] It is derived by invoking the piecewise function index identification as specified in Section TIFF2025134977000033.tif8115. 3.varScale=ChromaScaleCoef[idxYInv] varScale=ChromaScaleCoef[idxYInv] recSamples is derived as follows: If tu_cbf_cIdx[xCurr][yCurr] is equal to 1, the following applies: shiftC=11 recSamples[xCurr+i][yCurr+j]=ClipCidx1(predSamples[i][j]+Sign(resSamples[i][j]) *((Abs(resSamples[i][j])*varScale+(1<<(shiftC-1)))>>shiftC)) where i=0..nCurrSw-1, j=0..nCurrSh-1 [Outside 8] TIFF2025134977000034.tif9115 Otherwise (tu_cbf_cIdx[xCurr][yCurr] equals 0), recSamples[xCurr+i][yCurr+j]=ClipCidx1(predSamples[i][j]) [Outside 9] TIFF2025134977000035.tif9115
[0147] 8.5.6 Picture Reverse Mapping Process This section is relevant when the value of tile_group_reshaper_enabled_flag is equal to 1. The input is the reconstructed picture luma sample array S L and the output is the modified reconstructed picture luma sample array S' after the inverse mapping process. L The inverse mapping process for luma sample values is specified in 8.4.6.1. 8.5.6.1 Picture Inverse Mapping Process for Luma Sample Values The input to this process is the luma position (xP, yP) that specifies the luma sample location relative to the top-left luma sample of the current picture. The output of this process is the inverse-mapped luma sample value invLumaSample. The value of invLumaSample is derived by applying the following ordered steps: 1. The variable idxYInv is the luma sample value S. L It is derived by invoking the piecewise function index identification as specified in Section 8.5.6.2 with inputs [xP][yP]. 2. The value of reshapeLumaSample is derived as follows: shiftY=14 invLumaSample=InputPivot[idxYInv]+(InvScaleCoeff[idxYInv]*(S L [xP][yP]-ReshapePivot[idxYInv]) +(1<<(shiftY-1)))>>shiftY [Outside 10] TIFF2025134977000036.tif91153.clipRange=((reshaper_model_min_bin_idx>0) && (reshaper_model_max_bin_idx <MaxBinIdx)); If clipRange is equal to 1, the following applies: minVal=16<<(BitDepth Y -8) maxVal=235<<(BitDepth Y -8) invLumaSample=Clip3(minVal,maxVal,invLumaSample) Otherwise (clipRange equals 0), the following applies: invLumaSample=ClipCidx1(invLumaSample) 8.5.6.2 Identifying Piecewise Function Indices for the Luma Component The input to this process is the luma sample value S. The output of this process is an index idxS that identifies the piece to which sample S belongs. The variable idxS is derived as follows: for(idxS=0, idxFound=0; idxS<=MaxBinIdx; idxS++){ if((S <ReshapePivot[idxS+1]){ idxFound=1 break } } Note, an alternative implementation for finding the identification idxS is as follows: if(S <ReshapePivot[reshaper_model_min_bin_idx]) idxS=0 else if(S>=ReshapePivot[reshaper_model_max_bin_idx]) idxS=MaxBinIdx else idxS=findIdx(S,0,MaxBinIdx+1,ReshapePivot[]) function idx=findIdx(val,low,high,pivot[]){ if(high-low<=1) idx=low else { mid=(low+high)>>1 if(val <pivot[mid]) high=mid else low=mid idx=findIdx(val,low,high,pivot[]) } }
[0148] Several aspects will be described. [Aspect 1] 1. A method for reconstructing encoded video data with one or more processors, the method comprising: receiving an encoded bitstream (122) including one or more encoded shaped images in input codeword representation; receiving shaping metadata (207) for the one or more encoded shaping images in the encoded bitstream; generating a forward shaping function (282) based on the shaping metadata, the forward shaping function mapping pixels of an image from a first codeword representation to the input codeword representation; generating an inverse shaping function (265-3) based on the shaping metadata or the forward shaping function, the inverse shaping function mapping pixels of the shaped image from the input codeword representation to the first codeword representation; extracting from the coded bitstream a coded shaped image comprising one or more coded units; For an intra-coded coding unit (intra CU) in the coded shaped image: generating shaped reconstructed samples of the intra CU based on the shaped residual within the intra CU and the intra predicted shaped prediction samples (285); applying the inverse shaping function (265-3) to the shaped reconstructed samples of the intra CU to generate decoded samples of the intra CU in the first codeword representation; applying a loop filter (270) to the decoded samples of the intra CU to generate output samples of the intra CU; storing the output samples of the intra CU in a reference buffer; For inter-coded CUs (inter CUs) in the coded shaped image: applying the forward shaping function (282) to the inter-predicted samples stored in the reference buffer in the first codeword representation to generate shaped predicted samples for the inter-CU in the input codeword representation; generating shaped reconstructed samples for the inter CU based on a shaped residual for the inter CU and the shaped prediction samples for the inter CU; applying the inverse shaping function (265-3) to the shaped reconstructed samples of the inter-CU to generate decoded samples of the inter-CU in the first codeword representation; applying a loop filter (270) to the decoded samples of the inter-CU to generate output samples of the inter-CU; storing the output samples of the inter-CU in the reference buffer; generating a decoded image in the first codeword representation based on the output samples in the reference buffer; method. [Aspect 2] generating shaped reconstructed samples (RecSamples) of the intra CU; RecSample=(Res+IpredSample) Calculating where Res represents shaped residual samples in the intra CU in the input codeword representation, and IpredSample represents intra predicted shaped prediction samples in the input codeword representation. The method of embodiment 1. Aspect 3 generating shaped reconstructed samples (RecSamples) for the inter-CU; RecSample=(Res+Fwd(PredSample)) Calculating where Res represents the shaped residual for the inter CU in the input codeword representation, Fwd() represents the forward shaping function, and PredSample represents the inter predicted sample in the first codeword representation. The method of embodiment 1. Aspect 4 Generating output samples (RecSampleInDPB) to be stored in the reference buffer includes: RecSampleInDPB=LPF(Inv(RecSample)) Calculating where Inv() represents the inverse shaping function and LPF() represents the loop filter. The method of embodiment 2 or 3. Aspect 5 For chroma residual samples in the inter-coded CU (inter CU) in the input codeword representation, further: determining a chroma scaling factor based on luma pixel values in the input codeword representation and the shaping metadata; multiplying the chroma residual samples in the inter CU by the chroma scaling factor to generate scaled chroma residual samples in the inter CU in the first codeword representation; generating reconstructed chroma samples of the inter CU based on the scaled chroma residual of the inter CU and the chroma inter predicted samples stored in the reference buffer, thereby generating decoded chroma samples of the inter CU; applying the loop filter (270) to the decoded chroma samples of the inter-CU to generate output chroma samples of the inter-CU; and storing the output chroma samples of the inter-CU in the reference buffer. The method of embodiment 1. Aspect 6 6. The method of embodiment 5, wherein in an intra mode, the chroma scaling factor is based on an average of intra-predicted luma values. Aspect 7 6. The method of aspect 5, wherein in an inter mode, the chroma scaling factor is based on an average of inter-predicted luma values in the input codeword representation. Aspect 8 The formatting metadata: a first parameter indicating a number of bins used to represent the first codeword representation; A second parameter indicating the minimum bin index to be used in shaping; a first set of parameters indicating absolute delta codeword values for each bin in the input codeword representation; a second set of parameters indicating the sign of the delta codeword value for each bin in the input codeword representation; The method of embodiment 1. Aspect 9 2. The method of claim 1, wherein the forward shaping function is reconstructed as a piecewise linear function with linear segments derived by the shaping metadata. Aspect 10 1. A method for adaptively shaping a video sequence by a processor, the method comprising: accessing, by a processor, the input image in a first codeword representation; generating a forward shaping function that maps pixels of the input image to a second codeword representation, the second codeword representation allowing for more efficient compression than the first codeword representation, wherein generating the forward shaping function comprises: Dividing the input image into a plurality of pixel regions; assigning each pixel region to one of a plurality of codeword bins according to a first luminance characteristic of each pixel region; calculating a bin metric for each of the plurality of codeword bins according to a second luminance characteristic of each of the pixel regions assigned to each of the plurality of codeword bins; assigning a number of codewords in the second codeword representation to each of the plurality of codeword bins according to a bin metric and a rate-distortion optimization criterion for each of the plurality of codeword bins; generating the forward shaping function in response to an allocation of codewords in the second codeword representation to each of the plurality of codeword bins. method. Aspect 11 11. The method of embodiment 10, wherein the first luminance characteristic of a pixel region comprises an average luminance pixel value within the pixel region. Aspect 12 11. The method of embodiment 10, wherein the second luminance characteristic of a pixel region comprises a variance of luminance pixel values of the pixel region. Aspect 13 13. The method of embodiment 12, wherein calculating a bin metric for a codeword bin includes calculating an average of the variance of luminance pixel values for all pixel regions assigned to that codeword bin. Aspect 14 Assigning numbers of codewords in said second codeword representation to codeword bins according to their bin metrics: If no pixel region is assigned to that codeword bin, then do not assign a codeword to that codeword bin; assign the first number of codewords if the bin metric of that codeword bin is lower than the upper threshold; otherwise, assigning a second number of codewords to the codeword bin. The method of embodiment 10. Aspect 15 15. An apparatus having a processor and configured to perform the method of any one of aspects 1 to 14. Aspect 16 A non-transitory computer-readable storage medium storing computer-executable instructions for performing, by one or more processors, the method of any one of aspects 1 to 14.
Claims
1. 1. A method for reconstructing encoded video data with one or more processors, the method comprising: receiving an encoded bitstream including one or more encoded shaped images in shaped codeword representation; and receiving shaping parameters for the one or more coded shaped images in the coded bitstream, the shaping parameters including parameters for generating a forward shaping function based on the shaping parameters, the forward shaping function mapping image pixels from an input codeword representation to the shaped codeword representation, the shaping parameters comprising: a delta index parameter that determines an active maximum bin index to be used for shaping, the active maximum bin index being smaller than a predefined maximum bin index, the delta index parameter representing the difference between the active maximum bin index and the predefined maximum bin index; a minimum index parameter indicating the minimum bin index used in said shaping; an absolute delta codeword value for each active bin in the shaped codeword representation; a sign of the absolute delta codeword value for each active bin in the shaped codeword representation; The method further comprises: decoding the encoded bitstream based on the forward shaping function. method.
2. The method of claim 1 , wherein the forward shaping function is reconstructed as a piecewise linear function with linear segments driven by the shaping parameters.
3. 2. The method of claim 1, wherein determining the active maximum bin index used to represent the input codeword representation comprises calculating a difference between the predefined maximum bin index and the delta index parameter.
4. The method of claim 1 , wherein the predefined maximum bin index is one of 15, 31, or 63.
5. 1. A method for generating formatting parameters for an encoded bitstream, the method comprising: receiving a sequence of video pictures in input codeword representation; applying a forward shaping function to one or more pictures in the sequence of video pictures to generate a shaped picture in a shaped codeword representation, the forward shaping function mapping image pixels from the input codeword representation to the shaped codeword representation; generating shaping parameters for the shaped codeword representation; and generating an encoded bitstream based on at least the shaped picture, wherein the shaping parameters include: a delta index parameter that determines an active maximum bin index to be used for shaping, the active maximum bin index being smaller than a predefined maximum bin index, the delta index parameter representing the difference between the active maximum bin index and the predefined maximum bin index; a minimum index parameter indicating the minimum bin index used in said shaping; an absolute delta codeword value for each active bin in the shaped codeword representation; and a sign of the absolute delta codeword value for each active bin in the shaped codeword representation. method.
6. The method of claim 5 , wherein the forward shaping function comprises a piecewise linear function having linear segments driven by the shaping parameters.
7. 6. The method of claim 5, wherein determining the active maximum bin index used to represent the input codeword representation comprises calculating a difference between the predefined maximum bin index and the delta index parameter.
8. The method of claim 5 , wherein the predefined maximum bin index is one of 15, 31, or 63.
9. one or more non-transitory computer-readable media storing instructions for generating a bitstream and recording the bitstream on a recording medium, the instructions, when executed by the processor, causing the processor to perform a video encoding method to generate the bitstream; The encoding method is: receiving a sequence of video pictures in input codeword representation; applying a forward shaping function to one or more pictures in the sequence of video pictures to generate a shaped picture in a shaped codeword representation, the forward shaping function mapping image pixels from the input codeword representation to the shaped codeword representation; generating shaping parameters for the shaped codeword representation; and generating an encoded bitstream based on at least the shaped picture, wherein the shaping parameters include: a delta index parameter that determines an active maximum bin index to be used for shaping, the active maximum bin index being smaller than a predefined maximum bin index, the delta index parameter representing the difference between the active maximum bin index and the predefined maximum bin index; a minimum index parameter indicating the minimum bin index used in said shaping; an absolute delta codeword value for each active bin used in said shaping; and a sign of the absolute delta codeword value for each active bin used in the shaping. method.
Citation Information
Patent Citations
System for reshaping and coding high dynamic range and wide color gamut sequences
US20170085880A1
Content-adaptive perceptual quantizer for high dynamic range images
WO2016140954A1
In-loop block-based image reshaping in high dynamic range video coding
WO2016164235A1
Signal reshaping and coding for HDR and wide color gamut signals
WO2017011636A1
High dynamic range video coding architectures with multiple operating modes
WO2017019818A1