Image shaping in video coding using rate distortion optimization

The RDO-based signal shaping method optimizes codeword assignment in video encoding, addressing inefficiencies in higher bit depth encoding by improving both visual quality and compression metrics.

JP7706626B2Active Publication Date: 2025-07-11DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024189444
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-14
Filing Date
2024-10-29
Publication Date
2025-07-11
Estimated Expiration
2039-02-13

Smart Images

  • Figure 0007706626000036
    Figure 0007706626000036
  • Figure 0007706626000037
    Figure 0007706626000037
  • Figure 0007706626000038
    Figure 0007706626000038
Patent Text Reader

Abstract

To provide image reshaping in video coding using rate distortion optimization.SOLUTION: Given a sequence of images in a first codeword representation, a method, a process, and a system are presented for image reshaping using rate distortion optimization. Reshaping allows the images to be coded in a second codeword representation which allows more efficient compression than using a first codeword representation. A syntax method for signaling reshaping parameters is also presented.SELECTED DRAWING: Figure 2E
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 792,122, filed on January 14, 2019; U.S. Provisional Patent Application No. 62 / 782,659, filed on December 20, 2018; U.S. Provisional Patent Application No. 62 / 772,228, filed on November 28, 2018; U.S. Provisional Patent Application No. 62 / 739,402, filed on October 1, 2018; U.S. Provisional Patent Application No. 62 / 726,608, filed on September 4, 2018; U.S. Provisional Patent Application No. 62 / 691,366, filed on June 28, 2018; and U.S. Provisional Patent Application No. 62 / 630,385, filed on February 14, 2018, each of which is hereby incorporated by reference in its entirety.

[0002] Technique The present invention generally relates to image and video encoding. More particularly, certain embodiments of the present invention relate to image shaping in video encoding.

Background Art

[0003] In 2013, the MPEG group of the International Organization for Standardization (ISO), in conjunction with the International Telecommunication Union (ITU), published the first draft of the High Efficiency Video Coding (HEVC, also known as H.265) video coding standard (Non - Patent Document 4). More recently, the group published a call for evidence to support the development of a next - generation coding standard that provides improved coding performance over existing video coding technologies.

[0004] As used herein, the term "bit depth" represents the number of pixels used to represent one of the color components of an image. Traditionally, images were encoded at 8 bits per pixel per color component (e.g., 24 bits per pixel), but current architectures may now support higher bit depths such as 10 bits, 12 bits, etc.

[0005] In a traditional image pipeline, the captured image is quantized using a non-linear opto-electronic function (OETF), which converts linear scene light into a non-linear video signal (e.g., gamma-encoded RGB or YCbCr). Then, on the receiver, before being displayed on the display, the signal is processed by an electro-optical transfer function (EOTF) that converts video signal values into output screen color values. Such non-linear functions include the traditional "gamma" curves described in ITU-R Rec. BT.709 and BT.2020, the "PQ" (perceptual quantization) curve described in SMPTE ST2084, and the "Hybrid Log-gamma" or "HLG" curve described in Rec. ITU-R BT.2100. Summary of the Invention

[0006] As used herein, the term "forward reshaping" refers to the process of mapping from the original bit depth and the distribution or representation of the original codeword (e.g., gamma or PQ or HLG, etc.) of a digital image to an image with the same or different bit depth and different codeword distribution or representation, sample by sample from the samples of the digital image, or codeword by codeword. Reshaping enables improved compressibility or improved image quality at a fixed bitrate. For example, but not limited to, reshaping may be applied to 10-bit or 12-bit PQ-encoded HDR video to improve the encoding efficiency in a 10-bit video encoding architecture. At the receiver, after decompressing the reshaped signal, the receiver can apply an "inverse reshaping function" to restore the signal to the original codeword distribution. As recognized herein by the inventors, as development begins for the next generation of video coding standards, improved techniques for integrated reshaping and encoding of images are desired. The method of the present invention may be applicable to a variety of video content including, but not limited to, standard dynamic range (SDR) and / or high dynamic range (HDR) content.

[0007] The techniques described in this section are techniques that could have been pursued, but not necessarily techniques that were previously conceived or pursued. Therefore, unless otherwise indicated, no technique described in this section should be assumed to be eligible as prior art solely by virtue of being included in this section. Similarly, the problems identified with respect to one or more techniques should not be assumed to have been recognized in any prior art based on this section, unless otherwise stated. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Embodiments of the present invention are shown by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals refer to like elements.

[0009]

Fig. 1A

[0010]

Fig. 1B

[0011]

Fig. 2A

[0012]

Fig. 2B

[0013]

Fig. 2C

[0014]

Fig. 2D

[0015]

Fig. 2E

[0016]

Fig. 2F

[0017]

Fig. 3A

[0018]

Fig. 3B

[0019]

Fig. 4

[0020]

Fig. 5

[0021]

Fig. 6A

Fig. 6B

Fig. 6C

Fig. 6D

[0022]

Fig. 6E

DETAILED DESCRIPTION OF THE INVENTION

[0023] Signal shaping and encoding techniques for compressing images using rate-distortion optimization (RDO) are described in this paper. In the following description, for the purpose of explanation, numerous individual details are set forth to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without these specific details. On the other hand, well-known structures and devices are not described in exhaustive detail to avoid unnecessarily obscuring, burying, or obfuscating the present invention.

[0024] Summary The exemplary embodiments described herein relate to signal shaping and encoding for video. In an encoder, a processor receives an input image of a first codeword representation to be shaped into a second codeword representation, the second codeword representation allowing more efficient compression than the first codeword representation, and the processor further generates a forward shaping function that maps the pixels of the input image to the second codeword representation. To generate the forward shaping function, the encoder: divides the input image into a plurality of pixel regions, assigns each of the pixel regions to one of a plurality of codeword bins according to a first luminance characteristic of each pixel region, calculates a bin metric for each of the plurality of codeword bins according to a second luminance characteristic of each of the pixel regions assigned to each codeword bin, assigns some codewords in the second codeword representation to each codeword bin according to the bin metric of each codeword bin and a rate-distortion optimization criterion, and generates the forward shaping function in response to the assignment of the codewords in the second codeword representation to each of the plurality of codeword bins.

[0025] In another embodiment, in the decoder, the processor receives an encoded bitstream syntax element that characterizes the shaping model, the syntax element being a flag indicating the minimum codeword bin index value to be used in the shaping construction process, a flag indicating the maximum codeword bin index value to be used in the shaping construction process, a flag indicating the shaping model profile type, where the model profile type is associated with default bin relationship parameters including bin importance values, or one or more flags indicating delta bin importance values to be used to adjust the default bin importance values defined in the shaping model profile. The processor determines, based on the shaping model profile, a default bin importance value for each bin and an assignment list of the default number of codewords assigned to each bin according to the bin importance value. Then, for each codeword bin, the processor: determines its bin importance value by adding its default bin importance value to its delta bin importance value; determines the number of codewords to be assigned to the codeword bin based on the bin importance value and assignment list of the bin; generates a forward shaping function based on the number of codewords assigned to each codeword bin.

[0026] In another embodiment, in the decoder, the processor receives an encoded bitstream including one or more encoded shaped images in a first codeword representation and metadata related to shaping information for the encoded shaped images. The processor generates an inverse shaping function and a forward shaping function based on metadata related to the shaping information. Here, the inverse shaping function maps the pixels of the shaped image from the first coded word representation to the second coded word representation, and the forward shaping function maps the pixels of the image from the second coded word representation to the first coded word representation. The processor extracts a coded shaped image including one or more coded units from the coded bitstream. Here, for one or more coded units in the coded shaped image: For the shaped intra-coded coding unit (CU) in the coded shaped image, the processor: generates a first shaped reconstructed sample of the CU based on the shaped residual and the shaped prediction samples in the CU; generates a shaped loop filter output based on the first shaped reconstructed sample and the loop filter parameters; applies the inverse shaping function to the shaped loop filter output to generate the decoded samples of the coded unit in the second coded word representation; stores the decoded samples of the coded unit in the second coded word representation in the reference buffer; For the shaped inter-coded coding unit in the coded shaped image, the processor: applies the forward shaping function to the prediction samples stored in the reference buffer in the second coded word representation to generate a second shaped prediction sample; generates a second shaped reconstructed sample of the coded unit based on the shaped residual and the second shaped prediction sample in the coded CU; generates a shaped loop filter output based on the second shaped reconstructed sample and the loop filter parameters; applies the inverse shaping function to the shaped loop filter output to generate the samples of the coded unit in the second coded word representation; Store samples of the coding units in the second symbol word representation in the reference buffer. Finally, the processor generates a decoded image based on the samples stored in the reference buffer.

[0027] In another embodiment, in the decoder, the processor receives an encoded bitstream that includes one or more encoded reconstructed images in the input symbol word representation and shaping metadata (207) for the one or more encoded reconstructed images in the encoded bitstream. The processor generates a forward shaping function (282) based on the shaping metadata, and this forward shaping function maps the pixels of the image from the first symbol word representation to the input symbol word representation. The processor generates an inverse shaping function (265-3) based on the shaping metadata or the forward shaping function. Here, the inverse shaping function maps the pixels of the reconstructed image from the input symbol word representation to the first symbol word representation. The processor extracts an encoded reconstructed image including one or more encoded units from the encoded bitstream. Here: For intra-coded coding units (intra CUs) in the encoded reconstructed image, the processor: Generates the reconstructed samples (285) of the intra CU based on the shaped residuals and the intra-predicted shaped prediction samples in the intra CU; Applies the inverse shaping function (265-3) to the reconstructed samples of the intra CU to generate the decoded samples of the intra CU in the first symbol word representation; Applies a loop filter (270) to the decoded samples of the intra CU to generate the output samples of the intra CU; Stores the output samples of the intra CU in the reference buffer; For inter-coded CUs (inter CUs) in the encoded reconstructed image, the processor: Apply the forward shaping function (282) to the inter prediction samples stored in the reference buffer in the first coded word representation to generate shaped prediction samples for the inter CU in the input coded word representation; Generate the shaped reconstruction samples of the inter CU based on the shaped residual of the inter CU and the shaped prediction samples for the inter CU; Apply the inverse shaping function (265-3) to the shaped reconstruction samples of the inter CU to generate the decoded samples of the inter CU in the first coded word representation; Apply the loop filter (270) to the decoded samples of the inter CU to generate the output samples of the inter CU; Store the output samples of the inter CU in the reference buffer; Generate the decoded image in the first coded word representation based on the output samples in the reference buffer.

[0028] Example of Video Delivery Processing Pipeline FIG. 1A shows an exemplary process of a conventional video delivery pipeline (100) that shows various stages from video capture to video content display. A sequence of video frames (102) is captured or generated using an image generation block (105). The video frame (102) can be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frame (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production phase (110), the video data (107) is edited to provide a video production stream (112).

[0029] Next, the video data of the production stream (112) is provided to the processor in blocks (115) for post-production editing. Post-production editing in block (115) may include adjusting or modifying the color or luminance in individual regions of the image to improve image quality or achieve a particular look of the image according to the creative intent of the video creator. This is sometimes referred to as "color timing" or "color grading". Other edits (such as scene selection and sequencing, image cropping, addition of computer-generated visual effects, etc.) may be performed in block (115) to give the final version (117) of the production for distribution. During post-production editing (115), the video image is displayed on the reference display (125).

[0030] Following post-production (115), the video data of the final production (117) may be delivered to an encoding block (120) for delivery to downstream decoding and playback devices such as televisions, set-top boxes, movie theaters, etc. In some embodiments, the encoding block (120) may include audio and video encoders such as those defined by ATSC, DVB, DVD, Blu-Ray, and other delivery formats to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) representing the same or a near approximation of the signal (117). The receiver may be attached to a target display (140) having characteristics completely different from those of the reference display (125). In that case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137).

[0031] Signal Reshaping FIG. 1B shows an exemplary process for signal shaping according to Patent Document 2. When an input frame (117) is provided, a forward shaping block (150) analyzes the input and coding constraints and generates a codeword mapping function that maps the input frame (117) to a re-quantized output frame (152). For example, the input (117) may be encoded according to a certain electro-optical transfer function (EOTF) (e.g., gamma). In some embodiments, metadata may be used to convey information about the shaping process to a downstream device (e.g., a decoder). As used herein, the term "metadata" relates to any auxiliary information that is transmitted as part of an encoded bitstream and assists the decoder in rendering the decoded image. Such metadata includes, but is not limited to, color space or color gamut information, reference display parameters, and auxiliary signal parameters as described herein.

[0032] Following encoding (120) and decoding (130), the decoded frame (132) may be processed by a backward (or inverse) shaping function (160), which performs a conversion that returns the re-quantized frame (132) to the original EOTF domain (e.g., gamma) for further downstream processing such as the aforementioned display management process (135). In some embodiments, the backward shaping function (160) may be integrated with the de-quantizer within the decoder (130), for example, as part of the de-quantizer in an AVC or HEVC video decoder.

[0033] As used herein, the term "reshaper" may represent a forward reshaping function or an inverse reshaping function used when encoding and / or decoding a digital image. Examples of reshaping functions are discussed in Patent Document 2. In Patent Document 2, an in-loop block-based image reshaping method for high dynamic range video encoding was proposed. Its design allows for block-based reshaping within the encoding loop, but at the cost of increased complexity. Specifically, this design requires maintaining two sets of decoded image buffers. One is a set for the reshaped (or unreshaped) decoded pictures that can be used for both prediction without reshaping and output to the display, and the other is a set for the forward reshaped decoded pictures that are used only for prediction with reshaping. The forward reshaped decoded pictures can be calculated on the fly, but the cost of complexity is very high, especially for inter prediction (motion compensation using sub-pixel interpolation). Generally, display-picture-buffer (DPB) management is complex and requires very careful attention. Thus, as understood by the inventors, a simplified method for encoding video is desired.

[0034] Patent Document 6 presented further reshaping-based codec architectures, including an architecture with an external out-of-loop reshaper, an in-loop intra-only reshaper, an architecture with an in-loop reshaper for prediction residuals, and a hybrid architecture that combines both intra in-loop reshaping and inter residual reshaping. The main goal of those proposed reshaping architectures is to improve subjective visual quality. Thus, many of these approaches result in worse objective metrics, especially the well-known Peak Signal to Noise Ratio (PSNR) metric.

[0035] In the present invention, a new shaper based on Rate-Distortion Optimization (RDO) is proposed. In particular, when the target distortion metric is MSE (Mean Square Error), the proposed shaper improves both the subjective visual quality and the well-known objective metrics based on PSNR, Bjontegaard PSNR (BD-PSNR), or Bjontegaard rate (BD-Rate). Without loss of generality, note that any of the proposed reconstruction architectures can be applied to one or more of the luminance component, chroma component, or a combination of the luma and chroma components.

[0036] Reshaping Based on Rate-Distortion Optimization Considering a shaped video signal represented by the B-bit bit depth in a certain color component (e.g., B = 10 for Y, Cb, and / or Cr), there are a total of 2 B available codewords. Consider dividing the desired codeword range [0 2 B into N segments or bins. Let M k represent the number of codewords in the k-th segment or bin such that the distortion D between the source picture and the decoded or reconstructed picture is minimized when the target bitrate R is given. Without loss of generality, D may be expressed as a measure of the sum of squared errors (SSE) between the corresponding pixel values of the source input (Source(i,j)) and the reconstructed picture (Recon(i,j)). After Reshaping Mapping D = SSE = Σ Diff(i,j) i,j (1) 2 where Diff(i,j) = Source(i,j) - Recon(i,j) (2) The optimization shaping problem can be rewritten as follows: find M such that D is minimized when the bitrate R is given.

[0037] k kFind (k = 0, 1, …, N - 1). Here,

Number

[0038] Although various optimization methods can be used to find the solution, the optimal solution can be very complex for real-time encoding. In the present invention, a sub-optimal but more practical analytical solution is proposed.

[0039] Without loss of generality, consider an input signal represented by a bit depth of B bits (e.g., B = 10). Here, the codewords are uniformly divided into N bins (e.g., N = 32). By default, each bin has M a = 2 B / N codewords assigned to it (e.g., for N = 32 and B = 10, M a = 32). Next, a more efficient codeword assignment based on RDO is demonstrated through examples.

[0040] As used herein, the term "narrow range" [CW1, CW2] indicates a continuous range of codewords between codewords CW1 and CW2, which is a subset of the entire dynamic range [0 2 B -1]. For example, in one embodiment, the narrow range can be defined as [16 * 2 (B-8) , 235 * 2 (B-8) (e.g., for B = 10, the narrow range includes the values [64 940]). Assuming that the bit depth of the output signal is B o and the dynamic range of the input signal is within the narrow range, in what is denoted as "default shaping", the signal can be extended to the full range [0 2 Bo -1]. Then, each bin has M f = CEIL((2 Bo / (CW2 - CW1)) * M a ), or, for our example, when B o = B = 10, M f=CEIL((1024 / (940 - 64)) * 32) = 38 codewords. Here, CEIL(x) represents the ceiling function that maps x to the smallest integer greater than or equal to x. Without loss of generality, in the following example, for simplicity, assume B o = B.

[0041] For the same quantization parameter (QP), the effect of increasing the number of codewords in a bin is equivalent to allocating more bits to encode the signal in the bin. Thus, it is equivalent to reducing the SSE or improving the PSNR. However, a uniform increase in codeword allocation in each bin may not give better results than unshaped coding. This is because the PSNR gain may not exceed the increase in bitrate. That is, this is not a good trade-off for RDO. Ideally, we want to allocate more codewords only to the bins that give the best trade-off for RDO, i.e., to cause a significant reduction in SSE (increase in PSNR) at the cost of a small increase in bitrate.

[0042] In one embodiment, the RDO performance is improved through adaptive segmental shaping mapping. This method can be applied to any type of signal, including standard dynamic range (SDR) and high dynamic range (HDR) signals. Using the simple case described above as an example, the goal of the present invention is to assign M a or M f codewords to each codeword segment or codeword bin.

[0043] In the encoder, when N codeword bins for the input signal are given, the average luminance variance of each bin can be approximated as follows: · Initialize the sum of block variances (var bin (k)) and the counter (c bin (k)) to 0 for each bin. For example, for k = 0, 1, …, N - 1, var bin (k)=0, cbin Let (k) = 0. · Divide the picture into non - overlapping blocks of L*L (e.g., L = 16). · For each picture block, calculate the block's luma average and the luma variance of block i (e.g., Luma_mean(i) and Luma_var(i)). · Based on the average luma of the block, assign the block to one of N bins. In one embodiment, if Luma_mean(i) is within the k - th segment of the input dynamic range, the total bin luminance variance for the k - th bin is incremented by the luma variance of the newly assigned block, and the counter for that bin is incremented by 1. That is, if the i - th pixel region belongs to the k - th bin: var bin (k)=var bin (k)+Luma_var(i); (2) c bin (k)=c bin (k)+1 · For each bin, assuming the counter is not equal to 0, calculate the average luminance variance for that bin by dividing the sum of the block variances within the bin by the counter. Alternatively, if c bin (k) is not 0, var bin (k)=var bin (k) / c bin (k) (3) Let it be so.

[0044] Those skilled in the art will understand that alternative metrics other than luminance variance may be applied to characterize the sub - blocks. For example, the standard deviation of the luminance values, weighted luminance variance or luminance values, peak luminance, etc. may be used.

[0045] In one embodiment, the following pseudo - code shows an example of how an encoder can adjust bin assignment using the calculated metrics for each bin. For the k-th bin, if there are no pixels in the bin M k = 0; else if var bin (k) < TH U (4) M k = M f ; else M k = M a ; / / (Note: This is to ensure that each bin has at least M a codewords.) / / Alternatively, M a + 1 codewords may be assigned.) end Here, TH U represents a predetermined upper threshold value.

[0046] In another embodiment, the assignment may be performed as follows: For the k-th bin, if there are no pixels in the bin M k = 0; else if TH0 < var bin (k) < TH1(5) M k = M f ; else M k = M a ; end Here, TH0 and TH1 represent predetermined lower and upper threshold values.

[0047] In another embodiment For the k-th bin, if there are no pixels in the bin M k = 0; else if var bin (k) > TH L (6) M k = M f ; else M k =M a ; end Here, TH L represents a predetermined lower threshold value.

[0048] The above example shows how to select the number of codewords for each bin from two preselected numbers M f and M a . The threshold value (e.g., TH U or TH L ) can be determined based on optimizing the rate distortion, for example, through an exhaustive search. The threshold value may be adjusted based on the quantization parameter value (QP). In one embodiment, when B = 10, the threshold value can be in the range of 1,000 to 10,000.

[0049] In one embodiment, to speed up the processing, the threshold value may be determined from a fixed set of values, such as {2,000, 3,000, 4,000, 5,000, 6,000, 7,000}, using the Lagrangian optimization method. For example, for each TH(i) value in the set, a compression test can be performed with a fixed QP using a predefined training clip, and the value of the objective function J can be calculated. J is J(i)=D + λR (7) defined as. Then, the optimal threshold value can be defined as the TH(i) value in the set where J(i) is minimized.

[0050] In a more general example, a look-up table (LUT) can be predefined. For example, in Table 1, the first row represents possible bin metrics (e.g., var binDefine a set of thresholds that divide the entire range of (k) values into segments, and the second row defines the corresponding number of coded words (CW) assigned to each segment. In one embodiment, one rule for constructing such a LUT is as follows: If the bin variance is too large, more bits may need to be spent to reduce the SSE, thus, M a Coded word (CW) values smaller than can be assigned. When the bin variance is very small, M a Larger CW values than can be assigned.

[0051] Table 1: Exemplary LUT of coded word assignment based on bin variance threshold [Table 1]

[0052] Using Table 1, the mapping of the threshold to the coded word can be generated as follows. For the k-th bin, if there are no pixels in the bin M k = 0; else if var bin (k) < TH0 M k = CW0; else if TH0 < var bin (k) < TH1(8) M k = CW1; … else if TH p-1 < var bin (k) < TH p Mk = CW p ; … else if var bin (k) > TH q-1 Mk = CW q ; end

[0053] For example, when two thresholds and three codeword assignments are given, for B = 10, in one embodiment, TH0 = 3,000, CW0 = 38, TH1 = 10,000, CW1 = 32, and CW2 = 28.

[0054] In another embodiment, the two thresholds TH0 and TH1 can be selected as follows: a) Consider TH1 as a very large number (which may be infinite), and for example, using the RDO optimization of Equation (7), select TH0 from a set of pre-determined values. Given TH0, now define a second set of possible values for TH1, for example, the set {10,000, 15,000, 20,000, 25,000, 30,000}, and apply Equation (7) to identify the optimal value. This approach can be executed sequentially and iteratively using a limited number of thresholds or until convergence.

[0055] After assigning the codewords to the bins according to any of the previously defined schemes, M k the sum of the values may exceed the maximum of the available codewords (2 B ), or there may be unused codewords, which may be noted. If there are unused codewords, one may simply decide to do nothing, or assign the unused codewords to a specific bin. On the other hand, if the algorithm assigns more codewords than available, for example, by re-normalizing the CW values, one may want to re-adjust the M k values. Alternatively, one may also generate a forward shaping function using the existing Mk values, but then re-adjust the output values of the shaping function by scaling by (Σ k M k ) / 2 B . Examples of codeword re-assignment techniques are also described in Patent Document 7.

[0056] Figure 4 shows an exemplary process for assigning codewords to the shaping domain according to the RDO technique described above. In step 405, the desired shaped dynamic range is divided into N bins. After the input image is divided into non-overlapping blocks (step 410), for each block: · Step 415 calculates its luminance characteristics (e.g., mean and variance). · Step 420 assigns each image block to one of the N bins. · Step 425 calculates the average luminance variance in each bin Given the value calculated in step 425, in step 430, each bin is assigned one or more codewords according to one or more thresholds, using, for example, any of the codeword assignment algorithms shown in equations (4) to (8). Finally, in step (435), the final codeword assignment can be used to generate the forward shaping function and / or the inverse shaping function.

[0057] In one embodiment, without limitation, by way of example, the forward LUT (FLUT) can be constructed using the following C code. tot_cw = 2 B ; hist_lens = tot_cw / N; for (i = 0; i < N; i++) { double temp = (double)M[i] / (double)hist_lens; / / M[i] corresponds to M k corresponding to for (j = 0; j < hist_lens; j++) { CW_bins_LUT_all[i * hist_lens + j] = temp; } Y_LUT_all[0] = CW_bins_LUT_all[0]; for (i = 1; i < tot_cw; i++) { Y_LUT_all[i]=Y_LUT_all[i - 1]+CW_bins_LUT_all[i]; } for (i = 0; i < tot_cw; i++) { FLUT[i]=Clip3(0, tot_cw - 1, (Int)(Y_LUT_all[i]+0.5)); }

[0058] In one embodiment, the inverse LUT can be constructed as follows. low = FLUT[0]; high = FLUT[tot_cw - 1]; first = 0; last = tot_cw - 1; for (i = 1; i < tot_cw; i++) if (FLUT[0]<FLUT[i]) { first = i - 1; break; } for (i = tot_cw - 2; i >= 0; i--) if (FLUT[tot_cw - 1]>FLUT[i]) { last = i + 1; break; } for (i = 0; i < tot_cw; i++) if (i <= low) { ILUT[i]=first; } else if (i >= high) { ILUT[i]=last; } else { for (j = 0; j < tot_cw - 1; j++) if (FLUT[j]>=i) { ILUT[i]=j; break; }} }

[0059] Syntactically, the syntax proposed in previous applications, such as the partitioned polynomial mode or parametric model in Patent Documents 5 and 6, can be reused. Table 2 shows such an example for N = 32 for Equation (4). Table 2: Syntax of shaping using the first parametric model [Table 2] Here,[[]] reshaper_model_profile_type specifies the profile type used in the reshaper construction process. A given profile can provide information about default values used, such as the number of bins, the importance or priority value of the default bins, and the default codeword assignment (e.g., M a and / or M f values). reshaper_model_scale_idx specifies the index value of the scale factor (denoted as ScaleFactor) used in the reshaper construction process. The value of ScaleFactor makes it possible to improve the control of the shaping function for improving the overall coding efficiency. reshaper_model_min_bin_idx specifies the minimum bin index used in the reshaper construction process. The value of reshaper_model_min_bin_idx ranges from 0 to 31, inclusive. reshaper_model_max_bin_idx specifies the maximum bin index used in the reshaper construction process. The value of reshaper_model_max_bin_idx ranges from 0 to 31, inclusive. reshaper_model_bin_profile_delta[i] specifies the delta value used to adjust the profile of the i-th bin in the reshaper construction process. The value of reshaper_model_bin_profile_delta[i] ranges from 0 to 1, inclusive.

[0060] Table 3 shows another embodiment with an alternative, more efficient syntactic representation. Table 3: Syntax of Reconstruction Using the Second Parametric Model

Table 3

[0061] In another embodiment, as described in Table 4, the number of codewords per bin may be explicitly defined. Table 4: Syntax of shaping using the third model

Table 4

[0062] In one embodiment, assuming one of the above examples, for example, the sign word assignment according to Equation (4), an example of how to define the parameters in Table 2 includes the following: First, assign "bin importance" as follows: For the k-th bin, if M k =0; bin_importance = 0; else if M k ==M f bin_importance = 2; (9) else bin_importance = 1; end

[0063] As used herein, the term "bin importance" is a value assigned to each of the N coded word bins, indicating the importance of all the coded words within that bin in the shaping process relative to other bins.

[0064] In one embodiment, the default_bin_importance (default bin importance) from reshaper_model_min_bin_idx to reshaper_model_max_bin_idx may be set to 1. The value of reshaper_model_min_bin_idx is set to the smallest bin index having a non-zero M k The value of reshaper_model_max_bin_idx is set to the largest bin index having a non-zero M k For each bin within [reshaper_model_min_bin_idx reshaper_model_max_bin_idx], reshaper_model_bin_profile_delta is the difference between bin_importance and default_bin_importance.

[0065] An example of how to use the proposed parametric model to construct a Forward Reshaping LUT (FLUT) and an Inverse Reshaping LUT (ILUT) is shown below. 1) Divide the luminance range into N bins (e.g., N = 32). 2) Derive the bin-importance index for each bin from the syntax. For example, for the k-th bin, if reshaper_model_min_bin_idx <= k <= reshaper_model_max_bin_idx bin_importance[k] = default_bin_importance[k] + reshaper_model_bin_profile_delta[k]; else bin_importance[k] = 0; 3) Automatically pre-assign the codewords based on the bin importance: for the k-th bin, if bin_importance[k] == 0 M k = 0; else if bin_importance[k] == 2 M k = M f ; else M k = M a ; end 4) Construct the forward shaping LUT based on the codeword assignment for each bin by accumulating the assigned codewords for each bin. The sum should be below the total codeword budget (e.g., 1024 for a 10-bit full range). (See the first C code for example.) 5) Construct the inverse shaping LUT (see the first C code for example).

[0066] From a syntax perspective, alternative methods are also applicable. The key is, explicitly or implicitly, the number of codewords in each bin (e.g., M for k = 0, 1, 2, …, N - 1) kis to specify. In one embodiment, the number of coded words in each bin can be explicitly specified. In another embodiment, the coded words can be specified in different ways. For example, the number of coded words in a bin can be determined using the difference in the number of coded words in the current bin and the previous bin (e.g., M_Delta(k)=M(k)-M(k - 1)). In another embodiment, the most commonly used number of coded words (e.g., M M ) can be specified, and the number of coded words in each bin can be expressed as the difference in the number of coded words in each bin from this number (e.g., M_Delta(k)=M(k)-M M ).

[0067] In one embodiment, two shaping methods are supported. One is represented as the "default shaper", and M f is assigned to all bins. The other is represented as the "adaptive shaper", and the above-described adaptive shaper is applied. These two methods can be signaled to the decoder using a special flag, such as sps_reshaper_adaptive_flag, as in Patent Document 6 (e.g., sps_reshaper_adaptive_flag = 0 is used for the default shaper, and sps_reshaper_adaptive_flag = 1 is used for the adaptive shaper).

[0068] The present invention is applicable to any shaper proposed in Patent Document 6, such as an external shaper, an intra-only shaper within a loop, a residual shaper within a loop, or a hybrid shaper within a loop. As an example, FIGS. 2A and 2B show an exemplary architecture for hybrid in-loop shaping according to an embodiment of the present invention. In FIG. 2A, the architecture combines elements from both an intra-only shaping architecture within the loop (upper part of the figure) and a residual architecture within the loop (lower part of the figure). Under this architecture, for intra-slices, shaping is applied to the picture pixels, while for inter-slices, shaping is applied to the prediction residuals. In the encoder (200_E), two new blocks are added to a conventional block-based encoder (e.g., HEVC): a block (205) that estimates a forward shaping function (e.g., according to FIG. 4), a forward picture shaping block (210-1), and a forward residual shaping block (210-2). This applies forward shaping to one or more of the color components of the input video (117) or the prediction residuals. In some embodiments, these two operations may be performed as part of a single image shaping block. Parameters (207) related to the determination of the inverse shaping function in the decoder can be passed to the lossless encoder block (e.g., CABAC 220) of the video encoder so that they can be embedded in the encoded bitstream (122). In the intra mode, intra prediction (225-1), transform and quantization (T&Q), inverse transform and inverse quantization (Q -1 &T -1 ) all use the shaped picture. In both modes, the pictures stored in the DPB (215) are always in the inverse shaping mode, which requires an inverse picture shaping block (e.g., 265-1) or an inverse residual shaping block (e.g., 265-2) before the loop filters (270-1, 270-2). As shown in FIG. 2A, an intra / inter-slice switch allows switching between the two architectures depending on the type of slice being encoded. In another embodiment, in-loop filtering for intra-slices may be performed before inverse shaping.

[0069] In decoder (200_D), the following new normative blocks are added to a conventional block-based decoder: a block (250) that reconstructs a forward shaping function and a backward shaping function based on the encoded shaping function parameters (207) (shaper decoding); a block (265-1) that applies an inverse shaping function to the decoded data; and a block (265-2) that applies both a forward shaping function and an inverse shaping function to generate a decoded video signal (162). For example, in (265-2), the reconstructed value is given by Rec = ILUT(FLUT(Pred)+Res), where FLUT represents a forward shaping LUT and ILUT represents an inverse shaping LUT.

[0070] In some embodiments, the operations associated with blocks 250 and 265 may be combined into a single processing block. As shown in FIG. 2B, an intra / inter slice switch allows switching between two modes depending on the type of slice in the encoded video picture.

[0071] FIG. 3A shows an exemplary process (300_E) for encoding video using a shaping architecture (e.g., 200_E) according to an embodiment of the present invention. In the case of no active shaping (path 305), the encoding (335) proceeds as known in the prior art encoders (e.g., HEVC). When shaping is active (path 310), the encoder may have the option of applying a predetermined (default) shaping function (315) or adaptively determining a new shaping function (325) based on picture analysis (320) (as described, for example, in FIG. 4). After encoding the image using the shaping architecture (330), the remaining part of the encoding follows the same steps as the conventional encoding pipeline (335). When adaptive shaping is used (312), metadata related to the shaping function is generated as part of the "encode shaper" step (327).

[0072] Figure 3B shows an exemplary process (300_D) for decoding video using a shaping architecture (e.g., 200_D) according to an embodiment of the present invention. In the case where no enabled shaping exists (path 340), after decoding the picture (350) in the same manner as a conventional decoding pipeline, an output frame is generated (390). In the case where shaping is enabled (path 360), the decoder determines whether to apply a predetermined (default) shaping function (375) or adaptively determine a shaping function based on the received parameters (e.g., 207) (380). Following decoding using the shaping architecture (385), the remainder of the decoding follows a conventional decoding pipeline.

[0073] As described above in Patent Document 6 and this specification, the forward shaping LUT FwdLUT may be constructed by integration, while the inverse shaping LUT may be constructed based on backward mapping using the forward shaping LUT (FwdLUT). In an embodiment, the forward LUT may be constructed using piecewise linear interpolation. In the decoder, inverse shaping can be performed directly using the backward LUT or also by linear interpolation. The piecewise linear LUT is constructed based on input pivot points and output pivot points.

[0074] Let (X1, Y1) and (X2, Y2) be two input pivot points and the corresponding output values for each bin. The input value X between X1 and X2 can be interpolated by the following equation: Y = ((Y2 - Y1) / (X2 - X1)) * (X - X1) + Y1 In a fixed-point implementation, the above equation can be rewritten as: Y = ((m * X + 2 FP_PREC-1 ) >> FP_PREC) + c where m and c represent scalars and offsets for linear interpolation, and FP_PREC is a constant related to fixed-point precision.

[0075] As an example, the FwdLUT can be constructed as follows: Variables: lutSize = (1 << BitDepth Y ) is set. Variables: binNum = reshaper_model_number_bins_minus1 + 1 and binLen = lutSize / binNum are set. For the i-th bin, the two pivot points at both ends (e.g., X1 and X2) can be derived as X1 = i * binLen and X2 = (i + 1) * binLen. And: binsLUT[0] = 0; for (i = 0; i < reshaper_model_number_bins_minus1 + 1; i++) { binsLUT[(i + 1) * binLen] = binsLUT[i * binLen] + RspCW[i]; Y1 = binsLUT[i * binLen]; Y2 = binsLUT[(i + 1) * binLen]; scale = ((Y2 - Y1) * (1 << FP_PREC) + (1 << (log2(binLen) - 1))) >> (log2(binLen)); for (j = 1; j < binLen; j++) { binsLUT[i * binLen + j] = Y1 + ((scale * j + (1 << (FP_PREC - 1))) >> FP_PREC); } }

[0076] FP_PREC defines the fixed-point precision of the fractional part of the variable (e.g., FP_PREC = 14). In an embodiment, binsLUT[] may be calculated with a higher precision than the precision of FwdLUT. For example, the binsLUT[] values may be calculated as 32-bit integers, but FwdLUT may be the binsLUT values clipped to 16 bits.

[0077] Adaptive Threshold Derivation As described above, during shaping, the codeword assignment may be adjusted using one or more thresholds (e.g., TH, TH U , TH L , etc.). In certain embodiments, such thresholds may be adaptively generated based on content characteristics. FIG. 5 shows an exemplary process for deriving such thresholds according to one embodiment. 1) In step 505, the luminance range of the input image is divided into N bins (e.g., N = 32). For example, let N also be denoted as PIC_ANALYZE_CW_BINS. 2) In step 510, image analysis is performed to calculate the luminance characteristics for each bin. For example, the percentage of pixels within each bin (denoted as BinHist[b], where b = 1, 2,..., N) may be calculated. Here, BinHist[b]=100*(total number of pixels in bin b) / (total number of pixels in the picture) (10) As discussed above, another good metric for image characteristics is the average variance (or standard deviation) of the pixels in each bin, denoted as BinVar[b]. BinVar[b] may be calculated in "block mode" as described in the steps leading to equations (2) and (3). Alternatively, the block-based calculation can be refined with a pixel-based calculation. For example, the variance associated with the group of pixels surrounding the i-th pixel within an m×m neighborhood window (e.g., m = 5) centered on the i-th pixel is denoted as vf(i). For example,

Equation

Equation

Equation

[0078] 3) In step 515, the average bin variances (and their corresponding indices) are sorted, for example but not limited to, in descending order. For example, the sorted BinVar values may be stored in BinVarSortDsd[b], and the sorted bin indices may be stored in BinIdxSortDsd[b]. As an example, using C code, this process may be described as follows: for(int b = 0; b < PIC_ANALYZE_CW_BINS; b++ / / Initialization (unsorted) { BinVarSortDsd[b] = BinVar[b]; BinIdxSortDsd[b] = b; } / / Sort (see the code example in Appendix 1) bubbleSortDsd(BinVarSortDsd, BinIdxSortDsd, PIC_ANALYZE_CW_BINS); An exemplary plot of the sorted average bin variance factors is shown in FIG. 6A.

[0079] 4) Given the bin histogram values calculated in step 510, in step 520, a cumulative density function (CDF) is calculated and stored according to the order of the sorted average bin variances. For example, if the CDF is stored in the array BinVarSortDsdCDF[b], in one embodiment: BinVarSortDsdCDF[0] = BinHist[BinIdxSortDsd[0]]; for (int b = 1; b < PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b] = BinVarSortDsdCDF[b - 1] + BinHist[BinIdxSortDsd[b]]; } An exemplary plot (605) of the CDF calculated based on the data in FIG. 6A is shown in FIG. 6B. The pairs of CDF values versus sorted average bin variances {x = BinVarSortDsd[b], y = BinVarSortDsdCDF[b]} can be interpreted as "there are y% of pixels in the picture having a variance greater than or equal to x" or "there are (100 - y)% of pixels in the picture having a variance less than x". 5) Finally, in step 525, a CDF, BinVarSortDsdCDF[BinVarSortDsd[b]], is given as a function of the sorted average bin variance values, and a threshold can be defined based on the bin variance and the cumulative percentage.

[0080] Examples for determining a single threshold or two thresholds are shown in FIGS. 6C and 6D, respectively. When only one threshold (e.g., T H ) is used, as an example, T H may be defined as "the average variance such that k% of the pixels have T H ". Then, T H can be calculated by finding the intersection point (e.g., the BinVarSortDsd[b] value where BinVarSortDsdCDF = k%) of the CDF plot (605) at k% (e.g., 610). For example, as shown in FIG. 6C, for k = 50, T H = 2.5. Then, M f codewords are assigned for bins with BinVar[b] < TH, and M aIndividual codewords can be assigned. As a rough rule, it is preferable to assign a larger number of codewords to bins with smaller variances (e.g., for a 10-bit video signal with 32 bins, M f > 32 > M a ).

[0081] When using two thresholds, an example of selecting TH L and TH U is shown in FIG. 6D. For example, without loss of generality, TH L may be defined as the variance in which 80% of the pixels have vf ≧ TH L (in this example, TH L = 2.3), and TH U may be defined as the variance in which 10% of all pixels have vf ≧ TH U (in this example, TH U = 3.5). Given these thresholds, M L codewords can be assigned for bins with BinVar[b] < TH f , and M U codewords can be assigned for bins with BinVar[b] ≧ TH a . For bins with BinVar between TH L and TH U , the original number of codewords per bin (e.g., 32 for B = 10) may be used.

[0082] The above technique can be easily extended when having two or more thresholds. This relationship can also be used to adjust the number of codewords (M f , M a , etc.). As a rough rule, in bins with low variance, more codewords should be assigned to boost PSNR (and lower MSE); for bins with high variance, fewer codewords should be assigned to save bits.

[0083] In one embodiment, a set of parameters (e.g., TH L , TH U , M a, M f If, for example, parameters such as these are manually obtained through exhaustive manual parameter tuning, this automated method can be applied to design a decision tree for categorizing each content in order to automatically set optimal manual parameters. For example, content categories include: movies, TV, SDR, HDR, comics, nature, action, etc.

[0084] To reduce complexity, in-loop shaping can be restricted in various ways. When in-loop shaping is adopted in a video coding standard, these restrictions should be normative to ensure decoder simplification. For example, in one embodiment, loop filtering may be disabled for certain block coding sizes. For example, when nTbW * nTbH < TH, the intra and inter filter modes in the inter slice can be disabled. Here, the variable nTbW specifies the transform block width and the variable nTbH specifies the transform block height. For example, for TH = 64, blocks of sizes 4×4, 4×8, and 8×4 are disabled for both intra and inter mode filtering in the inter-coded slice (or tile).

[0085] Similarly, in another embodiment, chroma residual scaling based on loop filtering may be disabled in the intra mode in the inter-coded slice (or tile). Or it may also be disabled when it is effective to have separate loop filtering and chroma splitting trees.

[0086] Interaction with Other Encoding Tools Loop Filtering In Patent Document 6, it is described that the loop filter can operate in either the original pixel domain or the reshaped pixel domain. In certain embodiments, it is proposed that loop filtering be performed in the original pixel domain (after picture reshaping). For example, in the hybrid in-loop reshaping architectures (200_E and 200_D), for intra pictures, inverse reshaping (265-1) needs to be applied before the loop filter (270).

[0087] Figures 2C and 2D show alternative decoder architectures (200B_D and 200C_D) where inverse reshaping (265) is performed immediately before storing the decoded data in the decoded picture buffer (DPB) (260) after loop filtering (270). In the proposed embodiments, compared to the architecture of 200_D, the inverse residual reshaping formula for inter slices is modified and inverse reshaping is performed (e.g., via the InvLUT() function or a look-up table) after loop filtering (270). In this way, inverse reshaping is performed after loop filtering for both intra slices and inter slices, and for both intra-coded CUs and inter-coded CUs, the reconstructed pixels before loop filtering are in the reshaped domain. After inverse reshaping (265), all output samples stored in the reference DPB are in the original domain. Such architectures allow for both slice-based adaptation and CTU-based adaptation for in-loop reshaping.

[0088] As shown in Figures 2C and 2D, in certain embodiments, loop filtering (270) is performed in the reshaped domain for both intra-coded CUs and inter-coded CUs, and inverse picture reshaping (265) occurs only once, thus presenting a unified and simpler architecture for both intra-coded CUs and inter-coded CUs.

[0089] To decode an intra-coded CU (200B_D), intra prediction (225) is performed on the reshaped neighboring pixels. Given the residual Res and the predicted sample PredSample, the reconstructed sample (227) is derived as follows: RecSample = Res + PredSample (14) Given the reconstructed sample (227), loop filtering (270) and inverse picture reshaping (265) are applied to derive the RecSampleInDPB sample stored in the DPB (260). Here,[[]] RecSampleInDPB = InvLUT(LPF(RecSample))) = InvLUT(LPF(Res + PredSample))) (15) Here, InvLUT() represents the inverse reshaping function or inverse reshaping look-up table, and LPF() represents the loop filtering operation.

[0090] In conventional coding, the inter / intra mode decision is based on calculating a distortion function (dfunc()) between the original sample and the predicted sample. Examples of such functions include sum of squared errors (SSE), sum of absolute differences (SAD), and others. When using reshaping, on the encoder side (not shown), CU prediction and mode decision are performed in the reshaping domain. That is, for mode decision, distortion = dfunc(FwdLUT(SrcSample) - RecSample) (16) Here, FwdLUT() represents the forward reshaping function (or LUT), and SrcSample represents the original image sample.

[0091] For an inter-coded CU, on the decoder side (e.g., 200C_D), inter prediction is performed using the reference pictures in the non-reshaping domain in the DPB. Then, in the reconstruction block 275, the reconstructed pixel (267) is derived as follows: RecSample=(Res + FwdLUT(PredSample)) (17) A reconstructed sample (267) is provided, and loop filtering (270) and inverse picture shaping (265) are applied to derive the RecSampleInDPB sample stored in the DPB. Here, RecSampleInDPB = InvLUT(LPF(RecSample)) = InvLUT(LPF(Res + FwdLUT(PredSample)))) (18)

[0092] On the encoder side (not shown), intra prediction is performed in the shaping domain as follows. Res = FwdLUT(SrcSample) - PredSample (19a) Here, it is assumed that all neighboring samples (PredSample) used for prediction are already in the shaping domain. Inter prediction (e.g., using motion compensation) is performed in the non-shaping domain (i.e., directly using the reference picture from the DPB). That is, PredSample = MC(RecSampleinDPB) (19b) Here, MC() represents the motion compensation function. For fast mode decision where motion estimation and residuals are not generated, the distortion can be calculated using the following formula: distortion = dfunc(SrcSample - PredSample) However, for full mode decision where residuals are generated, the mode decision is performed in the shaping domain. That is, for full mode decision, distortion = dfunc(FwdLUT(SrcSample) - RecSample) (20)

[0093] Block-Level Adaptation As described above, the proposed in-loop shaper allows adaptation of shaping at the CU level, e.g., setting the variable CU_reshaper on or off as needed. Under the same architecture, for inter-coded CUs, when CU_reshaper = off, the reconstructed pixels need to be in the shaping domain even if the CU_reshaper flag is set off for this inter-coded CU. RecSample = FwdLUT(Res + PredSample) (21) Therefore, intra prediction always has neighboring pixels in the shaping domain. DPB pixels can be derived as follows: RecSampleInDPB = InvLUT(LPF(RecSample)))= = InvLUT(LPF(FwdLUT(Res + PredSample))) (22)

[0094] For intra-coded CUs, depending on the encoding process, two alternative methods are proposed: 1) All intra-coded CUs are encoded with CU_reshaper = on. In this case, since all pixels are already in the shaping domain, no additional processing is required. 2) Some intra-coded CUs can be encoded using CU_reshaper = off. In this case, for CU_reshaper = off, when applying intra prediction, inverse shaping needs to be applied to the neighboring pixels so that the intra prediction is performed in the original domain and the final reconstructed pixels need to be in the shaping domain. That is, RecSample = FwdLUT(Res + InvLUT(PredSample)) (23) And, RecSampleInDPB = InvLUT(LPF(RecSample)))= = InvLUT(LPF(FwdLUT(Res + InvLUT(PredSample))))) (24)

[0095] Generally, the proposed architectures can be used in various combinations, such as in-loop intra-only shaping, in-loop shaping only for prediction residuals, or a hybrid architecture that combines both in-loop intra shaping and inter residual shaping. For example, to reduce the latency in a hardware decode pipeline, for inter-slice decoding, intra prediction can be performed before inverse shaping (i.e., decode the intra CU within the inter-slice). An exemplary architecture (200D_D) of such an embodiment is shown in FIG. 2E. In the reconstruction module (285), for the inter CU from equation (17) (e.g., the Mux enables the output from 280 and 282), RecSample=(Res+FwdLUT(PredSample)) where FwdLUT(PredSample) represents the output of the inter predictor (280) and the subsequent forward shaping (282). Instead, for the intra CU (e.g., the Mux enables the output from 284), the output of the reconstruction module (285) is RecSample=(Res+IPredSample) where IPredSample represents the output of the intra prediction block (284). The inverse shaping block (265-3) Y CU =InvLUT[RecSample] performs.

[0096] Applying intra prediction in the shaping domain for inter slices is applicable to other embodiments including those shown in FIG. 2C (where inverse shaping is performed after loop filtering) and FIG. 2D. Special care needs to be taken in the combined inter / intra prediction mode (i.e., when during reconstruction, some samples are from inter-coded blocks and some are from intra-coded blocks). In all such embodiments, the inter prediction is in the original domain, while the intra prediction is in the shaping domain. When combining data from both inter-predicted and intra-predicted coding units, the prediction can be performed in either of the two domains. For example, when the combined inter / intra prediction mode is executed in the shaped domain: PredSampleCombined = PredSampeIntra + FwdLUT(PredSampleInter) RecSample = Res + PredSampleCombined That is, the inter-coded samples in the original domain are shaped before addition. Otherwise, when the combined inter / intra prediction mode is performed in the original domain: PredSampleCombined = InvLUT(PredSampeIntra) + PredSampleInter RecSample = Res + FwdLUT(PredSampleCombined) That is, the intra-predicted samples are inverse-shaped to be in the original domain.

[0097] Similar considerations are applicable to the corresponding encoding embodiments. This is because the encoder (e.g., 200_E) includes a decoder loop that matches the corresponding decoder. As discussed above, Equation (20) describes an embodiment where mode decision is performed in the shaping domain. In another embodiment, the mode decision may be performed in the original domain, i.e.: distortion = dfunc(SrcSample - InvLUT(RecSample))

[0098] For chroma QP offset or chroma residual scaling based on luma, the average CU luma value

Number

[0099] Chroma QP Derivation Similar to Patent Document 6, in order to balance the relationship between luma and chroma caused by the shaping curve, the same proposed chroma DQP derivation process may be applied. In an embodiment, a segmented chroma DQP value can be derived based on the sign word assignment for each bin. For example: For the k-th bin scale k =(M k / M a ); (25) chromaDQP = 6 * log2(scale k ); end

[0100] Encoder Optimization As described in Patent Document 6, when luma DQP is enabled, it is recommended to use pixel-based weighted distortion. When shaping is used, in one example, the required weights are adjusted based on the shaping function (f(x)). For example: W rsp = f'(x) 2 (26) Here, f'(x) represents the gradient of the shaping function f(x).

[0101] In another embodiment, a segmented weight can be directly derived based on the sign word assignment for each bin. For example; For the k-th bin, Wrsp (k)=(M k / M a ) 2 (27)

[0102] For the chroma component, the weight can be set to 1 or some scaling factor sf. To reduce chroma distortion, sf can be set greater than 1. To increase chroma distortion, sf can be set greater than 1. In certain embodiments, sf can be used to compensate for Equation (25). Since chromaDQP can only be set to an integer, sf can be used to accommodate the fractional part of chromaDQP. Thus: sf = 2 ((chromaDQP-INT(chromaDQP)) / 3)

[0103] In another embodiment, to control chroma distortion, the chromaQPOffset value can be explicitly set in the Picture Parameter Set (PPS) or the slice header.

[0104] The shaper curve or mapping function need not be fixed for the entire video sequence. For example, it can be adapted based on the quantization parameter (QP) or the target bitrate. In certain embodiments, when the bitrate is low, a more aggressive shaper curve can be used, and when the bitrate is relatively high, a less aggressive shaping can be used. For example, when 32 bins are given for a 10-bit sequence, each bin initially has 32 codewords. When the bitrate is relatively low, codewords between [28 40] can be used to select the codeword for each bin. When the bitrate is high, codewords between [31 33] can be selected for each bin, or simply an identity shaper curve can be used.

[0105] Given a slice (or tile), shaping at the slice (tile) level can be performed in a variety of ways that can trade off encoding efficiency with complexity. This can include: 1) disabling shaping for intra-slices only; 2) disabling shaping in certain inter-slices, such as at a particular temporal level (single or multiple), or in inter-slices not used for the reference picture, or in inter-slices considered to be of low importance reference pictures. Such slice adaptation can also be QP / rate dependent, so different adaptation rules can apply for different QPs or bitrates.

[0106] In the encoder, under the proposed algorithm, the variance is calculated for each bin (e.g., BinVar(b) in Equation (13)). Based on this information, codewords can be assigned based on each bin variance. In one embodiment, BinVar(b) may be inversely linearly mapped to the number of codewords within each bin b. In another embodiment, a non-linear mapping such as (BinVar(b)) 2 , sqrt(BinVar(b)), etc. may be used. Essentially, this approach allows the encoder to apply any codeword to each bin beyond the simpler mappings used previously. Here, the encoder uses two upper-range values M f and M a (see, for example, Figure 6C) or three upper-range values M f , 32 or M a (see, for example, Figure 6D) to assign codewords in each bin.

[0107] As an example, FIG. 6E shows two codeword assignment methods based on the BinVar(b) value. Plot 610 shows codeword assignment using two thresholds, while plot 620 shows codeword assignment using an inverse linear mapping. Here, the codeword assignment for a certain bin is inversely proportional to its BinVar(b) value. For example, in a certain embodiment, the following code may be applied to derive the number of codewords in a specific bin (bin_cw): alpha=(minCW-maxCW) / (maxVar-minVar); beta=(maxCW*maxVar-minCW*minVar) / (maxVar-minVar); bin_cw=round(alpha*bin_var+beta);, Here, minVar represents the minimum variance across all bins, maxVar represents the maximum variance across all bins, and minCW and maxCW represent the minimum and maximum number of codewords per bin determined by the shaping model.

[0108] Refinement of Chroma QP Offset Based on Luma In Patent Document 6, an additional chroma QP offset (denoted as chromaDQP or cQPO) and a luma-based chroma residual scaler (cScale) were defined to compensate for the interaction between luma and chroma. For example: chromaQP=QP_luma+chromaQPOffset+cQPO (28) Here, chromaQPOffset represents the chroma QP offset, and QP_luma represents the luma QP for the coding unit. As shown in Patent Document 6, in a certain embodiment,

Number

Equation

[0109] If a non-linear relationship is given between the QP value derived from luma (denoted as qPi) and the final chroma QP value (denoted as Qp C ) (for example, refer to Table 8-10 "Specification of Qp as a function of qPi for ChromaArrayType equal to 1" in Non-Patent Document 4), in one embodiment, cScale may be further adjusted as follows. C For example, like Table 8-10 in Non-Patent Document 4, denote the mapping between the adjusted luma and chroma QP values as f_QPi2QPc(). Then,

[0110] chromaQP_actual = f_QPi2QPc[chromaQP]= = f_QPi2QPc[QP_luma + chromaQPOffset + cQPO] (31) To scale the chroma residuals, it is necessary to calculate the scale based on the actual difference between the actual chroma coding QP before applying cQPO and after applying cQPO: QPcBase = f_QPi2QPc[QP_luma + chromaQPOffset]; QPcFinal = f_QPi2QPc[QP_luma + chromaQPOffset + cQPO]; (32) ​ cQPO_refine = QPcFinal - QpcBase; cScale = pow(2, -cQPO_refine / 6)

[0111] In another embodiment, chromaQPOffset can also be absorbed into cScale. For example, QPcBase = f_QPi2QPc[QP_luma]; QPcFinal = f_QPi2QPc[QP_luma + chromaQPOffset + cQPO]; (33) cTotalQPO_refine = QPcFinal - QpcBase; cScale = pow(2, -cTotalQPO_refine / 6)

[0112] As an example, in one embodiment, as described in Patent Document 6: Assume that CSCALE_FP_PREC = 16 represents the precision parameter. · Forward scaling: After the chroma residual is generated, before transformation and quantization: C_Res = C_orig - C_pred C_Res_scaled = (C_Res * cScale + (1 << (CSCALE_FP_PREC - 1))) >> CSCALE_FP_PREC · Inverse scaling: After chroma inverse quantization and inverse transformation, but before reconstruction: C_Res_inv = (C_Res_scaled << CSCALE_FP_PREC) / cScale C_Reco = C_Pred + C_Res_inv;

[0113] In an alternative embodiment, the operations for in-loop chroma shaping may be expressed as follows. On the encoder side, for the residual (CxRes = CxOrg - CxPred) of the chroma component Cx (e.g., Cb or Cr) of each CU or TU,

Number

[0114] The use of cScale is not limited to chroma residual scaling for in-loop filtering. The same method can also be applied to out-of-loop filtering. For out-of-loop filtering, cScale may be used for scaling chroma samples. The operations are the same as the in-loop approach.

[0115] On the encoder side, when calculating chroma RDOQ, a lambda modifier for chroma adjustment (either when using QP offset or when using chroma residual scaling) also needs to be calculated based on a refined offset: Modifier = pow(2, -cQPO_refine / 3); New_lambda = Old_lambda / Modifier (38)

[0116] As described in Equation (35), using cScale may require division in the decoder. To simplify the decoder implementation, the same functionality can be implemented using division in the encoder and applying a simpler multiplication in the decoder. For example, cScaleInv = (1 / cScale) Then, as an example, in the encoder cResScale = CxRes * cScale = CxRes / (1 / cScale) = CxRes / cScaleInv And in the decoder CxRes = cResScale / cScale = CxRes * (1 / cScale) = CxRes * cScaleInv Do this.

[0117] In one embodiment, each lumadependent chroma scaling factor may be calculated for the corresponding lumarange in a piece - wise linear (PWL) representation rather than for each lumacoded word value. Thus, the chroma scaling factor may be stored in a smaller LUT (e.g., 16 or 32 entries), such as cScaleInv[binIdx], instead of a 1024 - entry LUT (for a 10 - bit lumacoded word) (e.g., cScale[Y]). The scaling operations on the encoder side and the decoder side may be implemented in fixed - point integer arithmetic as follows. c' = sign(c)*((abs(c)*s + 2 CSCALE_FP_PREC-1 ) >> CSCALE_FP_PREC) where c is the chroma residual, s is the chroma residual scaling factor from cScaleInv[binIdx], binIdx is determined by the corresponding average luma value, and CSCALE_FP_PREC is a constant value related to precision.

[0118] In one embodiment, the forward shaping function may be represented using N equal segments (e.g., N = 8, 16, 32, etc.), but the inverse representation will include non-linear segments. From an implementation perspective, it is desirable to have a representation of the inverse shaping function using equal segments, but forcing such a representation may cause a loss of coding efficiency. As a compromise, in one embodiment, the inverse shaping function may be constructed using a "hybrid" PWL representation that combines both equal and non-equal segments. For example, when using 8 segments, the entire range may first be divided into 2 equal segments, and then each of these may be further subdivided into 4 non-equal segments. Alternatively, the entire range may be divided into 4 equal segments, and then each of these may be divided into 2 non-equal segments. Alternatively, the entire range may first be divided into several unequal segments, and then each unequal segment may be divided into a plurality of equal segments. Alternatively, the entire range may first be divided into 2 equal segments, then each equal segment may be divided into equal sub-segments, and the segment lengths in each group of sub-segments may not be the same.

[0119] For example, but not limited to, using 1024 codewords: a) 4 segments each having 150 codewords and 2 segments each having 212 codewords, or b) 8 segments each having 64 codewords and 4 segments each having 128 codewords can be had. The general purpose of such a combination of segments is to reduce the number of comparisons necessary to identify the PWL partition index when a code value is given, thereby simplifying the hardware and software implementations.

[0120] In one embodiment, for a more efficient implementation related to chroma residual scaling, the following variations may be enabled: · Disable chroma residual scaling when separate luma / chroma trees are used. · Disable chroma residual scaling for 2×2 chroma. · For intra- and inter-coded units, use the prediction signal instead of the reconstructed signal.

[0121] As an example, a decoder (200D_D) shown in FIG. 2E for processing the luma component is provided, and FIG. 2F shows an exemplary architecture (200D_DC) for processing the corresponding chroma samples.

[0122] As shown in FIG. 2F, compared with FIG. 2E, the following changes are made when processing chroma: · Forward and inverse shaping blocks (282 and 265-3) are not used. · There is a new chroma residual scaling block (288), which substantially replaces the inverse shaping block (265-3) for luma. · The reconstruction block (285-C) is modified to handle the color residual in the original domain as described in Equation (36): CxRec = CxPred + CxRes.

[0123] From Equation (34), on the decoder side, CxResScaled represents the extracted scaled chroma residual signal after inverse quantization and transformation (before block 288), CxRes = CxResScaled * C ScaleInv is assumed to represent the rescaled chroma residual generated by the chroma residual scaling block (288) used by the reconstruction unit (285-C) to calculate CxRec = CxPred + CxRes. Here, CxPred is generated by the intra (284) or inter (280) prediction block.

[0124] The values used for the transform unit (TU) may be shared by the Cb and Cr components and can be calculated as follows: · In the intra mode, calculate the average of the intra-predicted luma values; · In the inter-mode, calculate the average of the forward-shaped inter-predicted luma values. That is, the average luma value is calculated in the shaping domain. · In the case of combined merge and intra prediction, calculate the average of the combined predicted luma values. For example, the combined predicted luma values may be calculated according to Section 8.4.6.6 of Appendix 2. · In one embodiment, avgY' TU Based on, C ScaleInv A LUT can be applied to calculate. Alternatively, a piecewise linear (PWL) representation of the shaping function is given, and the index idx to which the value avgY' TU belongs to the inverse mapping PWL can be found. · Then, C ScaleInv = cScaleInv[idx] Exemplary implementations applicable to the Versatile Video Coding codec (Non-Patent Document 8) currently under development by ITU and ISO can be found in Appendix 2 (see, for example, Section 8.5.5.1.2).

[0125] Disabling luma-based chroma residual scaling for intra-slices using a binary tree may result in some loss of coding efficiency. To improve the effect of chroma shaping, the following methods may be used: 1. The chroma scaling factor may be kept the same for the entire frame depending on the average or median of the luma sample values. This removes the TU-level dependence on luma for chroma residual scaling. 2. The chroma scaling factor can be derived using the reconstructed luma values from neighboring CTUs. 3. The encoder can derive the chroma scaling factor based on the source luma pixels and transmit it in the bitstream at the CU / CTU level (e.g., as an index to the piecewise representation of the shaping function). Then, the decoder can extract the chroma scaling factor from the shaping function without depending on the luma data. · The scale factor for the CTU can be derived and transmitted only for intra slices, but can also be used for inter slices. The additional signaling cost occurs only for intra slices and thus does not affect the coding efficiency in random access. 4. Chroma can be shaped at the frame level as luma, and the luma shaping curve is derived from the luma shaping curve based on the correlation analysis between luma and chroma. This completely eliminates the chroma residual scaling.

[0126] delta_qp Application In AVC and HEVC, it is allowed for the parameter delta_qp to modify the QP value for the coding block. In one embodiment, the luma curve in the shaper can be used to derive the delta_qp value. Based on the codeword assignment for each bin, the piecewise luma DQP value can be derived. For example: For the k-th bin, scale k =(M k / M a ); (39) lumaDQP k =INT(6*log2(scale k )) Here, INT() can be CEIL(), ROUND(), or FLOOR(). The encoder can use a function of luma, such as average(luma), min(luma), max(luma), etc., to find the luma value for that block, and then use the corresponding lumaDQP value for that block. From Equation (27), in order to obtain the rate - distortion benefit, weighted distortion is used in mode decision, W rsp (k)=scale k 2 and can be set as such.

[0127] Considerations for Reshaping and Number of Bins In typical 10 - bit video coding, it is preferable to use at least 32 bins for the shaping mapping. However, in some embodiments, a smaller number of bins, such as 16 or even 8 bins, may be used to simplify the decoder implementation. Considering that the encoder may already be using 32 bins to analyze the sequence and derive the distributed codewords, the original 32 - bin codeword distribution can be reused, and within each of the 32 bins, by adding two corresponding 16 - bin pairs, a 16 - bin codeword can be derived. That is, For i = 0 to 15 CWIn16Bin[i]=CWIn32Bin[2i]+CWIn32Bin[2i + 1]

[0128] For the chroma residual scaling factor, the codeword can simply be divided by 2 to point to a 32 - bin chromaScalingFactorLUT. For example, In32Bin

[32] ={0 0 33 38 38 38 38 38 38 38 38 38 38 38 38 38 38 33 33 33 33 33 33 33 33 33 33 33 33 33 0 0} being given, the corresponding 16 - bin CW assignment is CWIn16Bin

[16] = {0 71 76 76 76 76 76 76 71 66 66 66 66 66 66 0} This approach can be extended to handle a smaller number of bins, for example, 8 bins. For i = 0 to 7 CWIn8Bin[i] = CWIn16Bin[2 * i] + CWIn16Bin[2 * i + 1]

[0129] When using a narrow range of valid codewords (e.g., [64, 940] for a 10-bit signal, [64, 235] for an 8-bit signal), it should be noted that the mapping to codewords where the first and last bins are reserved is not considered. For example, for a 10-bit signal, with 8 bins, each bin has 1024 / 8 = 128 codewords, and the first bin is [0, 127], but since the standard codeword range is [64, 940], only the codewords [64, 127] should be considered for the first bin. If the input video has a range narrower than the full range [0, 2 bitdepth -1], a special flag (e.g., video_full_range_flag = 0) may be used to notify the decoder to take special care not to generate incorrect codewords when processing the first and last bins. This applies to both luma and chroma formatting.

[0130] As an example, but not limited to, Appendix 2 provides exemplary syntax structures and related syntax elements for supporting formatting in an ISO / ITU Video Versatile Codec (VVC) (Non-Patent Document 8) according to an embodiment using the architectures shown in FIGS. 2C, 2E, and 2F. Here, the forward formatting function includes 16 segments. References Each of the references cited in this specification is hereby incorporated by reference in its entirety. [Non-Patent Document 1] "Exploratory Test Model for HDR extension of HEVC", K. Minoo et al., MPEG output document, JCTVC-W0092 (m37732), 2016, San Diego, USA [Patent Document 2] PCT Application PCT / US2016 / 025082, In-Loop Block-Based Image Reshaping in High Dynamic Range Video Coding, G-M. Su, filing date March 30, 2016, published as WO2016 / 164235 [Patent Document 3] U.S. Patent Application 15 / 410,563, Content-Adaptive Reshaping for High Codeword representation Images, T. Lu et al., filing date Jan. 19, 2017 [Non-Patent Document 4] ITU-T H.265, "High efficiency video coding," ITU, Dec. 2016 [Patent Document 5] PCT Application PCT / US2016 / 042229, Signal Reshaping and Coding for HDR and Wide Color Gamut Signals, P. Yin et al., filing date July 14, 2016, published as WO2017 / 011636 [Patent Document 6] PCT Patent Application PCT / US2018 / 040287, Integrated Image Reshaping and Video Coding, T. Lu et al., filing date June 29, 2018 [Patent Document 7] J. Froehlich et al., "Content-Adaptive Perceptual Quantizer for High Dynamic Range Images," U.S. Patent Application Publication No. 2018 / 0041759, Feb. 08, 2018 Non-Patent Document 8 B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 3)," JVET output document, JVET-L1001, v9, uploaded, Jan. 8, 2019

[0131] Exemplary Computer System Implementation Embodiments of the present invention can be implemented using a computer system, an electronic circuit, and a system composed of components, an integrated circuit (IC) such as a microcontroller, a field programmable gate array (FPGA), or other configurable or programmable logic device (PLC), a discrete time or digital signal processor (DSP), an application specific integrated circuit (ASIC), and / or a device including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or perform instructions regarding signal shaping and encoding of images as described herein. The computer and / or IC may calculate any of a variety of parameters or values related to the signal shaping and encoding processes described herein. Embodiments of images and videos can be implemented in hardware, software, firmware, and various combinations thereof.

[0132] Certain implementations of the present invention include a computer processor that executes software instructions that cause the processor to execute the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc. can implement methods related to signal shaping and encoding of images as described above by executing software instructions in a program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any non-transitory and tangible medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to execute the methods of the present invention. The program product according to the present invention may be in any of a wide variety of non-transitory and tangible forms. The program product may include, for example, physical media such as magnetic data storage media including floppy disks, hard disk drives, optical data storage media including CD-ROMs, DVDs, electronic data storage media including ROMs, flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0133] When a component (e.g., software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, the reference to the component (including the reference to "means") shall be construed to include any component that performs the function of the described component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structure that performs the function in the exemplary embodiments shown of the present invention, as an equivalent of that component.

[0134] Equivalents, Extensions, Alternatives, and Others Thus, exemplary embodiments related to efficient signal shaping and encoding of images are described. In the foregoing specification, embodiments of the invention have been described with reference to numerous individual details that may vary from implementation to implementation. Therefore, the only and exclusive indicator of what is the invention and what is intended to be the invention by the applicant is the specific form in which the set of claims issued for this application, including any subsequent amendments thereto, is allowed. The definitions explicitly set forth in this document for the terms included in such claims govern the meaning of such terms as used in those claims. Therefore, no limitation, element, characteristic, feature, advantage or attribute not explicitly recited in the claims should limit the scope of such claims in any way. Therefore, this specification and the drawings should be considered in an illustrative rather than a limiting sense.

[0135] Listed Exemplary Embodiments The present invention includes, but is not limited to, the following enumerated example embodiments (EEE) that describe the structure, features, and functions of some parts of the present invention, and can be embodied in any form described herein. 〔EEE1〕 A method for adaptively shaping a video sequence using a processor, the method comprising: accessing, by the processor, an input image in a first codeword representation; and generating a forward shaping function that maps the pixels of the input image to a second codeword representation, the second codeword representation allowing more efficient compression than the first codeword representation, and generating the forward shaping function comprising: dividing the input image into a plurality of pixel regions; assigning each pixel region to one of a plurality of codeword bins according to a first luminance characteristic of each pixel region; calculating a bin metric for each of the plurality of codeword bins according to a second luminance characteristic of each pixel region assigned to each codeword bin; According to the bin metric and rate - distortion optimization criteria of each symbol - word bin, allocate the number of symbol - words in the second symbol - word representation to each symbol - word bin; generating the forward shaping function in response to the allocation of symbol - words in the second symbol - word representation for each of the plurality of symbol - word bins, A method. 〔EEE2〕 The method according to EEE1, wherein the first luminance characteristic of the pixel region includes the average luminance - pixel value within that pixel region. 〔EEE3〕 The method according to EEE1, wherein the second luminance characteristic of the pixel region includes the variance of the luminance - pixel values of that pixel region. 〔EEE4〕 Calculating the bin metric for a symbol - word bin includes calculating the average of the variances of the luminance - pixel values for all pixel regions assigned to that symbol - word bin, the method according to EEE3. 〔EEE5〕 Allocating the number of symbol - words in the second symbol - word representation to a symbol - word bin according to its bin metric: if no pixel region is assigned to that symbol - word bin, do not assign a symbol - word to that symbol - word bin; if the bin metric of that symbol - word bin is lower than the upper threshold value, assign a first number of symbol - words; otherwise, include assigning a second number of symbol - words to that symbol - word bin, The method according to EEE1. 〔EEE6〕 For a first symbol - word representation with a depth of B bits, a second symbol - word representation with a depth of Bo bits, and N symbol - word bins, the first number of symbol - words is M f =CEIL((2 Bo / (CW2 - CW1))*M a ), and the second number of symbol - words is M a =2 B / N, where CW1 < CW2 represents two symbol - words in [0 2 B -1], the method according to EEE5. 〔EEE7〕 CW1 = 16 × 2 (B-8) and CW2 = 235 × 2 (B-8) The method described in EEE6, where 〔EEE8〕 Determining the upper threshold value involves: Defining a set of potential threshold values; For each threshold value in the set of threshold values: Generating a forward shaping function based on that threshold value; Encoding and decoding a set of input test frames according to the shaping function and bit rate R to generate an output set of decoded test frames; Calculating an overall rate - distortion optimization (RDO) metric based on the input test frames and the decoded test frames; Selecting, as the upper threshold value, the threshold value in the set of potential threshold values for which the RDO metric is minimum, The method described in EEE5. 〔EEE9〕 Calculating the RDO metric involves: Calculating J = D + λR, where D represents an indicator of the distortion between the pixel values of the input test frames and the corresponding pixel values in the decoded test frames, and λ represents the Lagrange multiplier, The method described in EEE8. 〔EEE10〕 D is an indicator of the sum of squared differences between the corresponding pixel values of the input test frames and the decoded test frames, the method described in EEE9. 〔EEE11〕 Allocating the number of codewords in the second codeword representation to the codeword bins according to their bin metrics is based on a codeword allocation lookup table, which defines two or more threshold values that segment the range of bin metric values, and gives the number of codewords to be allocated to the bins having bin metrics within each segment, the method described in EEE1. 〔EEE12〕 A method as described in EEE11, in which a default codeword assignment to bins is given, and bins with large bin metrics are assigned fewer codewords than the default codeword assignment, while bins with small bin metrics are assigned more codewords than the default codeword assignment. 〔EEE13〕 For a first codeword representation using B bits and N bins, the default codeword assignment per bin is M a =2 B as given by 2 / N, according to the method described in EEE12. 〔EEE14〕 further comprising the step of generating shaping information in response to the forward shaping function, wherein the shaping information is: a flag indicating the minimum codeword bin index value used in the shaping reconstruction process; a flag indicating the maximum codeword bin index value used in the shaping configuration process; a flag indicating the shaping model profile type, each model profile type being associated with default bin-related parameters, or one or more delta values used to adjust the default bin-related parameters including one or more of the above, according to the method described in EEE1. 〔EEE15〕 further comprising the step of assigning a bin importance value to each codeword bin, wherein the bin importance value is: 0 if no codeword is assigned to the codeword bin; 2 if the first value of the codeword is assigned to the codeword; 1 otherwise, according to the method described in EEE5. 〔EEE16〕 Determining the upper threshold value is: dividing the luminance range of pixel values in the input image into bins; and For each bin, determining a bin - histogram value and an average bin variance value, wherein for the bin, the bin - histogram value includes the number of pixels within that bin in the total number of pixels in the image, and the average bin variance value provides a metric of the average pixel variance within that bin; Sorting the average bin variance values to generate a sorted list of average bin variance values and a sorted list of average bin variance - value indices; Calculating a cumulative density function as a function of the sorted average bin variance values based on the sorted average bin variance values based on the bin - histogram values and the sorted list of average bin variance - value indices; Determining an upper threshold based on a criterion satisfied by the value of the cumulative density function, comprising: The method described in EEE5. 〔EEE17〕 Calculating the cumulative density function is: BinVarSortDsdCDF[0]=BinHist[BinIdxSortDsd[0]]; for(int b = 1; b < PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b]=BinVarSortDsdCDF[b - 1]+BinHist[BinIdxSortDsd[b]];} including calculating, where b is the bin number, PIC_ANALYZE_CW_BINS is the total number of bins, BinVarSortDsdCDF[b] is the output of the CDF function for bin b, BinHist[i] is the bin - histogram value for bin i, and BinIdxSortDsd[] represents the sorted list of average bin variance - value indices; The method described in EEE16. 〔EEE18〕 The method described in EEE16, wherein the upper threshold is determined as the average bin variance value at which the CDF output is k% under the criterion that the average bin variance is greater than or equal to the upper threshold for k% of the pixels in the input image. 〔EEE19〕 The method described in EEE18 where k = 50. [EEE20] A method for reconstructing a shaping function in a decoder, the method comprising: Receiving an encoded bitstream syntax element characterizing a shaping model, the syntax element being: A flag indicating the minimum codeword bin index value used in the shaping configuration process, A flag indicating the maximum codeword bin index value used in the shaping configuration process, A flag indicating a shaping model profile type, the model profile type being associated with default bin-related parameters including bin importance values, or A flag indicating one or more delta bin importance values used to adjust the default bin importance values defined in the shaping model profile including one or more of: Determining, based on the shaping model profile, a default bin importance value for each bin and an assignment list of the default number of codewords assigned to each bin according to the bin importance value; For each codeword bin: Determining its bin importance value by adding its default bin importance value to its delta bin importance value; Determining the number of codewords assigned to the codeword bin based on its bin importance value and the assignment list; Generating a forward shaping function based on the number of codewords assigned to each codeword bin. Method. [EEE21] Using the assignment list to determine the number of codewords M k assigned to the k-th codeword bin, further comprising: For the k-th bin: if bin_importance[k] == 0, M k is set to 0; Otherwise, if bin_importance[k] == 2 M k = M f and otherwise M k = M a including where M a and M f are elements of the assignment list, and bin_importance[k] represents the bin importance value of the k-th bin the method described in EEE20 〔EEE522〕 A method for reconstructing encoded data in a decoder having one or more processors, the method comprising: receiving an encoded bitstream (122) including one or more encoded shaped images in a first coded word representation and metadata (207) related to shaping information for the encoded shaped images; generating an inverse shaping function (250) based on the metadata related to the shaping information, the inverse shaping function mapping pixels of the shaped image from the first coded word representation to a second coded word representation; (250) generating a forward shaping function (250) based on the metadata related to the shaping information, the forward shaping function mapping pixels of an image from the second coded word representation to the first coded word representation; extracting an encoded shaped image including one or more encoded units from the encoded bitstream, for one or more encoded units in the encoded shaped image: for an intra-coded coded unit (CU) within the encoded shaped image: generating a first shaped reconstructed sample (227) of the CU based on the shaped residual and a first shaped prediction sample in the CU; Generate a shaped loop filter output based on the first shaped and reconstructed sample and loop filter parameters (270); Apply the inverse shaping function to the shaped loop filter output to generate the decoded sample of the coded unit in the second coded word representation (265); Storing the decoded sample of the coded unit in the second coded word representation in a reference buffer; For inter-coded coded units in the coded and shaped image: Apply the forward shaping function to the predicted sample stored in the reference buffer in the second coded word representation to generate a second shaped predicted sample; Generate a second shaped and reconstructed sample of the coded unit based on the shaped residual in the coded CU and the second shaped predicted sample; Generate a shaped loop filter output based on the second shaped and reconstructed sample and loop filter parameters; Apply the inverse shaping function to the shaped loop filter output to generate the sample of the coded unit in the second coded word representation; Storing the sample of the coded unit in the second coded word representation in a reference buffer; Generating a decoded image based on the samples stored in the reference buffer, Method. 〔EEE23〕 An apparatus comprising a processor and configured to execute the method according to any one of EEE1 to 22. 〔EEE24〕 A non-transitory computer-readable storage medium storing thereon computer-executable instructions for executing a method on one or more processors according to any one of EEE1 to 22.

[0136] Appendix 1 Exemplary implementation of bubble sort void bubbleSortDsd(double* array, int*idx, int n) { int i,j; bool swapped; for(i=0; i < n-1; i++) { swapped=false; for(j=0; j <n-i-1; j++) { if(array[j]<array[j+1]) { swap(&array[j],&array[j+1]); swap(&idx[j],&idx[j+1]); swapped=true; } } if(swapped==false) break; } }

[0137] Appendix 2 As an example, this appendix provides an exemplary syntax structure and related syntax elements according to an embodiment that supports shaping in the Versatile Video Codec (VVC) (Non-Patent Document 8) currently under joint development by ISO and ITU. New syntax elements in existing draft versions are highlighted or explicitly noted. Formula numbers such as (8-xxx) indicate placeholders that will be updated as necessary in the final specification.

[0138] 7.3.2.1 In the sequence parameter set RBSP syntax [Table 5-1] [Table 5-2] [Table 5-3]

[0139] 7.3.3.1 In the general tile group header syntax

Table 6-1

Table 6-2

Table 6-3

[0140] Add a new syntax table and tile reshaper model.

Table 7

[0141] In the general sequence parameter set RBSP semantics, add the following semantics. That sps_reshaper_enabled_flag is equal to 1 specifies that the reshaper is used in the coded video sequence (CVS). That sps_reshaper_enabled_flag is equal to 0 specifies that the reshaper is not used in the CVS.

[0142] In the tile group header syntax, add the following semantics. The fact that tile_group_reshaper_model_present_flag is equal to 1 indicates that it is specified to exist in the tile_group_reshaper_model() tile group header. The fact that tile_group_reshaper_model_present_flag is equal to 0 specifies that tile_group_reshaper_model() does not exist in the tile group header. If tile_group_reshaper_model_present_flag does not exist, it is presumed to be equal to 0. The fact that tile_group_reshaper_enabled_flag is equal to 1 specifies that the reshaper is enabled in the current tile group. The fact that tile_group_reshaper_enabled_flag is equal to 0 specifies that the reshaper is not enabled for the current tile group. If tile_group_resharper_enable_flag does not exist, it is presumed to be equal to 0. The fact that tile_group_reshaper_chroma_residual_scale_flag is equal to 1 specifies that chroma residual scaling is enabled for the current tile group. The fact that tile_group_reshaper_chroma_residual_scale_flag is equal to 0 specifies that chroma residual scaling is not enabled for the current tile group. If tile_group_reshaper_chroma_residual_scale_flag does not exist, it is presumed to be equal to 0.

[0143] Add the tile_group_reshaper_model() syntax reshaper_model_min_bin_idx specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of reshaper_model_min_bin_idx ranges from 0 to MaxBinIdx, inclusive. The value of MaxBinIdx is equal to 15. reshaper_model_delta_max_bin_idx specifies the value obtained by subtracting the maximum bin (or piece) index MaxBinIdx from the maximum bin index used in the reshaper configuration process. The value of reshaper_model_max_bin_idx is set equal to MaxBinIdx - reshaper_model_delta_max_bin_idx. One plus reshaper_model_bin_delta_abs_cw_prec_minus1 specifies the number of bits used for the representation of the syntax reshaper_model_bin_delta_abs_CW[i]. reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshaper_model_bin_delta_abs_CW[i] as follows. · When reshaper_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. · Otherwise (when reshaper_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is a negative value. If reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. The variable RspDeltaCW[i] = (1 - 2 * reshaper_model_bin_delta_sign_CW[i]) * reshaper_model_bin_delta_abs_CW[i]; The variable RspCW[i] is derived in the following steps: The variable OrgCW is set equal to (1 << BitDepthY) / (MaxBinIdx + 1). · When reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx RspCW[i] = OrgCW + RspDeltaCW[i] · Otherwise, RspCW[i] = 0 The value of RspCW[i] ranges from 32 to 2*OrgCW - 1 when the value of BitDepthYis is equal to 10. For variable InputPivot[i] in the range of i from 0 to MaxBinIdx + 1 including both ends, it is derived as follows. InputPivot[i] = i * OrgCW For variable ReshapePivot[i] in the range of i from 0 to MaxBinIdx + 1 including both ends, and variables ScaleCoef[i] and InvScaleCoeff[i] in the range of i from 0 to MaxBinIdx including both ends, they are derived as follows: shiftY = 14 ReshapePivot[0] = 0; for(i = 0; i <= MaxBinIdx; i++){ ReshapePivot[i + 1] = ReshapePivot[i] + RspCW[i] ScaleCoef[i] = (RspCW[i] * (1 << shiftY) + (1 << (Log2(OrgCW) - 1))) >> (Log2(OrgCW)) if(RspCW[i] == 0) InvScaleCoeff[i] = 0 else InvScaleCoeff[i] = OrgCW * (1 << shiftY) / RspCW[i] } For variable ChromaScaleCoef[i] in the range of i from 0 to MaxBinIdx including both ends, it is derived as follows: ChromaResidualScaleLut

[64] = {16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024} shiftC = 11 · When (RspCW[i] == 0), then ChromaScaleCoef[i] = (1 << shiftC) · Otherwise (RspCW[i] != 0), ChromaScaleCoef[i] = ChromaResidualScaleLut[Clip3(1, 64, RspCW[i] >> 1) - 1] Note: In an alternative implementation, the scaling of luma and chroma can be unified, and ChromaResidualScaleLut[] can be made unnecessary. In that case, chroma scaling can be implemented as follows: shiftC = 11 · When (RspCW[i] == 0), then ChromaScaleCoef[i] = (1 << shiftC) · Otherwise (RspCW[i] != 0), the following applies: BinCW = BitDepth Y > 10? (RspCW[i] >> (BitDepth Y - 10)) : BitDepth Y < 10? (RspCW[i] << (10 BitDepth Y )) : RspCW[i]; ChromaScaleCoef[i] = OrgCW * (1 << shiftC) / BinCW[i]

[0144] In the weighted sample prediction process for combined merge and intra prediction, add the following. The addition is highlighted. 8.4.6.6 Weighted Sample Prediction Process for Combined Merge and Intra Prediction The input to this process is as follows: · The width cbWidth of the current coding block · The height cbHeight of the current coding block · Two (cbWidth) × (cbHeight) arrays, preSamplesInter and preSamplesIntra · The intra prediction mode predModeIntra · The variable cIdx that specifies the color component index The output of this process is a (cbWidth) × (cbHeight) array predSamplesComb of predicted sample values. The variable bitDepth is derived as follows. · If cIdx is equal to 0, bitDepth is set equal to BitDepth Y is set equal to. · Otherwise, bitDepth is set equal to BitDepth C is set equal to. For x = 0..cbWidth - 1 and y = 0..cbHeight - 1, the predicted sample predSamplesComb[x][y] is derived as follows: · The weight w is derived as follows: · If predModeIntra is INTRA_ANGULAR50, w is specified in Table 8-10 with nPos equal to y and nSize equal to cbHeight. · Otherwise, when predModeIntra is INTRA_ANGULAR18, w is specified in Table 8-10 with nPos equal to x and nSize equal to cbWidth. · Otherwise, w is set to 4.

Table 8

Table 9

[0145] Add the following to the picture reconstruction process. 8.5.5 Picture Composition Process The input to this process is as follows: · The top - left sample of the current block, at the position (xCurr, yCurr) specified relative to the top - left sample of the current picture component · Variables nCurrSw and nCurrSh specifying the width and height of the current block respectively · Variable cIdx specifying the color component of the current block · (nCurrSw)×(nCurrSh) array predSamples specifying the predicted samples of the current block · (nCurrSw)×(nCurrSh) array resSamples specifying the residual samples of the current block Depending on the value of the color component cIdx, the following assignments are made: · When cIdx is equal to 0, recSamples corresponds to the reconstructed picture sample array S L and the function clipCidx1 is Clip1 Ycorresponds to · Otherwise, when cIdx is equal to 1, recSamples is the reconstructed chroma sample array S Cb corresponds to, and the function clipCidx1 is Clip1 C corresponds to · Otherwise (cIdx is equal to 2), recSamples is the reconstructed chroma sample array S Cr corresponds to, and the function clipCidx1 is Clip1 C corresponds to

Table 10

[0146] [Outer 2] TIFF0007706626000027.tif10115 This section defines picture reconstruction using the mapping process. Picture reconstruction using the mapping process for luma sample values is defined in 8.5.5.1.1. Picture reconstruction using the mapping process for chroma sample values is defined in 8.5.5.1.2. 8.5.5.1 Picture reconstruction using the mapping process for luma sample values The inputs to this process are as follows: · The (nCurrSw)×(nCurrSh) array predSamples specifying the luma prediction samples of the current block · Specify the luma residual samples of the current block as an (nCurrSw)×(nCurrSh) array resSamples · The output for this process is as follows: · A mapped luma prediction sample array predMapSamples of (nCurrSw)×(nCurrSh) · A reconstructed luma sample array recSamples of (nCurrSw)×(nCurrSh) predMapSamples is derived as follows: If (CuPredMode[xCurr][yCurr]==MODE_INTRA) || (CuPredMode[xCurr][yCurr]==MODE_INTER && mh_intra_flag[xCurr][yCurr]) predMapSamples[xCurr+i][yCurr+j]=predSamples[i][j] where i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outer 3] TIFF0007706626000028.tif9115 Otherwise ((CuPredMode[xCurr][yCurr]==MODE_INTER &&!mh_intra_flag[xCurr][yCurr])), the following applies: shiftY = 14 idxY = predSamples[i][j] >> Log2(OrgCW) predMapSamples[xCurr+i][yCurr+j]=ReshapePivot[idxY] +(ScaleCoeff[idxY]*(predSamples[i][j]-InputPivot[idxY]) +(1<<(shiftY-1))) >> shiftY where i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outer 4] TIFF0007706626000029.tif9115 recSamples is derived as follows: recSamples[xCurr + i][yCurr + j] = Clip1 Y (predMapSamples[xCurr + i][yCurr + j] + resSamples[i][j]]) Here, i = 0..nCurrSw - 1, j = 0..nCurrSh - 1 [Outside 5] TIFF0007706626000030.tif9115 8.5.5.1.2 Picture reconstruction using the mapping process for chroma sample values The input to this process is as follows: · A (nCurrSw x 2) x (nCurrSh x 2) mapped array predMapSamples that specifies the mapped luma prediction samples of the current block · A (nCurrSw) x (nCurrSh) array PredSamples that specifies the chroma prediction samples of the current block · A (nCurrSw) x (nCurrSh) array resSamples that specifies the chroma residual samples of the current block The output of this process is the reconstructed chroma sample array recSamples. recSamples is derived as follows: · If (!tile_group_reshaper_chroma_residual_scale_flag || ((nCurrSw) x (nCurrSh) <= 4)) recSamples[xCurr + i][yCurr + j] = Clip1 C (predSamples[i][j] + resSamples[i][j]) Here, i = 0..nCurrSw - 1, j = 0..nCurrSh - 1 [Outside 6] TIFF0007706626000031.tif9115 · Otherwise (tile_group_reshaper_chroma_residual_scale_flag && ((nCurrSw) x (nCurrSh) > 4)), the following applies: The variable varScale is derived as follows: 1. invAvgLuma = Clip1 Y ((Σ i Σ j predMapSamples[(xCurr << 1) + i][(yCurr << 1) + j] + nCurrSw * nCurrSh * 2) / (nCurrSw * nCurrSh * 4)) 2. The variable idxYInv is derived by calling the identification of the function index for each section as specified in [External 7] TIFF0007706626000032.tif8115 3. varScale = ChromaScaleCoef[idxYInv] varScale = ChromaScaleCoef[idxYInv] recSamples is derived as follows: · When tu_cbf_cIdx[xCurr][yCurr] is equal to 1, the following applies: shiftC = 11 recSamples[xCurr + i][yCurr + j] = ClipCidx1(predSamples[i][j] + Sign(resSamples[i][j]) * ((Abs(resSamples[i][j]) * varScale + (1 << (shiftC - 1))) >> shiftC)) where i = 0..nCurrSw - 1, j = 0..nCurrSh - 1 [External 8] TIFF0007706626000033.tif9115 · Otherwise (when tu_cbf_cIdx[xCurr][yCurr] is equal to 0), recSamples[xCurr + i][yCurr + j] = ClipCidx1(predSamples[i][j]) [External 9] TIFF0007706626000034.tif9115

[0147] 8.5.6 Picture inverse mapping process This section is invoked when the value of tile_group_reshaper_enabled_flag is equal to 1. The input is the reshaped picture luma sample array S L and the output is the modified reshaped picture luma sample array S' after the inverse mapping process L . The inverse mapping process for luma sample values is defined in 8.4.6.1. 8.5.6.1 Picture inverse mapping process for luma sample values The input to this process is the luma position (xP, yP) that specifies the luma sample position relative to the top - left luma sample of the current picture. The output of this process is the inverse - mapped luma sample value invLumaSample. The value of invLumaSample is derived by applying the following ordered steps: 1. The variable idxYInv is derived by calling the identification of the per - segment function index as specified in section 8.5.6.2 using the input of the luma sample value S L [xP][yP]. 2. The value of reshapeLumaSample is derived as follows: shiftY = 14 invLumaSample = InputPivot[idxYInv]+(InvScaleCoeff[idxYInv]*(S L [xP][yP]-ReshapePivot[idxYInv]) +(1<<(shiftY - 1)))>>shiftY [Outer 10] TIFF0007706626000035.tif91153.clipRange = ((reshaper_model_min_bin_idx>0) && (reshaper_model_max_bin_idx<MaxBinIdx)); When clipRange is equal to 1, the following applies: minVal = 16 << (BitDepth Y - 8) maxVal = 235 << (BitDepth Y - 8) invLumaSample = Clip3(minVal, maxVal, invLumaSample) Otherwise (when clipRange is equal to 0), the following applies: invLumaSample = ClipCidx1(invLumaSample) 8.5.6.2 Identification of the function index for each section of the luma component The input to this process is the luma sample value S. The output of this process is the index idxS that identifies the piece [segment] to which the sample S belongs. The variable idxS is derived as follows: for (idxS = 0, idxFound = 0; idxS <= MaxBinIdx; idxS++) { if ((S < ReshapePivot[idxS + 1]) { idxFound = 1 break } } Note that an alternative implementation for finding the identifying idxS is as follows: if (S < ReshapePivot[reshaper_model_min_bin_idx]) idxS = 0 else if (S >= ReshapePivot[reshaper_model_max_bin_idx]) idxS = MaxBinIdx else idxS = findIdx(S, 0, MaxBinIdx + 1, ReshapePivot[]) function idx = findIdx(val, low, high, pivot[]) { if (high - low <= 1) idx = low else { mid = (low + high) >> 1 if (val < pivot[mid]) high = mid else low = mid idx = findIdx(val, low, high, pivot[]) } }

[0148] Some aspects will be described. [Aspect 1] A method for reconstructing encoded video data by one or more processors, the method comprising: receiving an encoded bitstream (122) including one or more encoded reconstructed images in an input codeword representation; receiving reconstruction metadata (207) for the one or more encoded reconstructed images in the encoded bitstream; generating a forward reconstruction function (282) based on the reconstruction metadata, the forward reconstruction function mapping pixels of an image from a first codeword representation to the input codeword representation; generating an inverse reconstruction function (265-3) based on the reconstruction metadata or the forward reconstruction function, the inverse reconstruction function mapping pixels of the reconstructed image from the input codeword representation to the first codeword representation; extracting from the encoded bitstream an encoded reconstructed image including one or more encoded units; for an intra-coded coded unit (intra CU) in the encoded reconstructed image: generating a reconstructed sample of the intra CU based on the reconstructed residual and the intra-predicted reconstructed sample within the intra CU (285); applying the inverse reconstruction function (265-3) to the reconstructed sample of the intra CU to generate a decoded sample of the intra CU in the first codeword representation; Apply a loop filter (270) to the decoded samples of the intra CU to generate output samples of the intra CU; Store the output samples of the intra CU in a reference buffer; For an inter-coded CU (inter CU) in the coded and reconstructed image: Apply the forward reconstruction function (282) to the inter-prediction samples stored in the reference buffer in the first codeword representation to generate reconstructed prediction samples for the inter CU in the input codeword representation; Generate reconstructed and reconstructed samples of the inter CU based on the reconstructed residual in the inter CU and the reconstructed prediction samples for the inter CU; Apply the inverse reconstruction function (265-3) to the reconstructed and reconstructed samples of the inter CU to generate decoded samples of the inter CU in the first codeword representation; Apply a loop filter (270) to the decoded samples of the inter CU to generate output samples of the inter CU; Store the output samples of the inter CU in the reference buffer; Generating a decoded image in the first codeword representation based on output samples in the reference buffer. Method. 〔Aspect 2〕 Generating the reconstructed and reconstructed samples (RecSample) of the intra CU is: RecSample=(Res+IpredSample) including calculating where Res represents the reconstructed residual samples in the intra CU in the input codeword representation, and IpredSample represents the intra-predicted reconstructed prediction samples in the input codeword representation. The method according to Aspect 1. 〔Aspect 3〕 Generating the reshaped and reconstructed sample (RecSample) of the inter-CU: RecSample = (Res + Fwd(PredSample)) including calculating: where Res represents the reshaped residual in the inter-CU in the input coded word representation, Fwd() represents the forward reshaping function, and PredSample represents the inter-prediction sample in the first coded word representation. The method according to aspect 1. [Aspect 4] Generating the output sample (RecSampleInDPB) stored in the reference buffer: RecSampleInDPB = LPF(Inv(RecSample)) including calculating: where Inv() represents the inverse reshaping function and LPF() represents the loop filter. The method according to aspect 2 or 3. [Aspect 5] For the chroma residual samples in the inter-coded CU (inter-CU) in the input coded word representation, further: determining a chroma scaling factor based on the luma pixel values in the input coded word representation and the reshaping metadata; multiplying the chroma scaling factor by the chroma residual samples in the inter-CU to generate scaled chroma residual samples in the inter-CU in the first coded word representation; generating a reconstructed chroma sample of the inter-CU based on the scaled chroma residual in the inter-CU and the chroma inter-prediction samples stored in the reference buffer to generate a decoded chroma sample of the inter-CU; applying the loop filter (270) to the decoded chroma sample of the inter-CU to generate an output chroma sample of the inter-CU. storing the output chroma samples of the inter-CU in the reference buffer; The method according to aspect 1. [Aspect 6] The method according to aspect 5, wherein in the intra mode, the chroma scaling factor is based on the average of the intra-predicted luma values. [Aspect 7] The method according to aspect 5, wherein in the inter mode, the chroma scaling factor is based on the average of the inter-predicted luma values in the input coded word representation. [Aspect 8] The shaping metadata includes: a first parameter indicating the number of bins used to represent the first coded word representation; a second parameter indicating the minimum bin index used in shaping; a first set of parameters indicating the absolute delta coded word values for each bin in the input coded word representation; a second set of parameters indicating the signs of the delta coded word values for each bin in the input coded word representation. The method according to aspect 1. [Aspect 9] The method according to aspect 1, wherein the forward shaping function is reconstructed as a per-segment linear function having linear segments derived from the shaping metadata. [Aspect 10] A method for adaptively shaping a video sequence by a processor, the method comprising: accessing, by the processor, an input image in a first coded word representation; and generating a forward shaping function for mapping pixels of the input image to a second coded word representation, the second coded word representation allowing more efficient compression than the first coded word representation, and generating the forward shaping function includes: dividing the input image into a plurality of pixel regions; assigning each pixel region to one of a plurality of coded word bins according to a first luminance characteristic of each pixel region; Calculating a bin metric for each of the plurality of coded word bins according to a second luminance characteristic of each of the pixel regions assigned to each of the plurality of coded word bins; Allocating the number of coded words in the second coded word representation to each of the plurality of coded word bins according to the bin metric and rate - distortion optimization criteria of each of the plurality of coded word bins; Generating the forward shaping function in response to the allocation of the coded words in the second coded word representation for each of the plurality of coded word bins, including Method. 〔Aspect 11〕 The method according to aspect 10, wherein the first luminance characteristic of the pixel region includes an average luminance pixel value within the pixel region. 〔Aspect 12〕 The method according to aspect 10, wherein the second luminance characteristic of the pixel region includes the variance of the luminance pixel values of the pixel region. 〔Aspect 13〕 Calculating the bin metric for a coded word bin includes calculating an average of the variances of the luminance pixel values for all pixel regions assigned to the coded word bin, according to the method of aspect 12. 〔Aspect 14〕 Allocating the number of coded words in the second coded word representation to a coded word bin according to its bin metric includes: If no pixel region is assigned to the coded word bin, no coded word is assigned to the coded word bin; If the bin metric of the coded word bin is lower than an upper threshold value, a first number of coded words is assigned; Otherwise, including assigning a second number of coded words to the coded word bin, The method according to aspect 10. 〔Aspect 15〕 An apparatus having a processor and configured to execute the method according to any one of aspects 1 to 14. 〔Aspect 16〕 A non-transitory computer-readable storage medium storing computer-executable instructions for performing a method by one or more processors according to any one of aspects 1 to 14.

Claims

Claim 1 An apparatus for reconstructing encoded video data, the apparatus comprising: a processor; and an input for receiving an encoded bitstream including one or more encoded reconstructed pictures in a shaped codeword representation and shaping parameters for the one or more encoded reconstructed pictures in the encoded bitstream, the shaping parameters including parameters for generating a forward shaping function and a luma-based chroma residual scalar based on the shaping parameters, the forward shaping function being reconstructed as a per-segment linear function having a linear segment derived by the shaping parameters, the shaping parameters comprising: a delta index parameter for determining, using the processor, an active maximum bin index used for shaping, the active maximum bin index being smaller than a predefined maximum bin index, determining the active maximum bin index used for shaping including calculating a difference between the predefined maximum bin index and the delta index parameter; a minimum bin parameter indicating a minimum bin index used in the shaping; an absolute delta codeword value for each active bin in the shaped codeword representation; and the sign of the absolute delta codeword value for each active bin in the shaped codeword representation, the apparatus. Claim 2 An apparatus for generating shaping parameters for an encoded bitstream, the apparatus comprising: a processor; and an input for receiving a sequence of video pictures in an input codeword representation, wherein for a picture in the sequence of video pictures, the processor: applies a forward shaping function to luma pixel values to generate shaped luma pixel values in a shaped codeword representation, the forward shaping function being reconstructed as a per-segment linear function having a linear segment derived by shaping parameters; applies a luma-based scalar to chroma residuals to generate scaled chroma residuals; and generates shaping parameters for the shaped codeword representation. Generate an encoded bitstream based on at least the shaped luma pixel values and the scaled chroma residuals. The shaping parameters are: A delta index parameter that determines an active maximum bin index used for shaping, where the active maximum bin index is less than or equal to a predefined maximum bin index, and determining the active maximum bin index used for shaping includes calculating a difference between the predefined maximum bin index and the delta index parameter, the delta index parameter; A minimum bin parameter indicating a minimum bin index used in the shaping; An absolute delta codeword value for each active bin in the shaped codeword representation; And a sign of the absolute delta codeword value for each active bin in the shaped codeword representation. Device. **Claim 3** A method for transmitting a bitstream generated by a video encoding method, the video encoding method comprising: Receiving a sequence of video pictures in an input codeword representation; For a picture in the sequence of video pictures, applying a forward shaping function to luma pixel values to generate shaped luma pixel values in a shaped codeword representation; Applying a luma-based scalar to chroma residuals to generate scaled chroma residuals; Generating shaping parameters for the shaped codeword representation; Generating an encoded bitstream based on at least the shaped luma pixel values and the scaled chroma residuals, the shaping parameters being: A delta index parameter that determines an active maximum bin index used for shaping, where the active maximum bin index is less than a predefined maximum bin index, the delta index parameter; A minimum bin parameter indicating a minimum bin index used in the shaping; An absolute delta codeword value for each active bin in the shaped codeword representation; A method including signs of the absolute delta codeword values for each active bin in the shaped codeword representation. Method.

Citation Information

Patent Citations

  • Signal reshaping and coding for HDR and wide color gamut signals

    WO2017011636A1

  • High dynamic range video coding architectures with multiple operating modes

    WO2017019818A1

  • Signal reshaping for high dynamic range signals

    WO2017024042A2

  • Encoding and decoding reversible production-quality single-layer video signals

    WO2017165494A2