Image shaping in video coding using rate distortion optimization

The RDO-based signal shaping technique optimizes codeword assignment in video encoding to enhance compression efficiency and image quality, addressing inefficiencies in existing technologies by minimizing distortion and bitrate.

JP2026136335APending Publication Date: 2026-08-25DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026092453
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-01-14
Filing Date
2026-06-02
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing video encoding technologies struggle with inefficient compression and image quality trade-offs, particularly when dealing with higher bit depths and dynamic ranges, leading to suboptimal subjective and objective metrics like PSNR.

Method used

Implementing a Rate-Distortion Optimization (RDO) based signal shaping technique that adaptively assigns codewords to bins based on luminance variance, optimizing the number of codewords in each bin to minimize distortion and bitrate, using threshold-based methods and Lagrangian optimization for efficient compression.

Benefits of technology

Improves both subjective visual quality and objective metrics such as PSNR and Bjontegaard PSNR, while maintaining efficient compression, applicable to standard and high dynamic range signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136335000001_ABST
    Figure 2026136335000001_ABST
Patent Text Reader

Abstract

This provides image shaping in video coding using rate distortion optimization. [Solution] Given a sequence of images in a first codeword representation, a method, process, and system for image shaping using rate-distortion optimization are presented. The shaping enables encoding the images in a second codeword representation that allows for more efficient compression than using the first codeword representation. A syntax for signaling shaping parameters is also presented.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 792,122 filed on 14 January 2019, U.S. Provisional Patent Application No. 62 / 782,659 filed on 20 December 2018, U.S. Provisional Patent Application No. 62 / 772,228 filed on 28 November 2018, U.S. Provisional Patent Application No. 62 / 739,402 filed on 1 October 2018, U.S. Provisional Patent Application No. 62 / 726,608 filed on 4 September 2018, U.S. Provisional Patent Application No. 62 / 691,366 filed on 28 June 2018, and U.S. Provisional Patent Application No. 62 / 630,385 filed on 14 February 2018, each of which is incorporated herein by reference in its entirety.

[0002] technology This invention broadly relates to image and video coding. More specifically, one embodiment of the invention relates to image formatting in video coding. [Background technology]

[0003] In 2013, the MPEG group of the International Organization for Standardization (ISO), in collaboration with the International Telecommunication Union (ITU), published the first draft of the HEVC (also known as H.265) video coding standard (Non-Patent Literature 4). More recently, the group announced a call for evidence to support the development of next-generation coding standards that will provide improved coding performance compared to existing video coding technologies.

[0004] As used herein, the term “bit depth” refers to the number of pixels used to represent one of the color components of an image. Traditionally, images were encoded with 8 bits per pixel, per color component (for example, 24 bits per pixel), but current architectures may now support higher bit depths, such as 10 bits, 12 bits, and above.

[0005] In traditional image pipelines, captured images are quantized using a nonlinear opto-electronic function (OETF), which converts linear scene light into a nonlinear video signal (e.g., gamma-coded RGB or YCbCr). The signal is then processed on the receiver by an electro-optical transfer function (EOTF) that converts the video signal values ​​into output screen color values ​​before being displayed on a screen. Such nonlinear functions include the traditional “gamma” curve described in ITU-R Rec. BT.709 and BT.2020, the “PQ” (perceptual quantization) curve described in SMPTE ST2084, and the “HybridLog-gamma” or “HLG” curve described in Rec. ITU-R BT.2100. [Overview of the project]

[0006] As used herein, the term “forward reshaping” refers to the process of mapping a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma, PQ, or HLG) to an image of the same or different bit depth and different codeword distribution or representation, from sample to sample, or codeword to codeword. Reshaping enables improved compressibility or improved image quality at a fixed bitrate. For example, but not limited to, reshaping may be applied to 10-bit or 12-bit PQ encoded HDR video to improve encoding efficiency in a 10-bit video encoding architecture. After decompressing the reshaped signal, the receiver can apply a “de-shaping function” to restore the signal to its original codeword distribution. As recognized herein, as development for the next generation of video encoding standards begins, improved techniques for integrated image reshaping and encoding are desired. The method of the present invention may be applicable to a variety of video content, including but not limited to standard dynamic range (SDR) and / or high dynamic range (HDR) content.

[0007] The methods described in this section are methods that could have been pursued, but they are not necessarily methods that were conceived or pursued previously. Therefore, unless otherwise indicated, none of the methods described in this section should be assumed to be eligible as prior art simply because they are included in this section. Similarly, any problems identified with respect to one or more methods should not be assumed to have been recognized in any prior art based on this section, unless otherwise specified. [Brief explanation of the drawing]

[0008] Embodiments of the present invention are shown in the accompanying drawings as examples, not limiting, where similar reference numerals refer to similar elements.

[0009] [Figure 1A]This demonstrates an exemplary process for a video distribution pipeline;

[0010] [Figure 1B] This illustrates an exemplary process for data compression using conventional signal shaping techniques;

[0011] [Figure 2A] An exemplary architecture for an encoder using hybrid in-loop shaping according to one embodiment of the present invention is shown.

[0012] [Figure 2B] An exemplary architecture for a decoder using hybrid in-loop shaping according to one embodiment of the present invention is shown.

[0013] [Figure 2C] This document presents an exemplary architecture for intraCU decoding using formatting according to one embodiment.

[0014] [Figure 2D] This document presents an exemplary architecture for interCU decoding using formatting according to one embodiment.

[0015] [Figure 2E] An exemplary architecture for intra-CU decoding in an inter-encoded slice is shown, according to one embodiment for lumer or chromer processing.

[0016] [Figure 2F] An exemplary architecture for intra-CU decoding in an inter-encoded slice is shown, according to one embodiment for chroma processing.

[0017] [Figure 3A] An exemplary process for encoding video using a shaping architecture according to one embodiment of the present invention is shown.

[0018] [Figure 3B] An exemplary process for decoding video using a shaping architecture according to one embodiment of the present invention is shown.

[0019] [Figure 4] An exemplary process for reassigning codewords within a shaping domain according to one embodiment of the present invention is shown.

[0020] [Figure 5] An exemplary process for deriving a shaping threshold according to one embodiment of the present invention is shown.

[0021] [Figure 6A] Figure 5 shows the process and exemplary data plots for deriving a shaped threshold according to one embodiment of the present invention. [Figure 6B] Figure 5 shows the process and exemplary data plots for deriving a shaped threshold according to one embodiment of the present invention. [Figure 6C] Figure 5 shows the process and exemplary data plots for deriving a shaped threshold according to one embodiment of the present invention. [Figure 6D] Figure 5 shows the process and exemplary data plots for deriving a shaped threshold according to one embodiment of the present invention.

[0022] [Figure 6E] An example of codeword assignment according to bin distribution, based on one embodiment of the present invention, is shown. [Modes for carrying out the invention]

[0023] This paper describes signal shaping and encoding techniques for compressing images using rate-distortion optimization (RDO). Numerous specific details are included in the following description for illustrative purposes to provide a full understanding of the invention. However, it will be apparent that the invention can be carried out without these specific details. On the other hand, well-known structures and apparatus are not described in exhaustive detail to avoid unnecessarily obscuring, burying, or obscuring the invention.

[0024] overview The exemplary embodiments described herein relate to signal shaping and encoding for video. In an encoder, a processor receives an input image of a first codeword representation to be shaped into a second codeword representation, the second codeword representation allowing for more efficient compression than the first codeword representation, and the processor further generates a forward shaping function that maps the pixels of the input image to the second codeword representation. To generate the forward shaping function, the encoder: divides the input image into a plurality of pixel regions; assigns each of the pixel regions to one of a plurality of codeword bins according to a first luminance characteristic of each pixel region; calculates the bin metric of each of the plurality of codeword bins according to a second luminance characteristic of each of the pixel regions assigned to each codeword bin; assigns some codewords in the second codeword representation to each codeword bin according to the bin metric and rate-distortion optimization criterion of each codeword bin; and generates the forward shaping function in response to the assignment of codewords in the second codeword representation to each of the plurality of codeword bins.

[0025] In another embodiment, in the decoder, the processor receives an encoded bitstream syntax element characterizing a shaping model, the syntax element being a flag indicating the minimum codeword bin index value to be used in the shaping construction process, a flag indicating the maximum codeword bin index value to be used in the shaping construction process, and a flag indicating a shaping model profile type, the model profile type including one or more flags indicating a default bin relation parameter, including a bin importance value, or a delta bin importance value to be used to adjust the default bin importance value defined in the shaping model profile. Based on the shaping model profile, the processor determines a default bin importance value for each bin and an assignment list of default numbers to be assigned to each bin according to the bin importance value. Then, for each codeword bin, the processor: The bin importance value is determined by adding the default bin importance value to the delta bin importance value; Based on the bin importance value and assignment list of the bins, determine the number of codewords that should be assigned to the codeword bin; A forward shaping function is generated based on the number of codewords assigned to each codeword bin.

[0026] In another embodiment, the decoder receives an encoded bitstream which includes one or more encoded formatted images in a first codeword representation and metadata relating to formatting information about the encoded formatted images. The processor generates a de-shaping function and a forward-shaping function based on metadata related to the shaping information. Here, the de-shaping function maps the pixels of the shaping image from a first codeword representation to a second codeword representation, and the forward-shaping function maps the pixels of the image from the second codeword representation to the first codeword representation. The processor extracts an encoded shaping image from the encoded bitstream that contains one or more encoded units. Here, regarding one or more encoded units in the encoded shaping image: For the shaped intra-encoded coding unit (CU) in the encoded shaped image, the processor is: A first shaped reconstructed sample of the CU is generated based on the shaped residuals and shaped predicted sample in the CU; Based on the first shaped reconstructed sample and loop filter parameters, a shaped loop filter output is generated; The inverse shaping function is applied to the shaped loop filter output to generate decoded samples of the coded unit in the second codeword representation; The decoded sample of the coded unit in the second codeword representation is stored in the reference buffer; For the reshaped interencoded encoding unit in the encoded reshaped image, the processor: Apply a forward shaping function to the predicted samples stored in the reference buffer in the second codeword representation to generate the second shaped predicted samples; A second shaped reconstructed sample of the encoded unit is generated based on the shaped residual and the second shaped predicted sample in the encoded CU; Based on the second shaped reconstructed sample and loop filter parameters, a shaped loop filter output is generated; The de-shaping function is applied to the shaped loop filter output to generate samples of the coded unit in the second codeword representation; The sample of the encoded unit in the second codeword representation is stored in a reference buffer. Finally, the processor generates a decoded image based on the sample stored in the reference buffer.

[0027] In another embodiment, in the decoder, the processor receives an encoded bitstream which includes one or more encoded shaped images in the input codeword representation and shaping metadata (207) for one or more encoded shaped images in the encoded bitstream. The processor generates a forward shaping function (282) based on the shaping metadata, which maps the pixels of the image from the first codeword representation to the input codeword representation. The processor generates an inverse shaping function (265-3) based on the shaping metadata or the forward shaping function, which maps the pixels of the shaped image from the input codeword representation to the first codeword representation. The processor extracts an encoded shaped image from the encoded bitstream which includes one or more encoded units, where: For intra-encoded coding units (intraCUs) in encoded shaped images, the processor: Based on the shaped residuals and shaped predicted samples in the intraCU, a shaped reconstructed sample (285) of the intraCU is generated; The inverse reshaping function (265-3) is applied to the reshaped reconstructed sample of the intraCU to generate the decoded sample of the intraCU in the first codeword representation; A loop filter (270) is applied to the decoded samples of the intraCU to generate the output samples of the intraCU; The output sample of the intraCU is stored in the reference buffer; Regarding inter-encoded CUs (interCUs) in encoded shaped images, the processor: The forward shaping function (282) is applied to the inter-prediction samples stored in the reference buffer in the first codeword representation to generate shaped prediction samples for the inter-CU in the input codeword representation; A shaped reconstructed sample of the interCU is generated based on the shaped residuals of the interCU and the shaped predicted sample for the interCU; The inverse reshaping function (265-3) is applied to the reshaped reconstructed sample of interCU to generate the decoded sample of interCU in the first codeword representation; A loop filter (270) is applied to the decoded sample of the interCU to generate the output sample of the interCU; The output sample of the interCU is stored in the reference buffer; Based on the output samples in the reference buffer, a decoded image in a first codeword representation is generated.

[0028] Example of a video distribution processing pipeline Figure 1A illustrates an exemplary process of a conventional video distribution pipeline (100) showing the various stages from video capture to video content display. A sequence of video frames (102) is captured or generated using an image generation block (105). The video frames (102) may be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) may be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production phase (110), the video data (107) is edited to provide a video production stream (112).

[0029] Next, the video data from the production stream (112) is provided to the processor in block (115) for post-production editing. Post-production editing in block (115) may include adjusting or correcting the color or brightness in individual areas of the image to improve image quality or to achieve a particular appearance of the image according to the creative intent of the video creator. This is sometimes called “color timing” or “color grading.” Other edits (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, etc.) may be performed in block (115) to give the final version of the production (117) for distribution. During post-production editing (115), the video image is displayed on a reference display (125).

[0030] Following post-production (115), the video data of the final production (117) may be delivered to an encoding block (120) for delivery to a decoding and playback device such as a downstream television set, set-top box, or cinema. In some embodiments, the encoding block (120) may include audio and video encoders, such as those defined by ATSC, DVB, DVD, Blu-ray, and other distribution formats, to generate an encoded bitstream (122). At the receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to produce a decoded signal (132) that represents an identical or near approximation of the signal (117). The receiver may be mounted on a target display (140) which may have completely different characteristics from the reference display (125). In that case, a display management block (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137).

[0031] Signal Reshaping Figure 1B shows an exemplary process for signal shaping according to Patent Document 2. Given an input frame (117), a forward shaping block (150) analyzes the input and encoding constraints and generates a codeword mapping function that maps the input frame (117) to a requantized output frame (152). For example, the input (117) may be encoded according to some kind of electro-optical transfer function (EOTF) (e.g., gamma). In some embodiments, metadata may be used to convey information about the shaping process to downstream devices (e.g., decoders). As used herein, the term “metadata” refers to any auxiliary information transmitted as part of an encoded bitstream to assist the decoder in rendering a decoded image. Such metadata includes, but is not limited to, color space or color gamut information, reference display parameters, and auxiliary signal parameters, as described herein.

[0032] Following encoding (120) and decoding (130), the decoded frame (132) may be processed by a back-shaping function (160) that transforms the requantized frame (132) back to the original EOTF domain (e.g., gamma) for further downstream processing such as the display management process (135) described above. In some embodiments, the back-shaping function (160) may be integrated with the dequantizer in the decoder (130), for example, as part of the dequantizer in an AVC or HEVC video decoder.

[0033] As used herein, the term “reshaper” may refer to a forward or inverse reshaper function used when encoding and / or decoding a digital image. An example of a reshaper function is discussed in Patent Document 2, which proposes an in-loop block-based image reshaping method for high dynamic range video coding. Its design allows for block-based reshaping within the coding loop, but at the cost of increased complexity. Specifically, the design requires maintaining two sets of decoded image buffers: one set for inversely reshaped (or unreshaped) decoded pictures that can be used for both unreshaped prediction and output to the display, and the other set for forwardly reshaped decoded pictures that can be used only for prediction with reshaping. Forwardly reshaped decoded pictures can be computed on the fly, but the cost of complexity is very high, especially for inter-prediction (motion compensation with sub-pixel interpolation). Generally, display-picture-buffer (DPB) management is complex and requires very careful attention, so a simplified method for encoding video is desired, as the inventors understand.

[0034] Patent Document 6 presented further shaping-based codec architectures, including architectures with external out-of-loop shaping, intra-only shaping within the loop, architectures with in-loop shaping for predictive residuals, and hybrid architectures combining both intra-in-loop shaping and intra-residual shaping. The main goal of these proposed shaping architectures is to improve subjective visual quality. Therefore, many of these approaches result in worse objective metrics, particularly the well-known peak signal-to-noise ratio (PSNR) metric.

[0035] This invention proposes a novel shaper based on Rate-Distortion Optimization (RDO). In particular, when the target distortion metric is MSE (Mean Square Error), the proposed shaper improves both subjective visual quality and commonly used objective metrics based on PSNR, Bjontegaard PSNR (BD-PSNR), or Bjontegaard rate (BD-Rate). It should be noted that, without loss of generality, any of the proposed reconstruction architectures can be applied to one or more of the luminance component, chroma component, or a combination of the luminance and chroma components.

[0036] Shaping based on Rate-Distortion Optimization Consider a shaped video signal represented by the bit depth of B bits in a certain color component (for example, B=10 for Y, Cb and / or Cr), totaling 2 B There are 1 available codewords. Desired codeword range [0 2 B Consider dividing ] into N segments or bins. k This is the case where, given a target bitrate R, the distortion D between the source picture and the decoded or reconstructed picture is minimized. After shaping and mapping Let D represent the number of codewords in the k-th segment or bin. Without loss of generality, D may be expressed as a measure of the sum of squared errors (SSE) between the corresponding pixel values ​​of the source input (Source(i,j)) and the reconstructed picture (Recon(i,j)). D=SSE=Σ i,j Diff(i,j) 2 (1) Here, Diff(i,j)=Source(i,j)-Recon(i,j) That is the case.

[0037] The optimization problem can be rewritten as follows: Given a bitrate R, find M such that D is minimized. kFind (k = 0, 1, …, N - 1). Here,

Number

[0038] Various optimization methods can be used to find the solution, but the optimal solution can be very complex for real-time encoding. In the present invention, a sub-optimal but more practical analytical solution is proposed.

[0039] Without loss of generality, consider an input signal represented by a bit depth of B bits (e.g., B = 10). Here, the codewords are uniformly divided into N bins (e.g., N = 32). By default, each bin has M a = 2 B / N codewords (e.g., for N = 32 and B = 10, M a = 32) assigned to it. Next, a more efficient codeword assignment based on RDO is demonstrated through examples.

[0040] As used herein, the term "narrow range" [CW1, CW2] indicates a continuous range of codewords between codewords CW1 and CW2, which is a subset of the entire dynamic range [0 2 B - 1]. For example, in one embodiment, the narrow range can be defined as [16 * 2 (B-8) , 235 * 2 (B-8) (e.g., for B = 10, the narrow range includes the values [64 940]). Assuming that the bit depth of the output signal is B o , if the dynamic range of the input signal is within the narrow range, in what is denoted as "default shaping", the signal can be extended to the full range [0 2 Bo - 1]. Then, each bin has M f = CEIL((2 Bo / (CW2 - CW1)) * M a ) codewords, or, for our example, when B o = B = 10, M f=CEIL((1024 / (940-64))*32)=There are 38 codewords. Here, CEIL(x) represents the ceiling function that maps x to the smallest integer greater than or equal to x. Without loss of generality, in the example below, for simplicity, B o Let's assume that = B.

[0041] For the same quantization parameter (QP), increasing the number of codewords in a bin is equivalent to allocating more bits to encode the signal in that bin, and therefore equivalent to decreasing SSE or improving PSNR. However, a uniform increase in codeword allocation in each bin may not yield better results than unrefined encoding because the PSNR gain may not outweigh the increase in bit rate. In other words, this is not a good trade-off with respect to RDO. Ideally, we want to allocate more codewords only to bins that give the best trade-off with respect to RDO, i.e., to produce a significant decrease in SSE (increase in PSNR) at the cost of a small increase in bit rate.

[0042] In one embodiment, RDO performance is improved through adaptive piecewise shaping mapping. This method can be applied to any type of signal, including standard dynamic range (SDR) and high dynamic range (HDR) signals. Using the simple case described above as an example, the objective of the present invention is to achieve M for each codeword segment or codeword bin. a Individual or M f This involves assigning one of the codewords.

[0043] In an encoder, given N codeword bins for the input signal, the luminance variance of the average of each bin can be approximated as follows: • For each bin, the sum of block variances (var bin (k) and counter (c bin Initialize (k) to 0. For example, for k=0,1,...,N-1, var bin (k=0, cbin Let (k) = 0. • Divide the picture into L*L non-overlapping blocks (for example, L=16). For each picture block, calculate the lumen mean of the block and the lumen variance of block i (e.g., Luma_mean(i) and Luma_var(i)). Based on the average luma of a block, it is assigned to one of N bins. In one embodiment, if Luma_mean(i) is within the k-th segment of the input dynamic range, the total bin luminance variance for the k-th bin is incremented by the luma variance of the newly assigned block, and the counter for that bin is incremented by 1. That is, if the i-th pixel region belongs to the k-th bin: var bin (k=var bin (k)+Luma_var(i); (2) c bin (k=c bin (k)+1 For each bin, assuming the counter is not equal to 0, calculate the average luminance variance for that bin by dividing the sum of the block variances within that bin by the counter. Alternatively, c bin If (k) is not 0, var bin (k=var bin (k) / c bin (k) (3) Let's assume that.

[0044] Those skilled in the art will understand that alternative metrics other than luminance variance may be applied to characterize subblocks. For example, the standard deviation of luminance values, weighted luminance variance or luminance values, peak luminance, etc., may be used.

[0045] In one embodiment, the following pseudocode illustrates how an encoder may adjust bin assignments using a calculated metric for each bin. For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k) <TH U (4) M k =M f ; else M k =M a ; / / (Note: This means that each bin is at least M a To ensure that each codeword has its own unique identifier. / / Alternatively, M a (You may assign one more codeword.) end Here, TH U This represents a predetermined upper threshold.

[0046] In another embodiment, the assignment may be performed as follows: For the kth bin, if there are no pixels in the bin M k =0; else if TH0 bin (k) <TH1(5) M k =M f ; else M k =M a ; end Here, TH0 and TH1 represent predetermined lower and upper threshold values.

[0047] Another embodiment For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k>TH L (6) M k =M f ; ​else M k =M a ; end Here, TH L This represents a predetermined lower threshold.

[0048] The above example uses two pre-selected numbers M. f and M a This shows how to select the number of codewords for each bin from the threshold (for example, TH). U or TH L The threshold can be determined, for example, by optimizing rate distortion through an exhaustive search. The threshold may be adjusted based on the quantization parameter value (QP). In one embodiment, at B=10, the threshold may be in the range of 1,000 to 10,000.

[0049] In one embodiment, to expedite processing, the threshold may be determined using the Lagrangian optimization method from a fixed set of values, for example, {2,000, 3,000, 4,000, 5,000, 6,000, 7,000}. For example, for each TH(i) value in the set, a compression test can be performed with a fixed QP using a predefined training clip, and the value of the objective function J can be calculated. J is J(i) = D + λR (7) It is defined as follows. The optimal threshold can then be defined as the TH(i) value in the set that minimizes J(i).

[0050] In a more general example, a lookup table (LUT) can be predefined. For example, in Table 1, the first row is the possible bin metrics (e.g., var bin(k) defines a set of thresholds that divide the entire range of the value into segments, and the second row defines the corresponding number of codewords (CW) to be assigned in each segment. In one embodiment, one rule for constructing such a LUT is: if the bin variance is too large, it may be necessary to spend many bits to reduce the SSE, thus M a A smaller codeword (CW) value can be assigned. If the bin variance is very small, M a A larger CW value can be assigned.

[0051] Table 1: Exemplary LUTs for codeword assignment based on bin dispersion thresholds [Table 1]

[0052] Using Table 1, the mapping of thresholds to codewords can be generated as follows: For the kth bin, if there are no pixels in the bin M k =0; else if var bin (k) <TH0 M k =CW0; else if TH0 bin (k) <TH1(8) M k =CW1; ... else if TH p-1 bin (k) <TH p Mk=CW p ; ... else if var bin (k>TH q-1 Mk=CW q ; end

[0053] ​​For example, given two thresholds and three codeword assignments, for B=10, in one embodiment, TH0=3,000, CW0=38, TH1=10,000, CW1=32, and CW2=28.

[0054] In another embodiment, two thresholds TH0 and TH1 may be selected as follows: a) Consider TH1 to be a very large number (maybe even infinity), and select TH0 from a predetermined set of values ​​using, for example, the RDO optimization of equation (7). Given TH0, now define a second set of possible values ​​for TH1, for example the set {10,000, 15,000, 20,000, 25,000, 30,000}, and apply equation (7) to find the optimal value. This approach can be performed sequentially and iteratively with a limited number of thresholds, or until convergence occurs.

[0055] After assigning codewords to bins according to one of the previously defined methods, M k The sum of the values ​​is the maximum (2) available codewords. B It may be noted that this may exceed ) or there may be unused codewords. If there are unused codewords, one may simply decide to do nothing, or one may assign the unused codewords to a specific bin. On the other hand, if the algorithm assigns more codewords than are available, one may renormalize the CW value, for example, by M k You might want to readjust the values. Alternatively, you could generate a forward shaping function using existing Mk values, but then (Σ k M k ) / 2 B The output values ​​of the shaping function may be readjusted by scaling. An example of a codeword reassignment technique is also described in Patent Document 7.

[0056] Figure 4 shows an exemplary process for assigning codewords to the shaped domain according to the RDO technique described earlier. In step 405, the desired shaped dynamic range is divided into N bins. After the input image has been divided into non-overlapping blocks (step 410), for each block: Step 415 calculates its luminance properties (e.g., mean and variance). Step 420 assigns each image block to one of N bins. Step 425 calculates the mean luminance variance in each bin. Given the values ​​calculated in step 425, in step 430, each bin is assigned several codewords according to one or more thresholds, for example, using one of the codeword assignment algorithms shown in equations (4) to (8). Finally, in step (435), the final codeword assignments may be used to generate a forward shaping function and / or an inverse shaping function.

[0057] In one embodiment, without limitation, for example, a forward LUT (FLUT) can be constructed using the following C code: tot_cw=2 B ; hist_lens=tot_cw / N; for (i=0;i <N;i++) { double temp=(double) M[i] / (double)hist_lens; / / M[i] is M k Compatible for (j=0;j <hist_lens;j++) { CW_bins_LUT_all[i*hist_lens + j]=temp; } Y_LUT_all[0]=CW_bins_LUT_all[0]; for (i=1;i <tot_cw;i++) { Y_LUT_all[i]=Y_LUT_all[i-1]+CW_bins_LUT_all[i]; } for (i=0;i <tot_cw;i++) { FLUT[i]=Clip3(0,tot_cw-1,(Int)(Y_LUT_all[i]+0.5)); }

[0058] In one embodiment, the inverse LUT can be constructed as follows: low = FLUT[0]; high=FLUT[tot_cw-1]; first = 0; last=tot_cw-1; for (i=1;i <tot_cw;i++) if(FLUT[0] <FLUT[i]) { first = i - 1; break; } for (i=tot_cw-2;i>=0;i--) if(FLUT[tot_cw-1]>FLUT[i]) { last = i + 1; break; } for (i=0;i <tot_cw;i++) if(i<=low) { ILUT[i] = first; } else if(i>=high) { ILUT[i] = last; } else { for(j=0;j <tot_cw-1;j++) if(FLUT[j]>=i) { ILUT[i]=j; break; }} }

[0059] Syntactically, syntax proposed in previous applications, such as the piecewise polynomial mode or parametric model in Patent Documents 5 and 6, can be reused. Table 2 shows such an example for N=32 for equation (4). Table 2: Syntax for formatting using the first parametric model [Table 2] Here, reshaper_model_profile_type This specifies the profile type used in the shaper construction process. A given profile has a number of bins, a default bin importance or priority value, and a default codeword assignment (for example, M a and / or M f It may provide information about the default values ​​used, such as the value itself. reshaper_model_scale_idx This specifies the index value of the scale factor (denoted as ScaleFactor) used in the shaper construction process. The value of ScaleFactor allows for better control over the shaper function to improve overall coding efficiency. reshaper_model_min_bin_idx This specifies the minimum bin index used in the reshaper construction process. The value of reshaper_model_min_bin_idx should be in the range of 0 to 31, including both ends. reshaper_model_max_bin_idx This specifies the maximum bin index used in the reshaper construction process. The value of reshaper_model_max_bin_idx should be in the range of 0 to 31, including both ends. reshaper_model_bin_profile_delta[i] This specifies the delta value used to adjust the profile of the i-th bin during the reshaper construction process. The value of reshaper_model_bin_profile_delta[i] should be in the range of 0 to 1, including both ends.

[0060] Table 3 shows another embodiment with an alternative, more efficient syntax representation. Table 3: Syntax of reconstruction using the second parametric model [Table 3] Here, resharper_model_delta_max_bin_idx This is set to be equal to the maximum allowable bin index (e.g., 31) minus the maximum bin index used in the shaper construction process. reshaper_model_num_cw_minus1 Adding 1 to this value specifies the number of codewords to be transmitted. reshaper_model_delta_abs_CW[i] This specifies the i-th absolute delta codeword value. reshaper_model_delta_sign_CW[i] This specifies the sign for the i-th delta codeword. And: reshaper_model_delta_CW[i]=(1-2*reshaper_model_delta_sign_CW[i])* reshaper_model_delta_abs_CW [i]; reshaper_model_CW[i] = 32 + reshaper_model_delta_CW[i]. reshaper_model_bin_profile_delta[i]This specifies the delta value used to adjust the profile of the i-th bin in the shaper construction process. The value of reshaper_model_bin_profile_delta[i] is in the range of 0 to 1 when reshaper_model_num_cw_minus1 is equal to 0. The value of reshaper_model_bin_profile_delta[i] is in the range of 0 to 2 when reshaper_model_num_cw_minus1 is equal to 1. When reshaper_model_bin_profile_delta[i] is set to be equal to 0, CW=32; when reshaper_model_bin_profile_delta[i] is set to be equal to 1, CW=reshaper_model_CW[0]; and when reshaper_model_bin_profile_delta[i] is set to be equal to 2, CW=reshaper_model_CW[1]. In one embodiment, reshaper_model_num_cw_minus1 It is allowed to be greater than 1, and thereby, reshaper_model_num_cw_minus1 and reshaper_model_bin_profile_delta[i] This allows the signal to be transmitted via ue(v) for more efficient encoding.

[0061] In another embodiment, the number of codewords per bin may be explicitly defined, as shown in Table 4. Table 4: Syntax for formatting using the third model [Table 4] reshaper_model_number_bins_minus1 Adding 1 to this specifies the number of bins used for the luma component. In some embodiments, it may be more efficient for the number of bins to be a power of 2. Then the total number of bins is given by its log2 representation, for example log2_reshaper_model_number_bins_minus_minus1 It can be expressed using alternative parameters such as the following. For example, in the case of 32 bins, log2_reshaper_model_number_bins_minus1 = 4 reshaper_model_bin_delta_abs_cw_prec_minus1Adding 1 to it specifies the number of bits used for the representation of the syntax reshaper_model_bin_delta_abs_CW[i]. reshaper_model_bin_delta_abs_CW[i] Specifies the absolute delta codeword value for the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] Specifies the sign of reshaper_model_bin_delta_abs_CW[i] as follows: · When reshaper_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] has a positive value. · Otherwise (when reshaper_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] has a negative value. When reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. The variable RspDeltaCW[i]=(1 - 2*reshaper_model_bin_delta_sign_CW [i])*reshaper_model_bin_delta_abs_CW [i]; The variable OrgCW is set to (1<<BitDepthY) / (resharper_model_number_bins_minus_1 + 1); The variable RspCW[i] is derived as follows: if reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx then RspCW[i]=OrgCW + RspDeltaCW[i] else, RspCW[i]=0

[0062] In one embodiment, assuming one of the aforementioned examples, for example, the sign word assignment according to Equation (4), an example of how to define the parameters in Table 2 includes the following: First, assign "bin importance" as follows: For the k-th bin, if M k =0; bin_importance = 0; else if M k ==M f bin_importance = 2; (9) else bin_importance = 1; end

[0063] As used herein, the term "bin importance" is a value assigned to each of the N coded word bins, indicating the importance of all the coded words within that bin relative to the other bins in the shaping process.

[0064] In one embodiment, the default_bin_importance (default bin importance) from reshaper_model_min_bin_idx to reshaper_model_max_bin_idx may be set to 1. The value of reshaper_model_min_bin_idx is set to the smallest bin index having a non-zero M k The value of reshaper_model_max_bin_idx is set to the largest bin index having a non-zero M k For each bin within [reshaper_model_min_bin_idx reshaper_model_max_bin_idx], reshaper_model_bin_profile_delta is the difference between bin_importance and default_bin_importance.

[0065] An example of how to use the proposed parametric model to construct the Forward Reshaping LUT (FLUT) and the Inverse Reshaping LUT (ILUT) is shown below. 1) Divide the luminance range into N bins (e.g., N = 32). 2) Derive the bin - importance index for each bin from the syntax. For example, for the k - th bin, if reshaper_model_min_bin_idx <= k <= reshaper_model_max_bin_idx bin_importance[k]=default_bin_importance[k]+reshaper_model_bin_profile_delta[k]; else bin_importance[k]=0; 3) Automatically pre - assign the codewords based on the bin importance: for the k - th bin, if bin_importance[k]==0 M k =0; else if bin_importance[k]==2 M k =M f ; else M k =M a ; end 4) Construct the forward reshaping LUT based on the codeword assignment for each bin by accumulating the assigned codewords for each bin. The sum should be below the total codeword budget (e.g., 1024 for a 10 - bit full - range case). (See the first C code for example.) 5) Construct the inverse reshaping LUT (see the first C code for example).

[0066] From a syntax perspective, alternative methods are also applicable. The key is, explicitly or implicitly, the number of codewords in each bin (e.g., M for k = 0, 1, 2, …, N - 1) kThis involves specifying the number of codewords in each bin. In one embodiment, the number of codewords in each bin can be explicitly specified. In another embodiment, codewords can be specified in a different way. For example, the number of codewords in a bin can be determined using the difference between the number of codewords in the current bin and the previous bin (e.g., M_Delta(k)=M(k)-M(k-1)). In another embodiment, the number of most commonly used codewords (e.g., M M ) can be specified, and the number of codewords in each bin can be set to the difference in the number of codewords in each bin from this number (for example, M_Delta(k)=M(k)-M M It can be expressed as ).

[0067] In one embodiment, two shaping methods are supported. One is referred to as the "default shaping method," and M f This is assigned to all bins. The other is referred to as the "adaptive reshaper," which applies the adaptive reshaper described above. These two methods can be signaled to the decoder using a special flag, for example, sps_reshaper_adaptive_flag, as described in Patent Document 6 (for example, sps_reshaper_adaptive_flag=0 is used for the default reshaper and sps_reshaper_adaptive_flag=1 is used for the adaptive reshaper).

[0068] The present invention is applicable to any shaper proposed in Patent Document 6, such as an external shaper, an intra-only in-loop shaper, an intra-loop residual shaper, or an intra-loop hybrid shaper. As an example, Figures 2A and 2B show an exemplary architecture for hybrid intra-loop shaping according to an embodiment of the present invention. In Figure 2A, the architecture combines elements from both the intra-only in-loop shaping architecture (top of the figure) and the intra-loop residual architecture (bottom of the figure). Under this architecture, shaping is applied to picture pixels for intra-slices, while for inter-slices, shaping is applied to the predicted residual. In encoder (200_E), two new blocks are added to a conventional block-based encoder (e.g., HEVC): a block (205) for estimating the forward shaping function (e.g., according to Figure 4), a forward picture shaping block (210-1), and a forward residual shaping block (210-2). This applies forward shaping to one or more color components of the input video (117) or predicted residuals. In some embodiments, these two operations may be performed as part of a single image shaping block. Parameters (207) related to the determination of the inverse shaping function in the decoder may be passed to the lossless encoder block of the video encoder (e.g., CABAC 220) so that they can be embedded in the encoded bitstream (122). In intra mode, intra prediction (225-1), transform and quantization (T&Q), inverse transform and inverse quantization (Q -1 &T -1 All of these use a shaped picture. In both modes, the picture stored in DPB(215) is always in inverse shaping mode, which requires an inverse picture shaping block (e.g., 265-1) or an inverse residual shaping block (e.g., 265-2) before the loop filters (270-1, 270-2). As shown in Figure 2A, an intra / inter-slice switch allows switching between the two architectures depending on the type of slice being encoded. In another embodiment, in-loop filtering for intra-slices may be performed before inverse shaping.

[0069] In decoder (200_D), the following new normative blocks are added to the conventional block-based decoder: block (250) (shaper decode) which reconstructs the forward and backward shaping functions based on the encoded shaping function parameters (207); block (265-1) which applies the inverse shaping function to the decoded data; and block (265-2) which applies both the forward and inverse shaping functions to generate the decoded video signal (162). For example, in (265-2), the reconstructed value is given by Rec = ILUT(FLUT(Pred) + Res), where FLUT represents the forward reshaping LUT and ILUT represents the inverse reshaping LUT.

[0070] In some embodiments, the operations associated with blocks 250 and 265 may be combined into a single processing block. As shown in Figure 2B, an intra / inter-slice switch allows switching between two modes depending on the type of slice in the encoded video picture.

[0071] Figure 3A shows an exemplary process (300_E) for encoding video using a shaping architecture (e.g., 200_E) according to one embodiment of the present invention. If no shaping is enabled (path 305), encoding (335) proceeds as known with prior art encoders (e.g., HEVC). If shaping is enabled (path 310), the encoder may have the option of applying a predetermined (default) shaping function (315) or adaptively determining a new shaping function (325) based on picture analysis (320) (as shown in Figure 4, for example). After encoding the image using the shaping architecture (330), the remainder of the encoding follows the same steps as a conventional encoding pipeline (335). If adaptive shaping is used (312), metadata related to the shaping function is generated as part of the “encode shaping function” step (327).

[0072] Figure 3B shows an exemplary process (300_D) for decoding video using a shaping architecture (e.g., 200_D) according to one embodiment of the present invention. If no shaping is enabled (path 340), the picture is decoded (350) and an output frame is generated (390) in the same way as a conventional decoding pipeline. If shaping is enabled (path 360), the decoder decides whether to apply a predetermined (default) shaping function (375) or to adaptively determine a shaping function based on received parameters (e.g., 207) (380). Following decoding using the shaping architecture (385), the remainder of the decoding follows a conventional decoding pipeline.

[0073] As described in Patent Document 6 and above herein, the forward-shaping LUT (FwdLUT) may be constructed by integration, while the inverse-shaping LUT may be constructed based on a back-mapping using the forward-shaping LUT (FwdLUT). In some embodiments, the forward LUT may be constructed using piecewise linear interpolation. In the decoder, inverse shaping can be performed directly using the back-LUT, or again by linear interpolation. The piecewise linear LUT is constructed based on an input pivot point and an output pivot point.

[0074] Let (X1,Y1) and (X2,Y2) be the two input pivot points and the corresponding output values ​​for each bin. The input value X between X1 and X2 can be interpolated by the following equation: Y = ((Y2 - Y1) / (X2 - X1)) * (X - X1) + Y1 In a fixed-point implementation, the above expression can be rewritten as follows: Y=((m*X+2 FP_PREC-1 )>>FP_PREC)+c Here, m and c represent scalars and offsets for linear interpolation, and FP_PREC is a constant related to fixed-point precision.

[0075] For example, FwdLUT can be constructed as follows: variables: lutSize=(1< <BitDepth Y ) Let's assume the following. Variables: binNum=reshaper_model_number_bins_minus1+1 and binLen=lutSize / binNum Let's assume that. For the i-th bin, its two pivots (e.g., X1 and X2) can be derived as X1 = i * binLen and X2 = (i + 1) * binLen. And: binsLUT[0]=0; for(i=0; i <reshaper_model_number_bins_minus1+1; i++){ binsLUT[(i+1)*binLen]=binsLUT[i*binLen]+RspCW[i]; Y1 = binsLUT[i * binLen]; Y2 = binsLUT[(i+1)*binLen]; scale=((Y2-Y1)*(1< <FP_PREC)+(1<<(log2(binLen)-1)))> >(log2(binLen)); for(j=1;j <binLen;j++){ binsLUT[i*binLen+j]=Y1+((scale*j+(1<<(FP_PREC-1)))>>FP_PREC); } }

[0076] FP_PREC defines the fixed-point precision of the fractional part of the variable (e.g., FP_PREC=14). In some embodiments, binsLUT[] may be calculated with a higher precision than FwdLUT. For example, the binsLUT[] value may be calculated as a 32-bit integer, while FwdLUT may be a binsLUT value clipped to 16 bits.

[0077] Adaptive threshold derivation As mentioned above, during formatting, codeword assignment is done using one or more thresholds (e.g., TH, TH). U , TH L They may be adjusted using (etc.). In some embodiments, such thresholds may be adaptively generated based on content characteristics. Figure 5 shows an exemplary process for deriving such thresholds according to one embodiment. 1) In step 505, the luminance range of the input image is divided into N bins (for example, N=32). For example, N may also be denoted as PIC_ANALYZE_CW_BINS. 2) In step 510, image analysis is performed to calculate the luminance characteristics for each bin. For example, the percentage of pixels in each bin (BinHist[b], denoted as b=1,2,…,N) may be calculated. Here, BinHist[b] = 100 * (total number of pixels in bin b) / (total number of pixels in picture) (10) As discussed earlier, another good metric for image characteristics is the mean variance (or standard deviation) of pixels in each bin, denoted as BinVar[b]. BinVar[b] may be calculated in “block mode,” as described in the steps leading to equations (2) and (3). Alternatively, block-based calculations can be refined into pixel-based calculations. For example, let vf(i) be the variance associated with the group of pixels surrounding the i-th pixel in an m × m neighborhood window (e.g., m=5) centered on the i-th pixel. For example,

number

number

number

[0078] 3) In step 515, the mean bin variances (and their corresponding indices) are sorted in descending order, for example, but not limited to these. For example, the sorted BinVar values ​​may be stored in BinVarSortDsd[b], and the sorted bin indices may be stored in BinIdxSortDsd[b]. As an example, this process may be written using C code as follows: for(int b=0; b <PIC_ANALYZE_CW_BINS; b++ / / Initialization (unsorted) { BinVarSortDsd[b]=BinVar[b]; BinIdxSortDsd[b]=b; } / / Sort (see code example in Appendix 1) bubbleSortDsd(BinVarSortDsd,BinIdxSortDsd,PIC_ANALYZE_CW_BINS); Figure 6A shows an illustrative plot of sorted mean-bin variance factors.

[0079] 4) Given the bin histogram values ​​calculated in step 510, step 520 calculates and stores the cumulative density function (CDF) according to the sorted mean bin variance order. For example, if the CDF is stored in the array BinVarSortDsdCDF[b], in one embodiment: BinVarSortDsdCDF[0]=BinHist[BinIdxSortDsd[0]]; for(int b=1; b <PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b]=BinVarSortDsdCDF[b-1]+BinHist[BinIdxSortDsd[b]]; } An exemplary plot (605) of the calculated CDF based on the data in Figure 6A is shown in Figure 6B. A pair of CDF values ​​versus sorted mean bin variances {x=BinVarSortDsd[b], y=BinVarSortDsdCDF[b]} can be interpreted as either "y% of pixels in the picture have a variance greater than or equal to x" or "(100-y)% of pixels in the picture have a variance less than x". 5) Finally, in step 525, given CDF, BinVarSortDsdCDF[BinVarSortDsd[b]] as a function of sorted mean bin variances, a threshold can be defined based on the bin variance and cumulative proportion.

[0080] Examples for determining a single threshold or two thresholds are shown in Figures 6C and 6D, respectively. One threshold (e.g., T H When only ) is used, for example, T H "The k% of pixels is T H It may also be defined as "mean variance having". Then T H This can be calculated by finding the intersection of the CDF plot (605) at k% (e.g., 610) (for example, the BinVarSortDsd[b] value where BinVarSortDsdCDF=k%). For example, as shown in Figure 6C, for k=50, T H = 2.5. Next, BinVar[b] <THをもつビンのためにM f Assign a codeword to each bin, and for bins where BinVar[b]≧TH, M aA number of codewords can be assigned. As a general rule, it is preferable to assign more codewords to bins with smaller variance (for example, for a 10-bit video signal with 32 bins, M f >32>M a ).

[0081] When using two thresholds, TH L and TH U An example of selecting TH is shown in Figure 6D. For example, without loss of generality, L 80% of pixels have vf≧TH L It may also be defined as a variance having (TH in this example) L =2.3), TH U 10% of all pixels have vf ≥ TH U It may be defined as a variance having (TH in this example) U =3.5). Given these thresholds, BinVar[b] <TH L For a bottle with M f The codewords are defined as BinVar[b]≧TH U For a bottle with M a Individual codewords can be assigned. TH L and TH U For bins with BinVar between , the original number of codewords per bin (for example, 32 for B=10) may be used.

[0082] The above technique can be easily extended to cases where there are two or more thresholds. This relationship is given by the number of codewords (M) f M a It can also be used to adjust (etc.). As a general rule, for low-variance bins, more codewords should be allocated to boost the PSNR (and lower the MSE); for high-variance bins, fewer codewords should be allocated to save bits.

[0083] In one embodiment, a set of parameters for specific content (for example, TH L , TH U M a, M f If they are manually obtained, for example, through comprehensive manual parameter tuning, this automatic method can be applied to design a decision tree for categorizing each content in order to automatically set the optimal manual parameters. For example, content categories include: movies, TV, SDR, HDR, comics, nature, action, etc.

[0084] To reduce complexity, in-loop shaping can be restricted in various ways. When in-loop shaping is adopted in a video coding standard, these restrictions should be normative to ensure the simplification of the decoder. For example, in certain embodiments, in-loop shaping may be disabled for certain block coding sizes. For example, when nTbW*nTbH < TH, the intra and inter shaping modes in the inter-slice can be disabled. Here, the variable nTbW specifies the conversion block width, and the variable nTbH specifies the conversion block height. For example, for TH = 64, blocks of sizes 4×4, 4×8, and 8×4 are disabled for both intra and inter mode shaping in the inter-coded slice (or tile).

[0085] Similarly, in another embodiment, chroma residual scaling based on luma in the intra mode in the inter-coded slice (or tile) may be disabled. Or it may also be disabled when it is effective to have separate luma and chroma split trees.

[0086] Interaction with other coding tools Loop filtering Patent Document 6 describes that loop filtering can operate in either the original pixel domain or the reshaped pixel domain. In one embodiment, it is proposed that loop filtering is performed in the original pixel domain (after picture reshaping). For example, in hybrid in-loop reshaping architectures (200_E and 200_D), for intra-pictures, de-shaping (265-1) must be applied before loop filtering (270-1).

[0087] Figures 2C and 2D show alternative decoder architectures (200B_D and 200C_D) in which deformation (265) is performed immediately after loop filtering (270) and before storing the decoded data in a decoded picture buffer (DPB) (260). In the proposed embodiments, compared to the architecture of 200_D, the inverse residual shaping formula for inter-slices is modified and deformation is performed after loop filtering (270) (e.g., via the InvLUT() function or a lookup table). Thus, for both intra-slices and inter-slices, deformation is performed after loop filtering, and for both intra-encoded and inter-encoded CUs, the reconstructed pixels before loop filtering are in the shaping domain. After deformation (265), all output samples stored in the reference DPB are in the original domain. Such architectures allow both slice-based and CTU-based adaptations for in-loop shaping.

[0088] As shown in Figures 2C and 2D, in one embodiment, loop filtering (270) is performed in the shaping domain for both intra-encoded and inter-encoded CUs, and inverse picture shaping (265) occurs only once, thus resulting in a unified, simpler architecture for both intra-encoded and inter-encoded CUs.

[0089] To decode the intra-encoded CU (200B_D), an intra-prediction (225) is performed on the reshaped neighboring pixels. Given the residual Res and the predicted sample PredSample, the reconstructed sample (227) is derived as follows: RecSample = Res + PredSample (14) Given the reconstructed sample (227), loop filtering (270) and inverse picture shaping (265) are applied to derive the RecSampleInDPB sample, which is stored in DPB (260). Here, RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(Res+PredSample))) (15) Here, InvLUT() represents the deformation function or deformation lookup table, and LPF() represents the loop filtering operation.

[0090] In conventional encoding, inter / intra-mode determination is based on calculating a distortion function (dfunc()) between the original sample and the predicted sample. Examples of such functions include the sum of squared errors (SSE), the sum of absolute differences (SAD), and others. When using shaping, on the encoder side (not shown), CU prediction and mode determination are performed in the shaping domain. That is, for mode determination, distortion=dfunc(FwdLUT(SrcSample)-RecSample) (16) Here, FwdLUT() represents the forward formatting function (or LUT), and SrcSample represents the original image sample.

[0091] For the intercoded CU, on the decoder side (e.g., 200C_D), interpretation is performed using a reference picture of the unshaped domain in the DPB. Then, in reconstruction block 275, the reconstructed pixel (267) is derived as follows: RecSample=(Res+FwdLUT(PredSample)) (17) Given a reconstructed sample (267), loop filtering (270) and inverse picture shaping (265) are applied to derive the RecSampleInDPB sample to be stored in the DPB. Here, RecSampleInDPB=InvLUT(LPF(RecSample)))=InvLUT(LPF(Res+FwdLUT(PredSample)))) (18)

[0092] On the encoder side (not shown), intra prediction is performed in the shaping domain as follows: Res=FwdLUT(SrcSample)-PredSample (19a) Here, it is assumed that all neighboring samples (PredSample) used for prediction are already in the shaped domain. Interpretation (e.g., using motion compensation) is performed in the unshaped domain (i.e., using the reference picture directly from the DPB). That is, PredSample=MC(RecSampleinDPB) (19b) Here, MC() represents the motion compensation function. For fast mode determination where motion estimation and residuals are not generated, the strain can be calculated using the following equation: distortion=dfunc(SrcSample-PredSample) However, for full mode determination, in which residuals are generated, mode determination is performed in the shaping domain. That is, for full mode determination, distortion=dfunc(FwdLUT(SrcSample)-RecSample) (20)

[0093] Block-level adaptation As explained earlier, the proposed in-loop shaper allows shaping to be adapted at the CU level, for example by setting the variable CU_reshaper on or off as needed. Under the same architecture, for inter-encoded CUs, when CU_reshaper=off, the reconstructed pixels must be in the shaping domain, even if the CU_reshaper flag is set to off for this inter-encoded CU. RecSample=FwdLUT(Res+PredSample) (21) Therefore, intra-predictions always have neighboring pixels in the shaping domain. DPB pixels can be derived as follows: RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(FwdLUT(Res+PredSample))) (22)

[0094] For intra-encoded CUs, two alternative methods are proposed, depending on the encoding process: 1) All intra-encoded CUs are encoded with CU_reshaper=on. In this case, no additional processing is required as all pixels are already in the shaping domain. 2) Some intra-encoded CUs can be encoded using CU_reshaper=off. In this case, for CU_reshaper=off, when applying intra-prediction, de-shaping must be applied to neighboring pixels so that the intra-prediction is performed in the original domain and the final reconstructed pixel must be in the shaping domain. That is, RecSample=FwdLUT(Res+InvLUT(PredSample)) (23) and, RecSampleInDPB=InvLUT(LPF(RecSample)))= =InvLUT(LPF(FwdLUT(Res+InvLUT(PredSample))))) (24)

[0095] Generally, the proposed architectures can be used in various combinations, such as only in-loop intra shaping, in-loop shaping only for prediction residuals, or a hybrid architecture that combines both in-loop intra shaping and inter residual shaping. For example, to reduce the latency in the hardware decode pipeline, for inter-slice decoding, intra prediction can be performed before inverse shaping (i.e., decode intra CUs within an inter-slice). An exemplary architecture (200D_D) of such an embodiment is shown in FIG. 2E. In the reconstruction module (285), for inter CUs from Equation (17) (e.g., the Mux enables the output from 280 and 282), RecSample=(Res+FwdLUT(PredSample)) where FwdLUT(PredSample) represents the output of the inter predictor (280) and the subsequent forward shaping (282). Otherwise, for intra CUs (e.g., the Mux enables the output from 284), the output of the reconstruction module (285) is RecSample=(Res+IPredSample) where IPredSample represents the output of the intra prediction block (284). The inverse shaping block (265-3) Y CU =InvLUT[RecSample] performs.

[0096] Applying intra prediction in the shaping domain for inter slices is applicable to other embodiments including those shown in FIG. 2C (where inverse shaping is performed after loop filtering) and FIG. 2D. Special care needs to be taken in the combined inter / intra prediction mode (i.e., when, during reconstruction, some samples are from inter-coded blocks and some are from intra-coded blocks). This is because in all such embodiments, the inter prediction is in the original domain while the intra prediction is in the shaping domain. When combining data from both inter-predicted and intra-predicted coding units, the prediction can be performed in either of the two domains. For example, when the combined inter / intra prediction mode is executed in the shaped domain: PredSampleCombined = PredSampeIntra + FwdLUT(PredSampleInter) RecSample = Res + PredSampleCombined That is, the inter-coded samples in the original domain are shaped before addition. Otherwise, when the combined inter / intra prediction mode is performed in the original domain: PredSampleCombined = InvLUT(PredSampeIntra) + PredSampleInter RecSample = Res + FwdLUT(PredSampleCombined) That is, the intra-predicted samples are inverse-shaped to be in the original domain.

[0097] Similar considerations are applicable to the corresponding encoding embodiments. This is because the encoder (e.g., 200_E) includes a decoder loop that matches the corresponding decoder. As discussed earlier, Equation (20) describes embodiments where mode decision is executed in the shaping domain. In another embodiment, the mode decision may be executed in the original domain, i.e.: distortion=dfunc(SrcSample-InvLUT(RecSample))

[0098] For chroma QP offset or chroma residual scaling based on lumens, use the mean CU lumens value.

number

[0099] Chroma QP derivation Similar to Patent Document 6, the same proposed chromaDQP derivation process may be applied to balance the relationship between lumens and chromens resulting from the shaping curve. In one embodiment, piecewise chromaDQP values ​​can be derived based on codeword assignments for each bin. For example: About the kth bin scale k =(M k / M a ); (twenty five) chromaDQP=6*log2(scale k ); end

[0100] Encoder optimization As described in Patent Document 6, when lumaDQP is enabled, it is recommended to use pixel-based weighted distortion. When shaping is used, for example, the required weights are adjusted based on the shaping function (f(x)). For example: W rsp = f'(x) 2 (26) Here, f'(x) represents the gradient of the shape function f(x).

[0101] In another embodiment, piecewise weights can be directly derived based on the codeword assignment for each bin. For example; For the kth bin, Wrsp (k) = (M k / M a ) 2 (27)

[0102] For the chroma component, the weight can be set to 1 or some scaling factor sf. To reduce chroma distortion, sf can be set to a value greater than 1. To increase chroma distortion, sf can be set to a value greater than 1. In one embodiment, sf can be used to compensate for equation (25). Since chromaDQP can only be set to an integer, sf can be used to accommodate the fractional part of chromaDQP. Thus: sf=2 ((chromaDQP-INT(chromaDQP)) / 3)

[0103] In another embodiment, the chromaQPOffset value can be explicitly set in the Picture Parameter Set (PPS) or slice header to control chroma distortion.

[0104] The shaper curve or mapping function does not need to be fixed for the entire video sequence. For example, it can be adapted based on the quantization parameter (QP) or the target bitrate. In one embodiment, a more aggressive shaper curve can be used when the bitrate is low, and a less aggressive shaper can be used when the bitrate is relatively high. For example, given 32 bins in a 10-bit sequence, each bin initially has 32 codewords. When the bitrate is relatively low, codewords between

[2840] can be used to select a codeword for each bin. When the bitrate is high, codewords between

[3133] can be selected for each bin, or the identity shaper curve can simply be used.

[0105] Given a slice (or tile), shaping at the slice (tile) level can be performed in a variety of ways, with complexity being a trade-off for encoding efficiency. This includes: 1) disabling shaping only on intra-slices; and 2) disabling shaping on specific inter-slices, such as inter-slices at a particular temporal level (single or multiple), or on inter-slices not used for a reference picture, or on inter-slices considered to be of low importance as reference pictures. Such slice adaptations can also be QP / rate dependent, so different adaptation rules may apply for different QPs or bitrates.

[0106] In the encoder, under the proposed algorithm, the variance is calculated for each bin (e.g., BinVar(b) in equation (13)). Based on this information, codewords can be assigned based on each bin variance. In one embodiment, BinVar(b) may be inversely mapped to the number of codewords in each bin b. In another embodiment, to inversely map the number of codewords in bin b, (BinVar(b)) 2 Nonlinear mappings such as sqrt(BinVar(b)) may also be used. Essentially, this approach allows the encoder to apply an arbitrary codeword to each bin, beyond the simpler mappings used earlier. Here, the encoder takes two upper-range values ​​M f and M a (See, for example, Figure 6C) or three upper range values ​​M f , 32 or M a A codeword is assigned to each bin using (see, for example, Figure 6D).

[0107] As an example, FIG. 6E shows two codeword assignment methods based on the BinVar(b) value. Plot 610 shows codeword assignment using two thresholds, while plot 620 shows codeword assignment using inverse linear mapping. Here, the codeword assignment for a certain bin is inversely proportional to its BinVar(b) value. For example, in an embodiment, the following code may be applied to derive the number of codewords (bin_cw) within a specific bin: alpha=(minCW-maxCW) / (maxVar-minVar); beta=(maxCW*maxVar-minCW*minVar) / (maxVar-minVar); bin_cw=round(alpha*bin_var+beta);, Here, minVar represents the minimum variance across all bins, maxVar represents the maximum variance across all bins, and minCW, maxCW represent the minimum and maximum number of codewords per bin determined by the shaping model.

[0108] Refinement of Chroma QP offset based on Luma In Patent Document 6, an additional chroma QP offset (denoted as chromaDQP or cQPO) and a luma-based chroma residual scaler (cScale) were defined to compensate for the interaction between luma and chroma. For example: chromaQP=QP_luma+chromaQPOffset+cQPO (28) Here, chromaQPOffset represents the chroma QP offset, and QP_luma represents the luma QP for the coding unit. As shown in Patent Document 6, in an embodiment,

Number

number

[0109] The QP value derived from the chroma (denoted as qPi) and the final chroma QP value (Qp C Given a nonlinear relationship between (for example, Table 8-10 in Non-Patent Literature 4, "Qp as a function of qPi for ChromaArrayType equal to 1"), C (See "Specification"), in one embodiment, cScale may be further adjusted as follows:

[0110] For example, as shown in Table 8-10 of Non-Patent Document 4, the mapping between the adjusted luma and chroma QP values ​​is denoted as f_QPi2QPc(). Then, chromaQP_actual=f_QPi2QPc[chromaQP]= =f_QPi2QPc[QP_luma+chromaQPOffset+cQPO] (31) To scale the chroma residuals, the scale needs to be calculated based on the actual difference between the real chroma-encoded QPs before and after applying cQPO: QPcBase=f_QPi2QPc[QP_luma+chromaQPOffset]; QPcFinal=f_QPi2QPc[QP_luma+chromaQPOffset+cQPO]; (32) cQPO_refine=QPcFinal-QpcBase; cScale=pow(2,-cQPO_refine / 6)

[0111] In another embodiment, chromaQPOffset can also be absorbed into cScale. For example, QPcBase=f_QPi2QPc[QP_luma]; QPcFinal=f_QPi2QPc[QP_luma+chromaQPOffset+cQPO]; (33) cTotalQPO_refine=QPcFinal-QpcBase; cScale=pow(2,-cTotalQPO_refine / 6)

[0112] As an example, in one embodiment, as described in Patent Document 6: Let CSCALE_FP_PREC=16 represent the accuracy parameter. • Forward scaling: After chroma residuals are generated, before transformation and quantization: C_Res=C_orig-C_pred C_Res_scaled=C_Res*cScale+(1<<(CSCALE_FP_PREC-1)))>>CSCALE_FP_PREC • Inverse scaling: After chromatic inverse quantization and inverse transform, but before reconstruction: C_Res_inv=(C_Res_scaled< <CSCALE_FP_PREC) / cScale C_Reco = C_Pred + C_Res_inv;

[0113] In an alternative embodiment, the operation for in-loop chromatic shaping may be expressed as follows: On the encoder side, for each CU or TU chromatic component Cx (e.g., Cb or Cr) residual (CxRes = CxOrg - CxPred),

number

number

number

number

[0114] The use of cScale is not limited to chroma residual scaling for in-loop shaping. The same method can be applied to out-loop shaping. In out-loop shaping, cScale may be used for scaling chroma samples. The operation is the same as in the in-loop approach.

[0115] On the encoder side, when calculating the chroma RDOQ, the lambda modifier for chroma adjustment (whether using QP offsets or chroma residual scaling) also needs to be calculated based on the refined offset: Modifier=pow(2,-cQPO_refine / 3); New_lambda=Old_lambda / Modifier (38)

[0116] As shown in equation (35), using cScale may require division in the decoder. To simplify the decoder implementation, the same functionality could be implemented using division in the encoder, with a simpler multiplication applied in the decoder. For example, cScaleInv=(1 / cScale) For example, with an encoder cResScale=CxRes*cScale=CxRes / (1 / cScale)=CxRes / cScaleInv And with the decoder CxRes=cResScale / cScale=CxRes*(1 / cScale)=CxRes*cScaleInv Let's assume that.

[0117] In one embodiment, each rumor-dependent chroma scaling factor may be calculated for the corresponding rumor range in the piece-wise linear (PWL) representation, rather than for each rumor codeword value. Thus, the chroma scaling factor may be stored in a smaller LUT (e.g., 16 or 32 entries), such as cScaleInv[binIdx], instead of a 1024-entry LUT (for a 10-bit rumor codeword) (e.g., cScale[Y]). The encoder and decoder scaling operations may be implemented using fixed-point integer arithmetic as follows: c' = sign(c) * ((abs(c) * s + 2 CSCALE_FP_PREC-1 )>>CSCALE_FP_PREC) Here, c is the chroma residual, s is the chroma residual scaling factor from cScaleInv[binIdx], binIdx is determined by the corresponding mean lumen value, and CSCALE_FP_PREC is a constant value relating to precision.

[0118] In some embodiments, the forward shaping function may be represented using N equal segments (e.g., N=8, 16, 32, etc.), but the inverse representation would include nonlinear segments. From an implementation standpoint, it is desirable to have a representation of the inverse shaping function using equal segments, but forcing such a representation may result in a loss of coding efficiency. As a compromise, in some embodiments, the inverse shaping function may be constructed using a “mixed” PWL representation that combines both equal and unequal segments. For example, when using eight segments, the entire range may first be divided into two equal segments, and then each of these may be subdivided into four unequal segments. Alternatively, the entire range may be divided into four equal segments, and then one of each may be subdivided into two unequal segments. Alternatively, the entire range may first be divided into several unequal segments, and then each unequal segment may be subdivided into several equal segments. Alternatively, the entire range may first be divided into two equal segments, and then each equal segment may be subdivided into equal subsegments, where the segment lengths in each group of subsegments may not be the same.

[0119] For example, using 1024 codewords, we could have: a) four segments, each with 150 codewords and two segments, each with 212 codewords; or b) eight segments, each with 64 codewords and four segments, each with 128 codewords. The general purpose of such combinations of segments is to reduce the number of comparisons required to identify the PWL partition index given a code value, thereby simplifying hardware and software implementations.

[0120] In one embodiment, the following modifications may be possible for a more efficient implementation related to chroma residual scaling: • Disable chroma residual scaling when separate chroma / chroma trees are used. For 2x2 chroma grids, disable chroma residual scaling. • For intra- and inter-encoded units, use the prediction signal instead of the reconstruction signal.

[0121] As an example, given a decoder (200D_D) shown in Figure 2E for processing the chroma component, Figure 2F shows an exemplary architecture (200D_DC) for processing the corresponding chroma sample.

[0122] As shown in Figure 2F, the following changes are made when processing chroma compared to Figure 2E: • Forward and reverse shaping blocks (282 and 265-3) are not used. • A new chroma residual scaling block (288) is available, which effectively replaces the inverse shaping block (265-3) for lumens. • Reconstruction block (285-C) is modified to handle the chromophores of the original domain, as described in equation (36): CxRec=CxPred+CxRes.

[0123] From equation (34), on the decoder side, CxResScaled represents the extracted scaled chroma residual signal after inverse quantization and transformation (before block 288), CxRes = CxResScaled * C ScaleInv Let CxRec = CxPred + CxRes represent the rescaled chroma residuals generated by the chroma residual scaling block (288), which is used by the reconstruction unit (285-C) to calculate CxRec = CxPred + CxRes, where CxPred is generated by the intra(284) or inter(280) prediction block.

[0124] The value used for the conversion unit (TU) may be shared by the Cb and Cr components and can be calculated as follows: • If in intra mode, calculate the average of the intra-predicted luma values; If it is inter-mode, the mean of the forward-shaped inter-predicted lumens is calculated. That is, the mean lumens is calculated in the shaped domain. • For combined merge and intra predictions, calculate the average of the combined prediction ruma values. For example, the combined prediction ruma value may be calculated according to Section 8.4.6.6 of Appendix 2. · In one embodiment, avgY' TU Based on C ScaleInv A LUT can be applied to calculate it. Alternatively, given a piecewise linear (PWL) representation of the shaping function, the value avgY' can be calculated. TU This allows us to find the index idx that belongs to the reverse mapping PWL. Next, C ScaleInv =cScaleInv[idx] Exemplary implementations applicable to the Versatile Video Coding codec currently under development by the ITU and ISO (Non-Patent Document 8) can be found in Appendix 2 (see, for example, Section 8.5.5.1.2).

[0125] Disabling ruma-based chroma residual scaling for intra-slice using a double tree may result in some loss of encoding efficiency. To improve the effect of chroma shaping, the following methods may be used: 1. The chroma scaling factor may be kept the same across the entire frame, depending on the mean or median of the lumen sample values. This eliminates the dependence of the TU level on the lumen for chroma residual scaling. 2. Chroma scaling factors can be derived using reconstructed lumens from neighboring CTUs. 3. The encoder can derive a chroma scaling factor based on the source luma pixels and transmit it in the bitstream at the CU / CTU level (for example, as an index to the piecewise representation of the shaping function). The decoder can then extract the chroma scaling factor from the shaping function independently of the luma data. The scaling factor for CTU can be derived and transmitted only for intra-slice, but can also be used for inter-slice. Additional signaling costs occur only for intra-slice and therefore do not affect coding efficiency in random access. 4. Chroma can be shaped as luma at the frame level, and the luma shaping curve is derived from the luma shaping curve based on a correlation analysis between luma and chroma. This completely eliminates chroma residual scaling.

[0126] apply delta_qp In AVC and HEVC, the parameter delta_qp is allowed to modify the QP value for a coded block. In some embodiments, the delta_qp value can be derived using a lumer curve in the shaper. A piecewise lumer DQP value can be derived based on the codeword assignment for each bin. For example: For the kth bin, scale k =(M k / M a ); (39) lumaDQP k =INT(6*log2(scale k )) Here, INT() can be CEIL(), ROUND(), or FLOOR(). The encoder can use a function of luma, such as average(luma), min(luma), max(luma), etc., to find the luma value for that block, and then use the corresponding lumaDQP value for that block. From equation (27), in order to obtain the rate-distortion benefit, the weighted distortion is used in mode determination, W rsp (k) = scale k 2 It can be set as follows.

[0127] Considerations regarding shaping and number of bottles In a typical 10-bit video encoding, it is preferable to use at least 32 bins for shaping the mapping, but in some embodiments, fewer bins, such as 16 or even 8, may be used to simplify the decoder implementation. Given that the encoder may have already used 32 bins to analyze the sequence and derive the distributed codeword, the original 32-bin codeword distribution can be reused, and a 16-bin codeword can be derived by adding the corresponding two 16-bins within each of the 32 bins. That is, For i=0 to 15 CWIn16Bin[i]=CWIn32Bin[2i]+CWIn32Bin[2i+1]

[0128] For the chroma residual scaling factor, you can simply divide the codeword by 2 and point to a 32-bin chromaScalingFactorLUT. For example, In32Bin

[32] ={0 0 33 38 38 38 38 38 38 38 38 38 38 38 38 38 38 33 33 33 33 33 33 33 33 33 33 33 33 33 0 0} Given, the corresponding 16-bin CW allocation is: CWIn16Bin

[16] ={ 0 71 76 76 76 76 76 76 71 66 66 66 66 66 66 0} This approach can be extended to handle even fewer bins, such as eight. For i=0 to 7 CWIn8Bin[i]=CWIn16Bin[2i]+CWIn16Bin[2i+1]

[0129] When using a narrow range of valid codewords (e.g., [64,940] for a 10-bit signal, [64,235] for an 8-bit signal), care should be taken not to consider mappings of the first and last bins to reserved codewords. For example, for a 10-bit signal, with 8 bins, each bin has 1024 / 8 = 128 codewords, and the first bin is [0,127], but since the standard codeword range is [64,940], the first bin should only consider the codeword [64,127]. bitdepth A special flag (e.g., video_full_range_flag=0) may be used to inform the decoder that it should have a narrower range than -1 and that special care should be taken to avoid generating invalid codewords when processing the first and last bins. This can be applied to both lumer and chroma shaping.

[0130] As an example, but not limited to, Appendix 2 provides exemplary syntax structures and related syntax elements for supporting formatting in an ISO / ITU Video Versatile Codec (VVC) (Non-Patent Literature 8) using the architecture shown in Figures 2C, 2E, and 2F. Here, the forward formatting function includes 16 segments. References Each of the references cited herein is incorporated hereby by reference. [Non-Patent Document 1] "Exploratory Test Model for HDR extension of HEVC", K. Minoo et al., MPEG output document, JCTVC-W0092 (m37732), 2016, San Diego, USA [Patent Document 2] PCT application PCT / US2016 / 025082, In-Loop Block-Based Image Reshaping in High Dynamic Range Video Coding, GM. Su, filed March 30, 2016, also published as WO2016 / 164235. [Patent Document 3] U.S. Patent Application 15 / 410,563, Content-Adaptive Reshaping for High Codeword Representation Images, T. Lu et al., Filing Date: Jan. 19, 2017 [Non-Patent Document 4] ITU-T H.265, "High efficiency video coding," ITU, Dec. 2016 [Patent Document 5] PCT application PCT / US2016 / 042229, Signal Reshaping and Coding for HDR and Wide Color Gamut Signals, P. Yin et al., Filing date July 14, 2016, also published as WO2017 / 011636. [Patent Document 6] PCT Patent Application PCT / US2018 / 040287, Integrated Image Reshaping and Video Coding, T. Lu et al., Filing Date June 29, 2018 [Patent Document 7] J. Froehlich et al., "Content-Adaptive Perceptual Quantizer for High Dynamic Range Images," U.S. Patent Application Publication No. 2018 / 0041759, Feb. 08, 2018. [Non-Patent Document 8] B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 3)," JVET output document, JVET-L1001, v9, uploaded, Jan. 8, 2019

[0131] Exemplary Receiving System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprising electronic circuits and components, integrated circuits (ICs) such as microcontrollers, field-programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLCs), discrete-time or digital signal processors (DSPs), application-specific ICs (ASICs), and / or apparatus comprising one or more such systems, devices, or components. Computers and / or ICs may execute, control, or run instructions relating to the signal shaping and encoding of images as described herein. Computers and / or ICs may calculate any of the various parameters or values ​​relating to the signal shaping and encoding processes described herein. Embodiments of images and videos may be implemented in hardware, software, firmware, and various combinations thereof.

[0132] Certain implementations of the present invention include a computer processor that executes software instructions causing the processor to perform the method of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., can perform the methods related to the signal shaping and encoding of images as described above by executing software instructions in program memory accessible to the processor. The present invention may be provided in the form of a program product. The program product may include any non-temporary and tangible medium that, when executed by a data processor, carries a set of computer-readable signals including instructions causing the data processor to perform the method of the present invention. The program product according to the present invention may take any of a wide variety of non-temporary and tangible forms. The program product may include, for example, physical media such as magnetic data storage media including floppy diskettes, optical data storage media including hard disk drives, CD-ROMs, DVDs, and electronic data storage media including ROMs and flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0133] Where components (e.g., software modules, processors, assemblies, devices, circuits, etc.) are referred to above, unless otherwise indicated, references to such components (including references to “means”) should be interpreted as including any components that perform the function of the described component (e.g., functionally equivalent), including components that are not structurally equivalent to the disclosed structures that perform the function in the illustrated embodiments of the present invention.

[0134] Equivalents, extensions, substitutes, and others Thus, exemplary embodiments relating to the efficient signal shaping and encoding of images are described. In the preceding specification, embodiments of the invention have been described with reference to numerous individual details that may vary from implementation to implementation. Therefore, the sole and exclusive indicator of what constitutes an invention and what the applicant intends to be an invention is the specific form in which such claims are permitted, including any subsequent amendments, of the set of claims issued for this application. The definitions of terms included in such claims expressly provided herein govern the meaning of such terms as used in those claims. Therefore, no limitations, elements, characteristics, features, advantages, or attributes not expressly provided in the claims should in any way limit the scope of such claims. Accordingly, this specification and the drawings should be considered illustrative and not restrictive.

[0135] Bulleted Exemplary Embodiments The present invention includes, but is not limited to, the following Enumerated Example Embodiments (EEEs) describing the structure, features, and functions of some parts of the invention, and can be embodied in any form described herein. [EEE1] A method for adaptively shaping a video sequence using a processor, wherein the method is: The processor accesses the input image in a first codeword representation; The process includes the step of generating a forward shaping function that maps the pixels of the input image to a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation, and the generation of the forward shaping function is: The input image is divided into multiple pixel regions; Each pixel region is assigned to one of several codeword bins according to the first luminance characteristic of that pixel region; The bin metric for each of the plurality of codeword bins is calculated according to the second luminance characteristics of each pixel region assigned to each codeword bin; Assign the number of codewords in the second codeword representation to each codeword bin according to the bin metric and rate-distortion optimization criteria for each codeword bin; The process includes generating the forward shaping function in response to the allocation of codewords in the second codeword representation for each of the plurality of codeword bins, method. [EEE2] The method according to EEE1, wherein the first luminance characteristic of a pixel region includes the average luminance pixel value within that pixel region. [EEE3] The method according to EEE1, wherein the second luminance characteristic of a pixel region includes the variance of the luminance pixel value of that pixel region. [EEE4] The method described in EEE3 for calculating the bin metric for a codeword bin includes calculating the average of the variances of luminance pixel values ​​for all pixel regions assigned to that codeword bin. [EEE5] Assigning the number of codewords in the second codeword representation to a codeword bin according to its bin metric is: If no pixel region is assigned to that codeword bin, no codeword is assigned to that codeword bin; If the bin metric of that codeword bin is lower than the upper threshold, assign a first number of codewords; Otherwise, the process involves assigning a second number of codewords to that codeword bin. Method as described in EEE1. [EEE6] A first codeword representation with a depth of B bits, a second codeword representation with a depth of Bo bits, and N codeword bins, the first number of codewords is M f =CEIL((2 Bo / (CW2-CW1))*M a ) including, the second number of codewords is M a =2 B / N is included, where CW1 <CW2は、[0 2 B A method described in EEE5 for representing two codewords in -1]. [EEE7] CW1 = 16 × 2 (B-8) and CW2 = 235 × 2 (B-8) The method described in EEE6. [EEE8] Determining the aforementioned upper threshold: Define a set of potential thresholds; For each threshold in the aforementioned set of thresholds: A forward shaping function is generated based on that threshold; The set of input test frames is encoded and decoded according to the aforementioned formatting function and bitrate R to generate an output set of decoded test frames; Based on the input test frame and the decoded test frame, the overall rate-distortion optimization (RDO) metric is calculated; This includes selecting the threshold in the set of potential thresholds that has the smallest RDO metric as the upper limit threshold. Method as described in EEE5. [EEE9] Calculating the aforementioned RDO metric is: This involves calculating J = D + λR, where D represents an index of the distortion between the pixel values ​​of the input test frame and the corresponding pixel values ​​in the decoded test frame, and λ represents the Lagrange multiplier. Method as described in EEE8. [EEE10] D is an index of the sum of squared differences between the corresponding pixel values ​​of the input test frame and the decoded test frame, as described in the EEE9 method. [EEE11] The method according to EEE1, wherein assigning the number of codewords in the second codeword representation to a codeword bin according to its bin metric is based on a codeword assignment lookup table, the codeword assignment lookup table defines two or more thresholds that divide a range of bin metric values ​​into segments, and gives the number of codewords to be assigned to bins having a bin metric within each segment. [EEE12] The method described in EEE11, in which a default codeword assignment is given to the bins, and bins with a large bin metric are assigned fewer codewords than the default assignment, and bins with a small bin metric are assigned more codewords than the default assignment. [EEE13] For a first codeword representation using B bits and N bins, the default codeword assignment per bin is M a =2 B The method described in EEE12, given by / N. [EEE14] The process further includes the step of generating formatting information in response to the forward formatting function, wherein the formatting information is: A flag indicating the minimum codeword bin index value used in the reshaping and reconstruction process; A flag indicating the maximum codeword bin index value used in the aforementioned shaping process; A flag indicating a shaping model profile type, where each model profile type is associated with a default bin-related parameter, or One or more delta values ​​used to adjust the default bin-related parameters Including one or more of the following: Method as described in EEE1. [EEE15] The process further includes assigning a bin importance value to each codeword bin, wherein the bin importance value is: If no codeword is assigned to that codeword bin, the codeword is 0. If the codeword is assigned the first value of the codeword, then it is 2; Otherwise, it is 1. Method as described in EEE5. [EEE16] Determining the aforementioned upper threshold: The first step is to divide the luminance range of the pixel values ​​in the input image into bins; For each bin, the step is to determine the bin histogram value and the mean bin variance value, wherein for each bin, the bin histogram value includes the number of pixels in that bin out of the total number of pixels in the image, and the mean bin variance value gives a metric of the mean pixel variance of the pixels in that bin; The steps include sorting the mean bin variance values ​​to generate a sorted list of mean bin variance values ​​and a sorted list of mean bin variance value indices; A step of calculating a cumulative density function as a function of the sorted mean bin variances, based on the sorted mean bin variances based on the bin histogram values ​​and the sorted list of mean bin variance-value indices; The step includes determining an upper threshold based on criteria satisfied by the value of the cumulative density function, Method as described in EEE5. [EEE17] To calculate the aforementioned cumulative density function: BinVarSortDsdCDF[0]=BinHist[BinIdxSortDsd[0]]; for(int b=1; b <PIC_ANALYZE_CW_BINS; b++) { BinVarSortDsdCDF[b]=BinVarSortDsdCDF[b-1]+BinHist[BinIdxSortDsd[b]];} This includes calculating the following, where b is the bin number, PIC_ANALYZE_CW_BINS is the total number of bins, BinVarSortDsdCDF[b] is the output of the CDF function for bin b, BinHist[i] is the bin histogram value for bin i, and BinIdxSortDsd[] represents a sorted list of mean bin variance indexes. Method as described in EEE16. [EEE18] The method according to EEE16, wherein the upper threshold is determined as the mean bin variance value for which the CDF output is k%, based on the criterion that the mean bin variance for k% of pixels in the input image is greater than or equal to the upper threshold. [EEE19] The method described in EEE18, where k=50. [EEE20] A method for reconstructing a shaping function in a decoder, the method being: A step of receiving an encoded bitstream syntax element that characterizes the shaping model, wherein the syntax element is: A flag indicating the minimum codeword bin index value used in the formatting and construction process. A flag indicating the maximum codeword bin index value used in the formatting and construction process. A flag indicating a shaped model profile type, wherein the model profile type is associated with default bin-related parameters including bin importance values, or Flags indicating one or more delta bin importance values ​​used to adjust the default bin importance values ​​defined in the aforementioned shaping model profile. A stage including one or more of the following; A step of determining, based on the aforementioned shaping model profile, a default bin importance value for each bin and an assignment list of the default number of codewords to be assigned to each bin according to the bin importance value; For each codeword bin: The bin importance value is determined by adding the default bin importance value to the delta bin importance value; Based on the bin importance value of that bin and the assignment list, the number of codewords to be assigned to that codeword bin is determined; This includes the step of generating a forward shaping function based on the number of codewords assigned to each codeword bin. method. [EEE21] The number of codewords M assigned to the k-th codeword bin using the assignment list. k To decide further: For the kth bin: If bin_importance[k] == 0, M k Set = 0; Otherwise, if bin_importance[k]==2 M k =M f year, Otherwise, M k =M a This includes, Here, M a and M f is an element of the aforementioned assignment list, and bin_importance[k] represents the bin importance value of the k-th bin. Method as described in EEE20. [EEE22] A method for reconstructing encoded data in a decoder having one or more processors, wherein the method is: A step of receiving an encoded bitstream (122) which includes one or more encoded formatted images in a first codeword representation and metadata (207) relating to formatting information about the encoded formatted images; A step (250) of generating an inverse formatting function based on metadata related to the formatting information, wherein the inverse formatting function maps the pixels of the formatted image from the first codeword representation to the second codeword representation; (250) A step of generating a forward formatting function based on metadata relating to the formatting information, wherein the forward formatting function maps the pixels of the image from the second codeword representation to the first codeword representation; A step of extracting an encoded shaped image containing one or more encoded units from the encoded bitstream, wherein with respect to one or more encoded units in the encoded shaped image: Regarding the intra-encoded coding unit (CU) within the aforementioned encoded reshaped image: Based on the shaped residuals and the first shaped predicted sample in the CU, a first shaped reconstructed sample (227) of the CU is generated; Based on the first shaped and reconstructed sample and loop filter parameters, a shaped loop filter output is generated (270); The de-shaping function is applied to the shaped loop filter output to generate decoded samples of its coded units in the second codeword representation (265); The steps include storing the decoded sample of the encoding unit in the second codeword representation in a reference buffer; Regarding the intercoded coding units within the aforementioned encoded and shaped image: The forward shaping function is applied to the prediction samples stored in the reference buffer in the second codeword representation to generate a second shaped prediction sample; Based on the shaped residual in the encoded CU and the second shaped predicted sample, a second shaped reconstructed sample of the encoded unit is generated; Based on the second shaped and reconstructed sample and loop filter parameters, a shaped loop filter output is generated; The de-shaping function is applied to the shaped loop filter output to generate samples of its coded units in the second codeword representation; The steps include: storing a sample of the encoded unit in the second codeword representation in a reference buffer; The step of generating a decoded image based on the samples stored in the reference buffer is included. method. [EEE23] A device equipped with a processor and configured to perform any of the methods described in EEE1 to 22. [EEE24] A non-temporary computer-readable storage medium that stores computer-executable instructions thereon for executing a method on one or more processors in accordance with EEE1 to 22.

[0136] Appendix 1 Exemplary implementation of bubble sort void bubbleSortDsd(double* array, int*idx, int n) { int i,j; bool swapped; for(i=0; i < n-1; i++) { swapped=false; for(j=0; j <n-i-1; j++) { if(array[j] <array[j+1]) { swap(&array[j],&array[j+1]); swap(&idx[j],&idx[j+1]); swapped=true; } } if(swapped==false) break; } }

[0137] Appendix 2 As an example, this appendix provides exemplary syntax structures and related syntax elements in a formatting-supporting embodiment of the Versatile Video Codec (VVC) (Non-Patent Document 8), currently under joint development by ISO and ITU. New syntax elements in existing draft versions are highlighted or explicitly noted. Equation numbers such as (8-xxx) indicate placeholders that will be updated as needed in the final specification.

[0138] 7.3.2.1 Sequence parameter set RBSP syntax [Table 5-1] [Table 5-2] [Table 5-3]

[0139] 7.3.3.1 General Tile Group Header Syntax [Table 6-1] [Table 6-2] [Table 6-3]

[0140] New Syntax Table and Tile Shaping Machine models have been added. [Table 7]

[0141] The following semantics have been added to the general sequence parameter set RBSP semantics. If sps_reshaper_enabled_flag is equal to 1, it specifies that the reshaper will be used in the encoded video sequence (CVS). If sps_reshaper_enabled_flag is equal to 0, it specifies that the reshaper will not be used in the CVS.

[0142] The following semantics have been added to the tile group header syntax. A value of 1 for `tile_group_reshaper_model_present_flag` indicates that `tile_group_reshaper_model()` exists in the tile group header. A value of 0 for `tile_group_reshaper_model_present_flag` indicates that `tile_group_reshaper_model()` does not exist in the tile group header. If `tile_group_reshaper_model_present_flag` does not exist, it is assumed to be equal to 0. A value of `tile_group_reshaper_enabled_flag` equal to 1 indicates that the reshaper is enabled for the current tile group. A value of `tile_group_reshaper_enabled_flag` equal to 0 indicates that the reshaper is not enabled for the current tile group. If `tile_group_resharper_enable_flag` does not exist, it is assumed to be equal to 0. A value of `tile_group_reshaper_chroma_residual_scale_flag` equal to 1 indicates that chroma residual scaling is enabled for the current tile group. A value of `tile_group_reshaper_chroma_residual_scale_flag` equal to 0 indicates that chroma residual scaling is not enabled for the current tile group. If `tile_group_reshaper_chroma_residual_scale_flag` does not exist, it is assumed to be equal to 0.

[0143] Added the tile_group_reshaper_model() syntax. `reshaper_model_min_bin_idx` specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of `reshaper_model_min_bin_idx` is in the range of 0 to `MaxBinIdx`, including both ends. The value of `MaxBinIdx` is equal to 15. reshaper_model_delta_max_bin_idx specifies a value obtained by subtracting the maximum bin (or piece) index MaxBinIdx from the maximum bin index used in the reshaper configuration process. The value of reshaper_model_max_bin_idx is set equal to MaxBinIdx - reshaper_model_delta_max_bin_idx. One plus reshaper_model_bin_delta_abs_cw_prec_minus1 specifies the number of bits used for the representation of the syntax reshaper_model_bin_delta_abs_CW[i]. reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshaper_model_bin_delta_abs_CW[ i ] as follows. · When reshaper_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. · Otherwise (when reshaper_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is a negative value. If reshaper_model_bin_delta_sign_CW_flag[i] does not exist, it is assumed to be equal to 0. The variable RspDeltaCW[i] = (1 - 2 * reshaper_model_bin_delta_sign_CW[i]) * reshaper_model_bin_delta_abs_CW[i]; The variable RspCW[i] is derived in the following steps: The variable OrgCW is set equal to (1 << BitDepthY) / (MaxBinIdx + 1). If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx RspCW[i]=OrgCW+RspDeltaCW[i] • Otherwise, RspCW[i]=0 The value of RspCW[i] is in the range of 32 to 2*OrgCW-1, when the value of BitDepthYis is equal to 10. For i in the range from 0 to MaxBinIdx+1, including both ends, the variable InputPivot[i] is derived as follows: InputPivot[i] = i * OrgCW The variables ReshapePivot[i] for i in the range from 0 to MaxBinIdx+1, including both ends, and the variables ScaleCoef[i] and InvScaleCoeff[i] for i in the range from 0 to MaxBinIdx, including both ends, are derived as follows: shiftY=14 ReshapePivot[0] = 0; for(i=0; i<=MaxBinIdx; i++){ ReshapePivot[i+1]=ReshapePivot[i]+RspCW[i] ScaleCoef[i]=(RspCW[i]*(1< <shiftY)+(1<<(Log2(OrgCW)-1)))> >(Log2(OrgCW)) if(RspCW[i]==0) InvScaleCoeff[i]=0 else InvScaleCoeff[i]=OrgCW*(1< <shiftY) / RspCW[i] } For i in the range from 0 to MaxBinIdx, including both ends, the variable ChromaScaleCoef[i] is derived as follows: ChromaResidualScaleLut

[64] ={16384,16384,16384,16384,16384,16384,16384,8192,8192,8192,8192 ,5461,5461,5461,5461,4096,4096,4096,4096,3277,3277,3277,3277,2731,2731,2731,2731,2341,234 1,2341,2048,2048,2048,1820,1820,1820,1638,1638,1638,1638,1489,1489,1489,1489,1365,1365,1365,1365,1260,1260,1260,1260,1170,1170,1170,1170,1092,1092,1092,1092,1024,1024,1024}; shiftC=11 If (RspCW[i]==0) then ChromaScaleCoef[i]=(1< <shiftC) Otherwise, (RspCW[i]!=0) ChromaScaleCoef[i]=ChromaResidualScaleLut[Clip3(1,64,RspCW[i]>>1)-1] Note: In alternative implementations, chroma and chroma scaling may be unified, eliminating the need for ChromaResidualScaleLut[]. In that case, chroma scaling could be implemented as follows: shiftC=11 If (RspCW[i]==0) then ChromaScaleCoef[i]=(1< <shiftC) • Otherwise (RspCW[i]!=0), the following applies: BinCW = BitDepth Y >10?(RspCW[i]>>(BitDepth Y -10)) BitDepth Y <10?(RspCW[i]<<(10 BitDepth Y )):RspCW[i]; ChromaScaleCoef[i]=OrgCW*(1< <shiftC) / BinCW[i]

[0144] The following has been added to the weighted sample prediction process for combined merge and intra prediction. The additions are highlighted. 8.4.6.6 Weighted sample prediction process for combined merge and intra prediction The inputs to this process are as follows: • Current encoding block width cbWidth • Current encoding block height cbHeight • Two (cbWidth) × (cbHeight) arrays, preSamplesInter and preSamplesIntra • Intra predictive mode • Variable cIdx that specifies the color component index The output of this process is the sequence predSamplesComb, which is (cbWidth) × (cbHeight) of the predicted sample values. The variable bitDepth is derived as follows: If cIdx is equal to 0, then bitDepth is BitDepth Y It is set to be equal to. Otherwise, bitDepth is BitDepth C It is set to be equal to. The predicted samples predSamplesComb[x][y] for x=0..cbWidth-1 and y=0..cbHeight-1 are derived as follows: The weight w is derived as follows: If predModeIntra is INTRA_ANGULAR50, w is specified in tables 8-10 as nPos being equal to y and nSize being equal to cbHeight. Otherwise, if predModeIntra is INTRA_ANGULAR18, w is specified in Tables 8-10 with nPos equal to x and nSize equal to cbWidth. Otherwise, w is set to 4. [Table 8] The predicted sample predSamplesComb[x][y] is derived as follows: predSamplesComb[x][y]=(w*predSamplesIntra[x][y]+ (8-w)*predSamplesInter[x][y])>>3) (8-740) Table 8-10 Specifying w as a function of position nP and size nS [Table 9]

[0145] Add the following to the picture reconstruction process. 8.5.5 Picture Composition Process The inputs to this process are as follows: • The position (xCurr, yCurr) that specifies the top-left sample of the current block relative to the top-left sample of the current picture component. • Variables nCurrSw and nCurrSh that specify the width and height of the current block, respectively. • Variable cIdx specifies the color component of the current block. • Specifies the predicted samples for the current block: (nCurrSw) × (nCurrSh) array predSamples • Specify the residual samples of the current block: (nCurrSw) × (nCurrSh) array resSamples The following assignments are made depending on the value of the color component cIdx: If cIdx is equal to 0, recSamples is the reconstructed picture sample sequence S L The function clipCidx1 corresponds to Clip1 YIt corresponds to. Otherwise, if cIdx is equal to 1, recSamples is the reconstructed chroma sample sequence S Cb The function clipCidx1 corresponds to Clip1 C It corresponds to. Otherwise (cIdx is equal to 2), recSamples is the reconstructed chroma sample sequence S Cr The function clipCidx1 corresponds to Clip1 C It corresponds to. [Table 10] Otherwise, the (nCurrSw)×(nCurrrSh) block of the reconstructed sample sequence recSamples at position (xCurr,yCurr) is derived as follows: For i=0..nCurrSw-1 and j=0..nCurrSh-1, recSamples[xCurr+i][yCurr+j]=clipCidx1(predSamples[i][j]+resSamples[i][j]) [Outside 1] TIFF2026136335000027.tif7115

[0146] [Outside 2] TIFF2026136335000028.tif10115 This section specifies picture reconstruction using the mapping process. Picture reconstruction using the mapping process for lumar sample values ​​is specified in 8.5.5.1.1. Picture reconstruction using the mapping process for chroma sample values ​​is specified in 8.5.5.1.2. 8.5.5.1 Picture reconstruction using a mapping process for lumar sample values The inputs to this process are as follows: • Specify the rumor prediction samples for the current block: (nCurrSw) × (nCurrSh) array predSamples • Specify the lumen residual samples of the current block: (nCurrSw) × (nCurrSh) array resSamples The output for this process is as follows: • (nCurrSw) × (nCurrSh) mapped lumens prediction sample sequences predMapSamples Reconstructed luma sample sequences of (nCurrSw) × (nCurrSh) recSamples predMapSamples is derived as follows: If(CuPredMode[xCurr][yCurr]==MODE_INTRA)||(CuPredMode[xCurr][yCurr]==MODE_INTER && mh_intra_flag[xCurr][yCurr]) predMapSamples[xCurr+i][yCurr+j]=predSamples[i][j] Here, i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outside 3] TIFF2026136335000029.tif9115 Otherwise ((CuPredMode[xCurr][yCurr]==MODE_INTER && !mh_intra_flag[xCurr][yCurr])), the following applies: shiftY=14 idxY=predSamples[i][j]>>Log2(OrgCW) predMapSamples[xCurr+i][yCurr+j]=ReshapePivot[idxY] +(ScaleCoeff[idxY]*(predSamples[i][j]-InputPivot[idxY]) +(1<<(shiftY-1)))>> shiftY Here, i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outside 4] TIFF2026136335000030.tif9115recSamples is derived as follows: recSamples[xCurr+i][yCurr+j]=Clip1 Y (predMapSamples[xCurr+i][yCurr+j]+resSamples[i][j]]) Here, i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outside 5] TIFF2026136335000031.tif9115 8.5.5.1.2 Picture reconstruction using a mapping process for chroma sample values The inputs to this process are as follows: • Specify the mapped ruma prediction samples for the current block using the (nCurrSwx2) × (nCurrShx2) mapped array predMapSamples • PredSamples: An array of (nCurrSw) × (nCurrSh) specifying the chroma prediction samples for the current block. • Specify the chroma residual samples of the current block: the (nCurrSw) × (nCurrSh) array resSamples The output of this process is the reconstructed chroma sample sequence, recSamples. recSamples is derived as follows: ·If(!tile_group_reshaper_chroma_residual_scale_flag||((nCurrSw)x(nCurrSh)<=4)) recSamples[xCurr+i][yCurr+j]=Clip1 C (predSamples[i][j]+resSamples[i][j]) Here, i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outside 6] TIFF2026136335000032.tif9115 · Otherwise (tile_group_reshaper_chroma_residual_scale_flag &&((nCurrSw)x(nCurrSh)>4)), the following applies: The variable varScale is derived as follows: 1.invAvgLuma=Clip1 Y ((Σ i Σ j predMapSamples[(xCurr<<1)+i][(yCurr<<1)+j] +nCurrSw*nCurrSh*2) / (nCurrSw*nCurrSh*4)) 2. The variable idxYInv is generated using the input of the sample value invAvgLuma. [Outside 7] As specified in section TIFF2026136335000033.tif8115, it is derived by calling the identification of the function index for each section. 3.varScale=ChromaScaleCoef[idxYInv] varScale=ChromaScaleCoef[idxYInv] recSamples is derived as follows: If tu_cbf_cIdx[xCurr][yCurr] is equal to 1, then the following applies: shiftC=11 recSamples[xCurr+i][yCurr+j]=ClipCidx1(predSamples[i][j]+Sign(resSamples[i][j]) *((Abs(resSamples[i][j])*varScale+(1<<(shiftC-1)))>>shiftC)) Here, i = 0..nCurrSw-1, j = 0..nCurrSh-1 [Outside 8] TIFF2026136335000034.tif9115 · Otherwise (if tu_cbf_cIdx[xCurr][yCurr] is equal to 0), recSamples[xCurr+i][yCurr+j]=ClipCidx1(predSamples[i][j]) [Outside 9] TIFF2026136335000035.tif9115

[0147] 8.5.6 Picture Reverse Mapping Process This section is invoked when the value of tile_group_reshaper_enabled_flag is equal to 1. The input is a reconstructed picture lumen sample array S L The output is the corrected and reconstructed picture lumens sample sequence S' after the reverse mapping process. L The reverse mapping process for rumor sample values ​​is defined in 8.4.6.1. 8.5.6.1 Picture inverse mapping process for lumar sample values The input to this process is a lumen position (xP, yP) that specifies the lumen sample position relative to the lumen sample in the upper left corner of the current picture. The output of this process is the inversely mapped Luma sample value invLumaSample. The value of invLumaSample is derived by applying the following ordered steps: 1. The variable idxYInv is a sample value S. L It is derived by using the inputs [xP][yP] to call the identification of the function index for each section as specified in Section 8.5.6.2. 2. The value of reshapeLumaSample is derived as follows: shiftY=14 invLumaSample=InputPivot[idxYInv]+(InvScaleCoeff[idxYInv]*(S L [xP][yP]-ReshapePivot[idxYInv]) +(1<<(shiftY-1)))>> shiftY [Outside 10] TIFF2026136335000036.tif91153.clipRange=((reshaper_model_min_bin_idx>0) && (reshaper_model_max_bin_idx <MaxBinIdx)); If clipRange is equal to 1, the following applies: minVal=16<<(BitDepth Y -8) maxVal=235<<(BitDepth Y -8) invLumaSample=Clip3(minVal,maxVal,invLumaSample) Otherwise (clipRange is equal to 0), the following applies: invLumaSample=ClipCidx1(invLumaSample) 8.5.6.2 Identification of function indices for each segment of the luma component The input to this process is a lumen sample value S. The output of this process is an index idxS that identifies the piece (or section) to which sample S belongs. The variable idxS is derived as follows: for(idxS=0, idxFound=0; idxS<=MaxBinIdx; idxS++){ if((S <ReshapePivot[idxS+1]){ idxFound=1 break } } Note: Alternative implementations for finding the identifier idxS are as follows: if(S <ReshapePivot[reshaper_model_min_bin_idx]) idxS=0 else if(S>=ReshapePivot[reshaper_model_max_bin_idx]) idxS = MaxBinIdx else idxS=findIdx(S,0,MaxBinIdx+1,ReshapePivot[]) function idx=findIdx(val,low,high,pivot[]){ if(high-low<=1) idx=low else { mid=(low+high)>>1 if(val <pivot[mid]) high=mid else low=mid idx=findIdx(val,low,high,pivot[]) } }

[0148] Several aspects are described below. [Aspect 1] A method for reconstructing encoded video data using one or more processors, the method being: The steps include receiving an encoded bitstream (122) containing one or more encoded formatted images in the input codeword representation; The steps include receiving formatting metadata (207) for one or more encoded formatted images in the encoded bitstream; A step of generating a forward formatting function (282) based on the formatting metadata, wherein the forward formatting function maps the pixels of the image from a first codeword representation to the input codeword representation; A step of generating an inverse formatting function (265-3) based on the formatting metadata or the forward formatting function, wherein the inverse formatting function maps the pixels of the formatted image from the input codeword representation to the first codeword representation; The steps include: extracting an encoded shaped image containing one or more encoded units from the encoded bitstream; Regarding the intra-encoded encoding unit (intraCU) in the aforementioned encoded reshaped image: A shaped and reconstructed sample of the intraCU is generated based on the shaped residuals and intra-predicted shaped and predicted samples within the intraCU (285); The de-shaping function (265-3) is applied to the shaped and reconstructed sample of the intraCU to generate a decoded sample of the intraCU in the first codeword representation; A loop filter (270) is applied to the decoded sample of the intraCU to generate the output sample of the intraCU; The step of storing the output sample of the intraCU in a reference buffer; Regarding the intercoded CU (interCU) in the above-mentioned encoded reshaped image: The forward shaping function (282) is applied to the inter prediction samples stored in the reference buffer in the first codeword representation to generate shaped prediction samples for the interCU in the input codeword representation; Based on the shaped residuals in the interCU and the shaped predicted samples for the interCU, a shaped reconstructed sample of the interCU is generated; The inverse reshaping function (265-3) is applied to the reshaped and reconstructed sample of the interCU to generate the decoded sample of the interCU in the first codeword representation; A loop filter (270) is applied to the decoded sample of the interCU to generate the output sample of the interCU; The steps include: storing the output sample of the interCU in the reference buffer; The step of generating a decoded image in the first codeword representation based on the output samples in the reference buffer, method. [Aspect 2] To generate a reshaped and reconstructed sample (RecSample) of the aforementioned intraCU: RecSample=(Res+IpredSample) This includes calculating, Here, Res represents the shaped residual sample in the intraCU in the input codeword representation, and IpredSample represents the intra-predicted shaped predicted sample in the input codeword representation. The method described in Embodiment 1. [Aspect 3] To generate a reshaped and reconstructed sample (RecSample) of the aforementioned interCU: RecSample=(Res+Fwd(PredSample)) This includes calculating, Here, Res represents the shaped residual in the interCU in the input codeword representation, Fwd() represents the forward shaping function, and PredSample represents the interprediction sample in the first codeword representation. The method described in Embodiment 1. [Aspect 4] The output sample (RecSampleInDPB) stored in the aforementioned reference buffer can be generated: RecSampleInDPB=LPF(Inv(RecSample)) This includes calculating, Here, Inv() represents the inverse reshaping function, and LPF() represents the loop filter. The method according to embodiment 2 or 3. [Aspect 5] Regarding the chroma residual samples in the intercoded CU (interCU) in the input codeword representation, further: A step of determining the chroma pixel value in the input codeword representation and the chroma scaling factor based on the formatting metadata; The steps include: multiplying the chroma residual sample in the interCU by the chroma scaling factor to generate a scaled chroma residual sample in the interCU in the first codeword representation; A step of generating a reconfigured chroma sample of the interCU based on the scaled chroma residual in the interCU and the chroma interpredicted sample stored in the reference buffer, thereby generating a decoded chroma sample of the interCU; The steps include: applying the loop filter (270) to the decoded chroma sample of the interCU to generate an output chroma sample of the interCU; The step includes storing the output chroma sample of the interCU in the reference buffer, The method described in Embodiment 1. [Aspect 6] The method according to embodiment 5, wherein in intra-mode, the chroma scaling factor is based on the mean of the intra-predicted chroma values. [Aspect 7] The method according to embodiment 5, wherein, in inter-mode, the chroma scaling factor is based on the mean of the inter-predicted chroma values ​​in the input codeword representation. [Aspect 8] The aforementioned formatted metadata is: A first parameter indicating the number of bins used to represent the first codeword representation; The second parameter, indicating the minimum bin index used in shaping; A first set of parameters that represent the absolute delta codeword value for each bin in the aforementioned input codeword representation; A second set of parameters indicating the sign of the delta codeword value for each bin in the aforementioned input codeword representation, The method described in Embodiment 1. [Aspect 9] The method according to embodiment 1, wherein the forward formatting function is reconstructed as a linear function for each segment having a linear segment derived from the formatting metadata. [Aspect 10] A method for adaptively shaping a video sequence using a processor, the method being: The processor accesses the input image in a first codeword representation; The process includes the step of generating a forward shaping function that maps the pixels of the input image to a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation, and the generation of the forward shaping function is: The input image is divided into multiple pixel regions; Each pixel region is assigned to one of several codeword bins according to the first luminance characteristic of that pixel region; The bin metric for each of the plurality of codeword bins is calculated according to the second luminance characteristics of each of the pixel regions assigned to each of the plurality of codeword bins; The number of codewords in the second codeword representation is assigned to each of the plurality of codeword bins according to the bin metric and rate-distortion optimization criteria for each of the plurality of codeword bins; The process includes generating the forward shaping function in response to the allocation of codewords in the second codeword representation to each of the plurality of codeword bins, method. [Aspect 11] The method according to embodiment 10, wherein the first luminance characteristic of a pixel region includes the average luminance pixel value within that pixel region. [Aspect 12] The method according to embodiment 10, wherein the second luminance characteristic of the pixel region includes the dispersion of the luminance pixel value of the pixel region. [Aspect 13] The method according to embodiment 12, wherein calculating the bin metric for a codeword bin includes calculating the average of the variances of luminance pixel values ​​for all pixel regions assigned to that codeword bin. [Aspect 14] Assigning the number of codewords in the second codeword representation to a codeword bin according to its bin metric is: If no pixel region is assigned to that codeword bin, no codeword is assigned to that codeword bin; If the bin metric of that codeword bin is lower than the upper threshold, assign a first number of codewords; Otherwise, the process involves assigning a second number of codewords to that codeword bin. The method described in aspect 10. [Aspect 15] An apparatus having a processor and configured to perform the method described in any one of embodiments 1 to 14. [Aspect 16] A non-temporary computer-readable storage medium storing computer-executable instructions for performing a method by one or more processors as described in any one of embodiments 1 to 14.

Claims

1. A method for reconstructing encoded video data using one or more processors, the method being: The steps include: receiving an encoded bitstream containing one or more encoded formatted images in an input codeword representation; The steps include receiving formatting metadata for one or more encoded formatted images in the encoded bitstream, wherein the formatting metadata includes parameters for generating a forward formatting function based on the formatting metadata, the forward formatting function maps pixels of an image from a first codeword representation to the input codeword representation, and the formatting metadata is: The first parameter indicates the minimum bin index used in shaping; A second parameter for determining the active maximum bin index used in the shaping, wherein the active maximum bin index is less than or equal to a predefined maximum bin index, and determining the active maximum bin index includes calculating the difference between the predefined maximum bin index and the second parameter, and a delta index parameter; The absolute delta codeword value for each active bin in the aforementioned input codeword representation, and; The input codeword representation includes the sign of the absolute delta codeword value for each active bin, This method further, A step of generating a forward formatting function based on the formatting metadata, wherein the forward formatting function is reconstructed as a segment-specific linear function having linear segments derived from the formatting metadata; A step of generating an inverse formatting function based on the formatting metadata or the forward formatting function, wherein the inverse formatting function maps the pixels of the formatted image from the input codeword representation to the first codeword representation; The step includes decoding the encoded bitstream based on the forward formatting function and the inverse formatting function, method.

2. A method for generating formatting parameters for an encoded bitstream, the method being: The stage of receiving a sequence of video pictures in an input codeword representation; A step of applying a forward shaping function to one or more pictures in the sequence of video pictures to generate a shaped picture in a shaped codeword representation, wherein the forward shaping function maps the pixels of the image from the input codeword representation to the shaped codeword representation, and the forward shaping function is expressed as a segment-by-segment linear function having linear segments derived by shaping parameters; A step of generating the formatting parameters for the formatted codeword representation; The steps include at least the step of generating an encoded bitstream based on the shaped picture, wherein the shaping parameters are: A first parameter for determining the active maximum bin index used for shaping, wherein the active maximum bin index is less than or equal to a predefined maximum bin index; A second parameter indicating the minimum bin index used in the aforementioned shaping; The absolute delta codeword value for each active bin in the aforementioned formatted codeword representation, and; The code of the absolute delta codeword value for each active bin in the formatted codeword representation, method.

3. A method for transmitting a bitstream generated by an encoding device, the method comprising transmitting the bitstream and encoding the bitstream: The stage of receiving a sequence of video pictures in an input codeword representation; A step of applying a forward shaping function to one or more pictures in the sequence of video pictures to generate a shaped picture in a shaped codeword representation, wherein the forward shaping function maps the pixels of the image from the input codeword representation to the shaped codeword representation, and the forward shaping function is expressed as a segment-by-segment linear function having linear segments derived by shaping parameters; The steps include: generating formatting parameters for the formatted codeword representation; The steps include at least the step of generating an encoded bitstream based on the shaped picture, wherein the shaping parameters are: A first parameter for determining the active maximum bin index used for shaping, wherein the active maximum bin index is less than or equal to a predefined maximum bin index; A second parameter indicating the minimum bin index used in the aforementioned shaping; The absolute delta codeword value for each active bin in the aforementioned formatted code representation, and; The code of the absolute delta codeword value for each active bin in the formatted codeword representation, method.

4. A method for reconstructing encoded video data using one or more processors, the method being: The steps include receiving an encoded bitstream containing one or more encoded formatted images in a formatted codeword representation; The steps include receiving formatting metadata for one or more encoded formatted images in the encoded bitstream, wherein the formatting metadata includes parameters for generating a forward formatting function based on the formatting metadata, the forward formatting function is reconstructed as a segment-by-segment linear function having linear segments derived from the formatting metadata, and the formatting metadata is: The first parameter indicates the minimum bin index used in shaping; A second parameter for determining the active maximum bin index used in the shaping, wherein the active maximum bin index is less than or equal to a predefined maximum bin index, and determining the active maximum bin index includes calculating the difference between the predefined maximum bin index and the second parameter; The absolute delta codeword value for each active bin in the aforementioned formatted codeword representation, and; The code of the absolute delta codeword value for each active bin in the formatted codeword representation, method.