Integrated image shaping and video encoding
By introducing out-of-loop shaping, in-loop-only intra-frame shaping, and an in-loop shaping architecture for predicting residuals, the image coding process is optimized, solving the problems of low efficiency and high complexity in high bit-depth image coding, and achieving more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2018-06-29
- Publication Date
- 2026-05-12
AI Technical Summary
Existing video coding technologies are inefficient and complex when processing high bit depth images, especially in inter-frame prediction where the computational cost is too high.
By employing an out-of-loop shaping architecture, an in-loop-only intra-frame shaping architecture, and an in-loop shaping architecture for predicting residuals, combined with forward and inverse shaping functions, the image encoding and decoding process is optimized, reducing complexity and improving encoding efficiency.
By integrating signal shaping and coding techniques, the coding efficiency of high bit depth images is improved, and the complexity of video coding is reduced, especially in inter-frame prediction where computational costs are significantly reduced.
Smart Images

Figure CN116095314B_ABST
Abstract
Description
[0001] Case Separation Statement
[0002] This application is a divisional application of the invention patent application entitled "Integrated Image Shaping and Video Coding", which has PCT international application number PCT / US2018 / 040287, international application date of June 29, 2018, application number 201880012069.2 in the Chinese national phase.
[0003] Cross-reference to related applications
[0004] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 686,738, filed June 19, 2018; Serial No. 62 / 680,710, filed June 5, 2018; Serial No. 62 / 629,313, filed February 12, 2018; Serial No. 62 / 561,561, filed September 21, 2017; and Serial No. 62 / 526,577, filed June 29, 2017, each of which is incorporated herein by reference in its entirety. Technical Field
[0005] This invention generally relates to image and video coding. More specifically, embodiments of the invention relate to integrated image shaping and video coding. Background Technology
[0006] In 2013, the MPEG expert group within the International Organization for Standardization (ISO), together with the International Telecommunication Union (ITU), released an initial draft of the HEVC (also known as H.265) video coding standard. More recently, the MPEG expert group released evidence to support the development of a next-generation coding standard that offers improved coding performance compared to existing video coding technologies.
[0007] As used in this article, the term 'bit depth' refers to the number of pixels used to represent one of the color components of an image. Traditionally, images are encoded with 8 bits per pixel per color component (e.g., 24 bits per pixel); however, modern architectures can now support higher bit depths, such as 10 bits, 12 bits, or more.
[0008] In a conventional image pipeline, a nonlinear photoelectric function (OETF) is used to quantize the captured image. This OETF converts linear scene light into a nonlinear video signal (e.g., gamma-coded RGB or YCbCr). The signal is then processed at a receiver by an electro-optical conversion function (EOTF) before being displayed on a monitor. This EOTF converts the video signal values into output screen color values. Such nonlinear functions include the conventional “gamma” curve described in ITU-R Rec. BT. 709 and BT. 2020, and the “PQ” (perceptual quantization) curve described in SMPTE ST 2084 and Rec. ITU-R BT. 2100.
[0009] As used herein, the term "forward reshaping" refers to the process of mapping a digital image from its original bit depth and original codeword distribution or representation (e.g., gamma or PQ) to images with the same or different bit depths and different codeword distributions or representations, either sample-to-sample or codeword-to-codeword. Reshaping allows for improved compressibility or image quality at a fixed bit rate. For example, without limitation, reshaping can be applied to 10-bit or 12-bit PQ-coded HDR video to improve coding efficiency in a 10-bit video coding architecture. In the receiver, after decompressing the shaped signal, the receiver can apply an "inverse reshaping function" to restore the signal to its original codeword distribution. As understood herein by the inventors, improved techniques for reshaping and coding images are desired as development begins for next-generation video coding standards. The methods of this invention can be applied to a variety of video content, including but not limited to content in Standard Dynamic Range (SDR) and / or High Dynamic Range (HDR).
[0010] The methods described in this section are methods that can be sought, but are not necessarily methods that have been previously conceived or sought. Therefore, unless otherwise specified, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise specified, problems identified with respect to one or more methods should not be considered to be in any prior art based on this section. Summary of the Invention
[0011] A first aspect of this disclosure relates to a method for encoding an image using a processor, the method comprising: accessing an input image represented by a first codeword using the processor; generating a forward shaping function that maps pixels of the input image to a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation; generating an inverse shaping function based on the forward shaping function, wherein the inverse shaping function maps pixels from the second codeword representation to the first codeword representation; for an input pixel region in the input image; calculating a predicted region based on pixel data in a reference frame buffer or in a previously encoded spatial neighborhood; generating a shaped residual region based on the input pixel region, the predicted region, and the forward shaping function; generating a quantized residual region based on the shaped residual region; generating a dequantized residual region based on the quantized residual region; generating a reconstructed pixel region based on the dequantized residual region, the predicted region, the forward shaping function, and the inverse shaping function; and generating a reference pixel region to be stored on the reference frame buffer based on the reconstructed pixel region.
[0012] A second aspect of this disclosure relates to a method for decoding an encoded bitstream using a processor to generate an output image using a first codeword representation. The method may include: receiving an encoded image partially encoded using a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation; receiving shaping information of the encoded image; generating a forward shaping function based on the shaping information that maps pixels from the first codeword representation to the second codeword representation; generating an inverse shaping function based on the shaping information, wherein the inverse shaping function maps pixels from the second codeword representation to the first codeword representation; for a region of the encoded image; generating a decoded and shaped residual region; generating a prediction region based on pixels in a reference pixel buffer or in a previously decoded spatial neighborhood; generating a reconstructed pixel region based on the decoded and shaped residual region, the prediction region, the forward shaping function, and the inverse shaping function; generating an output pixel region of the output image based on the reconstructed pixel region; and storing the output pixel region in the reference pixel buffer.
[0013] A third aspect of this disclosure relates to a method for decoding an encoded bitstream using a processor to generate an output image represented by a first codeword, the method comprising: receiving an encoded image partially encoded using a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation; receiving shaping information of the encoded image; generating a shaping scaling function based on the shaping information; for regions of the encoded image; generating decoded and shaped residual regions; generating prediction regions based on pixels in a reference pixel buffer or in a previously decoded spatial neighborhood; generating reconstructed pixel regions based on the decoded and shaped residual regions, the prediction regions, and the shaping scaling function; generating output pixel regions of the output image based on the reconstructed pixel regions; and storing the output pixel regions in the reference pixel buffer.
[0014] A fourth aspect of this disclosure relates to a method for encoding an image using a processor, the method comprising: accessing an input image represented by a first codeword using the processor; selecting a shaping architecture from two or more candidate coding architectures for compressing the input image using a second codeword representation, wherein the second codeword representation allows for more efficient compression than the first codeword representation, wherein the two or more candidate coding architectures include an out-of-loop shaping architecture, an in-loop shaping architecture for intra-frame use only, and an in-loop architecture for predicting residuals; and compressing the input image according to the selected shaping architecture.
[0015] A fifth aspect of this disclosure relates to a method for decoding an encoded bitstream using a processor to generate an output image represented by a first codeword, the method comprising: receiving an encoded bitstream comprising one or more encoded images, wherein at least a portion of the encoded images is represented by a second codeword, wherein the second codeword representation allows for more efficient compression than the first codeword representation; determining a shaping decoder architecture based on metadata in the encoded bitstream, wherein the shaping decoder architecture includes one of an out-of-loop shaping architecture, an in-of-loop shaping architecture for intra-frame only, or an in-of-loop architecture for predicting residuals; receiving shaping information of the encoded images in the encoded bitstream; and decompressing the encoded images according to the shaping decoder architecture to generate the output image.
[0016] A sixth aspect of this disclosure relates to an apparatus for image reshaping, the apparatus comprising: one or more processors, and a memory having software instructions stored thereon, which, when executed by the one or more processors, cause the execution of a method according to this disclosure.
[0017] A seventh aspect of this disclosure relates to a non-transitory computer-readable storage medium that may have computer-executable instructions stored thereon for performing a method according to this disclosure. Attached Figure Description
[0018] Embodiments of the invention are shown in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:
[0019] Figure 1A An example process of a video transmission pipeline is described;
[0020] Figure 1B An example process for data compression using signal shaping according to existing technology is described;
[0021] Figure 2A An example architecture of an encoder using canonical out-of-loop shaping is depicted according to an embodiment of the present invention;
[0022] Figure 2B An example architecture of a decoder using canonical out-of-loop shaping is described according to an embodiment of the present invention;
[0023] Figure 2C An example architecture of an encoder using only intra-frame in-loop shaping according to an embodiment of the present invention is described.
[0024] Figure 2D An example architecture of a decoder using only intra-frame in-loop shaping according to an embodiment of the present invention is described.
[0025] Figure 2E An example architecture of an encoder for in-loop shaping for predicting residuals is described according to an embodiment of the present invention;
[0026] Figure 2F An example architecture of a decoder using in-loop shaping for predicting residuals is described according to an embodiment of the present invention;
[0027] Figure 2G An example architecture of an encoder using hybrid in-loop shaping according to an embodiment of the present invention is depicted;
[0028] Figure 2H An example architecture of a decoder using hybrid in-ring shaping according to an embodiment of the present invention is described;
[0029] Figure 3A An example process for encoding video using an out-of-ring shaping architecture according to an embodiment of the present invention is described;
[0030] Figure 3BAn example process for decoding video using an out-of-ring shaping architecture according to an embodiment of the present invention is described;
[0031] Figure 3C An example process for encoding video using an in-ring intra-frame-only shaping architecture according to an embodiment of the present invention is described;
[0032] Figure 3D An example process for decoding video using an in-ring intra-frame-only shaping architecture according to an embodiment of the present invention is described;
[0033] Figure 3E An example process for encoding video using an in-loop shaping architecture for predicting residuals, according to an embodiment of the present invention, is described.
[0034] Figure 3F An example process for decoding video using an in-loop shaping architecture for predicting residuals, according to an embodiment of the present invention, is described.
[0035] Figure 4A An example process for encoding video using any one of three shape-based architectures or a combination thereof, according to embodiments of the present invention, is described.
[0036] Figure 4B An example process for decoding video using any one of three shape-based architectures or a combination thereof, according to embodiments of the present invention, is described.
[0037] Figure 5A and Figure 5B The process of reconstructing the image using a shaping function in a video decoder according to an embodiment of the present invention is described;
[0038] Figure 6A and Figure 6B An example is depicted showing how the chroma QP offset value changes according to the luminance quantization parameter (QP) of the PQ encoded signal and the HLG encoded signal, according to an embodiment of the invention; and
[0039] Figure 7 An example of a pivot-based representation of an integer function according to an embodiment of the present invention is described. Detailed Implementation
[0040] This document describes standardized signal shaping and coding techniques for both out-of-loop and in-loop integration in image compression. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in detail to avoid unnecessarily obscuring, obscuring, or confusing the invention.
[0041] Overview
[0042] The example embodiments described herein relate to signal shaping and encoding for video integration. In the encoder, a processor receives an input image represented using a first codeword, which is represented by an input bit depth N and an input codeword mapping (e.g., gamma, PQ, etc.). The processor selects an encoder architecture (where the shaper is a component of the encoder) from two or more candidate encoder architectures for compressing the input image using a second codeword representation that allows for more efficient compression than the first codeword representation. The two or more candidate encoder architectures include an out-of-loop shaping architecture, an in-loop shaping architecture for intra-frame images only, or an in-loop architecture for predicting residuals. The processor compresses the input image according to the selected encoder architecture.
[0043] In another embodiment, a decoder for generating an output image represented using a first codeword receives an encoded bitstream, wherein at least a portion of the encoded image is compressed using a second codeword. The decoder also receives associated shaping information. A processor receives signaling indicating decoder architectures from two or more candidate decoder architectures for decompressing the input encoded bitstream, wherein the two or more candidate decoder architectures include an out-of-loop shaping architecture, an in-loop shaping architecture for intra-frame images only, or an in-loop architecture for predicting residuals, and the processor decompresses the encoded image according to the received shaping architecture to generate the output image.
[0044] In another embodiment, in an encoder for compressing an image based on an in-loop architecture for predicting residuals, a processor accesses an input image represented using a first codeword and generates a forward shaping function that maps the pixels of the input image from the first codeword representation to a second codeword representation. The processor then uses the forward shaping function to generate an inverse shaping function that maps pixels represented by the second codeword back to pixels represented by the first codeword. Then, for an input pixel region in the input image, the processor performs the following operations:
[0045] Calculate at least one prediction region based on pixel data in the reference frame buffer or in the previously encoded spatial neighborhood;
[0046] The shaped residual region is generated based on the input pixel region, the prediction region, and the forward shaping function.
[0047] Encoded (transformed and quantized) residual regions are generated based on shaped residual regions;
[0048] The decoded (inversely quantized and inversely transformed) residual region is generated based on the encoded residual region;
[0049] The reconstructed pixel region is generated based on the decoded residual region, the predicted region, the forward shaping function, and the inverse shaping function; and
[0050] The reference pixel region to be stored on the reference frame buffer is generated based on the reconstructed pixel region.
[0051] In another embodiment, in a decoder used to generate an output image represented using a first codeword based on an in-loop architecture for predicting residuals, the processor receives a portion of an encoded bitstream encoded using a second codeword representation. The processor also receives associated shaping information. Based on the shaping information, the processor generates a forward shaping function and a backward shaping function, wherein the forward shaping function maps pixels from the first codeword representation to the second codeword representation, and the backward shaping function maps pixels from the second codeword representation back to the first codeword representation. For a region of the encoded image, the processor performs the following operations:
[0052] Generate decoded and shaped residual regions based on coded images;
[0053] The prediction region is generated based on pixels in the reference pixel buffer or in the previously decoded spatial neighborhood.
[0054] The reconstructed pixel region is generated based on the decoded and shaped residual region, the predicted region, the forward shaping function, and the inverse shaping function.
[0055] The output pixel region is generated based on the reconstructed pixel region; and the output pixel region is stored in the reference pixel buffer.
[0056] Example video transmission and processing pipeline
[0057] Figure 1A An example process of a conventional video transmission pipeline (100) is depicted, illustrating the various stages from video capture to video content display. An image generation frame (105) is used to capture or generate a sequence of video frames (102). The video frames (102) can be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data (107). Alternatively, the video frames (102) can be captured on film by a film camera. The film is converted to a digital format to provide video data (107). In the production stage (110), the video data (107) is edited to provide a video production stream (112).
[0058] The video data (112) is then provided to the processor at frame (115) for post-production editing. Post-production editing at frame (115) may include adjusting or modifying the color or brightness in specific areas of the image to enhance image quality or achieve a specific look for the image according to the video creator's creative intent. This is sometimes referred to as "color timing" or "color grading". Other edits (e.g., scene selection and sorting, image cropping, adding computer-generated visual effects, etc.) may be performed at frame (115) to produce a final version (117) of the product for distribution. During post-production editing (115), the video image is viewed on a reference monitor (125).
[0059] After post-production (115), the video data of the final product (117) can be transferred to an encoding frame (120) for downstream transmission to decoding and playback devices such as televisions, set-top boxes, and cinemas. In some embodiments, the encoding frame (120) may include audio encoders and video encoders such as those defined by ATSC, DVB, DVD, Blu-ray, and other transmission formats to generate an encoded bitstream (122). In a receiver, the encoded bitstream (122) is decoded by a decoding unit (130) to generate a decoded signal (132) that is exactly the same as or closely approximated by the signal (117). The receiver may be attached to a target display (140), which may have characteristics completely different from those of a reference display (125). In this case, the display management frame (135) may be used to map the dynamic range of the decoded signal (132) to the characteristics of the target display (140) by generating a display mapping signal (137).
[0060] signal shaping
[0061] Figure 1BAn example process for signal shaping according to the prior art (Reference [1]) is described. Given an input frame (117), a forward shaping frame (150) analyzes the input constraints and encoding constraints and generates a codeword mapping function that maps the input frame (117) to a requantized output frame (152). For example, the input (117) can be encoded according to some electro-optical conversion function (EOTF) (e.g., gamma). In some embodiments, metadata can be used to transmit information about the shaping process to downstream devices (such as decoders). As used herein, the term “metadata” refers to any auxiliary information transmitted as part of the encoded bitstream and used by the decoder to render the decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and auxiliary signal parameters as described herein.
[0062] After encoding (120) and decoding (130), the decoded frame (132) can be processed by a backward (or inverse) shaping function (160) that converts the requantized frame (132) back to the original EOTF domain (e.g., gamma) for further downstream processing, such as the display management process (135) discussed earlier. In some embodiments, the backward shaping function (160) can be integrated with a dequantizer in the decoder (130), for example, as part of a dequantizer in an AVC or HEVC video decoder.
[0063] As used herein, the term "shaper" can refer to a forward shaping function or an inverse shaping function used when encoding and / or decoding a digital image. Examples of shaping functions are discussed in references [1] and [2]. For the purposes of this invention, it is assumed that those skilled in the art can determine suitable forward and inverse shaping functions based on the characteristics of the input video signal and the available bit depth of the encoding and decoding architectures.
[0064] In reference [1], a block-based in-loop image shaping method for high dynamic range video coding is proposed. This design allows block-based shaping within the coding loop, but at the cost of increased complexity. Specifically, the design requires maintaining two sets of decoded image buffers: one set for decoded images that have been inversely shaped (or not shaped), which can be used for prediction without shaping and for output to the display; and another set for decoded images that have been forward shaped, which are only used for prediction with shaping. Although the forward-shaped decoded images can be computed in real time, the complexity is very high, especially for inter-frame prediction (using motion compensation with subpixel interpolation). Typically, display image buffer (DPB) management is complex and requires great care; therefore, as the inventors understand, a simplified method for encoding video is desired.
[0065] The embodiments of the shaping-based codec architecture proposed in this paper can be categorized as follows: an architecture with an external out-of-loop shaper, an architecture with an in-loop intra-frame-only shaper, and an architecture with an in-loop shaper for predicting residuals (also referred to as an 'in-loop residual shaper'). Video encoders or decoders can support any one of these architectures or a combination thereof. Each of these architectures can also be applied individually or in combination with any of the other architectures. Each architecture can be applied to the luminance component, chrominance component, or a combination of luminance and one or more chrominance components.
[0066] In addition to these three architectures, the additional embodiments describe efficient signaling methods for metadata related to shaping, as well as several encoder-based optimization tools for improving coding efficiency when applying shaping.
[0067] Standard external ring shaper
[0068] Figure 2A and Figure 2B An architecture for a video encoder (200A_E) and a corresponding video decoder (200A_D) with a "normative" out-of-loop shaper is depicted. The term "normative" indicates a departure from previous designs where shaping was considered a preprocessing step and therefore outside the descriptions of encoding standards such as AVC and HEVC; in this embodiment, forward shaping and reverse shaping are part of the specification requirements. Figure 1B The architecture for testing bitstream consistency according to the standard after decoding (130) differs. Figure 2B In the reverse shaping frame (265) (for example, in Figure 1B After outputting at point 162 in the code, we tested for consistency.
[0069] In the encoder (200A_E), two new boxes are added to a conventional block-based encoder (e.g., HEVC): a box (205) for estimating the forward shaping function, and a forward image shaping box (210) for applying forward shaping to one or more color components of the input video (117). In some embodiments, these two operations may be performed as part of a single image shaping box. Parameters (207) associated with determining the inverse shaping function in the decoder may be passed to a lossless encoder block of the video encoder (e.g., CABAC 220) such that the parameters can be embedded in the encoded bitstream (122). The shaped image stored in the DPB (215) is used to perform intra-frame or inter-frame prediction (225), transform and quantization (T and Q), inverse transform and inverse quantization (Q). -1 and T -1And all operations related to loop filtering.
[0070] In the decoder (200A_D), two new canonical boxes are added to the conventional block-based decoder: a box (250) for reconstructing the inverse shaping function based on the encoded shaping function parameters (207), and a box (265) for applying the inverse shaping function to the decoded data (262) to generate the decoded video signal (162). In some embodiments, the operations associated with boxes 250 and 265 can be combined into a single processing box.
[0071] Figure 3A An example process (300A_E) for encoding video using an out-of-loop shaping architecture (200A_E) according to an embodiment of the invention is described. If shaping is not enabled (path 305), encoding is performed as is known in prior art encoders (e.g., HEVC). If shaping is enabled (path 310), the encoder may have the option to either apply a pre-defined (default) shaping function (315) or adaptively determine a new shaping function (325) based on image analysis (320) (e.g., as described in references [1] to [3]). After forward shaping (330), the remainder of the encoding follows a conventional encoding pipeline (335). If adaptive shaping (312) is employed, metadata associated with the inverse shaping function is generated as part of the “encoder shaper” step (327).
[0072] Figure 3B An example process (300A_D) for decoding video using an out-of-loop shaping architecture (200A_D) according to an embodiment of the present invention is described. If shaping is not enabled (path 355), an output frame is generated (390) after decoding the image (350), as in a conventional decoding pipeline. If shaping is enabled (path 360), in step (370), the decoder determines whether to apply a predetermined (default) shaping function (375) or adaptively determine an inverse shaping function based on received parameters (e.g., 207) (380). After inverse shaping (385), the remainder of the decoding follows a conventional decoding pipeline.
[0073] Standard intra-ring only intra-frame shaper
[0074] Figure 2C An example architecture for an encoder (200B_E) using only intra-frame in-loop shaping according to an embodiment of the invention is depicted. The design is very similar to the design proposed in reference [1]; however, to reduce complexity, especially when dealing with the use of DPB memories (215 and 260), only intra-frame images are encoded using this architecture.
[0075] The main difference between encoder 200B_E and out-of-loop shaping (200A_E) is that the DPB (215) stores the inversely shaped image instead of the shaped image. In other words, the decoded intra-frame image needs to be inversely shaped (by inverse shaping unit 265) before being stored in the DPB. The reason behind this approach is that if the intra-frame image is encoded using shaping, the improved performance of encoding the intra-frame image will propagate to (implicitly) improve the encoding of the inter-frame image, even if the inter-frame image is not encoded using shaping. In this way, one can utilize shaping without dealing with the complexity of in-loop shaping of the inter-frame image. Since inverse shaping (265) is part of the inner loop, it can be performed before the in-loop filter (270). The advantage of adding inverse shaping before the in-loop filter is that, in this case, the design of the in-loop filter can be optimized based on the characteristics of the original image rather than the characteristics of the forward-shaped image.
[0076] Figure 2D An example architecture for a decoder (200B_D) using only intra-frame in-ring shaping according to an embodiment of the present invention is described. Figure 2D As described in the diagram, the determination of the inverse shaping function (250) and the application of inverse shaping (265) are now performed before the in-loop filtering (270).
[0077] Figure 3C An example process (300B_E) for encoding video using an in-ring-only intra-frame shaping architecture according to an embodiment of the present invention is depicted. As depicted, Figure 3C The operation process and Figure 3A The operational flow shares many elements. Currently, by default, shaping is not applied to inter-frame coding. For intra-frame coded images, if shaping is enabled, the encoder again has the option to use the default shaping curve or apply adaptive shaping (312). If the image is shaped, inverse shaping (385) is part of the process, and associated parameters are encoded in step (327). Figure 3D The corresponding decoding process (300B_D) is described in the text.
[0078] like Figure 3D As described, shaping-related operations are enabled only for received intra-frame images and only when intra-frame shaping is applied on the encoder.
[0079] In-loop shaper for predicting residuals
[0080] In encoding, the term 'residual' refers to the difference between a prediction of a sample or data element and its original or decoded value. For example, given an original sample (denoted as Orig_sample) from an input video (117), intra-frame or inter-frame prediction (225) can generate a corresponding predicted sample (227) denoted as Pred_sample. Without shaping, the unshaped residual (Res_u) can be defined as:
[0081] Res_u=Orig_sample-Pred_sample. (1)
[0082] In some embodiments, applying shaping to the residual domain may be beneficial. Figure 2E An example architecture for an encoder (200C_E) using in-loop shaping for predicting residuals, according to an embodiment of the present invention, is depicted. Let Fwd() denote the forward shaping function, and let Inv() denote the corresponding inverse shaping function. In the embodiment, the shaped residual (232) can be defined as:
[0083] Res_r=Fwd(Orig_sample)-Fwd(Pred_sample). (2)
[0084] Accordingly, at the output (267) of the inverse shaper (265), the reconstructed sample (267) represented as Reco_sample can be expressed as:
[0085] Reco_sample=Inv(Res_d+Fwd(Pred_sample)), (3)
[0086] Wherein, Res_d represents the residual (234) after in-ring encoding and decoding in 200C_E, which is a close approximation of Res_r.
[0087] Note that although shaping is applied to the residual, the actual input video pixels are not shaped. Figure 2F The corresponding decoder (200C_D) is described. Note that, as Figure 2F As described in and based on equation (3), the decoder needs to access both the forward shaping function and the reverse shaping function, which can be extracted using the received metadata (207) and the “shaper decoding” box (250).
[0088] In this embodiment, to reduce complexity, equations (2) and (3) can be simplified. For example, assuming the forward shaping function can be approximated by a piecewise linear function and the absolute difference between Pred_sample and Orig_sample is relatively small, then equation (2) can be approximated as:
[0089] Res_r=a(Pred_sample)*(Orig_samlpe-Pred_sample), (4)
[0090] Where a(Pred_sample) represents the scaling factor based on the value of Pred_sample. According to equations (3) and (4), equation (3) can be approximated as:
[0091] Reco_sample=Pred_sample+(1 / a(Pred_sample))*Res_r, (5)
[0092] Therefore, in this embodiment, only the scaling factor a(Pred_sample) for the piecewise linear model needs to be sent to the decoder.
[0093] Figure 3E and Figure 3F An example process flow for encoding (300C_E) and decoding (300C_D) video using in-loop shaping of prediction residuals is described. The process is similar to... Figure 3A and Figure 3B The processes described in the text are very similar and therefore do not require further explanation.
[0094] Table 1 summarizes the key features of the three proposed architectures.
[0095] Table 1: Key features of the considered shaping architecture
[0096]
[0097] Figure 4A and Figure 4B Example encoding and decoding processes are described using a combination of the three proposed architectures for encoding and decoding. Figure 4A As described, if shaping is not enabled, the input video is encoded according to known video coding techniques (e.g., HEVC, etc.) without using any shaping. Otherwise, the encoder can select any of the three methods proposed in this paper based on the capabilities of the target receiver and / or the characteristics of the input. For example, in one embodiment, the encoder can switch between these methods at the scene level, where a 'scene' is represented as a sequence of consecutive frames with similar brightness characteristics. In another embodiment, high-level parameters are defined at the Sequence Parameter Set (SPS) level.
[0098] like Figure 4B As described, the decoder can invoke any decoding process in the corresponding decoding process based on the received shaped information signaling, in order to decode the incoming coded bit stream.
[0099] Intra-ring shaping
[0100] Figure 2G An example architecture (200D_E) for an encoder using a hybrid in-ring shaping architecture is depicted. This architecture combines elements from both the previously discussed in-ring intra-frame shaping architecture (200B_E) and the in-ring residual architecture (200C_E). Under this architecture, according to the in-ring intra-frame shaping coding architecture (e.g., Figure 2C The intra-frame slices are encoded using 200B_E, with one key difference: for intra-frame slices, inverse image shaping (265-1) is performed after loop filtering (270-1). In another embodiment, intra-loop filtering can be performed on the intra-frame slices after inverse shaping; however, experimental results show that this arrangement may result in worse coding efficiency compared to performing inverse shaping after loop filtering. The remaining operations remain the same as discussed above.
[0101] As discussed above, based on the in-ring residual coding architecture (e.g., Figure 2E The 200C_E in the code is used to encode inter-frame slices. For example... Figure 2G As described, intra / inter-frame slice switching allows switching between these two architectures depending on the type of slice to be encoded.
[0102] Figure 2H An example architecture (200D_D) for a decoder using hybrid in-ring shaping is depicted. Again, based on the in-ring intra-frame shaping decoder architecture (e.g., Figure 2D The intra-frame slice is decoded using 200B_D, where, again for the intra-frame slice, loop filtering (270-1) precedes inverse image shaping (265-1). This is based on the intra-loop residual coding architecture (e.g., Figure 2F The 200C_D in the code is used to decode inter-frame slices. For example... Figure 2H As described, intra / inter-frame slice switching allows switching between these two architectures based on the slice type in the encoded video picture.
[0103] By calling Figure 2G The encoding process described in 300D-E, Figure 4A This can be easily extended to include methods for encoding integers within mixed rings. Similarly, by calling... Figure 2H The decoding process described in 300D-D, Figure 4B It can be easily extended to include a hybrid ring-in-circle shaping decoding method.
[0104] Slice-level reshaping
[0105] Embodiments of the present invention allow for adaptation at various slice levels. For example, to reduce computation, shaping can be enabled only for intra-frame slices or only for inter-frame slices. In another embodiment, shaping can be enabled based on the value of a time ID (e.g., the HEVC variable TemporalId (reference
[11] ), where TemporalId = nuh_temporal_id_plus1-1). For example, if the TemporalId of the current slice is less than or equal to a predefined value, the slice_reshaper_enable_flag of the current slice can be set to 1; otherwise, the slice_reshaper_enable_flag will be 0. To avoid sending the slice_reshaper_enable_flag parameter for each slice, the sps_reshaper_temporal_id parameter can be specified at the SPS level, thus allowing the value of the slice_reshaper_enable_flag parameter to be inferred.
[0106] For slices with integer shaping enabled, the decoder needs to know which integer model to use. In one embodiment, the integer model defined at the SPS level can always be used. In another embodiment, the integer model defined in the slice header can always be used. If no integer model is defined in the current slice, the integer model used in the most recently decoded slice that used integer shaping can be applied. In yet another embodiment, the integer model can always be specified in the intra-slice, regardless of whether integer shaping is used for the intra-slice. In such an implementation, the parameters slice_reshaper_enable_flag and slice_reshaper_model_present_flag need to be deassociated. Table 5 illustrates an example of this slice syntax.
[0107] Signaling of plastic surgery information
[0108] Information related to forward and / or reverse shaping can reside in different information layers, such as the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Supplementary Information (SEI), or any other high-level syntax. As an example, and not a limitation, Table 2 provides examples of high-level syntax in the SPS used to signal whether shaping is enabled, whether shaping is adaptive, and which of the three architectures is being used.
[0109] Table 2: Examples of Integer Information in SPS
[0110]
[0111] Additional information can also be carried at another layer (e.g., in the slice header). The shaping function can be described by a lookup table (LUT), a piecewise polynomial, or other kind of parametric model. The type of shaping model used to transmit the shaping function can be signaled by additional syntax elements (e.g., the reshaping_model_type flag). For example, consider two systems with different representations: model_A (e.g., reshaping_model_type = 0) represents the shaping function as a set of piecewise polynomials (e.g., see reference [4]), while in model_B (e.g., reshaping_model_type = 1), the shaping function is adaptively obtained by assigning codewords to different brightness bands based on the image brightness characteristics and visual importance (e.g., see reference [3]). Table 3 provides examples of syntax elements in the slice header of an image to help the decoder determine the appropriate shaping model being used.
[0112] Table 3: Example Syntax for Shaping Signaling in Slice Headers
[0113]
[0114] The following three tables describe alternative examples of bitstream syntaxes for signal shaping at the sequence layer, slice layer, or code tree unit (CTU) layer.
[0115] Table 4: Examples of Integer Information in SPS
[0116]
[0117] Table 5: Example Syntax for Shaping Signaling in Slice Headers
[0118]
[0119] Table 6: Example Syntax for Shaping Signaling in CTU
[0120]
[0121] For Tables 4 to 6, the example semantics can be represented as:
[0122] A value of 1 for `sps_reshaper_enable_flag` indicates that a shaper is used in the encoded video sequence (CVS). A value of 0 for `sps_reshaper_enabled_flag` indicates that a shaper is not used in CVS.
[0123] A slice_reshaper_enable_flag value of 1 indicates that the current slice has the reshaper enabled. A slice_reshaper_enable_flag value of 0 indicates that the current slice has the reshaper disabled.
[0124] sps_reshaper_signal_type indicates the original codeword distribution or representation. By way of example and not limitation, sps_reshaper_signal_type equal to 0 specifies SDR (gamma); sps_reshaper_signal_type equal to 1 specifies PQ; and sps_reshaper_signal_type equal to 2 specifies HLG.
[0125] A reshaper_CTU_control_flag value of 1 indicates that the shaper is allowed to be adapted for each CTU. A reshaper_CTU_control_flag value of 0 indicates that the shaper is not allowed to be adapted for each CTU. If reshaper_CUT_control_flag does not exist, its value should be assumed to be 0.
[0126] A reshaper_CTU_flag value of 1 indicates that the shaper is used for the current CTU. A reshaper_CUT_flag value of 0 indicates that the shaper is not used for the current CTU. If reshaper_CTU_flag does not exist, its value should be assumed to be equal to slice_reshaper_enabled_flag.
[0127] A value of 1 for sps_reshaper_model_present_flag indicates that sps_reshaper_model() exists in SPS. A value of 0 for sps_reshaper_model_present_flag indicates that sps_reshaper_model() does not exist in SPS.
[0128] `slice_reshaper_model_present_flag` equal to 1 indicates that `slice_reshaper_model()` exists in the slice header. `slice_reshaper_model_present_flag` equal to 0 indicates that `slice_reshaper_model()` does not exist in the SPS. `sps_reshaper_chromaAdj` equal to 1 indicates that chroma QP adjustment was performed using chroma DQP. `sps_reshaper_chromaAdj` equal to 2 indicates that chroma QP adjustment was performed using chroma scaling.
[0129] `sps_reshaper_ILF_opt` indicates whether the in-loop filter is applied in the raw or integer domain for intra-slices and inter-slices. For example, using a two-bit syntax, where the least significant bit refers to the intra-slice:
[0130]
[0131]
[0132] In some embodiments, this parameter can be adjusted at the slice level. For example, in one embodiment, when slice_reshaper_enable_flag is set to 1, the slice may include slice_reshape_ILFOPT_flag. In another embodiment, in SPS, if sps_reshaper_ILF_opt is enabled, the sps_reshaper_ILF_Tid parameter may be included. If the TemporalID of the current slice is less than or equal to sps_reshaper_ILF_Tid and slice_reshaper_enable_flag is set to 1, then the in-loop filter is applied in the shaped domain. Otherwise, the in-loop filter is applied in the unshaped domain.
[0133] In Table 4, chroma QP adjustment is controlled at the SPS level. In this embodiment, chroma QP adjustment can also be controlled at the slice level. For example, in each slice, when slice_reshaper_enable_flag is set to 1, the syntax element slice_reshape_chromaAdj_flag can be added. In another embodiment, in SPS, if sps_reshaper_chromaAdj is enabled, the syntax element sps_reshaper_ChromaAdj_Tid can be added. If the TemporalID of the current slice is <= sps_reshaper_ChromaAdj_Tid and slice_reshaper_enable_flag is set to 1, then chroma adjustment is applied. Otherwise, chroma adjustment is not applied. Table 4B depicts example variations of Table 4 using the syntax described above.
[0134] Table 4B: Example Syntax for Integer Signaling Using Time ID in SPS
[0135]
[0136] `sps_reshaper_ILF_Tid` specifies the highest Temporal ID for applying the in-loop filter to the shaped slice in the shaped domain. `sps_reshaper_chromaAdj_Tid` specifies the highest Temporal ID for applying chroma adjustment to the shaped slice.
[0137] In another embodiment, the reshape model can be defined using a reshape model ID (e.g., reshape_model_id), for example, as part of the slice_reshape_model() function. The reshape model can be signaled at the SPS level, PPS level, or slice header level. If signaled in the SPS or PPS, the value of reshape_model_id can also be inferred from sps_seq_parameter_set_id or pps_pic_parameter_set_id. Table 5B below shows an example of how to use reshape_model_id for slices that do not carry slice_reshape_model() (e.g., slice_reshaper_model_present_flag equals 0); Table 5B is a variation of Table 5.
[0138] Table 5B: Example syntax for reshape signaling used in slice headers
[0139]
[0140] In the example syntax, the parameter `reshape_model_id` specifies the value of the `reshape_model` being used. The value of `reshape_model_id` should be in the range of 0 to 15.
[0141] As an example of using the proposed syntax, consider an HDR signal encoded with PQ EOTF, where shaping is used at the SPS level, no specific shaping is used at the slice level (shaping is used for all slices), and CTU adaptation is allowed only for inter-frame slices. Then:
[0142] sps_reshaper_signal_type=1(PQ);
[0143] sps_reshaper_model_present_flag=1;
[0144] / / Note: For inter-frame slices, you can manipulate slice_reshaper_enable_flag to enable and disable the reshaper.
[0145]
[0146] In another example, consider an SDR signal where shaping is applied only at the slice level and only for intra slices. CTU shaping adaptation is only allowed for inter slices. Then:
[0147]
[0148] At the CTU level, in an embodiment, CTU-level shaping can be enabled based on the lightness characteristic of the CTU. For example, for each CTU, the average lightness (e.g., CTU_avg_lum_value) can be calculated, compared with one or more thresholds, and based on the results of these comparisons, it is decided whether to turn on or off shaping. For example,
[0149] if CTU_avg_lum_value < THR1, or
[0150] if CTU_avg_lum_value > THR2, or
[0151] if THR3 < CTU_avg_lum_value < THR4,
[0152] then for this CTU, reshaper_CTU_Flag = 1.
[0153] In an embodiment, instead of using the average lightness, some other lightness characteristic of the CTU can be used, such as the minimum lightness, the maximum lightness, or the average lightness, variance, etc. The chroma-based characteristics of the CTU can also be applied, or the lightness characteristics and chroma characteristics can be combined with thresholds.
[0154] As described above (e.g., regarding Figure 3A 、 Figure 3B and Figure 3CIn the steps described above, the implementation may support default or static shaping functions, or adaptive shaping. A "default shaper" can be used to execute predefined shaping functions, thus reducing the complexity of analyzing each image or scene when obtaining the shaping curve. In this case, it is not necessary to signal the inverse shaping function at the scene, image, or slice level. The default shaper can be implemented by using a fixed mapping curve stored in the decoder to avoid any signaling, or it can be signaled once as part of a sequence-level parameter set. In another embodiment, a previously decoded adaptive shaping function can be reused for subsequent images in the encoding order. In another embodiment, the shaping curve can be signaled in a manner different from the previously decoded method. In other embodiments (e.g., for in-loop residual shaping that requires both the Inv() and Fwd() functions to perform inverse shaping), only one of the Inv() or Fwd() functions can be signaled in the bitstream, or alternatively, both functions can be signaled to reduce decoder complexity. Tables 7 and 8 provide two examples of signaling shaping information.
[0155] In Table 7, the integer functions are transmitted as a set of second-order polynomials. This is a simplified syntax of the Exploratory Test Model (ETM) (Reference [5]). Earlier variants can also be found in Reference [4].
[0156] Table 7: Example syntax for piecewise representation of integer functions (model_A)
[0157]
[0158] reshape_input_luma_bit_depth_minus8 specifies the sample bit depth of the input luminance component in the reshaping process.
[0159] `coeff_log2_offset_minus2` specifies the number of decimal places used in calculating the luminance component's shaped correlation coefficient. The value of `coeff_log2_offset_minus2` should be in the range of 0 to 3 (inclusive).
[0160] Incrementing `reshape_num_ranges_minus1` by 1 specifies the number of ranges in the piecewise integer function. If `reshape_num_ranges_minus1` does not exist, its value is presumed to be 0. `reshape_num_ranges_minus1` should be in the range of 0 to 7 (including the endpoints for the luminance component).
[0161] When `reshape_equal_ranges_flag` equals 1, the piecewise integer function is divided into NumberRanges segments of nearly equal length, and the length of each range is not explicitly signaled. When `reshape_equal_ranges_flag` equals 0, the length of each range is explicitly signaled.
[0162] The `reshape_global_offset_val` function is used to obtain the offset value used to specify the starting point of the 0th range.
[0163] reshape_range_val[i] is used to obtain the length of the i-th range of the luminance component.
[0164] `reshape_continuity_flag` specifies the continuity property of the shaping function for the luminance component. If `reshape_continuity_flag` equals 0, zero-order continuity is applied to the piecewise linear inverse shaping function between continuous pivot points. If `reshape_continuity_flag` equals 1, first-order smoothness is used to obtain the complete second-order polynomial inverse shaping function between continuous pivot points.
[0165] reshape_poly_coeff_order0_int[i] specifies the integer value of the coefficient of the 0th order polynomial of the i-th segment of the luminance component.
[0166] reshape_poly_coeff_order0_frac[i] specifies the small value of the coefficient of the 0th order polynomial in the i-th segment of the luminance component.
[0167] `reshape_poly_coeff_order1_int` specifies the integer values of the first-order polynomial coefficients for the luminance component.
[0168] reshape_poly_coeff_order1_frac specifies the small value of the first-order polynomial coefficients of the luminance component.
[0169] Table 8 describes example implementations of alternative parameterized representations of model_B (reference [3]) discussed above.
[0170] Table 8: Example syntax for the parameterized representation (model_B) of integer functions
[0171]
[0172] In Table 8, in the embodiments, the syntax parameter can be defined as: reshape_model_profile_type specifies the distribution type to be used during the shaper construction process.
[0173] `reshape_model_scale_idx` specifies the index value of the scaling factor (denoted as ScaleFactor) to be used during the shaper construction process. The value of ScaleFactor allows for improved control over the shape function, thereby improving overall coding efficiency. Discussions regarding the shape function reconstruction process (e.g., ...) Figure 5A and Figure 5B The description in [the document] provides additional details regarding the use of this ScaleFactor. By way of example and not limitation, the value of `reshape_model_scale_idx` should be in the range of 0 to 3 (inclusive). In the embodiment, the mapping between `scale_idx` and ScaleFactor, as shown in the table below, is given by the following formula:
[0174] ScaleFactor=1.0-0.05*reshape_model_scale_idx.
[0175] reshape_model_scale_idx ScaleFactor 0 1.0 1 0.95 2 0.9 3 0.85
[0176] In another embodiment, for a more efficient fixed-point implementation,
[0177] ScaleFactor=1-1 / 16*reshape_model_scale_idx.
[0178] reshape_model_scale_idx ScaleFactor 0 1.0 1 0.9375 2 0.875 3 0.8125
[0179] `reshape_model_min_bin_idx` specifies the minimum range index to use during the shaper construction process. The value of `reshape_model_min_bin_idx` should be in the range of 0 to 31, inclusive.
[0180] `reshape_model_max_bin_idx` specifies the maximum range index to use during the shaper construction process. The value of `reshape_model_max_bin_idx` should be in the range of 0 to 31, inclusive.
[0181] The `reshape_model_num_band` parameter specifies the number of bandwidths to be used during the shaper construction process. The value of `reshape_model_num_band` should be in the range of 0 to 15, inclusive.
[0182] `reshape_model_band_profile_delta[i]` specifies the increment value to be used to adjust the distribution of the i-th band during the shaper construction process. The value of `reshape_model_band_profile_delta[i]` should be in the range of 0 to 1, inclusive.
[0183] Compared to reference [3], the syntax in Table 8 is far more efficient by defining a set of “default distribution types” (e.g., highlights, midtones, and darks). In the embodiments, each type has a predefined visual band importance distribution. The predefined bands and corresponding distributions can be implemented as fixed values in the decoder, or they can be signaled using high-level syntax (e.g., sequence parameter sets). At the encoder, each image is first analyzed and classified into one of the distribution types. The distribution type is signaled via the syntax element “reshape_model_profile_type”. In adaptive shaping, in order to capture the full range of image dynamics, the default distribution is further adjusted by increments for each luminance band or subset of luminance bands. The increment values are based on the visual importance of the luminance bands and are signaled via the syntax element “reshape_model_band_profile_delta”.
[0184] In one embodiment, the increment value may take only the value 0 or 1. At the encoder, visual importance is determined by comparing the percentage of pixels in a frequency band across the entire image with the percentage of pixels in a frequency band within a "major frequency band," where a local histogram can be used to detect the major frequency band. If pixels within a frequency band are clustered in small local boxes, the frequency band is likely to be visually important within those boxes. The counts for the major frequency bands are summed and normalized to form meaningful comparisons, thus obtaining the increment value for each frequency band.
[0185] In the decoder, the shaper function must be called to reconstruct the LUT based on the method described in reference [3]. Therefore, the complexity is higher than that of a simpler piecewise approximation model that only requires evaluating a piecewise polynomial function to compute the LUT. The advantage of using the parameterized model syntax is that the bit rate of the shaper can be significantly reduced. For example, based on typical test content, the model depicted in Table 7 requires 200 to 300 bits to signal the shaper, while the parameterized model (as shown in Table 8) uses only about 40 bits.
[0186] In another embodiment, as depicted in Table 9, a forward-shaping lookup table can be obtained based on a parameterized model of dQP values. For example, in one embodiment,
[0187] dQP=clip3(min,max,scale*X+offset),
[0188] Where min and max represent the boundaries of dQP, scale and offset are two parameters of the model, and X represents parameters derived from signal brightness (e.g., the brightness value of a pixel, or, for a bounding box, a measure of its brightness (e.g., its minimum, maximum, average, variance, standard deviation, etc.)). For example, non-restrictively,
[0189] dQP=clip3(-3, 6, 0.015*X-7.5).
[0190] Table 9: Example syntax for the parameterized representation of integer functions (Model C)
[0191]
[0192] In this embodiment, the parameters in Table 9 can be defined as follows:
[0193] `full_range_input_flag` specifies the range of the input video signal. A `full_range_input_flag` of 0 corresponds to a standard dynamic range input video signal. A `full_range_input_flag` of 1 corresponds to a full range input video signal. When `full_range_input_flag` is not present, it is inferred to be 0.
[0194] Note: As used in this article, the term "full-range video" means that the valid codewords in the video are not "restricted". For example, for 10-bit full-range video, the valid codewords are between 0 and 1023, where 0 is mapped to the lowest brightness level. Conversely, for 10-bit "standard-range video", the valid codewords are between 64 and 940, and 64 is mapped to the lowest brightness level.
[0195] For example, the calculation of "full range" and "standard range" can be performed as follows:
[0196] For the normalized brightness value Ey' in
[01] , it is encoded in BD bits (e.g., BD = 10, 12, etc.):
[0197] Full range: Y = clip3(0, (1 << BD) - 1, Ey' * ((1 << BD) - 1)))
[0198] Standard range: Y = clip3(0, (1 << BD) - 1, round(1 << (BD - 8) * (219 * Ey' + 16)))
[0199] This syntax is similar to the “video_full_range_flag” syntax in the HEVCVUI parameters described in section E.2.1 of the HEVC(H.265) specification (reference
[11] ).
[0200] dQP_model_scale_int_prec specifies the number of bits used to represent dQP_model_scale_int. dQP_model_scale_int_prec equals 0, indicating that dQP_model_scale_int was not signaled and is inferred to be 0.
[0201] dQP_model_scale_int specifies an integer value for scaling the dQP model.
[0202] dQP_model_scale_frac_prec_minus16+16 specifies the number of bits used to represent dQP_model_scale_frac.
[0203] dQP_model_scale_frac specifies the minimum value for scaling the dQP model.
[0204] The variable dQPModelScaleAbs is obtained as follows:
[0205] dQPModelScaleAbs=dQP_model_scale_int<<(dQP_model_scale_frac_prec_minus16+16)+dQP_model_scale_frac
[0206] dQP_model_scale_sign specifies the sign of the dQP model scaling. When dQPModelScaleAbs equals 0, dQP_model_scale_sign is not signaled and is inferred to be 0.
[0207] Adding 3 to dQP_model_offset_int_prec_minus3 specifies the number of bits used to represent dQP_model_offset_int.
[0208] dQP_model_offset_int specifies an integer value for the dQP model offset.
[0209] Increment 1 by dQP_model_offset_frac_prec_minus1 to specify the number of bits used to represent dQP_model_offset_frac.
[0210] dQP_model_offset_frac specifies a small value for the dQP model offset.
[0211] The variable dQPModelOffsetAbs is obtained as follows:
[0212] dQPModelOffsetAbs=dQP_model_offset_int<<(dQP_model_offset_frac_prec_minus1+1)+dQP_model_offset_frac
[0213] dQP_model_offset_sign specifies the sign of the dQP model offset. When dQPModelOffsetAbs equals 0, dQP_model_offset_sign is not signaled and is inferred to be 0.
[0214] Adding 3 to dQP_model_abs_prec_minus3 specifies the number of bits used to represent dQP_model_max_abs and dQP_model_min_abs.
[0215] dQP_model_max_abs specifies the integer value of the maximum value of the dQP model.
[0216] dQP_model_max_sign specifies the sign of the maximum value of the dQP model. When dQP_model_max_abs equals 0, dQP_model_max_sign is not signaled and is inferred to be 0.
[0217] dQP_model_min_abs specifies the integer value of the minimum value of the dQP model.
[0218] dQP_model_min_sign specifies the sign of the minimum value of the dQP model. When dQP_model_min_abs equals 0, dQP_model_min_sign is not signaled and is inferred to be 0.
[0219] Decoding process of model C
[0220] Given the syntax elements in Table 9, the integer LUT can be obtained as follows.
[0221] The variable dQPModelScaleFP is obtained as follows:
[0222] dQPModelScaleFP=((1-2*dQP_model_scale_sign)*dQPModelScaleAbs)<<(dQP_model_offset_frac_prec_minus1+1).
[0223] The variable dQPModelOffsetFP is obtained as follows:
[0224] dQPModelOffsetFP=((1-2*dQP_model_offset_sign)*dQPModelOffsetAbs)<<(dQP_model_scale_frac_prec_minus16+16).
[0225] The variable dQPModelShift is obtained as follows:
[0226] dQPModelShift=(dQP_model_offset_frac_prec_minus1+1)+(dQP_model_scale_frac_prec_minus16+16).
[0227] The variable dQPModelMaxFP is obtained as follows:
[0228] dQPModelMaxFP=((1-2*dQP_model_max_sign)*dQP_model_max_abs)<<dQPModelShift
[0229] The variable dQPModelMinFP is obtained as follows:
[0230] dQPModelMinFP=((1-2*dQP_model_min_sign)*dQP_model_min_abs)<<dQPModelShift.
[0231] for Y = 0: maxY / / For example, for a 10-bit video, maxY = 1023
[0232]
[0233] If (full_range_input_flag == 0) / / If the input is a standard range video
[0234] For Y values outside the standard range (i.e., Y = [0:63] and [940:1023]), set slope[Y] = 0; CDF[0] = slope[0];
[0235] for Y = 0: maxY - 1
[0236] {
[0237] CDF[Y+1] = CDF[Y] + slope[Y]; / / CDF[Y] is the integral of slope[Y].
[0238] }
[0239] for Y = 0: maxY
[0240] {
[0241] FwdLUT[Y] = round(CDF[Y] * maxY / CDF[maxY]); / / Round and normalize to obtain FwdLUT
[0242] }
[0243] In another embodiment, as depicted in Table 10, the forward shaping function can be represented as a set of luminance pivot points (In_Y) and their corresponding codewords (Out_Y). To simplify encoding, the input luminance range is described using a linear piecewise representation based on a sequence of initial pivots and equally spaced subsequent pivots. Figure 7 The text describes an example of a forward shaping function for 10-bit input data.
[0244] Table 10: Example syntax for pivot-based representation (model D) for integer functions
[0245]
[0246] In this embodiment, the parameters in Table 10 can be defined as follows:
[0247] `full_range_input_flag` specifies the range of the input video signal. `full_range_input_flag` of 0 corresponds to a standard range input video signal. `full_range_input_flag` of 1 corresponds to a full range input video signal. When `full_range_input_flag` is not present, it is inferred to be 0.
[0248] bin_pivot_start specifies the pivot value (710) for the first equal-length interval. When full_range_input_flag is equal to 0, bin_pivot_start should be greater than or equal to the minimum standard range input and should be less than the maximum standard range input. (For example, for a 10-bit SDR input, bin_pivot_start(710) should be between 64 and 940).
[0249] bin_cw_start specifies the mapping value (715) of bin_pivot_start (710) (e.g., bin_cw_start = FwdLUT[bin_pivot_start]).
[0250] `log2_num_equal_bins_minus3` plus 3 specifies the number of equal-length intervals after the starting pivot (710). The variables `NumEqualBins` and `NumTotalBins` are defined as follows:
[0251] NumEqualBins=1<<(log2_num_equal_bins_minus3+3)
[0252] If full_range_input_flag == 0
[0253] NumTotalBins = NumEqualBins + 4
[0254] otherwise
[0255] NumTotalBins = NumEqualBins + 2
[0256] Note: Experimental results show that most forward integer functions can be represented using eight equal-length segments; however, complex integer functions may require more segments (e.g., 16 or more segments).
[0257] `equal_bin_pivot_delta` specifies the length of the equal-length interval (e.g., 720-1, 720-N). `NumEqualBins*equal_bin_pivot_delta` should be less than or equal to the valid input range. (For example, if `full_range_input_flag` is 0, then for a 10-bit input, the valid input range should be 940-64=876; if `full_range_input_flag` is 1, then for a 10-bit input, the valid input range should be 0 to 1023.)
[0258] bin_cw_in_first_equal_bin specifies the number of mapped codewords (725) in the first equal-length interval (720-1). Bin_cw_delta_abs_prec_minus4 plus 4 specifies the number of bits used to represent bin_cw_delta_abs[i] for each subsequent equal interval.
[0259] bin_cw_delta_abs[i] specifies the value of bin_cw_delta_abs[i] for each subsequent equal-length interval. bin_cw_delta[i] (e.g., 735) is the difference between the codeword (e.g., 740) in the current equal-length interval i (e.g., 720-N) and the codeword (e.g., 730) in the previous equal-length interval i-1.
[0260] `bin_cw_delta_sign[i]` specifies the sign of `bin_cw_delta_abs[i]`. When `bin_cw_delta_abs[i]` equals 0, `bin_cw_delta_sign[i]` is not signaled and is inferred to be 0. The variable `bin_cw_delta[i]` = (1 - 2 * `bin_cw_delta_sign[i]]` * `bin_cw_delta_abs[i]`
[0261] Decoding process of model D
[0262] Given the syntax elements of Table 10, for a 10-bit input, the integer LUT can be obtained as follows: Define constants:
[0263] minIN = minOUT = 0;
[0264] maxIN = maxOUT = 2^BD – 1 = 1023 for the 10-bit case / / BD = bit depth
[0265] minStdIN = 64 for the 10-bit case
[0266] maxStdIN = 940 for the 10-bit case
[0267] Step 1: For j = 0 to NumTotalBins, obtain the pivot value In_Y[j].
[0268]
[0269] Step 2: For j = 0 to NumTotalBins, obtain the mapping value Out_Y[j].
[0270]
[0271] bin_cw[j] = bin_cw[j-1] + bin_cw_delta[j-4]; / / bin_cw_delta[i] starts from idx 0.
[0272]
[0273] bin_cw[j] = bin_cw[j-1] + bin_cw_delta[j-3]; / / bin_cw_delta[i] starts from idx 0.
[0274] Step 3: Linear interpolation to obtain all LUT entries
[0275] Initialization: FwdLUT[]
[0276]
[0277] Typically, shaping can be turned on or off for each slice. For example, shaping can be enabled only for intra-frame slices and disabled for inter-frame slices. In another example, shaping can be disabled for inter-frame slices with the highest time level. (Note: As an example, as used in this article, the time sub-layer can match the definition of the time sub-layer in HEVC.) When defining the shaper model, in one example, the shaper model can be signaled only in the SPS, but in another example, the slice shaper model can be signaled in intra-frame slices. Alternatively, the shaper model can be signaled in the SPS and the slice shaper model can be allowed to update the SPS shaper model for all slices, or the slice shaper model can be allowed to update the SPS shaper model only for intra-frame slices. For inter-frame slices that follow intra-frame slices, either the SPS shaper model or the intra-frame slice shaper model can be applied.
[0278] As another example, Figure 5A and 5B The process of reconstructing the shape function in the decoder according to an embodiment is described. The process uses the method described herein and in reference [3], wherein the visual rating range is
[05] .
[0279] like Figure 5A As shown, first (step 510), the decoder extracts the `reshape_model_profile_type` variable and sets an appropriate initial frequency band distribution for each interval (steps 515, 520, and 525). For example, in the pseudocode:
[0280] If (reshape_model_profile_type == 0)R[b i ]=R bright [b i ];
[0281] Otherwise, if (reshape_model_profile_type == 1)R[b i ] = R dark [b i ];
[0282] Otherwise R[b i ]=R mid [b i ].
[0283] In step 530, the decoder uses the received reshape_model_band_profile_delta[bi] values to adjust the distribution of each frequency band, as shown below:
[0284] for(i=0:reshape_model_num_band-1)
[0285] {R[b i ]=R[b i ]+reshape_model_band_profile_delta[b i ]}.
[0286] In step 535, the decoder propagates the adjusted value to each binprofile as follows: if bin[j] belongs to frequency band b i , R_bin[j]=R[b i ].
[0287] In step 540, the interval distribution is modified as follows:
[0288] If (j > reshape_model_max_bin_idx) or (j < reshape_model_min_bin_idx)
[0289] Then {R_bin[j]=0}.
[0290] In parallel, in steps 545 and 550, the decoder can extract parameters to compute the scaling factor value and candidate codewords for each bin[j], as follows:
[0291] ScaleFactor=1.0-0.05*reshape_model_scale_idx
[0292] CW_dft[j] = the codeword in the interval if the default integer type is used.
[0293] CW_PQ[j]=TotalCW / TotalNumBins.
[0294] When calculating the ScaleFactor value, for fixed-point implementations, 1 / 16 = 0.0625 can be used instead of the scaling factor 0.05.
[0295] continue Figure 5B In step 560, the decoder begins pre-assigning codewords (CWs) for each interval based on the interval distribution, as shown below:
[0296] If R_bin[j] == 0, then CW[j] = 0
[0297] If R_bin[j] == 1, then CW[j] = CW_dft[j] / 2;
[0298] If R_bin[j] == 2, then CW[j] = min(CW_PQ[j], CW_dft[j]);
[0299] If R_bin[j] == 3, then CW[j] = (CW_PQ[j] + CW_dft[j]) / 2;
[0300] If R_bin[j] >= 4, CW[j] = max(CW_PQ[j], CW_dft[j]);
[0301] In step 565, the total number of codewords used is calculated and the codeword (CW) allocation is refined / completed, as shown below:
[0302] CW used =Sum(CW[j]):
[0303] If CW used >TotalCW, rebalance CW[j] = CW[j] / (CW used / TotalCW);
[0304] otherwise
[0305]
[0306] Finally, in step 565, the decoder: a) generates a forward integer function (e.g., FwdLUT) by accumulating the CW[j] value, b) multiplies the ScaleFactor value with the FwdLUT value to form the final FwdLUT (FFwdLUT), and c) generates the inverse integer function InvLUT based on FFwdLUT.
[0307] In the fixed-point implementation, the calculation of ScaleFactor and FFwdLUT can be expressed as:
[0308] ScaleFactor=(1<<SF_PREC)-reshape_model_scale_idx
[0309] FFwdLUT=(FwdLUT*ScaleFactor+(1<<(FP_PREC+SF_PREC-1)))>>(FP_PREC+SF_PREC),
[0310] Here, SF_PREC and FP_PREC are predefined precision-related variables (e.g., SF_PREC = 4 and FP_PREC = 14), and "c = a << n" means performing a binary shift operation on a, shifting it left by n bits (or c = a * (2 n And “c = a >> n” means performing a binary shift operation on a, shifting it right by n bits (or c = a / (2 n )).
[0311] Chromaticity QP Derivation
[0312] Chroma coding performance is closely related to luma coding performance. For example, in AVC and HEVC, tables are defined to specify the relationship between quantization parameters (QPs) of luma and chroma components, or between luminance and chroma. The specifications also allow for more flexible definition of the QP relationship between luma and chroma using one or more chroma QP offsets. When shaping is used, luma values are modified, and therefore, the relationship between luma and chroma can also be modified. To maintain and further improve coding efficiency during shaping, in embodiments, chroma QP offsets are obtained based on the shaping curve at the coding unit (CU) level. This needs to be done at both the decoder and encoder.
[0313] As used herein, the term “coding unit” (CU) refers to a coding box (e.g., a macro box, etc.). For example, without limitation, in HEVC, a CU is defined as “a luminance sample coding box, two corresponding chrominance sample coding boxes of an image having three sample arrays, or a sample coding box of a monochrome image or an image encoded using three separate color planes and a syntax structure for coding samples.”
[0314] In this embodiment, the chromaQP value can be obtained as follows:
[0315] 1) Based on the shaping curve, the equivalent luminance dQP mapping, dQPLUT, is obtained:
[0316] for CW = 0: MAX_CW_VALUE - 1
[0317] dQPLUT[CW]=-6*log2(slope[CW]);
[0318] Where slope[CW] represents the slope of the forward shaping curve at each CW (codeword) point, and MAX_CW_VALUE is the maximum codeword value for a given bit depth. For example, for a 10-bit signal, MAX_CW_VALUE = 1024(2 10 ).
[0319] Then, for each coding unit (CU):
[0320] 2) Calculate the average brightness of the coding unit, expressed as AvgY:
[0321] 3) The chromaDQP value is calculated based on dQPLUT[], AvgY, integer schema, inverse integer function Inv(), and slice type, as shown in Table 11 below:
[0322] Table 11: Example ChromaDQP values based on the shaping architecture
[0323]
[0324] 4) Calculate chromaQP as:
[0325] chromaQP == QP_luma + chromaQPOffset + chromaDQP, where chromaQPOffset represents the chroma QP offset, and QP_luma represents the luminance QP of the coding unit. Note that the chroma QP offset value can be different for each chroma component (e.g., Cb and Cr), and the chroma QP offset value is transmitted to the decoder as part of the encoded bitstream.
[0326] In this embodiment, dQPLUT[] can be implemented as a predefined LUT. Assume all codewords are divided into N intervals (e.g., N = 32), and each interval contains M = MAX_CW_VALUE / N codewords (e.g., M = 1024 / 32 = 32). When assigning new codewords to each interval, they can limit the number of codewords to 1 to 2*M, so they can pre-compute dQPLUT[1…2*M] and store the result as a LUT. This approach avoids any floating-point calculations or approximations of fixed-point calculations. The method also saves encoding / decoding time. For each interval, a fixed chromaQPOffset is used for all codewords in that interval. The DQP value is set to equal dQPLUT[L], where L is the number of codewords in that interval, and 1 ≤ L ≤ 2*M.
[0327] The dQPLUT value can be pre-calculated as follows:
[0328] for i = 1: 2 * M
[0329] slope[i] = i / M;
[0330] dQPLUT[i] = -6 * log2(slope[i]);
[0331] End
[0332] When calculating dQPLUT[x], different quantization schemes can be used to obtain an integer QP value, such as: round(), ceil(), floor(), or a mixture of them. For example, a threshold TH can be set, and if Y < TH, floor() is used to quantize the dQP value, otherwise, when Y ≥ TH, ceil() is used to quantize the dQP value. The use of such quantization schemes and the corresponding parameters can be predefined in the codec or can be signaled in the bitstream for adaptation. An example syntax that allows mixing a quantization scheme with a threshold as discussed above is as follows:
[0333]
[0334] The quant_scheme_signal_table() function can be defined at different levels of the integer syntax (e.g., sequence level, slice level, etc.) according to the adaptation needs.
[0335] In another embodiment, the chromaDQP value can be calculated by applying a scaling factor to the residual signal in each coding unit (or more specifically, the transform unit). This scaling factor can be a luminance-dependent value and can be calculated as follows: a) numerically, for example, as the first derivative (slope) of a forward integer LUT (see equation (6) in the next section for example), or b) calculated as:
[0336]
[0337] When using dQP(x) to calculate Slope(x), dQP can maintain floating-point precision without integer quantization. Alternatively, various different quantization schemes can be used to calculate the quantized integer dQP value. In some embodiments, this scaling can be performed at the pixel level rather than at the block level, where each chroma residual can be scaled by a different scaling factor obtained using the co-located luminance prediction value of that chroma sample. Therefore,
[0338] Table 12: Example chroma dQP values using scaling for the in-loop integer architecture
[0339]
[0340] For example, if CSCALE_FP_PREC = 16
[0341] • Forward scaling: After generating the chroma residuals, but before performing transformation and quantization:
[0342] -C_Res=C_orig-C_pred
[0343] -C_Res_scaled=C_Res*S+(1<<(CSCALE_FP_PREC-1)))>>CSCALE_FP_PREC
[0344] • Inverse scaling: After inverse chroma quantization and inverse transform, but before reconstruction:
[0345] -C_Res_inv=(C_Res_scaled<<CSCALE_FP_PREC) / S
[0346] -C_Reco = C_Pred + C_Res_inv;
[0347] Where S is either S_cu or S_px.
[0348] Note: In Table 12, the average brightness (AvgY) of the box is calculated before applying inverse shaping when calculating Scu. Alternatively, inverse shaping can be applied before calculating the average brightness, for example, Scu = SlopeLUT[Avg(Inv[Y])]. This alternative calculation order also applies to the values in Table 11; that is, calculating Inv(AvgY) can be replaced by calculating the Avg(Inv[Y]) value. The latter approach may be considered more accurate, but it increases computational complexity.
[0349] Encoder optimization for shaping
[0350] This section discusses several techniques for improving encoder coding efficiency by jointly optimizing shaping and encoder parameters when shaping is part of the normalization decoding process (as described in one of the three candidate architectures). Generally, encoder optimization and shaping have their own limitations when addressing coding problems in different places. In traditional imaging and coding systems, there are two types of quantization: a) sample quantization in the baseband signal (e.g., gamma or PQ coding), and b) transform-related quantization (part of compression). Shaping lies in between. Image-based shaping is typically updated based on the image and only allows mapping of sample values based on their brightness level, without considering any spatial information. In block-based codecs (such as HEVC), transform quantization (e.g., for brightness) is applied within a spatial frame and can be adjusted spatially, so encoder optimization methods must apply the same set of parameters to the entire frame containing samples with different brightness values. As understood by the inventors and described herein, joint shaping and encoder optimization can further improve coding efficiency.
[0351] Inter-frame / intra-frame mode decision
[0352] In traditional coding, inter-frame / intra-frame mode decisions are based on calculating a distortion function (dfiunc()) between the original and predicted samples. Examples of such functions include the sum of squared errors (SSE), the sum of absolute differences (SAD), etc. In implementations, shaped pixel values can be used to apply such distortion metrics. For example, if the original dfunct() uses Orig_sample(i) and Pred_sample(i), then when applying the shaping, dfunct() can use its corresponding shaped values Fwd(Orig_sample(i)) and Fwd(Pred_sample(i)). This approach allows for more accurate inter-frame / intra-frame mode decisions, thereby improving coding efficiency.
[0353] LumaDQP using plastic surgery
[0354] In the JCTVC HDR Common Test Conditions (CTC) document (Reference [6]), lumaDQP and chromaQPoffsets are two encoder settings used to modify the quantization (QP) parameters of the luma and chroma components to improve HDR coding efficiency. In this invention, several new encoder algorithms are proposed to further improve the original proposal. For each lumaDQP adapter unit (e.g., 64×64CTU), the dQP value is calculated based on the average input luma value of the unit (as shown in Table 3 of Reference [6]). The final quantization parameter QP for each coding unit within the lumaDQP adapter unit should be adjusted by subtracting this dQP. The dQP mapping table is configurable in the encoder input configuration. This input configuration is represented as dQP inp .
[0355] As discussed in references [6] and [7], in existing coding schemes, the same lumaDQP LUT dQP inp This applies to both intra-frame and inter-frame images. Intra-frame and inter-frame images may have different attributes and quality characteristics. In this invention, a method for adjusting lumaDQP settings based on image encoding type is proposed. Therefore, two dQP mapping tables are configurable in the encoder input configuration and are represented as dQP... inpIntra and dQP inpInter .
[0356] As discussed earlier, when using the in-loop intra-frame shaping method, since no shaping is performed on the inter-frame images, it is important to apply some lumaDQP settings to the inter-frame coded images to achieve similar quality, as if the inter-frame images were shaped using the same shaper used for the intra-frame images. In one embodiment, the lumaDQP settings for the inter-frame images should match the characteristics of the shaping curve used for the intra-frame images.
[0357] make
[0358] Slope(x)=Fwd'(x)=(Fwd(x+dx)-Fwd(x-dx)) / (2dx), (6)
[0359] Let represent the first derivative of the forward shaping function, then in this embodiment, represents the automatically obtained dQP. auto The value of (x) can be calculated as follows:
[0360] If Slope(x) = 0, then dQP auto (x) = 0, otherwise
[0361] dQP auto (x) = 6log2(Slope(x)), (7)
[0362] Among them, dQP auto (x) can be limited to a reasonable range, for example, [-66].
[0363] If lumaDQP is enabled for intra-frame images that utilize shaping (i.e., external DQP is set), inpIntra Therefore, the lumaDQP used for inter-frame images should take this into account. In an embodiment, this can be achieved by using the dQP obtained from the shaper. auto (Equation (7)) and dQP for intra-frame images inpIntra The final inter-frame dQP is calculated by adding the values. final In another embodiment, to utilize intra-frame quality propagation, dQP used for inter-frame images can be... final Set to dQP auto Or set it in small increments (by setting dQP) inpInter Add it to dQP auto .
[0364] In this embodiment, when shaping is enabled, the following general rules for setting the luminance dQP value may be applied:
[0365] (1) A luminance dQP mapping table can be set independently for intra-frame and inter-frame images (based on image encoding type);
[0366] (2) If the image within the coding loop is in the shaping domain (e.g., intra-frame images in an intra-ring shaping architecture or all images in an out-of-ring shaping architecture), then the input luminance to incremental QP mapping dQP also needs to be performed. inp Transform to Integer Domain dQP rsp .Right now
[0367] dQP rsp (x)=dQP inp [Inv(x)]. (8)
[0368] (3) If the image in the coding loop is in the unshaped domain (e.g., reverse shaped or unshaped, such as inter-frame images in the intra-loop shaped architecture or all images in the intra-loop residual shaped architecture), then the input luminance to incremental QP mapping does not need to be converted and can be used directly.
[0369] (4) Automatic inter-frame increment QP derivation is only valid for in-loop intra-frame shaping architectures. In this case, the actual increment QP used for inter-frame images is the sum of the automatically obtained and input values:
[0370] dQP final [x] = dQP inp [x]+dQP auto [x], (9)
[0371] And dQP final [x] can be limited to a reasonable range, such as [-1212];
[0372] (5) The luminance can be updated to the dQP mapping table in each image or when the shaped LUT changes. The actual dQP adaptation (obtaining the corresponding dQP for quantization of the box based on the average luminance value of the box) can occur at the CU level (configurable by the encoder).
[0373] Table 13 summarizes the dQP settings for each of the three proposed architectures.
[0374] Table 13: dQP Settings
[0375]
[0376] Rate-Distortion Optimization (RDO)
[0377] In the JEM6.0 software (Reference [8]), when lumaDQP is enabled, weighted distortion based on RDO (Rate Distortion Optimization) pixels is used. The weight table is fixed based on the luminance value. In the embodiment, the weight table should be adaptively adjusted based on the lumaDQP settings calculated as presented in the previous section. The two weights for the sum of squared errors (SSE) and the sum of absolute differences (SAD) are proposed as follows:
[0378]
[0379]
[0380] The weights calculated using equation (10a) or equation (10b) are based on the total weights of the final dQP, which includes both the input lumaDQP and the dQP obtained from the forward shaping function. For example, based on equation (9), equation (10a) can be written as:
[0381]
[0382] The total weights can be separated into weights calculated from the input lumaDQP:
[0383]
[0384] And the weighting from plastic surgery:
[0385]
[0386] When the total weights are calculated using the total dQP by first calculating the weights from the integer part, the integer dQP is obtained due to the clipping operation. autoHowever, this results in a loss of accuracy. In contrast, directly using the slope function to calculate the weights from the integer part maintains higher weight accuracy and is therefore more advantageous.
[0387] The weights obtained from the input lumaDQP are represented as W. dQP Let f'(x) denote the first derivative (or slope) of the forward shaping curve. In this embodiment, the total weight takes into account both the dQP value and the shape of the shaping curve, therefore the total weight value can be expressed as:
[0388] weight total =Clip3(0.0, 30.0, W) dQP *f′(x) 2 (11)
[0389] A similar approach can also be applied to the chromaticity component. For example, in the embodiment, for chromaticity, dQP[x] can be defined according to Table 13.
[0390] Interaction with other coding tools
[0391] When enabling shaping, this section provides several examples of suggested changes required for other coding tools. Interactions may exist for any existing or future coding tools that could be included in next-generation video coding standards. The examples given below are not limiting. Typically, it is necessary to identify the video signal domains (shaped, unshaped, and inversely shaped) during the coding steps, and the operations that process the video signal at each step need to take the shaping effect into account.
[0392] Cross-component linear model prediction
[0393] In CCLM (Cross-Component Linear Model Prediction) (Reference [8]), the brightness reconstruction signal rec can be used. L The predicted chromaticity sample pred is obtained by using '(i,j). c (i, j):
[0394] pred c (i, j) = α·rec L ′(i,j)+β。 (12)
[0395] When shaping is enabled, in an embodiment, it may be necessary to determine whether the reconstructed luminance signal is in a shaped domain (e.g., an out-of-loop shaper or an in-loop intra-frame shaper) or an unshaped domain (e.g., an in-loop residual shaper). In one embodiment, the reconstructed luminance signal can be used implicitly as is without any additional signaling or operation. In other embodiments, if the reconstructed signal is in the unshaped domain, the reconstructed luminance signal can be converted to also be in the unshaped domain, as follows:
[0396] pred c (i, j) = α·Inv(rec L ′(i,j))+β。 (13)
[0397] In other embodiments, bitstream syntax elements can be added to signal which field is desired (shaped or unshaped). This can be determined by the RDO process or based on the decoded information, thus saving the overhead required for explicit signaling. Appropriate operations can then be performed on the reconstructed signal based on this decision.
[0398] Shaper with residual prediction tools
[0399] The HEVC range extension profile includes a residual prediction tool. Based on the luminance residual signal from the encoder side, the chrominance residual signal is predicted as follows:
[0400] Δr C (x, y) = r C (x, y) - (α×r′) L (x, y))>>3, (14)
[0401] Furthermore, the chroma residual signal is compensated at the decoder side as follows:
[0402] r′ C (x, y) = Δr′ C (x, y) + (α × r′) L (x, y)) >> 3, (15)
[0403] Where, r c Let r′ represent the chromaticity residual sample at position (x, y). L Represents the reconstructed residual sample of the luminance component, Δr c This represents the prediction signal using inter-color prediction, Δr′ C This indicates that for Δr c The reconstructed signal after encoding and decoding, and r C This represents the reconstructed chromaticity residual.
[0404] When shaping is enabled, it may be necessary to consider which luminance residual to use for chrominance residual prediction. In one embodiment, the "residual" can be used as is (it can be shaped or unshaped based on the shaper architecture). In another embodiment, the luminance residual can be forced into a domain (such as in an unshaped domain) and appropriate mapping can be performed. In yet another embodiment, appropriate processing can be obtained by the decoder, or appropriate processing can be explicitly signaled as described above.
[0405] Shaper with adaptive amplitude limiting
[0406] Adaptive limiting (reference [8]) is a novel tool introduced to signal the range of raw data about content dynamics and to perform adaptive limiting instead of fixed limiting (based on internal bit depth information) at each step of the compression workflow where limiting occurs (e.g., in transform / quantization, loop filtering, output).
[0407] T clip =Clip BD (T,bitdepth,C)=Clip3(min c max c ,T). (16)
[0408] Where x = Clip3(min, max, c) means:
[0409]
[0410] and
[0411] • C is component ID (usually Y, Cb, or Cr).
[0412] ·min c It is the lower limit of the current slice used in component ID C.
[0413] ·max c It is the upper limit of the current slice used by component ID C.
[0414] When enabling shaping, in this embodiment, it may be necessary to determine the domain in which the data stream currently resides and perform clipping correctly. For example, if clipping is being processed within integer domain data, the original clipping boundaries need to be transformed to the integer domain:
[0415] T clip =Clip BD (T, bitdepth, C)==Clip3(Fwd(min C ), Fwd(max) C ), T). (17)
[0416] Typically, each sizing step needs to be handled correctly regarding the shaping architecture.
[0417] Shaper and loop filter
[0418] In HEVC and JEM 6.0 software, loop filters such as ALF and SAO require both reconstructed luminance samples and uncompressed "raw" luminance samples to estimate optimal filter parameters. When shaping is enabled, in embodiments, the domain to which filter optimization is desired can be specified (explicitly or implicitly). In one embodiment, filter parameters can be estimated over the shaped domain (relative to the shaped original when the reconstruction is in the shaped domain). In other embodiments, filter parameters can be estimated over the unshaped domain (relative to the original when the reconstruction is in the unshaped or inversely shaped domain).
[0419] For example, depending on the in-loop shaping architecture, the in-loop filter optimization (ILFOPT) options and operations can be described in Tables 14 and 15.
[0420] Table 14. Loop Filtering Optimization in In-Ring In-Frame-Only Shaping Architecture and In-Ring Hybrid Shaping
[0421]
[0422]
[0423]
[0424] Table 15. Loop Filtering Optimization in In-Loop Residual Shaping Architecture
[0425]
[0426] While most of the detailed discussion in this paper concerns the methods performed on the lightness component, those skilled in the art will understand that similar methods can be performed on the chroma color component and chroma-related parameters such as chromaQPOffset (see, for example, reference [9]).
[0427] In-ring shaping and Region of Interest (ROI)
[0428] Given an image, as used herein, the term 'region of interest' (ROI) refers to an image region considered to be of particular interest. In this section, novel embodiments are proposed that support only in-ring shaping of the ROI. That is, in these embodiments, shaping can be applied only inside the ROI, rather than outside it. In another embodiment, different shaping curves can be applied both inside and outside the ROI.
[0429] The use of Regions of Interest (ROIs) is driven by the need to balance bitrate with image quality. Consider, for example, a video sequence of a sunset. In the upper half of the image, the sun can be positioned against a relatively uniformly colored sky (so pixels in the sky background can have very low variance). Conversely, the lower half of the image can depict a moving wave. From the viewer's perspective, the upper half might be perceived as far more important than the lower half. On the other hand, moving waves are difficult to compress due to their higher pixel variance, requiring more bits per pixel; however, it might be desirable to allocate more bits to the sun portion than to the wave portion. In this case, the upper half can be represented as the Region of Interest.
[0430] ROI Description
[0431] Today, most codecs (e.g., AVC, HEVC, etc.) are block-based. To simplify implementation, regions can be specified in units of boxes. Without limitation, using HEVC as an example, a region can be defined as multiple coding units (CUs) or coding tree units (CTUs). One or more ROIs can be specified. Multiple ROIs can be different or overlapping. ROIs do not necessarily have to be rectangular. Syntax for ROIs can be provided at any level of interest (such as slice level, picture level, video stream level, etc.). In an embodiment, the ROI is first specified in the Sequence Parameter Set (SPS). Then, small variations in the ROI can be allowed in the slice header. Table 16 depicts an example of the syntax where one ROI is specified as multiple CTUs within a rectangular region. Table 17 describes the syntax for a modified ROI at the slice level.
[0432] Table 16: SPS Syntax for ROI
[0433]
[0434] Table 17: Slice Header Syntax for ROI
[0435]
[0436] When sps_reshaper_active_ROI_flag equals 1, it indicates that an ROI exists in the coded video sequence (CVS). When sps_reshaper_active_ROI_flag equals 0, it indicates that an ROI does not exist in the CVS.
[0437] The parameters `reshaper_active_ROI_in_CTUsize_left`, `reshaper_active_ROI_in_CTUsize_right`, `reshaper_active_ROI_in_CTUsize_top`, and `reshaper_active_ROI_in_CTUsize_bottom` each specify the image sample within the ROI based on the rectangular region defined in the image coordinates. For the left and top regions, the coordinates are equal to `offset * CTUsize`, while for the right and bottom regions, the coordinates are equal to `offset * CTUsize - 1`.
[0438] A reshape_model_ROI_modification_flag value of 1 indicates that the ROI has been modified in the current slice. A reshape_model_ROI_modification_flag value of 0 indicates that the ROI has not been modified in the current slice.
[0439] The parameters reshaper_ROI_mod_offset_left, reshaper_ROI_mod_offset_right, reshaper_ROI_mod_offset_top, and reshaper_ROI_mod_offset_bottom each specify the left / right / top / bottom offset values relative to reshaper_active_ROI_in_CTUsize_left, reshaper_active_ROI_in_CTUsize_right, reshaper_active_ROI_in_CTUsize_top, and reshaper_active_ROI_in_CTUsize_bottom, respectively.
[0440] For multiple ROIs, the example syntax for a single ROI in Tables 16 and 17 can be extended using an index (or ID) for each ROI, similar to the scheme in HEVC used to define multiple pan-scan rectangles using SEI messages (see HEVC specification, reference
[11] , section D.2.4).
[0441] ROI processing in intra-frame shaping only within the ring
[0442] For intra-frame shaping only, the ROI portion of the image is shaped first, and then encoding is applied to it. Since shaping is only applied to the ROI, boundaries between the ROI and non-ROI portions of the image may be visible. This is due to loop filters (e.g., Figure 2C or Figure 2DThe 270 in the image may cross boundaries, therefore special attention must be paid to the ROI for loop filter optimization (ILFOPT). In an embodiment, it is proposed that the loop filter be applied only if the entire decoded image is in the same domain. That is, the entire image is either entirely in the shaped domain or entirely in the unshaped domain. In one embodiment, on the decoder side, if loop filtering is applied in the unshaped domain, inverse shaping should first be applied to the ROI portion of the decoded image, and then the loop filter should be applied. Next, the decoded image is stored in the DPB. In another embodiment, if loop filtering is applied in the shaped domain, shaping should first be applied to the non-ROI portion of the decoded image, and then the loop filter should be applied, and then the entire image should be inversely shaped. Next, the decoded image is stored in the DPB. In yet another embodiment, if loop filtering is applied in the shaped domain, inverse shaping can first be applied to the ROI portion of the decoded image, then the entire image can be shaped, then the loop filter can be applied, and then the entire image can be inversely shaped. Next, the decoded image is stored in the DPB. Table 18 summarizes these three methods. From a computational perspective, method "A" is simpler. In the embodiment, enabling ROI can be used to specify the order in which inverse shaping and loop filtering (LF) are performed. For example, if ROI is actively used (e.g., SPS syntax flag = true), then inverse shaping ( Figure 2C and Figure 2D After box 265 in the middle, execute LF( Figure 2C and Figure 2D (See box 270 in the image). If no ROI is used explicitly, LF is performed before reverse reshaping.
[0443] Table 18. Loop Filtering (LF) Options for Using ROI
[0444]
[0445]
[0446] ROI processing in in-loop prediction residual shaping
[0447] For in-ring (predicted) residual shaping architecture (e.g., see...) Figure 2F In the 200C_D), at the decoder, using equation (3), the processing can be expressed as:
[0448] If (the current CTU belongs to the ROI)
[0449] Reco_sample = Inv(Res_d + Fwd(Pred_sample)), (see equation (3))
[0450] otherwise
[0451] Reco_sample=Res_d+Pred_sample
[0452] Finish
[0453] ROI and encoder considerations
[0454] In the encoder, it is necessary to check whether each CTU belongs to the ROI. For example, for in-loop prediction residual shaping, a simple check based on equation (3) can be performed as follows:
[0455] If (the current CTU belongs to the ROI)
[0456] Weighted distortion from RDO is applied to the luminance. The weights are obtained based on equation (10).
[0457] otherwise
[0458] Applying unweighted distortion in RDO to luminance
[0459] Finish
[0460] A sample coding workflow that considers ROI during shaping may include the following steps:
[0461] -For intra-frame images:
[0462] - Apply forward shaping to the ROI region of the original image.
[0463] - Encode intra-frames
[0464] - Apply inverse shaping to the ROI region of the reconstructed image before the loop filter (LF).
[0465] - Perform loop filtering in the unshaped domain as follows (e.g., see method "C" in Table 18), which includes the following steps:
[0466] • Apply forward shaping to the non-ROI regions of the original image (so that the entire original image is shaped for use as a loop filter reference).
[0467] • Apply forward shaping to the entire image region of the reconstructed image.
[0468] • Obtain the loop filter parameters and apply loop filtering.
[0469] • Apply inverse shaping to the entire image region of the reconstructed image and store it in the DPB. On the encoder side, since LF requires an uncompressed reference image for filter parameter estimation, the processing of the LF reference for each method is shown in Table 19:
[0470] Table 19. Treatment of LF reference for ROI
[0471]
[0472] -For inter-frame images:
[0473] - When encoding inter-frames, for each CU within the ROI, prediction residual shaping and weighted distortion are applied to the lumen; for each CU outside the ROI, no shaping is applied.
[0474] - Perform loop filtering optimization as before (as if no ROI was used) (Option 1):
[0475] • Perform forward shaping on the entire image area of the original image.
[0476] • Perform forward shaping on the entire image region of the reconstructed image.
[0477] • Obtain the loop filter parameters and apply loop filtering.
[0478] • Apply inverse shaping to the entire image region of the reconstructed image and store it in the DPB.
[0479] Reshaping HLG encoded content
[0480] The term HybridLog-Gamma, or HLG, refers to another transfer function defined in Rec.BT.2100 for mapping high dynamic range signals. HLG was developed to maintain backward compatibility with conventional standard dynamic range signals encoded using conventional gamma functions. When comparing codeword distribution between PQ-coded content and HLG-coded content, PQ mapping tends to allocate more codewords in dark and bright areas, while most HLG content codewords appear to be allocated in the middle range. HLG luma shaping can be performed using two methods. In one embodiment, the HLG content can be simply converted to PQ content, and then all the PQ-related shaping techniques discussed above can be applied. For example, the following steps can be applied:
[0481] 1) Map HLG lightness (e.g., Y) to PQ lightness. Let the transformation function or LUT be expressed as HLG2PQLUT(Y).
[0482] 2) Analyze the PQ brightness values and obtain a forward shaping function or LUT based on PQ. Represent it as PQAdpFLUT(Y).
[0483] 3) Combine these two functions or LUTs into a single function or LUT: HLGAdpFLUT[i] = PQAdpFLUT
[0484] [HLG2PQLUT[i]].
[0485] Because the HLG codeword distribution is completely different from the PQ codeword distribution, this method may produce suboptimal shaping results. In another embodiment, the HLG shaping function is obtained directly from the HLG samples. The same framework used for PQ signals can be applied, but the CW_Bins_Dft table can be modified to reflect the characteristics of the HLG signal. In this embodiment, using the midtone distribution for HLG signals, several CW_Bins_Dft tables can be designed according to user preferences. For example, when highlight preservation is preferred, for alpha = 1.4,
[0486] g_DftHLGCWBin0={8, 14, 17, 19, 21, 23, 24, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 36, 37, 38, 39, 39, 40, 41, 41, 42, 43, 43, 44, 44, 30}.
[0487] When it is preferable to retain the midtones (or midrange):
[0488] g_DftHLGCWBin1={12, 16, 16, 20, 24, 28, 32, 32, 32, 32, 36, 36, 40, 44, 48, 52, 56, 52, 48, 44, 40, 36, 36, 32, 32, 32, 26, 26, 20, 16, 16, 12}.
[0489] When skin tone is preferred to be preserved:
[0490] g_DftHLGCWBin2={12,16,16,24,28,32,56,64,64,64,64,56,48,40,32,32,32,32,32,32,28,28,24,24,20,20,20,20,20,16,16,12};
[0491] From the perspective of bitstream syntax, in order to distinguish between PQ-based and HLG-based shaping, a new parameter denoted as sps_reshaper_signal_type is added, where the value sps_reshaper_signal_type indicates the type of signal being shaped (e.g., 0 for gamma-based SDR signal, 1 for PQ-coded signal, and 2 for HLG-coded signal).
[0492] Tables 20 and 21 show examples of syntax tables for HDR shaping in the SPS and slice headers for both PQ and HLG, which have all the features discussed above (e.g., ROI, Loop Filter Optimization (ILFOPT), and ChromaDQPAdjustment).
[0493] Table 20: Example SPS syntax for shaping
[0494]
[0495] sps_in_loop_filter_opt_flag equal to 1 specifies the in-loop filter optimization to be performed in the integer domain of the coded video sequence (CVS).
[0496] `sps_in_loop_filter_opt_flag` equal to 0 specifies that in-loop filter optimization should be performed in the unshaped domain of the CVS. `sps_luma_based_chroma_qp_offset_flag` equal to 1 specifies that a luma-based chroma QP offset is obtained (e.g., according to Table 11 or Table 12) and applied to the chroma coding of each CU in the coded video sequence (CVS). `sps_luma_based_chroma_qp_offset_flag` equal to 0 specifies that luma-based chroma QP offset is not enabled in the CVS.
[0497] Table 21: Example syntax for shaping at the slice level
[0498]
[0499] Improve color quality
[0500] Proponents of HLG-based encoding argue that it provides better backward compatibility with SDR signaling. Therefore, theoretically, HLG-based signals can use the same encoding settings as traditional SDR signals. However, when viewing HLG-encoded signals in HDR mode, some color artifacts can still be observed, especially in achromatic areas such as white and gray. In practice, these artifacts can be reduced by adjusting the chromaQPOffset value during encoding. It is recommended that a less intense chromaQP adjustment be applied to HLG content than that used when encoding PQ signals. For example, in reference
[10] , a model that assigns QP offsets to Cb and Cr based on luminance QP and assigns factors based on the captured color primary and the represented color primary is described as follows:
[0501] QPoffsetCb=Clip3(-12,0,Round(c cb *(k*QP+l))), (18a)
[0502] QPoffsetCr=Clip3(-12,0,Round(c cr*(k*QP+l))), (18b)
[0503] Where, if the captured color primary is the same as the representing color primary, then c cb =1, if the captured color primary is equal to the P3D65 primary and the represented color primary is equal to the Rec.ITU-R BT.2020 primary, then c cb =1.04, and if the captured color primary is equal to the Rec.ITU-R BT.709 primary and the indicated primary is equal to the Rec.ITU-R BT.2020 primary, then c cb =1.14. Similarly, if the captured color primary is the same as the representing color primary, then c cr =1, if the captured color primary is equal to the P3D65 primary and the represented color primary is equal to the Rec.ITU-R BT.2020 primary, then c cr =1.39, and if the captured color primary is equal to the Rec.ITU-R BT.709 primary and the indicated primary is equal to the Rec.ITU-R BT.2020 primary, then c cr =1.78. Finally, k = -0.46 and l = 0.26.
[0504] In the embodiments, it is suggested to use the same model but with different parameters to produce a less pronounced chromaQPOffset variation. For example, without limitation, in the embodiments, for Cb in equation (18a), c cb =1, k=-0.2 and l=7, and for Cr in equation (18b), c cr =1, k=-0.2, and l=7. Figure 6A and Figure 6B Examples illustrating how the chromaQPOffset value changes based on the luminance quantization parameter (QP) for PQ (Rec.709) and HLG are shown. The changes in PQ-related values are more significant than those for HLG-related values. Figure 6A Corresponding to Cb (equation (18a)), while Figure 6B Corresponding to Cr (Equation (18b)).
[0505] References
[0506] Each of the references listed in this article is included in its full text by way of citation.
[0507] [1] PCT application PCT / US2016 / 025082, filed on March 30, 2016, for In-Loop Block-Based Image Reshaping in High Dynamic Range Video Coding, was also published by GM.Su as WO 2016 / 164235.
[0508] [2] D. Baylon, Z. Gu, A. Luthra, K. Minoo, P. Yin, F. Pu, T. Lu, T. Chen, W. Husak, Y. He, L. Kerofsky, Y. Ye, B. Yi, “Response to Call for Evidence for HDR and WCG Video Coding: Arris, Dolby and InterDigital”, document m36264, July 2015, Warsaw, Poland.
[0509] [3] U.S. Patent Application No. 15 / 410,563, filed on January 19, 2017, by T. Lu et al., entitled Content-Adaptive Reshaping for HighCodeword Representation Images.
[0510] [4] PCT application PCT / US2016 / 042229, filed on July 14, 2016, for Signal Reshaping and Coding for HDR and WideColor Gamut Signals, was also published by P. Yin et al. as WO 2017 / 011636.
[0511] [5] K. Minoo et al., “Exploratory Test Model for HDR extension of HEVC”, MPEG Output Document, JCTVC-W0092(m37732), 2016, San Diego, USA.
[0512] [6] E. Francois, J. Sole, J. P. Yin, “Common Test Conditions for HDR / WCG video coding experiments”, JCTVC document Z1020, Geneva, January 2017.
[0513] [7] A. Segall, E. Francois and D. Rusanovskyy, “JVET common test conditions and evaluation procedures for HDR / WCG Video”, JVET-E1020, ITU-T meeting, Geneva, January 2017.
[0514] [8] JEM 6.0 software: https: / / jvet.hhi.fraunhofer.de / svn / svnHMJEMSoftware / tags / HM-16.6- JEM-6.0
[0515] [9] The U.S. Provisional Patent Application Serial No. 62 / 406,483, filed on October 11, 2016, entitled “Adaptive Chroma Quantization in Video Coding for Multiple Color Imaging Formats”, was also filed as U.S. Patent Application Serial No. 15 / 728,939 and published as U.S. Patent Application Publication US 2018 / 0103253.
[0516]
[10] J. Samuelsson et al. (editors), “Conversion and coding practices for HDR / WCG Y'CbCr4:2:0 Video with PQ Transfer Characteristics”, JCTVC-Y1017, ITU-T / ISO meeting, Chengdu, October 2016.
[0517]
[11] ITU-T H.265, “High efficiency video coding”, ITU, version 4.0, (12 / 2016).
[0518] Example computer system implementation
[0519] Embodiments of the present invention can be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs), or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means comprising one or more such systems, devices, or components. The computer and / or IC can execute, control, or implement instructions related to integrated signal shaping and image encoding, such as those described herein. The computer and / or IC can calculate any of the various parameters or values associated with the signal shaping and encoding processes described herein. Image and video embodiments can be implemented in hardware, software, firmware, and various combinations thereof.
[0520] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement the methods related to integrated signal shaping and image encoding as described above by executing software instructions in a program memory accessible to the processor. The present invention can also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. The program product according to the present invention can take any of a variety of forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0521] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise stated, references to this component (including references to "device") should be interpreted as including equivalents (e.g., functionally equivalents) of any component that performs the function of the described component, including components that are not structurally equivalent to the disclosed structures that perform the functions in the illustrated exemplary embodiments of the invention.
[0522] Equivalents, extensions, substitutes and others
[0523] Example embodiments related to efficient integrated signal shaping and image coding are described thus. In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the claims that are the sole and unique indications of the invention, and intended by the applicant to be the specific form published in this application, including any subsequent modifications, are the claims. Any definitions of terms expressly set forth herein for inclusion in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, characteristics, features, advantages, or attributes not expressly referenced in the claims shall not in any way limit the scope of such claims. Therefore, the description and drawings should be considered illustrative rather than restrictive.
Claims
1. An apparatus for encoding an image, the apparatus comprising: The input terminal is used to access the input image represented by the first codeword; and processor, wherein the processor: Generate a forward shaping function that maps the pixels of the input image to the second codeword representation; An inverse shaping function is generated based on the forward shaping function, wherein the inverse shaping function maps pixels from the second codeword representation to the first codeword representation; Based on the input pixel region in the input image, the forward shaping function, and the inverse shaping function, the coded pixel region of the input image is generated; Generate integer metadata representing the forward integer function based on a piecewise linear representation; and An output bitstream is generated based on the coded pixel region of the input image and the shaped metadata. Specifically, for the input pixel region in the input image, in order to generate the coded pixel region, the processor: The prediction region is calculated based on pixel data in the reference frame buffer or in the previously encoded spatial neighborhood. A shaped residual region is generated based on the input pixel region, the prediction region, and the forward shaping function, wherein the shaped residual samples in the shaped residual region are obtained at least in part by forward shaping the corresponding prediction samples in the prediction region. A quantized residual region is generated based on the shaped residual region; The dequantized residual region is generated based on the quantized residual region; The reconstructed pixel region is generated based on the dequantized residual region, the predicted region, the forward shaping function, and the inverse shaping function; and A reference pixel region to be stored in the reference frame buffer is generated based on the reconstructed pixel region; and Generating the reconstructed pixel region includes calculating: Recon_sample ( i ) = Inv ( Res_d ( i ) + Fwd ( Pred_sample ( i ))), in, Fwd () represents the forward shaping function. Inv () represents the inverse shaping function. Recon_sample ( i ) represents the pixels of the reconstructed pixel region. Pred_sample ( i ) represents the pixels of the predicted region. Res_d ( i ) represents the representative region of the dequantized residual region. Res_r(i) The pixel is close to the pixel, and Res_r ( i ) represents the pixel of the shaped residual region.
2. The apparatus of claim 1, wherein, In order to generate the coded pixel region of the input image, the processor applies in-loop shaping.
3. The apparatus of claim 1, wherein, Generating the quantized residual region includes: The forward encoding transform is applied to the shaped residual region to generate transformed data; and A forward encoder quantizer is applied to the transformed data to generate quantized data.
4. The apparatus of claim 3, wherein, Generating the dequantized residual region includes: Applying an inverse encoder-quantizer to the quantized data to generate inverse-quantized data; and The inverse encoding transformation is applied to the inverse-quantized data to generate the dequantized residual region.
5. The apparatus of claim 1, wherein, Generating the reference pixel region to be stored on the reference frame buffer includes applying a loop filter to the reconstructed pixel region.
6. The apparatus of claim 1, wherein, Generating the shaped residual region includes calculating: in, Orig_sample ( i ) represents the pixels of the input image region.
7. The apparatus of claim 1, wherein, For each segment of the piecewise linear representation of the forward integer function, the integer metadata includes the absolute value of the increment and a flag for the absolute value of the increment.
8. The apparatus of claim 7, wherein, For bin_ce_delta_abs[i], which represents the absolute value of the increment of segment i, its value represents the difference between the codeword allocated in segment i and the codeword allocated in segment i-1 (bin_ce_delta_abs[i-1]).
9. An apparatus for decoding an encoded bitstream to generate an output image represented using a first codeword, the apparatus comprising: The input end is used to receive encoded images that are partially encoded using the second codeword representation. and processor, wherein the processor: Receive the reshaping metadata of the encoded image; Based on the integer metadata, a forward integer function is generated that maps pixels from the first codeword representation to the second codeword representation; An inverse integer function is generated based on the integer metadata, wherein the inverse integer function maps pixels from the second codeword representation to the first codeword representation; and The encoded region of the encoded image is decoded based on the forward shaping function and the inverse shaping function to generate the output pixel region; In order to decode the encoded region, the processor: Generate the decoded and shaped residual region; The prediction region is generated based on pixels in the reference pixel buffer or in the previously decoded spatial neighborhood. A reconstructed pixel region is generated based on the decoded and shaped residual region, the prediction region, the forward shaping function, and the inverse shaping function, wherein the reconstructed samples in the reconstructed pixel region are obtained at least in part by forward shaping the corresponding prediction samples in the prediction region. The output pixel region of the output image is generated based on the reconstructed pixel region; and The output pixel region is stored in the reference pixel buffer; Generating the reconstructed pixel region includes calculating: in, Reco_sample ( i ) represents the pixels of the reconstructed pixel region. Res_d ( i ) represents the pixels in the decoded and shaped residual region. Inv () represents the inverse shaping function. Fwd () represents the forward shaping function, and Pred_sample ( i ) represents the pixels of the predicted region.
10. The apparatus of claim 9, wherein, The processor decodes the encoded region of the encoded image based on in-loop shaping.
11. The apparatus of claim 9, wherein, The integer metadata represents the forward integer function based on a piecewise linear representation, and for each segment of the piecewise linear representation of the forward integer function, the integer metadata includes the absolute value of the increment and a flag of the absolute value of the increment.
12. The apparatus of claim 11, wherein, For bin_ce_delta_abs[i], which represents the absolute value of the increment of segment i, its value represents the difference between the codeword allocated in segment i and the codeword allocated in segment i-1 (bin_ce_delta_abs[i-1]).