Integrated image reconstruction and video coding
Integrated signal reconstruction architectures in video coding address inefficiencies in existing technologies by employing out-of-loop and in-loop methods, enhancing coding efficiency and reducing complexity for high bit depth and dynamic range formats.
Patent Information
- Application Number
- JP2025171062
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-06-19
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-21
AI Technical Summary
Existing video coding technologies struggle with inefficient compression and reconstruction methods, particularly when dealing with higher bit depths and dynamic range formats, leading to increased complexity and suboptimal coding performance.
The implementation of integrated out-of-loop and in-loop signal reconstruction architectures, including normative out-of-loop reshapers, intra-only in-loop reshapers, and in-loop reshapers for prediction residuals, along with efficient signaling methods and encoder-based optimization tools, to enhance coding efficiency and reduce complexity.
These architectures improve coding efficiency and reduce complexity by allowing adaptive and flexible reconstruction methods, enabling better compression and quality in video coding, especially for high bit depth and dynamic range formats.
Smart Images

Figure 2026010049000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 686,738, filed June 19, 2018, U.S. Provisional Patent Application No. 62 / 680,710, filed June 5, 2018, U.S. Provisional Patent Application No. 62 / 629,313, filed February 12, 2018, U.S. Provisional Patent Application No. 62 / 561,561, filed September 21, 2017, and U.S. Provisional Patent Application No. 62 / 526,577, filed June 29, 2017. Each of these applications is incorporated herein by reference in its entirety.
[0002] This invention relates generally to image and video coding. More particularly, one embodiment of the present invention relates to integrated image reshaping and video coding. [Background technology]
[0003] In 2013, the MPEG group within the International Organization for Standardization (ISO), in collaboration with the International Telecommunications Union (ITU), published the first draft of the HEVC (also known as H.265) video coding standard. More recently, the group issued a call for evidence to support the development of a next-generation coding standard that would provide improved coding performance over existing video coding technologies.
[0004] As used herein, the term "bit depth" refers to the number of bits per color component of an image. This refers to the number of pixels used to represent a color. Traditionally, images have been coded with 8 bits per color component per pixel (e.g., 24 bits per pixel). However, modern architectures can now support higher bit depths, such as 10 bits, 12 bits, or even more.
[0005] In a traditional image pipeline, a captured image is quantized using a nonlinear optical-to-electrical transfer function (OETF) that converts linear scene light into a nonlinear video signal (e.g., gamma-coded RGB or YCbCr). At the receiver, this signal is then processed by an electro-optical transfer function (EOTF) that converts the video signal values into output screen color values before being displayed on a display. Such nonlinear functions include the traditional "gamma" curve specified in ITU-R Rec. BT.709 and BT.2020, as well as the SMPTE ST 2084 and the "PQ" (perceptual quantization) curves described in Rec. ITU-R BT.2100. Summary of the Invention
[0006] As used herein, the term "forward reshaping" refers to digital The term "reconstruction" refers to the process of sample-to-sample or codeword-to-codeword mapping of a digital image from an original bit depth and original codeword distribution or codeword representation (e.g., gamma or PQ) of an image to an image with the same or a different bit depth and a different codeword distribution or codeword representation. Reconstruction allows for improved compression ratios or improved image quality at a fixed bit rate. For example, but not limited to, in a 10-bit video coding architecture, reconstruction can be applied to 10-bit or 12-bit PQ-coded HDR video to improve coding efficiency. At the receiver, after decompressing the reconstructed signal, the receiver can apply an "inverse reshaping function" to restore the signal to the original codeword distribution. The inventors have recognized that improved techniques for integrated reconstruction and coding of images are needed as next-generation video coding standards begin to be developed. The method of the present invention is applicable to a variety of video content. The content may include, but is not limited to, content in standard dynamic range (SDR) and / or high dynamic range (HDR).
[0007] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise specified, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, it should not be assumed, based on this section, that problems identified with one or more approaches have been recognized in any prior art, unless otherwise specified.
[0008] An embodiment of the present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings, in which like reference symbols represent similar elements and in which: [Brief explanation of the drawings]
[0009] [Figure 1A] FIG. 1 illustrates an example process for a video distribution pipeline. [Figure 1B] FIG. 1 illustrates an exemplary process of data compression with signal reconstruction according to the prior art. [Figure 2A] FIG. 2 illustrates an exemplary architecture of an encoder with normative out-of-loop reconstruction according to one embodiment of the present invention. [Figure 2B] FIG. 1 illustrates an exemplary architecture of a decoder with normative out-loop reconstruction according to one embodiment of the present invention. [Figure 2C] FIG. 1 illustrates an exemplary architecture of an encoder with normative intra-only in-loop reconstruction according to one embodiment of the present invention. [Figure 2D] FIG. 1 illustrates an exemplary architecture of a decoder with normative intra-only in-loop reconstruction according to one embodiment of the present invention. [Figure 2E] FIG. 2 illustrates an exemplary architecture of an encoder with in-loop reconstruction of prediction residuals according to one embodiment of the present invention. [Figure 2F] FIG. 2 illustrates an exemplary architecture of a decoder with in-loop reconstruction of prediction residuals according to one embodiment of the present invention; [Figure 2G] FIG. 2 illustrates an exemplary architecture of an encoder with hybrid in-loop reconstruction according to one embodiment of the present invention. [Figure 2H] FIG. 2 illustrates an exemplary architecture of a decoder with hybrid in-loop reconstruction according to one embodiment of the present invention. [Figure 3A] FIG. 2 illustrates an exemplary process for encoding video using an out-of-loop reconstruction architecture according to one embodiment of the present invention. [Figure 3B] FIG. 2 illustrates an exemplary process for decoding video using an out-of-loop reconstruction architecture according to one embodiment of the present invention. [Figure 3C] FIG. 2 illustrates an exemplary process for encoding video using an intra-only in-loop reconstruction architecture in accordance with one embodiment of the present invention. [Figure 3D] FIG. 2 illustrates an exemplary process for decoding video using an intra-only in-loop reconstruction architecture according to one embodiment of the present invention. [Figure 3E] FIG. 2 shows an exemplary process for encoding video using an in-loop reconstruction architecture of prediction residuals in accordance with one embodiment of the present invention. [Figure 3F] FIG. 2 shows an exemplary process for decoding video using an in-loop reconstruction architecture of prediction residuals in accordance with one embodiment of the present invention. [Figure 4A] FIG. 2 illustrates an exemplary process for encoding video using any one or combination of three reconstruction-based architectures, according to one embodiment of the present invention. [Figure 4B] FIG. 2 illustrates an exemplary process for decoding video using any one or combination of three reconstruction-based architectures, according to one embodiment of the present invention. [Figure 5A] FIG. 2 is a diagram illustrating a reconstruction function restoration process in a video decoder according to an embodiment of the present invention. [Figure 5B] FIG. 2 is a diagram illustrating a reconstruction function restoration process in a video decoder according to an embodiment of the present invention. [Figure 6A] 4A and 4B illustrate examples of how chroma QP offset values vary according to luma quantization parameters (QP) for PQ and HLG encoded signals, according to one embodiment of the present invention. [Figure 6B] 4A and 4B illustrate examples of how chroma QP offset values vary according to luma quantization parameters (QP) for PQ and HLG encoded signals, according to one embodiment of the present invention. [Figure 7] FIG. 2 illustrates an example of a pivot-based representation of a reconstruction function, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Described herein are techniques for integrated out-of-loop and in-loop signal reconstruction coding for compressing images. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices have not been described in exhaustive detail so as not to unnecessarily obscure, obscure, or obfuscate the present invention.
[0011] [overview]
[0003] Example embodiments described herein relate to integrated signal reconstruction and encoding of video. In the encoder, a processor receives an input image in a first codeword representation represented by an input bit depth N and an input codeword mapping (e.g., gamma, PQ, etc.). The processor selects one encoder architecture (having a reshaper as a component part of the encoder) from two or more candidate encoder architectures that compress the input image using a second codeword representation that allows more efficient compression than the first codeword representation. These two or more candidate encoder architectures include an out-of-loop reconstruction architecture, an in-loop reconstruction architecture for intra-pictures only, or an in-loop architecture for prediction residuals. The processor then compresses the input image according to the selected encoder architecture.
[0012] In another embodiment, a decoder that generates an output image in a first codeword representation receives an encoded bitstream in which at least a portion of an encoded image is compressed with a second codeword representation. The decoder also receives associated reconstruction information. A processor receives signaling indicating a decoder architecture from among two or more candidate decoder architectures for decompressing the input encoded bitstream. The two or more candidate decoder architectures include an out-of-loop reconstruction architecture, an in-loop reconstruction architecture for intra-pictures only, or an in-loop architecture for prediction residuals. The processor also decompresses the encoded image in accordance with the received reconstruction architecture to generate an output image.
[0013] In another embodiment, in an encoder for compressing images according to an in-loop architecture for prediction residuals, a processor accesses an input image in a first codeword representation and generates a forward reconstruction function that maps pixels of the input image from the first codeword representation to a second codeword representation. The processor generates a backward reconstruction function based on the forward reconstruction function that maps pixels from the second codeword representation to the first codeword representation. Then, for an input pixel region in the input image, the processor: Calculating at least one prediction region based on pixel data in a reference frame buffer or a pre-encoded neighboring space; generating a reconstructed residual region based on the input pixel region, the prediction region, and the forward reconstruction function; generating a coded (transformed and quantized) residual region based on the reconstructed residual region; generating a decoded (dequantized and transformed) residual region based on the coded residual region; generating a reconstructed pixel region based on the decoded residual region, the prediction region, the forward reconstruction function, and the backward reconstruction function; A reference pixel region is generated based on the reconstructed pixel region to be stored in the reference frame buffer.
[0014] In another embodiment, in a decoder that generates an output image in a first codeword representation according to an in-loop architecture for prediction residuals, the processor receives an encoded bitstream partially encoded in a second codeword representation. The processor also receives associated reconstruction information. Based on the reconstruction information, the processor generates a forward reconstruction function that maps pixels from the first codeword representation to the second codeword representation and a backward reconstruction function that maps pixels from the second codeword representation to the first codeword representation. For a region of the encoded image, the processor: generating a decoded reconstructed residual domain based on the coded image; generating a prediction region based on pixels in a reference pixel buffer or a pre-decoded neighborhood; generating a reconstructed pixel region based on the decoded reconstructed residual region, the prediction region, the forward reconstruction function, and the backward reconstruction function; generating an output pixel region based on the reconstructed pixel region; The output pixel region is stored in the reference pixel buffer.
[0015] [Example video streaming processing pipeline] 1A illustrates an example process for a conventional video distribution pipeline 100, showing various stages from video capture to display of video content. A sequence of video frames 102 is captured or generated using an image generation block 105. The video frames 102 can be captured digitally (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video data 107. Alternatively, the video frames 102 may be captured on film by a film camera. The film is converted to a digital format to provide the video data 107. In a production phase 110, the video data 107 is edited to provide a video production stream 112.
[0016] The video data in the production stream 112 is then provided to a processor in block 115 for post-production. Post-production in block 115 may involve adjusting or changing the color or brightness in specific areas of the image to enhance the image quality or achieve a particular look for the image, according to the creative intent of the video creator. This is known as "color timing" or "color grading." Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, etc.) can be performed in block 115 to obtain the final version of the production for distribution 117. In post-production 115, the video images are viewed on a reference display 125.
[0017] Following post-production 115, the final production video data 117 is sent to downstream decoding and playback devices such as television sets, set-top boxes, movie theaters, etc. The signal 117 may be delivered to an encoding block 120 for distribution to a broadcasting system. In some embodiments, encoding block 120 may comprise an audio encoder and a video encoder that generate an encoded bitstream 122, such as those specified by ATSC, DVB, DVD, Blu-Ray, and other distribution formats. At the receiver, encoded bitstream 122 is decoded by a decoding unit 130 to generate a decoded signal 132 that represents the same or a close approximation of signal 117. The receiver may be attached to a target display 140 that may have characteristics quite different from those of the reference display 125. In this case, a display management block 135 may be used to generate a display-mapped signal 137 to map the dynamic range of decoded signal 132 to the characteristics of the target display 140.
[0018] signal reconstruction FIG. 1B illustrates an exemplary process for signal reconstruction according to prior art reference [1]. Given an input frame 117, a forward reconstruction block 150 analyzes the input and coding constraints and generates a codeword mapping function that maps the input frame 117 to a requantized output frame 152. For example, the input 117 can be encoded according to a certain electro-optical transfer function (EOTF) (e.g., gamma). In some embodiments, information about the reconstruction process can be conveyed to downstream devices (such as a decoder) using metadata. As used herein, the term "metadata" refers to any additional information transmitted as part of an encoded bitstream that assists a decoder in rendering a decoded image. Such metadata may include, but is not limited to, color space or gamut information, reference display parameters, and additional signal parameters, as described herein.
[0019] Following encoding 120 and decoding 130, the decoded frames 132 may be processed by a backward (or inverse) reconstruction function 160, which converts the requantized frames 132 back to the EOTF domain (e.g., gamma) for further downstream processing, such as the aforementioned display management processing 135. In some embodiments, the backward reconstruction function 160 may be integrated with a dequantizer in the decoder 130, for example, as part of the dequantizer in an AVC video decoder or an HEVC video decoder.
[0020] In this specification, the term "reshaper" may refer to a forward reconstruction function or an inverse reconstruction function used when encoding and / or decoding a digital image. Examples of reconstruction functions are discussed in references [1] and [2]. In this invention, it is assumed that a person skilled in the art can derive appropriate forward and inverse reconstruction functions according to the characteristics of the input video signal and the available bit depth of the encoding and decoding architectures.
[0021] Reference [1] proposes an in-loop block-based image reconstruction method for high dynamic range video coding. The design allows block-based reconstruction inside the coding loop, but at the cost of increased complexity. Specifically, the design requires maintaining two sets of decoded picture buffers. One set is for backward reconstructed (or non-reconstructed) decoded pictures, which can be used for both prediction without reconstruction and output to the display. The other set is for forward reconstructed decoded pictures, which are used only for prediction with reconstruction. Forward reconstructed decoded pictures can be calculated on the fly, but the complexity cost is very high, especially in the case of inter prediction (motion compensation with sub-pixel interpolation). In general, management of the Display Picture Buffer (DPB) is complex and requires careful attention. Therefore, the inventors have recognized that a simplified method for encoding video is desirable.
[0022] Embodiments of the reconstruction-based codec architectures presented herein can be divided into architectures with an external out-of-loop reshaper, architectures with an intra-only in-loop reshaper, and architectures with an in-loop reshaper for prediction residuals, also called "in-loop residual reshaper" for short. A video encoder or decoder can support any one of these architectures or a combination of them. Each of these architectures can be applied on its own or in combination with any other one. Each architecture can be applied for the luma component, the chroma component, or a combination of the luma component and one or more chroma components.
[0023] In addition to these three architectures, additional embodiments describe efficient signaling methods for metadata related to the reconstruction, and several encoder-based optimization tools that improve coding efficiency when reconstruction is applied.
[0024] Normative Out-of-Loop Reshaper 2A and 2B show the architecture of a video encoder 200A_E and a corresponding video decoder 200A_D with a "normative" out-of-loop reshaper. The term "normative" refers to the normative description of coding standards such as AVC, HEVC, etc., because in previous designs, reconstruction was considered a pre-processing step. This represents that, unlike the architecture of Figure 1B where bitstream adaptability according to the standard is checked after decoding 130, in Figure 2B adaptability is checked after backward reconstruction block 265 (e.g., at output 162 in Figure 1B).
[0025] In encoder 200A_E, two new blocks are added to a conventional block-based encoder (e.g., HEVC): a block 205 that estimates a forward reconstruction function, and a forward picture reconstruction block 210 that applies forward reconstruction to one or more color components of input video 117. In some embodiments, these two operations may be performed as part of a single image reconstruction block. Parameters 207 related to determining the backward reconstruction function at the decoder can be embedded in the coded bitstream 122 by passing them to a lossless encoder block of the video encoder (e.g., CABAC 220). These include intra- or inter-prediction 225, transform and quantization (T&Q), and inverse transform and quantization (Q&Q). -1 &T -1 ) as well as all operations related to loop filtering are performed using the reconstructed picture stored in the DPB 215.
[0026] In decoder 200A_D, two new normative blocks are added to a conventional block-based decoder: block 250, which recovers the backward reconstruction function based on the encoding reconstruction function parameters 207, and block 265, which applies the backward reconstruction function to decoded data 262 to generate decoded video signal 162. In some embodiments, the operations associated with blocks 250 and 265 may be combined into a single processing block.
[0027] 3A illustrates an exemplary process 300A_E for encoding video using an out-of-loop reconstruction architecture 200A_E according to one embodiment of the present invention. If reconstruction is not enabled (path 305), encoding proceeds as known in prior art encoders (e.g., HEVC). If reconstruction is enabled (path 310), encoding proceeds as known in prior art encoders (e.g., HEVC). The encoder has the option of applying a predetermined (default) reconstruction function 315, or adaptively determining a new reconstruction function 325 based on picture analysis 320 (as described, for example, in references [1]-[3]). Following forward reconstruction 330, the remainder of the encoding follows a conventional encoding pipeline (335). If adaptive reconstruction 312 is used, metadata related to the backward reconstruction function is generated as part of the "Encode Reshaper" step 327.
[0028] 3B illustrates an exemplary process 300A_D for decoding video using the out-of-loop reconstruction architecture 200A_D according to one embodiment of the present invention. If reconstruction is not enabled (path 355), then after decoding a picture (350), an output frame is generated (390), as in a conventional decoding pipeline. If reconstruction is enabled (path 360), then in step 370, the decoder decides whether to apply a predetermined (default) reconstruction function (375) or adaptively determine a backward reconstruction function (380) based on received parameters (e.g., 207). Following backward reconstruction 385, the remainder of the decoding follows a conventional decoding pipeline.
[0029] Normative intra-only in-loop reshaper 2C shows an exemplary architecture of an encoder 200B_E using normative intra-only in-loop reconstruction according to one embodiment of the present invention. This design is quite similar to the design proposed in reference [1], but only intra pictures are coded using this architecture to reduce the complexity, especially in the use of DPB memories 215 and 260.
[0030] Compared to the out-of-loop reconstruction 200A_E, the main difference in the encoder 200B_E is that the DPB 215 stores backward reconstructed pictures instead of reconstructed pictures. In other words, decoded intra pictures need to be backward reconstructed (by the backward reconstruction unit 265) before being stored in the DPB. The reasoning behind this approach is that if intra pictures are coded with reconstruction, the improved coding performance of the intra pictures propagates, and even if the inter pictures are coded without reconstruction, the coding of the inter pictures is also (implicitly) improved. In this way, reconstruction can be utilized without dealing with the complexity of in-loop reconstruction of inter pictures. Since the backward reconstruction 265 is part of the inner loop, it can be performed before the in-loop filter 270. The advantage of adding backward reconstruction before the in-loop filter is that in this case, the design of the in-loop filter can be optimized based on the characteristics of the original picture rather than the forward reconstructed picture.
[0031] 2D shows an example architecture of a decoder 200B_D using normative intra-only in-loop reconstruction according to one embodiment of the present invention, where determining an inverse reconstruction function (250) and applying inverse reconstruction (265) are performed before in-loop filtering 270.
[0032] Figure 3C shows an exemplary process 300B_E for encoding video using an intra-only in-loop reconstruction architecture according to one embodiment of the present invention. As shown, the operational flow in Figure 3C shares many elements with the operational flow in Figure 3A. In this example, by default, no reconstruction is applied to inter-coding. For intra-coded pictures, if reconstruction is enabled, the encoder again has the option of using the default reconstruction curve or applying adaptive reconstruction (312). If the picture is reconstructed, backward reconstruction 385 is included in the process, and the associated parameters are coded in step 327. A corresponding decoding process 300B_D is shown in Figure 3D.
[0033] As shown in FIG. 3D, reconstruction-related operations are enabled only for received intra pictures, and only if intra reconstruction is applied at the encoder.
[0034] In-loop reshaper for prediction residuals In coding, the term "residual" refers to the difference between a predicted value of a sample or data element and its original or decoded value. For example, given an original sample from the input video 117, denoted as Orig_sample, intra or inter prediction 225 can generate a corresponding predicted sample 227, denoted as Pred_sample. When no reconstruction is performed, the unreconstructed residual Res_u can be defined as: Res_u = Orig_sample - Pred_sample. (1)
[0035] In some embodiments, it may be beneficial to apply reconstruction in the residual domain. Figure 2E shows an example architecture of an encoder 200C_E using in-loop reconstruction on prediction residuals, according to one embodiment of the present invention. We will denote the forward reconstruction function as Fwd() and the corresponding backward reconstruction function as Inv(). In one embodiment, the reconstructed residual 232 can be defined as: Res_r = Fwd(Orig_sample) - Fwd(Pred_sample). (2)
[0036] Correspondingly, at the output 267 of the reverse reshaper 265, the reconstructed sample, denoted as Reco_sample 267, can be expressed as: Reco_sample= Inv(Res_d + Fwd(Pred_sample)), (3) Here, Res_d represents the residual 234, which is a close approximation of Res_r after in-loop encoding and decoding in 200C_E.
[0037] Note that the reconstruction is applied to the residual, but the actual input video pixels are not reconstructed. Figure 2F shows the corresponding decoder 200C_D. Note that, as shown in Figure 2F, based on equation (3), the decoder needs access to both the forward and backward reconstruction functions, which can be extracted using the received metadata 207 and the "reshaper decoding" block 250.
[0038] In one embodiment, to reduce complexity, equations (2) and (3) can be simplified. For example, assuming that the forward reconstruction function can be approximated by a piecewise linear function and the absolute difference between Pred_sample and Orig_sample is relatively small, equation (2) can be approximated as follows: Res_r = a(Pred_sample) * (Orig_sample -Pred_sample), (4) Here, a(Pred_sample) represents a scaling factor based on the value of Pred_sample. From equations (3) and (4), equation (3) can be approximated as follows: Reco_sample= Pred_sample + (1 / a(Pred_sample))*Res_r, (5) Therefore, in one embodiment, only the scaling factor a(Pred_sample) of the piecewise linear model needs to be communicated to the decoder.
[0039] 3E and 3F show an example process flow 300C_E for encoding video using in-loop reconstruction of prediction residuals and an example process flow 300C_D for decoding video. These processes are quite similar to the processes shown in FIGS. 3A and 3B and therefore do not require further explanation.
[0040] Table 1 summarizes the key features of the three proposed architectures.
[0041] [Table 1]
[0042] 4A and 4B show an example encoding process flow for encoding and an example decoding process flow for decoding using a combination of the three proposed architectures. As shown in FIG. 4A, when reconstruction is not enabled, the input video is encoded according to a known video coding technique (e.g., HEVC, etc.) without reconstruction. When reconstruction is enabled, the encoder selects one of the three proposed architectures depending on the capabilities and / or input characteristics of the target receiver. Any one of the main proposed methods can be selected. For example, in one embodiment, the encoder can switch between these methods at the scene level, where "scene" refers to a sequence of consecutive frames with similar luminance characteristics. In another embodiment, high-level parameters are defined at the sequence parameter set (SPS) level.
[0043] As shown in FIG. 4B, depending on the received signal of reconstruction information, the decoder can invoke one of the corresponding decoding processes to decode the incoming encoded bitstream.
[0044] Hybrid In-Loop Reconfiguration FIG. 2G shows an example architecture 200D_E of an encoder using a hybrid in-loop reconstruction architecture. This architecture combines elements from both the intra-only in-loop reconstruction architecture 200B_E and the in-loop residual architecture 200C_E described above. Under this architecture, intra slices are coded according to the in-loop intra reconstruction coding architecture (e.g., 200B_E in FIG. 2C) with one difference: for intra slices, backward picture reconstruction 265-1 is performed after loop filtering 270-1. In another embodiment, in-loop filtering of intra slices can be performed after backward reconstruction. However, experimental results indicate that such a configuration may result in worse coding efficiency than when backward reconstruction is performed after loop filtering. The remaining operations are the same as those described above.
[0045] Inter slices are coded according to the in-loop residual coding architecture (e.g., 200C_E in FIG. 2E), as previously described. As shown in FIG. 2G, an intra / inter slice switch allows switching between the two architectures depending on the slice type being coded.
[0046] Figure 2H shows an example architecture 200D_D of a decoder using hybrid in-loop reconstruction. Again, intra slices are decoded according to an in-loop intra reconstruction decoder architecture (e.g., 200B_D in Figure 2D). Again, for intra slices, loop filtering 270-1 occurs before backward picture reconstruction 265-1. Inter slices are decoded according to an in-loop residual decoding architecture (e.g., 200C_D in Figure 2F). As shown in Figure 2H, an intra / inter slice switch allows switching between the two architectures depending on the slice type in a picture of coded video.
[0047] Figure 4A can be easily extended to include a method for hybrid in-loop reconstruction encoding by invoking the encoding process 300D-E shown in Figure 2G. Similarly, Figure 4B can be easily extended to include a method for hybrid in-loop reconstruction decoding by invoking the decoding process 300D-D shown in Figure 2H.
[0048] Slice-level reconstruction Embodiments of the present invention allow for various slice-level adaptive adjustments. For example, to reduce computational complexity, reconstruction can be enabled only for intra-slices or only for inter-slices. In another embodiment, reconstruction can be enabled based on the value of a temporal ID (e.g., the HEVC variable TemporalId (reference
[11] , where TemporalId=nuh_temporal_id_plus1-1). For example, if the TemporalId of the current slice is less than or equal to a default value, then the slice_reshaper_enable_flag of the current slice can be set to 1. otherwise slice_reshaper_enable_flag is 0. To avoid sending the slice_reshaper_enable_flag parameter for each slice, the sps_reshaper_temporal_id parameter can be specified at the SPS level, and its value can thus be deduced.
[0049] For slices with reconstruction enabled, the decoder needs to know which reconstruction model to use. In one embodiment, the decoder can always use the reconstruction model defined at the SPS level. In another embodiment, the decoder can always use the reconstruction model defined in the slice header. If a reconstruction model is not defined for the current slice, the decoder can apply the reconstruction model used in the most recently decoded slice that used reconstruction. In another embodiment, the reconstruction model can always be specified for intra-slices, regardless of whether reconstruction is used for intra-slices or not. In such an implementation, the parameters slice_reshaper_enable_flag and slice_reshaper_model_present_flag need to be separated. An example of such slice syntax is shown in Table 5.
[0050] Signaling of reconfiguration information Information related to forward and / or reverse reconstruction can be present in various information layers, such as a picture parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, supplemental information (SEI), or any other high-level syntax. By way of example and not limitation, Table 2 provides an example of high-level syntax in an SPS for signaling whether reconfiguration is enabled, whether reconfiguration is adaptive, and which of the three architectures is being used.
[0051] [Table 2]
[0052] Additional information can also be carried in some other layer, e.g., the slice header. The reconstruction function can be described by a look-up table (LUT), a piecewise polynomial, or other kinds of parametric models. The type of reconstruction model used to convey the reconstruction function can be signaled by an additional syntax element, e.g., the reshaping_model_type flag. For example, consider a system using two different representations: model_A (e.g., reshaping_model_type=0) represents the reconstruction function as a set of piecewise polynomials (see, e.g., reference [4]), while in model_B (e.g., reshaping_model_type=1), the reconstruction function is adaptively derived by assigning codewords to different luminance bands based on the image luminance characteristics and visual importance (e.g., (See, for example, reference [3].) Table 3 provides an example of syntax elements in a video slice header that assist the decoder in determining the appropriate reconstruction model to use.
[0053] [Table 3]
[0054] The following three tables list bitstream syntax alternatives for signal reconstruction at the sequence layer, slice layer, or coding tree unit (CTU) layer.
[0055] [Table 4]
[0056] [Table 5]
[0057] [Table 6]
[0058] For Tables 4 to 6, the example semantics can be expressed as follows: sps_reshaper_enable_flag equal to 1 specifies that the reshaper is used for Coded Video Sequences (CVS). sps_reshaper_enabled_flag equal to 0 specifies that the reshaper is not used for CVS. slice_reshaper_enable_flag equal to 1 specifies that the reshaper is enabled for the current slice. slice_reshaper_enable_flag equal to 0 specifies that the reshaper is not enabled for the current slice. sps_reshaper_signal_type indicates the distribution or representation of the original codeword. By way of example and not limitation, sps_reshaper_signal_type equal to 0 specifies SDR (gamma), sps_reshaper_signal_type equal to 1 specifies PQ, and sps_reshaper_signal_type equal to 2 specifies HLG. reshaper_CTU_control_flag equal to 1 indicates that it is possible to adapt the reshaper to each CTU. reshaper_CTU_control_flag equal to 0 indicates that it is not possible to adapt the reshaper to each CTU. When reshaper_CUT_control_flag is not present, its value shall be inferred to be 0. reshaper_CTU_flag equal to 1 specifies that the reshaper is used for the current CTU. reshaper_CUT_flag equal to 0 specifies that the reshaper is used for the current CTU. Specifies that reshaper_CTU_flag is not used for the current CTU. When reshaper_CTU_flag is not present, its value shall be inferred to be equal to slice_reshaper_enabled_flag. sps_reshaper_model_present_flag equal to 1 indicates that sps_reshaper_model() is present in SPS. sps_reshaper_model_present_flag equal to 0 indicates that sps_reshaper_model() is present in SPS. Indicates that aper_model() does not exist in the SPS. slice_reshaper_model_present_flag equal to 1 indicates that slice_reshaper_model() is present in the slice header. slice_reshaper_model_present_flag equal to 0 indicates that slice_reshaper_model() is not present in the slice header. sps_reshaper_chromaAdj equal to 1 indicates that the chroma QP adjustment is done using chromaDQP. sps_reshaper_chromaAdj equal to 2 indicates that the chroma QP adjustment is done using chroma scaling. sps_reshaper_ILF_opt indicates whether the in-loop filter should be applied in the original domain or in the reconstructed domain for intra- and inter-slices. For example, using the 2-bit syntax where the least significant bit relates to intra-slices, it is as in the following table:
[0059] [Table 7]
[0060] In some embodiments, this parameter can be adjusted at the slice level. For example, in one embodiment, a slice may include slice_reshape_ILFOPT_flag when slice_reshaper_enable_flag is set to 1. In another embodiment, an SPS may include a sps_reshaper_ILF_Tid parameter when sps_reshaper_ILF_opt is enabled. If the TemporalID of the current slice is less than or equal to sps_reshaper_ILF_Tid and slice_reshaper_enable_flag is set to 1, the in-loop filter is applied in the reconstructed domain. Otherwise, the in-loop filter is applied in the non-reconstructed domain.
[0061] In Table 4, the saturation QP adjustment is controlled at the SPS level. In one embodiment, the saturation QP adjustment can also be controlled at the slice level. For example, in each slice, when slice_reshaper_enable_flag is set to 1, the syntax element slice_reshape_chromaAdj_flag can be added. In another embodiment, in the SPS, when sps_reshaper_chromaAdj is enabled, the syntax element sps_reshaper_ChromaAdj_Tid can be added. If the TemporalID of the current slice is ≦ sps_reshaper_ChromaAdj_Tid and slice_reshaper_enable_flag is set to 1, the saturation adjustment is applied. Otherwise, the saturation adjustment is not applied. Table 4B shows an example variation of Table 4 using the above syntax.
[0062] [Table 8]
[0063] sps_reshaper_ILF_Tid specifies the highest TemporalID for which the in-loop filter is applied for the reconstructed slices in the reconstruction domain. sps_reshaper_chromaAdj_Tid specifies the highest TemporalID for which chroma adjustment is applied for the reconstructed slice.
[0064] In another embodiment, the reconstruction model can be defined using a reconstruction model ID, e.g., reshape_model_id, e.g., as part of the slice_reshape_model() function. The reconstruction model can be signaled at the SPS level, PPS level, or slice header level. If signaled in the SPS or PPS, the value of reshape_model_id can also be inferred from sps_seq_parameter_set_id or pps_pic_parameter_set_id. An example of how reshape_model_id is used for slices that do not have slice_reshape_model() (e.g., slice_reshaper_model_present_flag equals 0) is shown in Table 5B below, which is a variation of Table 5.
[0065] [Table 9]
[0066] In the example syntax, the parameter reshape_model_id specifies the value of reshape_model being used. This value of reshape_model_id shall be in the range 0 to 15.
[0067] As an example of the use of the proposed syntax, consider an HDR signal coded using PQ EOTF where reconstruction is used at the SPS level, no specific reconstruction is used at the slice level (reconstruction is used for all slices), and CTU adaptation is enabled only for inter-slice. sps_reshaper_signal_type=1(PQ); sps_reshaper_model_present_flag=1; / / Note: slice_reshaper_enable_flag can be manipulated to enable and disable the reshaper for inter-slices. slice_reshaper_enable_flag=1; if(CTUAdp) { if(I_slice) slice_reshaper_model_present_flag=0; reshaper_CTU_control_flag=0; else slice_reshaper_model_present_flag=0; reshaper_CTU_control_flag=1; } else { slice_reshaper_model_present_flag=0; reshaper_CTU_control_flag=0; }
[0068] As another example, consider an SDR signal where reconstruction is applied only at the slice level, and only for intra-slices. CTU reconstruction adaptation is enabled only for inter-slices. Then: sps_reshaper_signal_type=0(SDR); sps_reshaper_model_present_flag = 0; slice_reshaper_enable_flag = 1; if (I_slice) { slice_reshaper_model_present_flag = 1; reshaper_CTU_control_flag = 0; } else { slice_reshaper_model_present_flag = 0; if (CTUAdp) reshape_CTU_control_flag = 1; else reshaper_CTU_control_flag = 0; }
[0069] At the CTU level, in one embodiment, CTU-level reconstruction can be enabled based on the luminance characteristics of the CTU. For example, for each CTU, an average luminance (e.g., CTU_avg_lum_value) can be calculated, this average luminance can be compared with one or more thresholds, and based on the results of those comparisons, it can be determined whether to turn on or off the reconstruction. For example, if CTU_avg_lum_value < THR1, or if CTU_avg_lum_value > THR2, or if THR3 < CTU_avg_lum_value < THR4, for this CTU, reshaper_CTU_Flag = 1. In one embodiment, instead of using the average luminance, other luminance characteristics of the CTU such as the minimum luminance, the maximum luminance, or the average luminance and variance can be used. Chroma-based characteristics of the CTU can also be applied, or luminance characteristics and chroma characteristics and thresholds can be combined.
[0070] As described above (e.g., for the steps in Figures 3A, 3B, and 3C), embodiments can support both default or static reconstruction functions, or adaptive reconstruction. A "default reshaper" can be used to implement a default reconstruction function, thus reducing the complexity of analyzing each picture or scene when deriving a reconstruction curve. In this case, there is no need to signal the backward reconstruction function at the scene level, picture level, or slice level. The default reshaper can be implemented by using a fixed mapping curve stored in the decoder to avoid signaling, or it can be signaled once as part of a sequence-level parameter set. In another embodiment, a previously decoded adaptive reconstruction function can be reused for a later picture in coding order. In another embodiment, the reconstruction curve can be signaled as a difference from the previously decoded one. In other embodiments (e.g., for in-loop residual reconstruction where both the Inv() and Fwd() functions are needed to perform backward reconstruction), only one of the Inv() or Fwd() functions can be signaled in the bitstream, or alternatively, both can be signaled to reduce decoder complexity. Tables 7 and 8 provide two examples of signaling reconstruction information.
[0071] In Table 7, the reconstruction function is presented as a set of quadratic polynomials. The reconstruction function is a simplified syntax of the Exploratory Experimental Model (ETM) (reference [5]). A variation of this can also be found in reference [4].
[0072] [Table 10]
[0073] reshape_input_luma_bit_depth_minus8 specifies the sample bit depth of the input luma component for the reconstruction process. coeff_log2_offset_minus2 specifies the number of fractional bits for the calculation of the reconstruction-related coefficients of the luminance component. The value of coeff_log2_offset_minus2 shall be in the range 0 to 3, inclusive. reshape_num_ranges_minus1 plus 1 specifies the number of ranges in the piecewise reconstruction function. When not present, the value of reshape_num_ranges_minus1 is inferred to be 0. reshape_num_ranges_minusl shall be in the range 0 to 7, inclusive, for the luma component. reshape_equal_ranges_flag equal to 1 specifies that the piecewise reconstruction function is divided into NumberRanges parts of approximately equal length, and the length of each range is not explicitly signaled. reshape_equal_ranges_flag equal to 0 specifies that the length of each range is explicitly signaled. reshape_global_offset_val is used to obtain the offset value used to specify the start point of the 0th range. reshape_range_val[i] is used to obtain the length of the ith range of the luminance component. reshape_continuity_flag specifies the continuity properties of the reconstruction function for the luma component. If reshape_continuity_flag is equal to 0, zero-order continuity is applied to the piecewise linear inverse reconstruction function between consecutive pivot points. If reshape_continuity_flag is equal to 1, a full second-order polynomial inverse reconstruction function between consecutive pivot points is derived using first-order smoothing. reshape_poly_coeff_order0_int[i] specifies the integer value of the zeroth order polynomial coefficient of the i-th part of the luminance component. reshape_poly_coeff_order0_frac[i] specifies the fractional value of the zeroth order polynomial coefficient of the i-th part of the luminance component. reshape_poly_coeff_orderl_int is a first-order polynomial of the luminance component. Specifies the integer value of the nominal coefficient. reshape_poly_coeff_orderl_frac specifies the fractional value of the first-order polynomial coefficient for the luminance component.
[0074] Table 8 lists an exemplary embodiment of an alternative parametric representation according to model_B (reference [3]) mentioned above.
[0075] [Table 11]
[0076] In Table 8, in one embodiment, the syntax parameters can be defined as follows: reshape_model_profile_type specifies the profile type used in the reshaper construction process. reshape_model_scale_idx specifies the index value of the scale factor (represented as ScaleFactor) used in the reshaper construction process. The value of ScaleFactor allows for improved control of the reconstruction function, which improves overall coding efficiency. Further details regarding the use of this ScaleFactor are provided in connection with the discussion of the reconstruction function recovery process (e.g., as shown in Figures 5A and 5B). By way of example and not limitation, the value of reshape_model_scale_idx shall be in the range 0 to 3, inclusive. In one embodiment, the mapping relationship between scale_idx and ScaleFactor is as shown in the table below: ScaleFactor = 1.0 - 0.05* reshape_model_scale_idx. is given by
[0077] [Table 12]
[0078] In another embodiment, for a more efficient fixed-point implementation, the mapping relationship is: ScaleFactor = 1 - 1 / 16* reshape_model_scale_idx is given by
[0079] [Table 13]
[0080] reshape_model_min_bin_idx specifies the minimum bin index used in the reshaper construction process. The value of reshape_model_min_bin_idx must be in the range of 0 to 31. reshape_model_max_bin_idx specifies the maximum bin index used in the reshaper construction process. The value of reshape_model_max_bin_idx must be in the range of 0 to 31. reshape_model_num_band specifies the number of bands used in the reshaper construction process. The value of reshape_model_num_band must be in the range of 0 to 15. reshape_model_band_profile_delta[i] specifies the delta value used to adjust the profile of the i-th band in the reshaper construction process. The value of reshape_model_band_profile_delta[i] shall be in the range of 0 to 1.
[0081] Compared to reference [3], the syntax in Table 8 is much more efficient by defining a set of "default profile types," e.g., highlights, midtones, and darks. In one embodiment, each type has a default visual band importance profile. The default bands and corresponding profiles can be implemented in the decoder as fixed values or signaled using high-level syntax (e.g., sequence parameter sets). At the encoder, each image is first analyzed and categorized into one of the profiling types. The profile type is signaled by the syntax element "reshape_model_profile_type." In adaptive reconstruction, the default profiling is further adjusted by a delta for each or a subset of luma bands to capture the full range of image dynamics. The delta value is derived based on the visual importance of the luma band and is signaled by the syntax element "reshape_model_band_profile_delta."
[0082] In one embodiment, the delta value can only take on values of 0 or 1. In the encoder, visual importance is determined by comparing the percentage of band pixels in the overall image with the percentage of band pixels in a "dominant band," where dominant bands can be detected using local histograms. If the pixels in a band are concentrated in a small local block, then that band is most likely to be visually important in that block. The dominant band counts are summed and normalized to make a meaningful comparison to obtain the delta value for each band.
[0083] At the decoder, the reshaper function restoration process must be invoked to derive the reconstruction LUT based on the method described in [3]. Therefore, the complexity is higher compared to the simpler piecewise approximation model, which only requires evaluating a piecewise polynomial function to compute the LUT. The advantage of using the parametric model syntax is that The advantage of using a reshaper is that it can significantly reduce the bit rate. For example, based on typical test content, the model shown in Table 7 requires 200-300 bits to signal the reshaper, while a parametric model (as in Table 8) uses only about 40 bits.
[0084] In another embodiment, the forward reconstruction lookup table can be derived according to a parametric model of the dQP values, as shown in Table 9. For example, in one embodiment, dQP = clip3(min, max, scale*X+offset) where min and max represent the bounds of dQP, scale and offset are two parameters of the model, and X represents a parameter derived based on signal intensity (e.g., pixel intensity values, or in the case of a block, a metric of the block intensity, such as its minimum, maximum, mean, variance, standard deviation, etc.). For example, but not limited to, dQP = clip3(-3, 6, 0.015*X - 7.5) is.
[0085] [Table 14]
[0086] In one embodiment, the parameters in Table 9 can be defined as follows: full_range_input_flag specifies the range of the input video signal. A full_range_input_flag of 0 corresponds to a standard dynamic range input video signal. A full_range_input_flag of 1 corresponds to a full range input video signal. If not present, full_range_input_flag is inferred to be 0.
[0087] Note: In this specification, the term "full range video" means that the valid code words in the video are not "restricted." For example, for 10-bit full range video, the valid code words are 0 to 1023, with 0 mapping to the lowest luminance level. In contrast, for 10-bit "standard range video," the valid code words are 64 to 940, with 64 mapping to the lowest luminance level.
[0088] For example, the calculation of "full range" and "standard range" can be done as follows: When the normalized luminance value Ey' in
[0001] is coded with BD bits (for example, BD=10, 12, etc.), the following results: Full range: Y=clip3(0,(1≪BD)-1,Ey'*((1≪BD)-1))) Standard range: Y=clip3(0,(1≪BD)-1,round(l≪(BD-8)*(219*Ey'+16)))
[0089] This syntax is the "video_fu" HEVC VUI parameter described in section E.2.1 of the HEVC (H.265) specification (reference
[11] ). ll_range_flag" syntax. dQP_model_scale_int_prec specifies the number of bits used to represent dQP_model_scale_int. dQP_model_scale_int_prec equal to 0 indicates that dQP_model_scale_int is not signaled and is inferred to be 0. dQP_model_scale_int specifies the integer value of the dQP model scale. dQP_model_scale_frac_prec_minus16 plus 16 specifies the number of bits used to represent dQP_model_scale_frac. dQP_model_scale_frac specifies the decimal value of the dQP model scale.
[0090] The variable dQPModelScaleAbs is derived as follows: dQPModelScaleAbs = dQP_model_scale_int << (dQP_model_scale_frac_prec_minus16 + 16) + dQP_model_scale_frac dQP_model_scale_sign specifies the sign of the dQP model scale. When dQPModelScaleAbs is equal to 0, dQP_model_scale_sign is not signaled and is inferred to be 0. dQP_model_offset_int_prec_minus3 plus 3 specifies the number of bits used to represent dQP_model_offset_int. dQP_model_offset_int specifies the integer value of the dQP model offset. dQP_model_offset_frac_prec_minusl plus 1 specifies the number of bits used to represent dQP_model_offset_frac. dQP_model_offset_frac specifies the fractional value of the dQP model offset.
[0091] The variable dQPModelOffsetAbs is derived as follows: dQPModelOffsetAbs = dQP_model_offset_int << (dQP_model_offset_frac_prec_minus1 + 1) + dQP_model_offset_frac dQP_model_offset_sign specifies the sign of the dQP model offset. When dQPModelOffsetAbs is equal to 0, dQP_model_offset_sign is not signaled and is inferred to be equal to 0. dQP_model_abs_prec_minus3 plus 3 specifies the number of bits used to represent dQP_model_max_abs and dQP_model_min_abs. dQP_model_max_abs specifies the integer value of the dQP model max. dQP_model_max_sign specifies the sign of the dQP model max. When dQP_model_max_abs is equal to 0, dQP_model_max_sign is not signaled and is inferred to be equal to 0. dQP_model_min_abs specifies the integer value of the dQP model min. dQP_model_min_sign specifies the sign of the dQP model min. When dQP_model_min_abs is equal to 0, dQP_model_min_sign is not signaled and is inferred to be equal to 0.
[0092] Model C Decryption Process Given the syntax elements in Table 9, the reconstruction LUT can be derived as follows: The variable dQPModelScaleFP is derived as follows: dQPModelScaleFP = ((1- 2*dQP_model_scale_sign) * dQPModelScaleAbs ) << (dQP_model_offset_frac_prec_minus1 + 1). The variable dQPModelOffsetFP is derived as follows: dQPModelOffsetFP = ((1-2* dQP_model_offset_sign) * dQPModelOffsetAbs ) << (d QP_model_scale_frac_prec_minus16 + 16). The variable dQPModelShift is derived as follows: dQPModelShift = (dQP_model_offset_frac_prec_minus1 + 1) + (dQP_model_scale_f rac_prec_minus16 + 16). The variable dQPModelMaxFP is derived as follows: dQPModelMaxFP = ((1- 2*dQP_model_max_sign) * dQP_model_max_abs) << dQPModelS hift. The variable dQPModelMinFP is derived as follows. dQPModelMinFP = ((1- 2*dQP_model_min_sign) * dQP_model_min_abs) << dQPModelS hift. for Y=0:maxY / / For example, for 10-bit video, maxY=1023 { dQP[Y]=clip3(dQPModelMinFP, dQPModelMaxFP, dQPModelScaleFP*Y+dQPModelOffsetFP); slope[Y]=exp2((dQP[Y]+3) / 6); / / Perform fixed-point exp2, where exp2(x)=2^(x); } If(full_range_input_flag==0) / / If the input is a standard range video For Y outside the standard range (i.e., Y=[0:63] and [940:1023]): , set slope[Y]=0; CDF[0]=slope[0]; for Y=0:maxY-1 { CDF[Y+1]=CDF[Y]+slope[Y]; / / CDF[Y] is the integral of slope[Y] } for Y=0:maxY { FwdLUT[Y]=round(CDF[Y]*maxY / CDF[maxY]); / / FwdLU after rounding and normalization Get a T }
[0093] In another embodiment, the forward reconstruction function can be represented as a collection of intensity pivot points (In_Y) and their corresponding code words (Out_Y), as shown in Table 10. To simplify encoding, the input intensity range is described by a starting pivot followed by a sequence of equally spaced pivots using a linear piecewise representation. An example representation of the forward reconstruction function for 10-bit input data is shown in Figure 7.
[0094] [Table 15]
[0095] In one embodiment, the parameters in Table 10 may be defined as follows: full_range_input_flag specifies the range of the input video signal. A full_range_input_flag of 0 corresponds to a standard range input video signal. A full_range_input_flag of 1 corresponds to a full range input video signal. When not present, full_range_input_flag is inferred to be 0. bin_pivot_start specifies the pivot value (710) of the first equal-length bin. When full_range_input_flag is equal to 0, bin_pivot_start shall be greater than or equal to the minimum standard range input and less than the maximum standard range input. (For example, for a 10-bit SDR input, bin_pivot_start (710) shall be between 64 and 940.) bin_cw_start specifies the mapped value (715) of bin_pivot_start (710) (e.g., bin_cw_start=FwdLUT[bin_pivot_start]). log2_num_equal_bins_minus3 plus 3 specifies the number of equal length bins following the starting pivot (710). The variables NumEqualBins and NumTotalBins are defined by the following formulas: NumEqualBins = 1<<(log2_num_equal_bins_minus3+3) if(full_range_input_flag==0) NumTotalBins=NumEqualBins+4 else NumTotalBins=NumEqualBins+2 NOTE: Experimental results have shown that most forward reconstruction functions can be represented using eight equal-length segments, but complex reconstruction functions may require more segments (e.g., 16 or more). equal_bin_pivot_delta specifies the length of the equal-length bins (e.g., 720-1, 720-N). NumEqualBins*equal_bin_pivot_delta must be less than or equal to the valid input range. (For example, if full_range_input_flag is 0, the valid input range for 10-bit input is 940-64=876. If full_range_input_flag is 1, the valid input range for 10-bit input is 0 to 1023.) bin_cw_in_first_equal_bin specifies the number of codewords (725) to be mapped in the first equal length bin (720-1). bin_cw_delta_abs_prec_minus4 plus 4 specifies the number of bits used to represent bin_cw_delta_abs[i] of each subsequent equal length bin. bin_cw_delta_abs[i] specifies the value of bin_cw_delta_abs[i] for each subsequent equal length bin. bin_cw_delta[i] (e.g., 735) is the difference of the codeword (e.g., 740) in the current equal length bin i (e.g., 720-N) compared to the codeword (e.g., 730) in the previous equal length bin i-1. bin_cw_delta_sign[i] specifies the sign of bin_cw_delta_abs[i]. When bin_cw_delta_abs[i] is equal to 0, bin_cw_delta_sign[i] is not signaled and is inferred to be 0. The variable bin_cw_delta[i]=(1-2*bin_cw_delta_sign[i])*bin_cw_delta_abs[i].
[0096] Model D Decryption Process Given the syntax elements in Table 10, the reconstruction LUT can be derived as follows for a 10-bit input: Constant definitions: minIN=minOUT=0; maxIN=maxOUT=2^BD-1=1023 for 10-bit / / BD=bit depth minStdIN=64 for 10-bit maxStdIN=940 for 10 bits Step 1: For j=0, derive the pivot value In_Y[j]: NumTotalBins In_Y[0]=0; In_Y[NumTotalBins]=maxIN; if(full_range_input_flag==0) { In_Y[1]=minStdIN; In_Y[2]=bin_pivot_start; for(j=3:NumTotalBins-2) In_Y[j]=In_Y[j-1]+equal_bin_pivot_delta; In_Y[NumTotalBins-1]=maxStdIN; } else { In_Y[1]=bin_pivot_start; for j=2:NumTotalBins-1 In_Y[j]=In_Y[j-1]+equal_bin_pivot_delta; } Step 2: Derive the mapped value Out_Y[j] for j=0: NumTotalBins Out_Y[0]=0; Out_Y[NumTotalBins]=maxOUT; if(full_range_input_flag==0) { Out_Y[1]=0; Out_Y[2]=bin_cw_start; Out_Y[3]=bin_cw_start+bin_cw_in_first_equal_bin; bin_cw[3]=bin_cw_in_first_equal_bin; for j=(4:NumTotalBins-2) bin_cw[j]=bin_cw[j-1]+bin_cw_delta[j-4]; / / bin_cw_delta[i] starts from idx 0 for j=(4:NumTotalBins-2) Out_Y[j]=Out_Y[j-1]+bin_cw[j]; Out_Y[NumTotalBins-1]=maxOUT; } else { Out_Y[1]=bin_cw_start; Out_Y[2]=bin_cw_start+bin_cw_in_first_equal_bin; bin_cw[2]=bin_cw_in_first_equal_bin; for j=(3:NumTotalBins-1) bin_cw[j]=bin_cw[j-1]+bin_cw_delta[j-3]; / / bin_cw_delta[i ] starts at idx 0 for j=3:NumTotalBins-1 Out_Y[j]=Out_Y[j-1]+bin_cw[j]; } Step 3: Linear interpolation to get all LUT entries Init:FwdLUT[] for(j=0:NumTotalBins-1) { InS=In_Y[j]; InE=In_Y[j+1]; OutS=Out_Y[j]; OutE=Out_Y[j+1]; for(i=In_Y[j]:In_Y[j+1]-1) { FwdLUT[i]=OutS+round((OutE-OutS)*(i-InS) / (InE-InS)); } } FwdLUT[In_Y[NumTotalBins]]=Out_Y[NumTotalBins];
[0097] In general, reconstruction can be switched on or off for each slice. For example, only intra-slice reconstruction can be enabled, while inter-slice reconstruction can be disabled. In another example, reconstruction of the inter-slice with the highest temporal level can be disabled. (Note: As an example, in this specification, a temporal sublayer may be consistent with the definition of a temporal sublayer in HEVC.) In the reshaper model, in one example, only the reshaper model in SPS may be signaled, while in another example, the slice reshaper model in intra-slice may be signaled. Alternatively, the reshaper model in SPS may be signaled, allowing the slice reshaper model to update the SPS reshaper model for all slices, or allowing the slice reshaper model to update the SPS reshaper model only for intra-slices. For inter-slices following an intra-slice, either the SPS reshaper model or the intra-slice reshaper model can be applied.
[0098] As another example, Figures 5A and 5B show the reconstruction function recovery process in a decoder according to one embodiment, using the method described herein and in reference [3] using the visual rating range in
[0005] .
[0099] 5A, first (step 510), the decoder extracts the reshape_model_profile_type variable and sets the appropriate initial band profile for each bin (steps 515, 520, and 525). For example, in pseudocode: if(reshape_model_profile_type==0)R[b i ]=R bright [b i ]; elseif(reshape_model_profile_type==1)R[b i ]=R dark [bi ]; else R[b i ]=R mid [b i ].
[0100] In step 530, the decoder calculates the received reshape_model_band_profile_delta[b i ] value to define each band profile. Adjust the rule. for(i=0:reshape_model_num_band-l) {R[b i ]=R[b i ]+reshape_model_band_profile_delta[b i ]}.
[0101] In step 535, the decoder applies the adjusted values to each bin profile as follows: bin[j] is band b i If it belongs to R_bin[j]=R[b i ].
[0102] In step 540, the bin profile is modified as follows: if(j>reshape_model_max_bin_idx) or (j <reshape_model_min_bin_idx) {R_bin[j]=0}.
[0103] In parallel, in steps 545 and 550, the decoder can extract parameters to calculate scale factor values and candidate codewords for each bin[j] as follows: ScaleFactor=1.0-0.05*reshape_model_scale_idx CW_dft[j] = codeword in bin if using default reconstruction CW_PQ[j]=TotalCW / TotalNumBins.
[0104] When calculating the ScaleFactor value for a fixed-point implementation, instead of using the scaler 0.05, one can use 1 / 16=0.0625.
[0105] Proceeding to FIG. 5B, in step 560, the decoder begins pre-allocating codewords (CWs) for each bin based on the bin profile, as follows: If(R_bin[j]==0) CW[j]=0 If(R_bin[j]==1) CW[j]=CW_dft[j] / 2; If(R_bin[j]==2) CW[j]=min(CW_PQ[j], CW_dft[j]); If(R_bin[j]==3) CW[j]=(CW_PQ[j]+CW_dft[j]) / 2; If(R_bin[j]>=4) CW[j]=max(CW_PQ[j], CW_dft[j]);
[0106] In step 565, the decoder calculates all codewords used and refines / finishes the codeword (CW) allocation as follows: CW used =Sum(CW[j]): if(CW used >TotalCW) CW[j]=CW[j] / (CW used / TotalCW); else { CW_remain=TotalCW-CW used ; CW_remain is assigned to the bin with the largest R_bin[j]; }
[0107] Finally, in step 565, the decoder performs the following: a) by accumulating the CW[j] values b) generating a forward reconstruction function (eg, FwdLUT); b) multiplying the ScaleFactor value by the FwdLUT value to form a final FwdLUT (FFwdLUT); and c) generating an inverse reconstruction function InvLUT based on the FFwdLUT.
[0108] In a fixed-point implementation, the calculation of ScaleFactor and FFwdLUT can be expressed as follows: ScaleFactor = (1<< SF_PREC) - reshape_model_scale_idx FFwdLUT = (FwdLUT * ScaleFactor + (1 << (FP_PREC + SF_PREC - 1))) >> (FP_PREC + SF_PREC), Here, SF_PREC and FP_PREC are variables related to the default precision (for example, SF_PREC=4 and FP_PREC=14), and "c=a<<n" represents an operation to binary-shift a to the left by n bits (or c=a*(2 n )), "c=a≫n" means to change a to n represents the binary right shift operation by 2 bits (or c = a / (2 n )).
[0109] Derivation of saturation QP Chroma coding performance is closely related to luma coding performance. For example, AVC and HEVC define tables that specify the relationship between the quantization parameters (QP) of luma and chroma components or between luma and chroma. These specifications also allow the use of one or more chroma QP offsets, which provide additional flexibility in defining the QP relationship between luma and chroma. When reconstruction is used, the luma values are changed, and therefore the relationship between luma and chroma may be changed as well. To maintain and further improve coding efficiency during reconstruction, in one embodiment, at the coding unit (CU) level, a chroma QP offset is derived based on the reconstruction curve. This operation needs to be performed in both the decoder and the encoder.
[0110] As used herein, the term "coding unit" (CU) refers to a coding block (e.g., a macroblock, etc.). By way of example and not limitation, in HEVC, a CU is defined as a " "One coded block of luma samples for a video having a three sample array, two corresponding coded blocks of chroma samples, or a coded block of samples for a video that is coded using three separate color planes and syntax structures used to code monochrome video or samples."
[0111] In one embodiment, the chroma quantization parameter (QP) (chromaQP) value can be derived as follows: 1) Derive the equivalent luminance dQP mapping dQPLUT based on the reconstruction curve: for CW=0:MAX_CW_VALUE-1 dQPLUT[CW]=-6*log2(slope[CW]); where slope[CW] represents the slope of the forward reconstruction curve at each CW (codeword) point, and MAX_CW_VALUE is the maximum codeword value for a given bit depth, e.g., for a 10-bit signal, MAX_CW_VALUE=1024(2 10 ) Then, for each coding unit (CU): 2) Calculate the average luminance of the coding unit, denoted as AvgY: 3) Calculate the chromaDQP value based on dQPLUT[], AvgY, reconstruction architecture, inverse reconstruction function Inv(), and slice type as shown in Table 11 below.
[0112] [Table 16]
[0113] 4) Calculate chromaQP as follows: chromaQP = QP_luma + chromaQPOffset + chromaDQP; where chromaQPOffset represents the chroma QP offset of the coding unit, and QP_luma represents the luma QP of the coding unit. Note that the value of the chroma QP offset may be different for each chroma component (e.g., Cb and Cr), and the chroma QP offset value is communicated to the decoder as part of the encoded bitstream.
[0114] In one embodiment, dQPLUT[] can be implemented as a default LUT. Assume that all codewords are divided into N (e.g., N=32) bins, and each bin contains M=MAX_CW_VALUE / N (e.g., M=1024 / 32=32) codewords. When assigning new codewords to each bin, the number of codewords can be limited to 1 to 2*M, and therefore dQPLUT[1...2*M] can be pre-calculated and stored as a LUT. This approach can avoid approximations in floating-point or fixed-point calculations. This approach can also save encoding / decoding time. For each bin, one fixed chromaQPOffset is used for all codewords in this bin. The DQP value is set equal to dQPLUT[L], where L is the number of codewords in this bin, and 1≦L≦2*M.
[0115] The dQPLUT value can be pre-calculated as follows: for i=l:2*M slope[i]=i / M; dQPLUT[i]=-6*log2(slope[i]); end
[0116] When calculating dQPLUT[x], integer QP values can be obtained using various quantization methods such as round(), ceil(), floor(), or a mixture of them. For example, a threshold TH is set. When Y < TH, the dQP value can be quantized using floor(), and when Y ≥ TH, the dQP value can be quantized using ceil(). The usage of such quantization methods and corresponding parameters can be predefined in the codec or signaled in the bitstream for adaptation. An exemplary syntax that enables the mixing of such a quantization method with one threshold is shown as follows.
[0117]
Table 17
[0118] The quant_scheme_signal_table() function can be defined at various levels of the reconstruction syntax (e.g., sequence level, slice level, etc.) according to the need for adaptation.
[0119] In another embodiment, the chromaDQP value can be calculated by applying a scaling factor to the residual signal in each coding unit (more specifically, the transform unit). This scaling factor can be a luminance-dependent value, a) for example, as the first derivative (gradient) of the forward reconstruction LUT (see, for example, Equation (6) in the next section) or b)
Equation
[0120] [Table 18]
[0121] For example, if CSCALE_FP_PREC=16, Forward scaling: after the chroma residual is generated, but before transformation and quantization: - C_Res = C_orig - C_pred - C_Res_scaled = C_Res * S + (1 << (CSCALE_FP_PREC - 1))) >> CSCALE_FP_PREC Inverse scaling: after chroma dequantization and inverse transform, but before reconstruction: - C_Res_inv = (C_Res_scaled << CSCALE_FP_PREC) / S - C_Reco = C_Pred + C_Res_inv; Here, S is either S_cu or S_px. Note: In Table 12, in calculating Scu, the average luminance (AvgY) of the block is calculated before applying the inverse reconstruction. Alternatively, the inverse reconstruction can be applied before calculating the average luminance, e.g., Scu = SlopeLUT[Avg(Inv[Y])]. This alternative calculation order also applies to calculating the values in Table 11. That is, calculating Inv(AvgY) can be replaced by calculating the Avg(Inv[Y]) value. The latter approach can be considered more accurate, but increases computational complexity.
[0122] Encoder optimization for reconstruction In this section, we discuss several techniques for improving coding efficiency in an encoder by jointly optimizing reconstruction and encoder parameters when reconstruction is part of the normative decoding process (described in one of three candidate architectures). Generally, encoder optimization and reconstruction address the coding problem at different locations, each with its own limitations. In traditional imaging and coding systems, there are two types of quantization: a) sample quantization in the baseband signal (e.g., gamma or PQ coding) and b) transform-related quantization (part of compression). Reconstruction lies somewhere in between. Picture-based reconstruction is typically updated per picture and only allows sample value mapping based on its luminance level, without considering spatial information. In block-based codecs (e.g., HEVC), transform quantization (e.g., luma transform quantization) is applied within spatial blocks and can be spatially adjusted; therefore, encoder optimization methods must apply the same set of parameters for the entire block, which contains samples with different luminance values. As recognized by the inventors and described herein, combining reconstruction and encoder optimization can further improve coding efficiency.
[0123] Inter / Intra mode decision In traditional coding, the inter / intra mode decision is based on calculating a distortion function (dfunc()) between the original samples and the predicted samples. Examples of such functions include the sum of squared errors (SSE), sum of absolute differences (SAD), etc. In one embodiment, Such distortion metrics can be used with reconstructed pixel values. For example, when reconstruction is applied, if the original dfunct() uses Orig_sample(i) and Pred_sample(i), then dfunct() can use their corresponding reconstructed values Fwd(Orig_sample(i)) and Fwd(Pred_sample(i)). This approach allows for more accurate inter / intra decisions and therefore improves coding efficiency.
[0124] LumaDQP with reconstruction In the JCTVC HDR Common Test Criteria (CTC) document (reference [6]), lumaDQP and chromaQP offset are two encoder settings used to modify the quantization (QP) parameters of luma and chroma components to improve HDR coding efficiency. In this invention, several new encoder algorithms are proposed to further improve the original proposal. For each lumaDQP adaptation unit (e.g., 64x64 CTU), a dQP value is calculated based on the average input luma value of that unit (as in Table 3 of reference [6]). The final quantization parameter QP used for each coding unit within this lumaDQP adaptation unit should be adjusted by subtracting this dQP. The dQP mapping table is configurable in the encoder input configuration. This input configuration is used to calculate the dQP inp It is written as:
[0125] As discussed in references [6] and [7], existing coding methods use the same lumaDQP LUT for both intra and inter pictures. inp Intra-picture and inter-picture have different properties and quality characteristics. In this invention, we propose to adapt the lumaDQP setting value based on the video coding type. Therefore, two dQP mapping tables are configurable in the encoder input configuration, and the dQP inpIntra and dQP inpInter It is written as:
[0126] As mentioned above, when using the in-loop intra reconstruction method, since reconstruction is not performed on inter pictures, it is important that certain lumaDQP settings are applied to inter-coded video to achieve similar quality as if the inter pictures were reconstructed by the same reshaper used for intra pictures. In one embodiment, the lumaDQP settings for inter pictures should match the characteristics of the reconstruction curve used in intra pictures. The first derivative of the forward reconstruction function is Slope(x) = Fwd'(x) = (Fwd(x+dx)- Fwd(x-dx)) / (2dx), (6) In this case, in one embodiment, this expression can be expressed as the automatically derived dQP auto (x) value can be calculated as follows: If Slope(x) = 0, then dQP auto (x) = 0, otherwise dQP auto (x) = 6log2(Slope(x)), (7) where dQP auto (x) can be clipped to a reasonable range, e.g., [-6 6].
[0127] Luma DQP is enabled for intra pictures with reconstruction (i.e., external dQP inpIntra is set), the lumaDQP of the inter picture should take that into account. final is the dQP derived from the reshaper auto(Equation (7)) and dQP for intra pictures inpIntra In another embodiment, the inter picture dQP can be calculated by adding a set value to take advantage of intra quality propagation. final , dQP auto Set in It can also be determined by (dQP inpInter dQP in small increments) auto It can also be added to.
[0128] In one embodiment, when reconstruction is enabled, the following general formula for setting the luminance dQP value is used: The objective rules can be applied. (1) The luma dQP mapping table can be set independently for intra-pictures and inter-pictures (based on the video coding type). (2) If the picture in the coding loop is in the reconstruction domain (e.g., an intra-picture in an in-loop intra-reconstruction architecture or all pictures in an out-of-loop reconstruction architecture), the input luma dQP to delta QP mapping inp Similarly, the reconstruction area d QP rsp That is, dQP rsp (x) = dQP inp [Inv(x)] (8) is. (3) If the pictures in the coding loop are in the non-reconstructed domain (e.g., backward reconstructed or not, e.g., interpictures in an in-loop intra reconstruction architecture or all pictures in an in-loop residual reconstruction architecture), the input luma to delta QP mapping does not need to be converted and can be used directly. (4) Automatic inter delta QP derivation is only valid for in-loop intra reconstruction architectures. In such cases, the actual delta QP used for an inter picture is the sum of the automatically derived inputs, i.e., dQP final [x] = dQPinp [x] + dQP auto [x] (9) and dQP final [x] is clipped to a reasonable range, e.g. [-12 12]. It is possible. (5) The luminance to dQP mapping table can be updated at every image or when there is a change in the reconstruction LUT. The actual dQP adaptation (from the average luminance value of a block to get the corresponding dQP for the quantization of this block) can be done at the CU level (encoder configurable).
[0129] Table 13 summarizes the dQP settings for each of the three proposed architectures.
[0130] [Table 19]
[0131] Rate-Distortion Optimization (RDO) In the JEM6.0 software (reference [8]), when lumaDQP is enabled, RDO (rate distortion optimization) pixel-based weighted distortion is used. The weight table is fixed based on the luminance value. In one embodiment, the weight table is adaptively adjusted based on the lumaDQP setting value and should be calculated as proposed in the previous section. Two weights, sum of squared errors (SSE) and sum of absolute differences (SAD), are proposed as follows:
number
[0132] These weights calculated by Equation (10a) or Equation (10b) are total weights based on the final dQP, including both the input lumaDQP and the dQP derived from the forward reconstruction function. For example, based on Equation (9), Equation (10a) can be written as follows:
number
number
number
[0133] The weight derived from the input lumaDQP is W dQP Let us express it as f'(x) Let denote the first derivative (or gradient) of the forward reconstruction curve. In one embodiment, the total weight takes into account both the dQP value and the shape of the reconstruction curve, and thus the total weight value can be expressed as:
number
[0134] A similar approach can be applied to the chroma component as well. For example, in one embodiment, for chroma, dQP[x] can be defined according to Table 13.
[0135] Interaction with other coding tools When reconstruction is enabled, this section provides some examples of proposed changes required in other coding tools. Interactions may exist for any possible existing coding tool or future coding tool that will be included in the next generation video coding standard. The examples given below are not limiting. In general, the video signal domains (reconstructed, non-reconstructed, backward reconstructed) during the coding steps need to be identified, and the operations that handle the video signal at each step need to take the reconstruction effect into account.
[0136] Cross-component linear model prediction In CCLM (Cross-Component Linear Model Prediction) (reference [8]), the predicted saturation samples pred C (i,j) is the luminance restoration signal rec L '(i,j) can be used to derive it.
number
[0137] When reconstruction is enabled, in one embodiment, it may be necessary to distinguish whether the luma restored signal is in the reconstruction domain (e.g., out-of-loop reshaper or in-loop intra-reshaper) or in the non-reconstruction domain (e.g., in-loop residual reshaper). In one embodiment, the luma restored signal can be used implicitly as is without any additional signaling or operation. In other embodiments, if the restored signal is in the non-reconstruction domain, the luma restored signal can also be transformed to be in the non-reconstruction domain as follows:
number
[0138] In other embodiments, a bitstream syntax element can be added that signals which region (reconstructed or non-reconstructed) is desired, which can be determined by the RDO process, or the decision can be derived based on the decoded information, thus avoiding the overhead required by explicit signaling. Based on this decision, a corresponding operation can be performed on the recovered signal.
[0139] Reshaper with Residual Prediction Tool The HEVC Range Extension profile includes a residual prediction tool: the chroma residual signal is predicted from the luma residual signal at the encoder side as follows:
number
number
[0140] When reconstruction is enabled, it may be necessary to consider which luma residual to use for chroma residual prediction. In one embodiment, the "residual" can be used as is (it may or may not be reconstructed based on the reshaper architecture). In another embodiment, the luma residual can be forced into one domain (such as the non-reconstructed domain) and an appropriate mapping can be performed. In another embodiment, the appropriate handling can be derived by the decoder or explicitly signaled as described above.
[0141] Reshaper with adaptive clipping. Adaptive clipping (reference [8]) is a new tool introduced to signal the original data range with respect to content dynamics and perform adaptive clipping (based on internal bit depth information) instead of fixed clipping at each step in the compression workflow where clipping occurs (e.g., transform / quantization, loop filtering, output). It is a rule.
number
number
[0142] When reconstruction is enabled, in one embodiment, it may be necessary to locate the region the data flow currently resides in to perform the clipping accurately. For example, when processing clipping in the reconstruction domain data, the original clipping boundary needs to be transformed to the reconstruction domain.
number
[0143] Reshaper and Loop Filtering In HEVC and JEM6.0 software, loop filters such as ALF and SAO need to estimate optimal filter parameters using reconstructed luma samples and uncompressed "original" luma samples. When reconstruction is enabled, in one embodiment, you can specify (explicitly or implicitly) the region where you want to perform filter optimization. In one embodiment, you can estimate filter parameters on the reconstructed domain (relative to the reconstructed original domain when reconstruction is in the reconstructed domain). In other embodiments, you can estimate filter parameters on the non-reconstructed domain (relative to the original domain when reconstruction is in the non-reconstructed domain or backward reconstruction domain). For example, depending on the in-loop reconstruction architecture, the in-loop filter optimization (ILFOPT) options and operations can be described by Table 14 and Table 15.
[0144] Table 14. In-loop filtering optimization for in-loop intra-only reconstruction architecture and in-loop hybrid reconstruction [Table 20]
[0145] [Table 21]
[0146] While most of the detailed discussion herein refers to methods performed on the luma component, those skilled in the art will appreciate that similar methods can also be performed on chroma color components and chroma-related parameters such as chomaQPOffset (see, e.g., reference [9]).
[0147] In-the-loop reconstruction and region of interest (ROI) As used herein, the term "region of interest" (ROI) in reference to an image refers to a region of an image that is considered to be of particular interest. In this section, novel embodiments are presented that support in-loop reconstruction of the region of interest only. That is, in one embodiment, reconstruction may be applied only inside the ROI, but not outside. In another embodiment, different reconstruction curves may be applied inside and outside the region of interest.
[0148] The use of ROIs is motivated by the need to balance bitrate and image quality. For example, consider a video sequence of a sunset. The top half of the image might have a sun over a sky of relatively uniform color (thus, the pixels of the sky background might have very low variance). In contrast, the bottom half of the image might depict moving waves. From the viewer's perspective, the top might be considered much more important than the bottom. On the other hand, the moving waves are more difficult to compress and require more bits per pixel due to the higher variance of their pixels. However, one might want to allocate more bits to the sun than to the waves. In this case, the top half can be represented as the region of interest.
[0149] ROI description Today, most codecs (e.g., AVC, HEVC, etc.) are block-based. To simplify implementation, regions can be specified in blocks. Using HEVC as an example and not by way of limitation, regions can be defined as multiples of coding units (CUs) or coding tree units (CTUs). A single ROI or multiples of ROIs can be specified. Multiple ROIs can be separate or overlapping. ROIs do not have to be rectangular. ROI syntax can be provided at any level of interest, such as the slice level, the video level, or the video stream level. In one embodiment, the ROI is first specified in the sequence parameter set (SPS). In that case, slight modifications of the ROI can be allowed in the slice header. Table 16 shows an example of syntax in which a single ROI is specified as multiple CTUs in a rectangular region. Table 17 describes the modified ROI syntax at the slice level.
[0150] [Table 22]
[0151] [Table 23]
[0152] sps_reshaper_active_ROI_flag equal to 1 specifies that the ROI is present in the Coded Video Sequence (CVS). sps_reshaper_active_ROI_flag equal to 0 specifies that the ROI is not present in the CVS. reshaper_active_ROI_in_CTUsize_left, reshaper_active_ROI_in_CTUsize_right, reshaper_active_ROI_in_CTUsize_top, and reshaper_active_ROI_in_CTUsize_bottom each specify a sample of the image within the ROI by a rectangular area specified in image coordinates equal to offset*CTUsize for left and top, and offset*CTUsize-1 for right and bottom. reshape_model_ROI_modification_flag equal to 1 specifies that the ROI is modified in the current slice. reshape_model_ROI_modification_flag equal to 0 specifies that the ROI is not modified in the current slice. reshaper_ROI_mod_offset_left, reshaper_ROI_mod_offset_right, reshaper_ROI_mod_offset_top, and reshaper_ROI_mod_offset_bottom specify the left / right / top / bottom offset values from reshaper_active_ROI_in_CTUsize_left, reshaper_active_ROI_in_CTUsize_right, reshaper_active_ROI_in_CTUsize_top, and reshaper_active_ROI_in_CTUsize_bottom, respectively.
[0153] In the case of multiple ROIs, the example syntax in Tables 16 and 17 for a single ROI can be extended with an index (or ID) for each ROI, similar to the method used in HEVC to define multiple pan-scan rectangles using SEI messaging (see HEVC Specification
[11] , Section D.2.4).
[0154] ROI processing in intra-only in-loop reconstruction In the case of intra-only reconstruction, the ROI portion of the image is first reconstructed, and then coding is applied. Because the reconstruction is only applied to the ROI, the boundary between the ROI and non-ROI portions of the image may be visible. The loop filter (e.g., 270 in FIG. 2C or FIG. 2D) can cross the boundary, so loop filter optimization (ILF) is required. In the case of OPT, special attention must be paid to the ROI. In one embodiment, it is proposed that the loop filter be applied only when the entire decoded image is in the same region. That is, the entire image is either entirely in the reconstructed domain or entirely in the non-reconstructed domain. In one embodiment, at the decoder side, if loop filtering is applied to the non-reconstructed domain, backward reconstruction should be applied to the ROI section of the decoded image first, followed by the loop filter. The decoded image is then stored in the DPB. In another embodiment, if loop filtering is applied to the reconstructed domain, reconstruction should be applied to the non-ROI portion of the decoded image first, followed by the loop filter, and then backward reconstruction of the entire image. The decoded image is then stored in the DPB. In yet another embodiment, if loop filtering is applied to the reconstructed domain, backward reconstruction can be applied to the ROI portion of the decoded image first, followed by reconstructing the entire image, followed by applying the loop filter, and then backward reconstruction of the entire image. The decoded image is then stored in the DPB. These three approaches are summarized in Table 18. From a computational standpoint, method "A" is simpler. In one embodiment, ROI activation can be used to specify the execution order of backward reconstruction relative to loop filtering (LF). For example, if ROI is actively used (e.g., SPS syntax flag = true), LF (block 270 in Figures 2C and 2D) is executed after backward reconstruction (block 265 in Figures 2C and 2D). If ROI is not actively used, LF is executed before backward reconstruction.
[0155] [Table 24]
[0156] ROI processing in in-loop prediction residual reconstruction. For an in-loop (prediction) residual reconstruction architecture (see, for example, 200C_D in FIG. 2F), at the decoder, using equation (3), the process can be expressed as follows: If(currentCTU belongs to ROI) Reco_sample=lnv(Res_d+Fwd(Pred_sample)), (see formula (3)) else Reco_sample=Res_d+Pred_sample end
[0157] ROI and Encoder Considerations For each CTU, the encoder needs to check whether the CTU belongs to the ROI. For example, for in-loop prediction residual reconstruction, a simple check based on equation (3) can be done as follows: If(currentCTU belongs to ROI) Apply weighted distortion in RDO of luminance. The weights are derived based on Equation (10) else Applying unweighted distortion in luminance RDO end
[0158] An example encoding workflow that takes the ROI into account during reconstruction may include the following steps. -For intra-picture: - Apply forward reconstruction to the ROI area of the original image -Encode intraframes - Apply backward reconstruction to the ROI area of the restored video before the loop filter (LF) - Perform loop filtering in the unreconstructed domain (see, for example, method "C" in Table 18) as follows: The steps include: Apply forward reconstruction to the non-ROI area of the source image (to reconstruct the entire source image to obtain the loop filter criteria) Apply forward reconstruction to the entire image area of the restored image Derive loop filter parameters and apply loop filtering Apply backward reconstruction to the entire image area of the restored image and store it in the DPB.
[0159] On the encoder side, the LF needs to have an uncompressed reference image for filter parameter estimation, so the handling of the LF reference for each method is as shown in Table 19.
[0160] [Table 25]
[0161] -For Interpicture: When encoding interframes, for each CU within the ROI, apply prediction residual reconstruction and weighted distortion to luminance; for each CU outside the ROI, do not apply reconstruction. -Loop filtering optimization (option 1) is performed as usual (as if no ROI is used): - Forward reconstruction of the entire image area of the original image - Forward reconstruction of the entire image area of the restored image Derive loop filter parameters and apply loop filtering Apply backward reconstruction to the entire image area of the restored image and store it in the DPB.
[0162] Reconstruction of HLG encoded content The term HybridLog-Gamma, or HLG, refers to another transfer function defined in Rec.BT.2100 that maps high dynamic range signals. HLG was developed to maintain backward compatibility with conventional standard dynamic range signals coded using conventional gamma functions. Comparing the codeword distribution between PQ and HLG coded content, PQ mapping tends to allocate more codewords to dark and light areas, while HLG content codewords tend to be distributed more evenly across the image. The majority tends to be distributed in the middle range. Two approaches can be used for HLG luminance reconstruction. In one embodiment, the HLG content can simply be converted to PQ content, and then all of the reconstruction techniques related to PQ mentioned above can be applied. For example, the following steps can be applied: 1) Map HLG luminance (e.g., Y) to PQ luminance. Let us denote the function or LUT for this conversion as HLG2PQLUT(Y). 2) Analyze the PQ luminance values and derive a PQ-based forward reconstruction function or LUT, which we denote as PQAdpFLUT(Y). 3) Combine these two functions or LUTs into a single function or LUT, i.e., HLGAdpFLUT[i]=PQAdpFLUT[HLG2PQLUT[i]].
[0163] Since the HLG codeword distribution is significantly different from the PQ codeword distribution, such an approach may produce suboptimal reconstruction results. In another embodiment, the HLG reconstruction function is derived directly from the HLG samples. The same framework used for PQ signals can be applied, but the CW_Bins_Dft tables may be modified to reflect the characteristics of the HLG signal. In one embodiment, using a midtone profile for the HLG signal, several CW_Bins_Dft tables are designed according to user preferences. For example, when it is preferred to preserve highlights, for alpha=1.4, g_DftHLGCWBin0 = [ 8, 14, 17, 19, 21, 23, 24, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 36, 37, 38, 39, 39, 40, 41, 41, 42, 43, 43, 44, 44, 30 ]. When it is desired to preserve the midtones (or mid-range), g_DftHLGCWBin1= [ 12, 16, 16, 20, 24, 28, 32, 32, 32, 32, 36, 36, 40, 44, 48, 52, 56, 52, 48, 44, 40, 36, 36, 32, 32, 32, 26, 26, 20, 16, 16, 12 ]. When it is desirable to maintain skin tone, g_DftHLGCWBin2= [12, 16, 16, 24, 28, 32, 56, 64, 64, 64, 64, 56, 48, 40, 32, 32, 32, 32, 32, 32, 28, 28, 24, 24, 20, 20, 20, 20, 20, 16, 16, 12]; is.
[0164] From a bitstream syntax point of view, to distinguish between PQ-based reconstruction and HLG-based reconstruction, a new parameter, denoted as sps_reshaper_signal_type, is added, where the sps_reshaper_signal_type value indicates the reconstructed signal type (e.g., 0 is a gamma-based SDR signal, 1 is a PQ-coded signal, and 2 is an HLG-coded signal).
[0165] Examples of syntax tables for HDR reconstruction in SPS and slice headers for both PQ and HLG with all the features described above (e.g., ROI, in-loop filter optimization (ILFOPT), and ChromaDQPA adjustment) are shown in Tables 20 and 21.
[0166] [Table 26]
[0167] sps_in_loop_filter_opt_flag equal to 1 specifies that the in-loop filter optimization is performed in the reconstruction domain in the Coded Video Sequence (CVS). sps_in_loop_filter_opt_flag equal to 0 specifies that the in-loop filter optimization is performed within the non-reconstruction region in CVS. sps_luma_based_chroma_qp_offset_flag equal to 1 specifies that a luma-based chroma QP offset is derived (e.g., according to Table 11 or 12) and applied to the chroma coding of each CU in the Coded Video Sequence (CVS). sps_luma_based_chroma_qp_offset_flag equal to 0 specifies that a luma-based chroma QP offset is not enabled in the CVS.
[0168] [Table 27]
[0169] Improved saturation quality Proponents of HLG-based encoding claim that it offers better backward compatibility with SDR signaling. Therefore, theoretically, HLG-based signals can use the same encoding settings as legacy SDR signals. However, when viewing HLG-encoded signals in HDR mode, some color artifacts may still be observed, especially in achromatic areas (such as white and gray). In one embodiment, such artifacts can be reduced by adjusting the chromaQPOffset value during encoding. It is proposed to apply a more conservative chromaQP adjustment for HLG content compared to that used when encoding PQ signals. For example, reference
[10] describes the following model for assigning Cb and Cr QP offsets based on the luma QP and a factor based on the capture and display primaries:
number
[0170] In one embodiment, it is proposed to use the same model but with different parameters that do not actively change chromaQPOffset. By way of example and not limitation, in one embodiment, for Cb in equation (18a), cb = 1, k = -0.2, and l = 7, and for Cr in equation (18b), cr = 1, k = -0.2, and l = 7. Figures 6A and 6B show examples of how the chromaQPOffset values vary according to the luma quantization parameter (QP) of PQ (Rec. 709), and that the HLG.PQ-related values vary more dramatically than the HLG-related values. Figure 6A corresponds to Cb (Equation (18a)), while Figure 6B corresponds to Cr (Equation (18b)).
[0171] References Each of the references cited herein is incorporated herein in its entirety by reference. [1] International application PCT / US201 filed by GM. Su on March 30, 2016 No. 6 / 025082, "In-Loop Block-Based Image Reshaping in High Dynamic Range Video Coding" (also published as WO 2016 / 164235) [2] D. Baylon, Z. Gu, A. Luthra, K. Minoo, P. Yin, F. Pu, T. Lu, T. Chen, W. Husak, Y. He, L. Kerofsky, Y. Ye, B. Yi, “Response to Call for Evidence for HDR and WCG Video Coding: Arris, Dolby and InterDigital,” Doc. m36264, July (2015), Warsaw, Poland. [3] U.S. Patent Application No. 15 / 410, filed January 19, 2017, by T. Lu et al. [4] P. Yin et al., International Application No. PCT / US2016 / 042229, "Signal Reshaping and Coding for HDR and Wide Color Gamut Signals," filed July 14, 2016 (also published as WO 2017 / 011636). [5] K. Minoo et al. “Exploratory Test Model for HDR extension of HEVC” MPEG output document, JCTVC-W0092 (m37732), 2016, San Diego, USA [6] E. Francois, J. Sole, J. Stroem, P. Yin “Common Test Conditions for HDR / WCG video coding experiments” JCTVC doc. Z1020, Geneva, Jan. 2017 [7] A. Segall, E. Francois, and D. Rusanovskyy, “JVET common test conditions and evaluation procedures for HDR / WCG Video” (JVET-E1020, ITU-T meeting, Geneva, January 2017) [8] JEM6.0 Software: https: / / jvet.hhi.fraunhofer.de / svn / svn HMJEMSoftware / tags / HM-16.6-JEM-6.0 [9] U.S. Provisional Patent Application No. 62 / 4 filed October 11, 2016 by T. Lu et al. No. 06,483, "Adaptive Chroma Quantization in Video Coding for Multiple Color Imaging Formats" (also filed as U.S. Patent Application No. 15 / 728,939, U.S. Patent No. (Published as Patent Application Publication No. 2018 / 0103253)
[10] J. Samuelsson et al. (Eds) "Conversion and coding practices for HDR / WCG Y'CbCr 4:2:0 Video with PQ Transfer Characteristics" JCTVC-Y1017, ITU-T / ISO meeting, Chengdu, Oct. 2016
[11] ITU-T H.265 “High efficiency video coding” ITU, version 4.0, (12 / 2016)
[0172] Exemplary Computer System Implementation Embodiments of the present invention may be implemented using computer systems, systems comprised of electronic circuitry and components, integrated circuit (IC) devices such as microcontrollers, field programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete time or digital signal processors (DSPs), application specific ICs (ASICs), and / or apparatuses comprising one or more of such systems, devices, or components. These computers and / or ICs may be used to perform integrated image signal reconstruction and processing as described herein. The image and video embodiments may be implemented in hardware, software, firmware, or various combinations thereof.
[0173] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors in a display, encoder, set-top box, transcoder, etc., can implement the above-described methods related to integrated image signal reconstruction and encoding by executing software instructions in a program memory accessible to the processor. The present invention can also be provided in the form of a program product. The program product can include any non-transitory medium that carries a set of computer-readable signals that include instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. Program products according to the present invention can be in any of a wide variety of forms. Program products can include physical media, such as magnetic data storage media including floppy diskettes and hard disk drives, optical data storage media including CD-ROMs and DVDs, ROMs, and electronic data storage media including flash RAM. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0174] In the above, when a component (e.g., a software module, a processor, an assembly, a device, a circuit, etc.) is referred to, unless otherwise indicated, the reference to that component (including the reference to "means") is to be interpreted as including, as equivalents of that component, any component that performs the function of the described component (i.e., functionally equivalent components), including components that are not structurally equivalent to the disclosed structures that perform that function in the illustrated exemplary embodiments of the invention.
[0175] [Equivalents, extensions, alternatives and others] Exemplary embodiments of integrated, efficient image signal reconstruction and coding have been described above. In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the sole and exclusive indication of what is, and what is intended by the applicant to be, the invention is the following set of claims in their particular form as derived from this application. These claims include any and all subsequent amendments. Any definitions expressly set forth herein for terms contained in such claims shall determine the meaning of those terms used in the claims. Accordingly, no limitation, element, property, feature, advantage, or attribute not expressly recited in a claim should in any way limit the scope of such claims. The specification and drawings, therefore, are to be regarded in an illustrative and not a restrictive sense.
Claims
1. 1. A computer program comprising computer instructions for encoding an image, the computer instructions, when executed by a processor, causing the processor to: accessing an input image in input codeword representation; applying a forward reconstruction function to the input image to map it to a reconstructed codeword representation; generating an inverse reconstruction function based on parameters of the forward reconstruction function, the inverse reconstruction function mapping pixel values from the reconstructed codeword representation to the input codeword representation; generating an encoded pixel region of the input image based on an input pixel region in the input image, the forward reconstruction function, and the inverse reconstruction function; generating reconstruction metadata characterizing the forward reconstruction function based on a piecewise linear representation; generating an output bitstream based on the encoded pixel regions of the input image and the reconstruction metadata; A computer program that executes the following:
2. The computer program product of claim 1 , wherein generating the coded pixel region of the input image comprises applying in-loop reconstruction to the input image.
3. The step of generating an encoded pixel region for an input pixel region in the input image comprises: calculating a prediction region based on pixel data in a reference frame buffer or a pre-encoded neighboring space; generating a reconstructed residual region based on the input pixel region, the prediction region, and the forward reconstruction function, wherein reconstructed residual samples in the reconstructed residual region are obtained based, at least in part, on forward reconstruction of respective prediction samples in the prediction region; generating a quantized residual domain based on the reconstructed residual domain; generating a dequantized residual region based on the quantized residual region; generating a reconstructed pixel region based on the dequantized residual region, the prediction region, the forward reconstruction function, and the backward reconstruction function; generating a reference pixel region based on the reconstructed pixel region, the reference pixel region being stored in a reference frame buffer; 2. The computer program of claim 1, comprising:
4. The step of generating the quantized residual domain comprises: applying a forward-coding transform to the reconstructed residual domain to generate transformed data; applying a forward-coding quantizer to the transformed data to generate quantized data; 4. The computer program of claim 3, comprising:
5. The step of generating a dequantized residual domain comprises: applying a de-encoding quantizer to the quantized data to generate de-quantized data; applying an inverse coding transform to the dequantized data to generate the dequantized residual domain; 5. The computer program of claim 4, comprising:
6. The computer program product of claim 3 , wherein generating the reference pixel region to be stored in the reference frame buffer comprises applying a loop filter to the reconstructed pixel region.
7. 1. A computer program comprising computer instructions for decoding a bitstream, the computer instructions, when executed by a processor, causing the processor to: receiving an encoded image encoded with the reconstructed codeword representation; receiving reconstruction metadata of the encoded image; generating parameters of a forward reconstruction function based on the reconstruction metadata, the forward reconstruction function mapping pixels from a first codeword representation to a reconstructed codeword representation; generating parameters of an inverse reconstruction function based on the reconstruction metadata, the inverse reconstruction function mapping pixels from the reconstructed codeword representation to the first codeword representation; decoding a coded region of the coded image based on the forward reconstruction function and the inverse reconstruction function to generate an output pixel region of an output image of the first codeword representation; A computer program that executes
8. The computer program product of claim 7 , wherein decoding the coded regions of the coded image is based on in-loop reconstruction.
9. Decoding the encoded region includes: generating a decoded reconstructed residual domain; generating a prediction region based on pixels in a reference pixel buffer or a pre-decoded neighborhood; generating a reconstructed pixel region based on the decoded reconstructed residual region, the prediction region, the forward reconstruction function, and the backward reconstruction function, wherein reconstructed samples in the reconstructed pixel region are obtained at least in part based on forward reconstruction of respective prediction samples in the prediction region; generating an output pixel area for an output image based on the reconstructed pixel area; storing the output pixel region in the reference pixel buffer; The computer program of claim 7 further comprising:
10. the metadata characterizes the forward reconstruction function based on a piecewise linear representation; for each piece of the piecewise linear representation of the forward reconstruction function, the reconstruction metadata includes a delta magnitude and a sign of the delta magnitude; 8. A computer program according to claim 7.