Picture coding method of a coding device, storage medium, and data transmission method
By deriving the quantization parameter offset to optimize the quantization process, image compression efficiency is improved, the transmission and storage cost issues of high-resolution and high-quality images are solved, and efficient image data processing is achieved.
Patent Information
- Application Number
- CN202310308727.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-01
- Filing Date
- 2019-03-05
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2039-03-05
AI Technical Summary
Existing technologies incur high transmission and storage costs when processing high-resolution and high-quality images, necessitating improvements in image compression and quantization efficiency.
The quantization process is optimized by deriving the quantization parameter offset (QP offset), and combined with entropy decoding and inverse quantization modules, an efficient image encoding and decoding method is generated.
It improves image compression and quantization efficiency, enabling efficient image data transmission and storage.
Smart Images

Figure CN116320456B_ABST
Abstract
Description
[0001] The application is a divisional application of the invention patent application No. 201980006064.3 (International Application No. PCT / KR2019 / 002520, filed on March 5, 2019, entitled "Image encoding apparatus based on quantization parameter derivation and method therefor"). TECHNICAL FIELD
[0002] The present invention relates to an image encoding technology. More particularly, the present invention relates to an image encoding apparatus based on quantization parameter derivation and method therefor in an image encoding system. BACKGROUND
[0003] Recently, demand for high-resolution, high-quality images such as high definition (HD) images and ultra-high definition (UHD) images has increased in various fields. As image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to existing image data. Accordingly, when transmitting image data using a medium such as a wired / wireless broadband line, or when storing, transmission and storage costs can increase.
[0004] Therefore, an efficient image compression technology is needed for efficiently transmitting, storing, and reproducing information of high-resolution and high-quality images. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] The present invention provides a method and apparatus for enhancing video encoding efficiency.
[0007] The present invention also provides a method and apparatus for increasing quantization efficiency.
[0008] The present invention also provides a method and apparatus for efficiently deriving a quantization parameter.
[0009] TECHNICAL SOLUTION
[0010] According to an embodiment of the present invention, a picture decoding method performed by a decoding apparatus is provided. The method includes decoding image information including information on a quantization parameter (QP), deriving an expected average luma value of a current block from neighboring available samples, deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and the information on the QP, deriving the luma QP based on the QP offset, performing inverse quantization on a quantization group including the current block based on the derived luma QP, generating a residual sample of the current block based on the inverse quantization, generating a prediction sample of the current block based on the image information, and generating a reconstructed sample of the current block based on the residual sample of the current block and the prediction sample of the current block.
[0011] According to an embodiment of the present application, a decoding device for decoding a picture is provided. The decoding device includes: an entropy decoding module configured to decode image information including information about a quantization parameter (QP); an inverse quantization module configured to derive an expected average luma value of a current block from neighboring available samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and the information about the QP, derive the luma QP based on the QP offset, and perform inverse quantization on a quantization group including the current block based on the derived luma QP; an inverse transform module configured to generate residual samples of the current block based on the inverse quantization; a prediction module configured to generate prediction samples of the current block based on the image information; and a reconstruction module configured to generate reconstructed samples of the current block based on the residual samples of the current block and the prediction samples of the current block.
[0012] According to an embodiment of the present application, a picture encoding method performed by an encoding device is provided. The method includes: deriving an expected average luma value of a current block from neighboring available samples; deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and information about the QP; deriving the luma QP based on the QP offset; performing quantization on a quantization group including the current block based on the derived luma QP; and encoding image information including the information about the QP.
[0013] According to an embodiment of the present application, an encoding device for encoding a picture is provided. The encoding device includes: a quantization module configured to derive an expected average luma value of a current block from neighboring available samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and information about the QP, derive the luma QP based on the QP offset, and perform quantization on a quantization group including the current block based on the derived luma QP; and an entropy encoding module configured to encode image information including the information about the QP.
[0014] Advantageous Effects
[0015] According to the present application, overall image / video compression efficiency can be increased.
[0016] According to the present application, quantization efficiency can be increased.
[0017] According to the present application, a quantization parameter can be efficiently derived. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 FIG. 1 is a schematic diagram illustrating a configuration of an encoding device according to an embodiment.
[0019] Figure 2is a schematic diagram showing a configuration of a decoding apparatus according to an embodiment.
[0020] Figure 3 shows an example of a chromaticity diagram.
[0021] Figure 4 shows an example of mapping of linear light values for SDR and HDR representations.
[0022] Figure 5 is a flowchart showing a process of reconstructing a picture according to an embodiment.
[0023] Figure 6 is a flowchart showing a process of reconstructing a picture according to another embodiment.
[0024] Figure 7 is a flowchart showing an operation of an encoding apparatus according to an embodiment.
[0025] Figure 8 is a block diagram showing a configuration of an encoding apparatus according to an embodiment.
[0026] Figure 9 is a flowchart showing an operation of a decoding apparatus according to an embodiment.
[0027] Figure 10 is a block diagram showing a configuration of a decoding apparatus according to an embodiment. DETAILED DESCRIPTION
[0028] According to an embodiment of the present application, there is provided a picture decoding method performed by a decoding apparatus. The method includes decoding image information including information on a quantization parameter (QP), deriving an expected average luminance value of a current block from neighboring available samples, deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luminance value and the information on the QP, deriving the luma QP based on the QP offset, performing inverse quantization on a quantization group including the current block based on the derived luma QP, generating residual samples of the current block based on the inverse quantization, generating prediction samples of the current block based on the image information, and generating reconstructed samples of the current block based on the residual samples of the current block and the prediction samples of the current block.
[0029] MODES OF THE INVENTION
[0030] The present application can be modified in various forms, and specific embodiments thereof will be described and illustrated in the drawings. However, these embodiments are not intended to limit the present application. The terms used in the following description are merely used to describe specific embodiments and are not intended to limit the present application. Singular expressions include plural expressions as long as they are clearly different from the context. Terms such as "include" and "have" are intended to indicate that there is a feature, number, step, operation, element, component, or a combination thereof described in the following description, and it should be understood that the possibility of adding one or more different features, numbers, steps, operations, elements, components, or a combination thereof is not excluded.
[0031] On the other hand, the elements in the drawings described in the present application are independently drawn for the convenience of illustrating different specific functions in the image encoding / decoding apparatus, and are not intended to mean that the elements are embodied by independent hardware or independent software. For example, two or more of the elements can be combined to form a single element, or one element can be divided into a plurality of elements. Embodiments in which the elements are combined and / or divided belong to the present application without departing from the concept of the present application.
[0032] Hereinafter, exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings. In addition, the same reference numbers will be used throughout these drawings to refer to the same elements and a repeated description on the same elements will be omitted.
[0033] The following description can be applied to the technical field of processing video or image. For example, the method or embodiments disclosed in the following description can be applied to various video encoding standards, such as the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), the next generation video / image encoding standard after VVC, or the previous generation video / image encoding standard before VVC (for example, the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), etc.).
[0034] In this specification, video can mean a set of images according to time. A picture generally refers to a unit representing one image in a specific time period, and a slice is a unit constituting a part of a picture in encoding. One picture can be composed of a plurality of slices, and if necessary, a picture and a slice can be used in combination. In addition, in some cases, the term "image" can mean a concept including still images and video (a set of still images according to time lapse). In addition, "video" does not necessarily mean only a set of still images according to time, but can be interpreted as a concept including the meaning of still images in some embodiments.
[0035] A pixel or pel can mean the smallest unit of a picture (or image). In addition, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chroma component.
[0036] A unit represents a basic unit of image processing. The unit can include at least one of a specific region of a picture and information related to the region. The unit can be used in combination with terms such as a block or a region. In general, an M x N block can represent a set of samples or transform coefficients consisting of M columns and N rows.
[0037] Figure 1 The configuration of an encoding apparatus according to an embodiment is schematically illustrated.
[0038] Hereinafter, an encoding / decoding apparatus can include a video encoding / decoding apparatus and / or an image encoding / decoding apparatus. The video encoding / decoding apparatus can be used as a concept including the image encoding / decoding apparatus, and the image encoding / decoding apparatus can be used as a concept including the video encoding / decoding apparatus.
[0039] Referring to Figure 1 The encoding apparatus 100 can include a picture partitioning module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filtering module 150, and a memory 160. The residual processing module (residual processing unit) 120 can include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, an inverse quantization module 125, and an inverse transform module 126.
[0040] The picture partitioning module 105 can partition an input picture into at least one processing unit.
[0041] In one example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a largest coding unit (LCU) according to a quadtree binary tree (QTBT) structure. For example, one coding unit can be divided into a plurality of coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure.
[0042] In this case, for example, the quadtree structure is applied first, and the binary tree structure and the ternary tree structure can be applied later. Alternatively, the binary tree structure / ternary tree structure can be applied first. The encoding process according to the present application can be performed based on a final coding unit that is not further divided. In this case, the largest coding unit can be directly used as the final coding unit based on coding efficiency, etc., according to image characteristics, or the coding unit can be recursively divided into a lower depth coding unit and can be used as the final coding unit. Here, the encoding process can include processes such as prediction, conversion, and restoration (to be described later).
[0043] As another example, the processing unit can include a coding unit (CU) prediction module (PU) or a transform unit (TU). The coding unit can be split from a largest coding unit (LCU) along a quad-tree structure into coding units of a deeper depth. In this case, the largest coding unit can be directly used as a final coding unit based on coding efficiency or the like according to image characteristics, or the coding unit can be recursively divided into coding units of a lower depth and can be used as a final coding unit. When a minimum coding unit (SCU) is set, the coding unit can not be divided into coding units smaller than the minimum coding unit.
[0044] Herein, the term "final coding unit" means a coding unit that partitions or divides a prediction module or a transform unit. The prediction module is a unit that is partitioned from the coding unit, and can be a unit for which sample prediction is performed. At this time, the prediction module can be divided into sub-blocks. The transform unit can be divided from the coding unit along a quad-tree structure, and can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0045] Hereinafter, the coding unit can be referred to as a coding block (CB), the prediction module can be referred to as a prediction block (PB), and the transform unit can be referred to as a transform block (TB). The prediction block or the prediction module can refer to a specific region in the form of a block in a picture, and can include an array of predicted samples. In addition, the transform block or the transform unit can refer to a specific region in the form of a block within a picture, and can include an array of transform coefficients or residual samples.
[0046] The prediction module 110 predicts a current block or a residual block, and generates a block including a prediction of a predicted sample of the current block. The unit of prediction performed in the prediction module 110 can be a coding block, a transform block, or a prediction block.
[0047] The prediction module 110 predicts a current block or a residual block, and generates a block including a prediction of a predicted sample of the current block. The unit of prediction performed in the prediction module 110 can be a coding block, a transform block, or a prediction block.
[0048] The prediction module 110 can determine whether to apply intra prediction or inter prediction to a current block. For example, the prediction module 110 can determine whether to apply intra prediction or inter prediction in units of a CU.
[0049] In the case of intra prediction, the prediction module 110 can derive a predicted sample of the current block based on a reference sample outside the current block in a picture to which the current block belongs (hereinafter referred to as a current picture).
[0050] In this case, the prediction module 110 can derive a predicted sample based on an average or interpolation of neighboring reference samples of the current block (case (i)), or can derive a predicted sample based on a reference sample existing in a certain (prediction) direction with respect to the predicted sample among the samples (case (ii)).
[0051] Case (i) can be referred to as a non-directional mode or a non-angular mode, and case (ii) can be referred to as a directional mode or an angular mode. In intra prediction, the prediction modes can have, for example, 65 directional prediction modes and at least two non-directional modes. The non-directional modes can include a DC prediction mode and a Planar mode. The prediction module 110 can determine a prediction mode applied to a current block using a prediction mode applied to a neighboring block.
[0052] In case of inter prediction, the prediction module 110 can derive prediction samples of a current block based on samples on a reference picture specified by a motion vector. The prediction module 110 can derive the prediction samples of the current block by applying one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the skip mode and the merge mode, the prediction module 110 can use motion information of a neighboring block as motion information of the current block.
[0053] In the skip mode, unlike the merge mode, a difference (residual) between predicted samples and original samples is not transmitted. In the MVP mode, a motion vector of the current block can be derived by using a motion vector of a neighboring block as a motion vector predictor to be used as a motion vector predictor of the current block.
[0054] In case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in a current picture and temporal neighboring blocks existing in a reference picture. The reference picture including the temporal neighboring blocks can be referred to as a colPic. The motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.
[0055] When using motion information of a temporal neighboring block in the skip mode and the merge mode, a highest picture on a reference picture list can be used as a reference picture. The reference pictures included in a picture order count (POC) can be ordered based on a POC difference between a current picture and a corresponding reference picture. The POC corresponds to a display order of a picture and can be distinguished from an encoding order.
[0056] The subtractor 121 generates residual samples that are differences between original samples and predicted samples. When the skip mode is applied, the residual samples can not be generated as described above.
[0057] The transform module 122 transforms the residual samples based on a transform block to generate transform coefficients. The transform module (transform unit) 122 can perform a transform according to a size of the transform block and a prediction mode applied to an encoding block or a prediction block spatially overlapping the transform block.
[0058] For example, if intra prediction is applied to a coding block or a prediction block overlapping with a transform block and the transform block is a 4x4 residual array, the residual samples are transformed into a discrete sine transform (DST). In other cases, a DCT (discrete cosine transform) transform core can be used to transform the residual samples.
[0059] The quantization module (quantization unit) 123 can quantize the transform coefficients to generate quantized transform coefficients.
[0060] The rearrangement module 124 rearranges the quantized transform coefficients. The rearrangement module 124 can rearrange the block-shaped quantized transform coefficients into a one-dimensional vector form through a scanning method of the coefficients. The rearrangement module 124 can be a part of the quantization module 123, but is alternatively configured to describe the rearrangement module 124.
[0061] The entropy encoding module 130 can perform entropy encoding on the quantized transform coefficients. For example, the entropy encoding can include an encoding method such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC). The entropy encoding module 130 can encode information (e.g., values of syntax elements, etc.) required for video restoration together with or separately from the quantized transform coefficients according to the entropy encoding or a predetermined method.
[0062] The encoded information can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The bitstream can be transmitted via a network or stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0063] The inverse quantization module 125 inverse-quantizes the quantized values (quantized transform coefficients) obtained from the quantization module 123, and the inverse transform module 126 inverse-quantizes the inverse-quantized values obtained from the inverse quantization module 125 to generate residual samples.
[0064] The adder 140 combines the residual samples and the predicted samples to reconstruct a picture. The residual samples and the predicted samples are added in units of blocks, and thus a reconstructed block can be generated. Here, the adder 140 can be a part of the prediction module 110, and in addition, the adder 140 can be referred to as a reconstruction module or a reconstructed block generation unit.
[0065] For the reconstructed picture, the filtering module 150 can apply a deblocking filter and / or a sample adaptive offset. With deblocking filtering and / or a sample adaptive offset, artifacts in block boundaries or distortion in quantization processing in the reconstructed picture can be corrected. The sample adaptive offset can be applied on a sample-by-sample basis, and can be applied after the process of deblocking filtering is completed. The filtering module 150 can apply an ALF (adaptive loop filter) to the recovered picture. The ALF can be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset are applied.
[0066] The memory 160 can store the recovered picture (decoded picture) or information required for encoding / decoding. Here, the reconstructed picture can be a reconstructed picture for which the filtering module 150 has completed the filtering process. The stored recovered picture can be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 can store a (reference) picture for inter prediction. At this time, the picture for inter prediction can be designated by a reference picture set or a reference picture list.
[0067] Figure 2 The configuration of the decoding apparatus according to an embodiment is schematically illustrated.
[0068] Referring to Figure 2 The decoding apparatus 200 can include an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filtering module 250, and a memory 260. Here, the residual processing module 220 can include a rearrangement module 221, an inverse quantization module 222, and an inverse transform module 223. In addition, although not shown, the video decoding apparatus 200 can include a receiver for receiving a bitstream including video information. The receiving unit can be a separate module or can be included in the entropy decoding module 210.
[0069] When a bitstream including video / image information is input, the (video) decoding apparatus 200 can recover video / image / picture in response to the processing of processing video / image information in the (video) encoding apparatus 100.
[0070] For example, the video decoding apparatus 200 can perform video decoding using the processing units applied in the video encoding apparatus. Accordingly, the processing unit block for video decoding can be, for example, a coding unit, in another example, a coding unit, a prediction module, or a transform unit. The coding unit can be divided from a largest coding unit along a quad tree structure, a binary tree structure, and / or a ternary tree structure.
[0071] The prediction module and the transform unit can be further used according to circumstances. The prediction block can be a block derived or partitioned from the coding unit. The transform unit can be divided from the coding unit along a quad tree structure, and can be a unit of deriving a transform factor or a unit of deriving a residual signal from the transform factor.
[0072] The entropy decoding module 210 can parse a bitstream and output information required for video recovery or picture recovery. For example, the entropy decoding module 210 decodes information in the bitstream based on an encoding method such as exponential Golomb encoding, CAVLC, or CABAC, and calculates values of syntax elements required for video recovery and coefficient values regarding quantization of a residual.
[0073] More specifically, the CABAC entropy decoding method includes receiving bins corresponding to each syntax element in a bitstream, determining a context model based on a decoding target syntax element and decoding information of a decoding target, predicting a probability of occurrence of a bin according to the determined context model, and performing arithmetic decoding of the bin to generate a symbol corresponding to a value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using information of a decoded symbol / bin for a context model of a next symbol / bin.
[0074] Predicted information about decoded information in the entropy decoding module 210 is provided to the prediction module 230, and residual values on which entropy decoding is performed in the entropy decoding module 210 can be input to the rearrangement module 221.
[0075] The rearrangement module 221 can rearrange quantized transform coefficients in a two-dimensional block form. The rearrangement module 221 can perform rearrangement in response to coefficient scanning performed in an encoding apparatus. The rearrangement module 221 can be a part of the inverse quantization module 222, but is described as an alternative configuration.
[0076] The inverse quantization module 222 can inverse quantize quantized transform coefficients based on a (de)quantization parameter and output transform coefficients. At this time, information for deriving a quantization parameter can be signaled from an encoding apparatus.
[0077] The inverse transform module 223 can inverse transform transform coefficients to derive residual samples.
[0078] The prediction module 230 can predict a current block and can generate a predicted block including predicted samples of the current block. A unit of prediction performed by the prediction module 230 can be a coding block, a transform block, or a prediction block.
[0079] The prediction module 230 can determine whether to apply intra prediction or inter prediction based on the prediction information. In this case, the unit of determining whether to apply intra prediction or inter prediction can be different from the unit of generating the prediction samples. Also, the unit of generating the prediction samples in the inter prediction and the intra prediction can also be different. For example, it can be determined whether to apply the inter prediction or the intra prediction in units of CUs. Also, for example, in the inter prediction, the prediction mode can be determined in units of PUs to generate the prediction samples. In the intra prediction, the prediction mode can be determined in units of PUs, and the prediction samples can be generated in units of TUs.
[0080] In the case of the intra prediction, the prediction module 230 can derive the prediction samples of the current block based on the neighboring reference samples in the current picture. The prediction module 230 can apply a directional mode or a non-directional mode based on the neighboring reference samples of the current block to derive the prediction samples of the current block. In this case, the intra prediction mode of the neighboring block can be used to determine the prediction mode to be applied to the current block.
[0081] In the case of the inter prediction, the prediction module 230 can derive the prediction samples of the current block based on the samples on the reference picture specified by the motion vector on the reference picture. The prediction module 230 can derive the prediction samples of the current block by applying a skip mode, a merge mode, or an MVP mode. At this time, the motion information (e.g., information about the motion vector, the reference picture index, etc.) required for the inter prediction of the current block provided in the encoding apparatus can be acquired or derived based on the prediction information.
[0082] In the skip mode and the merge mode, the motion information of the neighboring block can be used as the motion information of the current block. In this case, the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0083] The prediction module 230 can construct a merge candidate list using the motion information of the available neighboring blocks, and use the information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding apparatus. The motion information can include the motion vector and the reference picture. When the motion information of the temporal neighboring block is used in the skip mode and the merge mode, the highest picture on the reference picture list can be used as the reference picture.
[0084] In the skip mode, unlike the merge mode, the difference (residual) between the predicted samples and the original samples is not transmitted.
[0085] In the MVP mode, the motion vector of the neighboring block can be used as the motion vector predictor to derive the motion vector of the current block. In this case, the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0086] For example, when the merge mode is applied, a motion vector of a reconstructed spatial neighboring block and / or a motion vector corresponding to the Col block (temporal neighboring block) can be used to generate a merge candidate list. In the merge mode, a motion vector of a candidate block selected in the merge candidate list is used as a motion vector of the current block. The information about the prediction can include a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. At this time, the prediction module 230 can derive the motion vector of the current block using the merge index.
[0087] As another example, when the motion vector prediction (MVP) mode is applied, a motion vector predictor candidate list is generated using a motion vector of a reconstructed spatial neighboring block and / or a motion vector corresponding to the Col block (temporal neighboring block). That is, the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the neighboring block Col can be used as a motion vector candidate. The information about the prediction can include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list.
[0088] At this time, the prediction module 230 can select a predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. The prediction unit of the encoding apparatus can obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can output the same as a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction module 230 can obtain the motion vector difference included in the information about the prediction, and derive the motion vector of the current block by addition of the motion vector difference and the motion vector predictor. The prediction module can also acquire or derive a reference picture index indicating a reference picture, etc. from the information about the prediction.
[0089] The adder 240 can add the residual samples and the predicted samples to reconstruct the current block or the current picture. The adder 240 can add the residual samples and the predicted samples block by block to reconstruct the current picture. When the skip mode is applied, since the residual is not transmitted, the predicted samples can be the restored samples. Here, the adder 240 is described as an alternative configuration, but the adder 240 can be a part of the prediction module 230. Also, the adder 240 can be referred to as a reconstruction module or a reconstructed block generation unit.
[0090] The filter module 250 can apply a sample adaptive offset of deblocking filtering and / or ALF to the reconstructed picture. At this time, the sample adaptive offset can be applied sample by sample, and can be applied after the deblocking filtering. The ALF can be applied after the deblocking filtering and / or the sample adaptive offset.
[0091] The memory 260 can store a reconstructed picture (decoded picture) or information required for decoding. Here, the reconstructed picture can be a reconstructed picture for which the filtering module 250 has completed a filtering process. For example, the memory 260 can store a picture used for inter prediction. At this time, the picture used for inter prediction can be designated by a reference picture set or a reference picture list. The reconstructed picture can be used as a reference picture for another picture. In addition, the memory 260 can output the reconstructed picture according to an output order.
[0092] Figure 3 An example of a chromaticity diagram is shown.
[0093] This embodiment relates to video coding, and more specifically, to techniques for optimizing video coding according to given conditions such as a defined or intended luminance transfer function, a dynamic range of a video, or luminance values of a coding block.
[0094] As used herein, the term luminance transfer function can refer to an optical-to-electrical transfer function (OETF) or an electrical-to-optical transfer function (EOTF). It should be noted that the optical-to-electrical transfer function can be referred to as an inverse electrical-to-optical transfer function, and the electrical-to-optical transfer function can be referred to as an inverse optical-to-electrical transfer function, even though the two transfer functions are not exactly inverses of each other.
[0095] The techniques described herein can be used to compensate for non-optimal video coding performance that occurs when the mapping of luminance values to digital code words is not considered with equal importance. For example, in practice, the OETF can allow more bits for dark regions than for bright regions (and vice versa). In this case, a video encoder / decoder designed based on the assumption that all digital code words are coded with equal importance typically cannot perform video coding in an optimal manner.
[0096] It should be noted that although the techniques of this disclosure are described with respect to the ITU-T H.264 standard and the ITU-T H.265 standard, the techniques of this disclosure are generally applicable to any video coding standard.
[0097] Video compression techniques have been deployed on a variety of devices, including digital televisions, desktop and laptop computers, tablet computers, digital recording devices, digital media players, video gaming devices, smartphones, and the like. Digital video can be coded according to a video coding standard, such as ITU-T H.264 and High-Efficiency Video Coding (HEVC). Video coding standards can allow a particular format (i.e., YUV 420) to be coded.
[0098] Conventional digital video cameras initially generate raw data corresponding to signals generated by their respective image sensors. For example, a digital image capture device records an image as a set of linearly related luminance values. However, human vision does not recognize changes in luminance values in a linear manner. That is, for example, an image region associated with a luminance value of 100 cd / m2 is not necessarily perceived as twice as bright as an image region associated with a luminance value of 200 cd / m2. Accordingly, a luminance transfer function (e.g., an optical-to-electrical transfer function (OETF) or an electrical-to-optical transfer function (EOTF)) can be used to convert linear luminance data to data that can be perceived in a meaningful manner. An optical-to-electrical transfer function (OETF) can map absolute linear luminance values in a non-linear manner to digital code words. The resulting digital code words can be converted to a video format supported by a video encoding standard.
[0099] Conventional video encoding / display systems, such as conventional television video distribution environments, have provided a standard dynamic range (SDR) that typically supports a luminance range of about 0.1 to 100 cd / m 2 (often referred to as "nits"). This range is significantly less than the range encountered in real life. For example, a light bulb can have a luminance of over 10000 cd / m 2 , a surface illuminated by sunlight can have a luminance of hundreds of thousands of cd / m 2 , and the night sky can be 0.005 cd / m 2 or less.
[0100] In recent years, LCD and OLED displays have been widely used, and the technology of these devices allows for higher luminance and wide color space reproduction. The achievable and desired luminance and dynamic range of various displays can be significantly different from those of conventional (SDR) capture and creation devices. For example, a content creation system can be able to create or capture content with a contrast ratio of 1000000: 1. Televisions and other video distribution environments are expected to give a viewing experience that is closer to a real life experience to provide a stronger "live" feeling to users. Instead of the existing SDR luminance range (which can be from 0.1 to 100 cd / m 2 ), a higher luminance range (from 0.005 to 10000 cd / m 2 ) can be considered. For example, when displaying HDR content on a display that supports a minimum luminance of 0.01 cd / m 2 and a maximum luminance of 2000 cd / m 2 , there can be a dynamic range of 200000: 1. It is noted that the dynamic range of a scene can be described as the ratio of the maximum light intensity to the minimum light intensity.
[0101] Additionally, Ultra High Definition Television (UHDTV) aims to provide a "realistic" feel to the user. Increasing resolution alone can not be sufficient to fully achieve this goal without additionally creating, capturing, and displaying content with higher peak luminance and greater contrast values than today's TVs. Additionally, greater realism requires rendering colors that are richer than those provided by the color gamut commonly used today (e.g., BT.709). Thus, new content will not only have several orders of magnitude greater luminance and contrast, but also a significantly wider color gamut (e.g., BT.2020 or even wider in the future). Figure 1 Various color gamut ranges are depicted.
[0102] Beyond this recent technical evolution extension, it is now possible to implement HDR (High Dynamic Range) image / video rendering for both production and consumer sides using appropriate transfer functions (OETF / EOTF).
[0103] Figure 4 An example of mapping of linear light values for SDR and HDR representations is shown.
[0104] A transfer function can be described as a mapping between input and output in the real (floating point) range [0.0, 1.0]. One example of a luminance transform function corresponding to HDR data includes the so-called SMPTE (Society of Motion Picture and Television Engineers) High Dynamic Range (HDR) transfer function, which can be referred to as SMPTE ST 2084. Another example of a luminance transform function corresponding to HDR data includes a hybrid logarithmic gamma transfer function for HDR signals (also referred to as ITU-R BT.2100). In particular, the SMPTE HDR transfer function includes an EOTF and an inverse EOTF. The SMPTE ST 2084 inverse EOTF is described according to the following set of mathematical expressions.
[0105] [mathematical expression 1]
[0106] L c = R / 10000
[0107] [mathematical expression 2]
[0108] V = ((c1 + c 2* L c n ) / (1 + c 3* L c n )) m
[0109] In Math Expression 1 and Math Expression 2, cl = c3 - c2 + 1 = 3424 / 4096 = 0.8359375, c2 = 32*2413 / 4096 = 18.8515625, c3 = 32*2392 / 4096 = 18.6875, m = 128*2523 / 4096 = 78.84375, and n = 0.25*2610 / 4096 = 0.1593017578125.
[0110] The SMPTE ST 2084 EOTF can be described according to the following set of mathematical expressions.
[0111] [mathematical expression 3]
[0112] L c = ((max[(V 1 / m -c1), 0]) / (c2 - c3*V 1 / m ))1 / n
[0113] [mathematical expression 4]
[0114] R = 10000*Lc
[0115] In the above equation, R is a luminance value having an expected range of 0 to 10000 cd / m 2 . For example, L c equal to 1 is intended to correspond to a luminance level of 10000 cd / m 2 . R can indicate an absolute linear luminance value. Further, in the above equation, V can be referred to as a luminance value (or a perceptual curve value). Since an OETF can map a perceptual curve value to a digital codeword, V can be mapped to a 2 N bit codeword. An example of a function that can be used to map V to an N bit codeword can be defined as:
[0116] [mathematical expression 5]
[0117] Digital value = INT((2 N -1)*V)
[0118] In Math Expression 5, INT(x) generates an integer value by rounding down decimal values less than 0.5 and rounding up decimal values greater than or equal to 0.5.
[0119] As an illustrative example, Figure 2 an 8 bit SDR system capable of representing 0.1 to 100 cd / m 2 with a BT.709 style transfer function (green curve) can be compared to an SDR system capable of representing 0.005 to 10000 cd / m 2The graph lines in this figure are approximate graph lines. They do not capture the exact form of the curve, but are shown for illustrative purposes only. In the graph, integer code levels are along the horizontal axis, and linear light values (scaled to log10) are along the vertical axis. This exemplary mapping includes a traditional code level range scale to accommodate both under-gamma (negative samples below the [0.0, 1.0] real value range) and over-gamma (samples above the real value 1.0). Due to design properties, the 10-bit HDR transfer function shown here assigns approximately twice as many code levels [119 to 509] as the traditional 8-bit SDR transfer function assigns [16 to 235] in the SDR range, while providing a similar number of new code levels [510 to 940] to extend luminance. For intensities below 0.01 cd / m 2 New code levels [64 to 118] are assigned for darker intensities.
[0120] In a sense, the 10-bit HDR system shown here distributes 2 bits extra than traditional consumer 8-bit "SDR" video, by assigning approximately 1 extra bit of precision in the traditional SDR intensity range, while applying another extra bit to extend the curve to intensities greater than 100 cd / m 2
[0121] While current video coding standards can encode video data without considering the luminance transfer function, the performance of the video coding standards can be affected by the luminance transfer function, as the distribution of codewords can depend on the luminance transfer function. For example, a video coding standard can be based on the assumption that each codeword is typically mapped to equal importance in terms of human visual sensitivity (HVS). However, this can not always be the case in reality. There are many transfer functions available, and each transfer function has its own mapping rule, which is not universal. Thus, this can result in non-optimal performance of a video encoder such as HEVC. For example, and as described in more detail below, techniques in HEVC and existing video compression systems based on quantization parameter values can not perform optimally, as they quantize the entire range of codewords with equal importance regardless of the luminance values.
[0122] Furthermore, some examples regarding standards that support HDR video processing / coding are described in Table 1 below.
[0123] [Table 1]
[0124]
[0125] Figure 5 is a flowchart showing a process of reconstructing a picture according to an embodiment.
[0126] Video content typically includes video sequences composed of a group of pictures / frames (GOP). Individual video frames or pictures can include multiple slices, where a slice includes multiple video blocks. A video block can be defined as the largest array of pixel values (also referred to as samples) that can be predictively coded. Video encoders / decoders apply predictive coding to video blocks and their subdivisions. ITU-T H.264 specifies a macroblock including 16x16 luma samples. ITU-T H.265 (or commonly known as HEVC) specifies a similar coding tree unit (CTU) structure, where a picture can be split into equal-sized CTUs, and each CTU can include a coding block (CB) having 16x16, 32x32, or 64x64 luma samples. In JEM, which is an exploration model beyond HEVC, a CTU can include a coding block having 128x128, 128x64, 128x32, 64x64, or 16x16, etc. Here, a coding block, a prediction block, and a transform block can be identical to each other. Specifically, a coding block (prediction block and transform block) can be a square or a non-square block.
[0127] According to embodiments, a decoding device can receive a bitstream (S500), perform entropy decoding (S510), perform inverse quantization (S520), determine whether to perform inverse transform (S530), perform inverse transform (S540), perform prediction (S550), and generate reconstructed samples (S560). More specific descriptions regarding embodiments are shown below.
[0128] As described above, a prediction syntax element can associate its coding block with corresponding reference samples. For example, for intra-predictive coding, an intra-prediction mode can specify the position of reference samples. In ITU-T H.265, possible intra-prediction modes for luma components include a planar prediction mode (predMode: 0), a DC prediction (predMode: 1), and a plurality of angular prediction modes (predMode: 2-N, where N can be 34 or 65 or more). One or more syntax elements can identify one of the intra-prediction modes. For inter-predictive coding, a motion vector (MV) identifies reference samples in a picture other than the picture to be coded, thereby exploiting temporal redundancy in the video. For example, a current coding block can be predicted from a reference block located in a previously coded frame, and a motion vector can be used to indicate the position of the reference block. For example, a motion vector and associated data can describe a horizontal component of the motion vector, a vertical component of the motion vector, a resolution of the motion vector (e.g., quarter-pixel precision), a prediction direction, and / or a reference picture index value. Furthermore, coding standards such as, for example, HEVC can support motion vector prediction. Motion vector prediction allows a motion vector to be specified using motion vectors of neighboring blocks.
[0129] A video encoder can generate residual data by subtracting a predicted video block from a source video block. The predicted video block can be intra-predicted or inter (motion vector) predicted. The residual data is obtained in the pixel domain. Transform coefficients are obtained by applying a transform (e.g., a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform) to the residual block to produce a set of residual transform coefficients. The transform coefficient generator can output the residual transform coefficients to a coefficient quantization unit.
[0130] The quantization parameter (QP) derivation process is summarized as follows.
[0131] The first step is to derive the luma QP. a) Find a predicted luma quantization parameter (qP Y_PRED ) based on the quantization parameter of the previous encoding (available), b) obtain a cu_delta_QP offset that indicates the difference between the predicted QP (obtained in a) and the actual QP, and c) determine the luma QP value based on the bit depth, the predicted QP, and the cu_delta_QP.
[0132] The second step is to derive the chroma QP. a) Derive the chroma QP from the luma QP, b) find the chroma QP offset from the PPS level offsets (i.e., pps_cb_qp_offset, pps_cr_qp_offset) and the slice level chroma QP offsets (i.e., slice_cb_qp_offset, slice_cr_qp_offset).
[0133] The following are the details of the above process.
[0134] The predicted luma quantization parameter qP Y_PRED is derived as follows.
[0135] [mathematical expression 6]
[0136] The predicted luma quantization parameter qP Y_PRED is derived as follows. Y_A + qP Y_B + 1) » 1
[0137] where the variables qP Y_A and qP Y_B indicate quantization parameters from previous quantization groups. Here, qP Y_A is set equal to the luma quantization parameter of the coding unit that contains the luma coding block that covers (xQg-1, yQg). Here, the luma position (xQg, yQg) specifies the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. where the variables qP Y_A and qP Y_B indicate quantization parameters from previous quantization groups. Specifically, qP Y_AThe luminance quantization parameter is set to be equal to the luminance coded unit containing the luminance coded block covering (xQg-1, yQg). Here, the luminance position (xQg, yQg) specifies the top-left luminance sample of the current quantization group relative to the top-left luminance sample of the current frame. Variable qP Y_B The luminance quantization parameter Qp is set to be equal to the luminance coded unit containing the luminance coded block covering (xQg, yQg-1). Y When qP Y_A or qP Y_B When unavailable, it is set to equal to qP. Y_PREV Here, qP Y_PREV The luminance quantization parameter Qp is set to be equal to the last encoded unit in the previous quantization group in decoding order. Y .
[0138] Once qP is determined Y_PRED The brightness quantization parameter is then updated by adding it to CuQpDeltaVal.
[0139] [Mathematical Expression 7]
[0140] Qp Y =((qP) Y_PRED +CuQpDeltaVal+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0141] The CuQpDeltaVal value is sent via a bitstream using two syntax elements, such as cu_qp_delta_abs and cu_qp_delta_sign_flag. QpBdOffset Y Specifies the value of the range offset for the luminance quantization parameter, and it depends on bit_depth_luma_minus8 (i.e., the bit depth of luminance - 8).
[0142] [Mathematical Expression 8]
[0143] QpBdOffset Y = 6 * bit_depth_luma_minus8
[0144] Finally, the brightness quantization parameter Qp′ is derived as follows. Y .
[0145] [Mathematical Expression 9]
[0146] Brightness quantization parameter Qp′ Y =Qp Y +QpBdOffsetY
[0147] The PPS-level offsets (pps_cb_qp_offset, pps_cr_qp_offset) and slice-level offsets (slice_cb_qp_offset, slice_cr_qp_offset) are considered to derive chroma QPs from luma QP as follows.
[0148] [mathematical expression 10]
[0149] qPi Cb = Clip3( -QpBdOffset C , 57, Qp Y + pps_cb_qp_offset + slice_cb_qp_offset )
[0150] qPi Cr = Clip3( -QpBdOffset C , 57, Qp Y + pps_cr_qp_offset + slice_cr_qp_offset )
[0151] The above qPi Cb and qPi Cr are further updated to qP Cb and qP Cr respectively based on the following Tables 2 and 3.
[0152] [Table 2]
[0153] qPi Cb ]]> <30 30 31 32 33 34 35 36 37 38 39 40 41 42 43 >43 qP Cb ]]> = qPi Cb ]] 29 30 31 32 33 33 34 34 35 35 36 36 37 37 = qPi Cb -6]]>
[0154] Table 2 represents the mapping from qPi Cb to qP Cb .
[0155] [Table 3]
[0156] qPi Cr ]]> <30 30 31 32 33 34 35 36 37 38 39 40 41 42 43 >43 qP Cr ]]> = qPi Cr ]] 29 30 31 32 33 33 34 34 35 35 36 36 37 37 = qPi Cr -6]]>
[0157] Table 3 represents the mapping from qPi Cr to qP Cr .
[0158] Finally, the chroma quantization parameters Qp' Cb and Qp' Cr for Cb and Cr components are derived as follows.
[0159] [mathematical expression 11]
[0160] Qp' Cb = qP Cb + QpBdOffsetC
[0161] Qp' Cr = qP Cr + QpBdOffset C
[0162] QpBdOffset C specifies a value of chroma quantization parameter range offset, and it depends on bit_depth_chroma_minus8 (i.e., bit depth of chroma - 8) as follows.
[0163] [mathematical expression 12]
[0164] QpBdOffset C = 6 * bit_depth_chroma_minus8
[0165] Further, Table 4 represents definitions of syntax elements used in the present specification.
[0166] [table 4]
[0167]
[0168] Figure 6 is a flowchart showing a process of reconstructing a picture according to another embodiment.
[0169] Since S600, S610, and S630 to S670 correspond to S500 to S560 of Figure 5 , detailed descriptions duplicated from the above description will be omitted or simplified.
[0170] According to the embodiment, the decoding device can receive a bitstream (S600), perform entropy decoding (S610), perform inverse quantization (S620), determine whether to perform inverse transform (S640), perform inverse transform (S650), perform prediction (S660), and generate reconstructed samples (S670). Also, additionally, the decoding device can derive a QP offset based on the entropy decoding (S620) and perform inverse quantization based on the derivation of the QP offset.
[0171] S620 can be specified as shown below. Hereinafter, the QP offset can be denoted by "Luma_avg_qp".
[0172] Once qPY_PRED is determined, the luma quantization parameter can be updated by adding Luma_avg_qp as follows.
[0173] [mathematical expression 13]
[0174] QpY = ((qPY PRED + CuQpDeltaVal + Luma_avg_qp + 52 + 2 * QpBdOffsetY) % (52 + QpBdOffsetY)) - QpBdOffsetY
[0175] In one example, Luma_avg_qp can be derived (or inferred) from the luminance values of the neighboring pixels (or blocks) that have been decoded and available. Luma_avg_qp can be determined from the neighboring pixel values based on a predefined derivation rule. For example, Luma_avg_qp can be derived as follows.
[0176] [mathematical expression 14]
[0177] Luma_avg_qp = A * (avg_luma - M) + B.
[0178] In mathematical expression 14, avg_luma: expected average luminance value obtained from the available (decoded) neighboring pixels (or blocks),
[0179] M: a predefined value that can depend on the bit depth,
[0180] A: a scale factor that maps the pixel value difference to the QP difference (can be predefined or signaled in the bitstream). It indicates the slope of the qp mapping, and
[0181] B: an offset value, can be predefined or signaled in the bitstream.
[0182] The derivation of Luma_avg_qp from the avg_luma value can not be limited by the above equation (one of the multiple). In another example, Luma_avg_qp can be obtained from a table mapping as follows.
[0183] [mathematical expression 15]
[0184] Luma_avg_qp = Mapping_Table_from_luma_to_QP[avg_luma]
[0185] where avg_luma is the input of the table, and the output of the table is Luma_avg_qp. To reduce the table size, the input value (avg_luma) range can be further reduced as follows.
[0186] [mathematical expression 16]
[0187] Luma_avg_qp = Mapping_Table_from_luma_to_QP[avg_luma / D]
[0188] where D is a predefined constant value to reduce the input value range.
[0189] In an embodiment, Luma_avg_qp can be derived based on information on QP. The decoding device can obtain the information on QP from a bitstream. In one example, the information on QP can include init_qp_minus26, slice_qp_delta, slice_cb_qp_offset, slice_cr_qp_offset, cu_qp_delta_abs, and cu_qp_delta_sign_flag. Further, the information on QP is not limited to the example listed above.
[0190] Figure 7 is a flowchart illustrating an operation of an encoding device according to an embodiment, Figure 8 is a block diagram illustrating a configuration of an encoding device according to an embodiment.
[0191] Figure 7 The various steps disclosed in Figure 1 may be performed by the encoding device 100 disclosed in Figure 1 . More specifically, S700 to S730 can be performed by the quantization module 123 illustrated in Figure 1 , and S740 can be performed by the entropy encoding module 130 illustrated in Figure 6 . In addition, the operations according to S700 to S740 are based on some descriptions described above in Figure 1 . Thus, the detailed description that is repetitive of the above-described contents in Figure 6 and
[0192] As illustrated in Figure 8 , the encoding device according to an embodiment can include the quantization module 123 and the entropy encoding module 130. However, in some cases, Figure 8 all the components illustrated in Figure 8 may not be essential components, and the encoding device can be implemented by more or less components than those illustrated in
[0193] The quantization module 123 and the entropy encoding module 130 in the decoding device according to an embodiment can be implemented as separate chips, or at least two or more components can be implemented through one chip.
[0194] The encoding device according to an embodiment can derive an expected average luminance value of a current block from neighboring available samples (S700). More specifically, the quantization module 123 of the encoding device can derive the expected average luminance value of the current block from the neighboring available samples.
[0195] The encoding apparatus according to an embodiment can derive a QP offset for deriving a luma QP based on the expected average luma value and the information on the QP (S710). More specifically, the quantization module 123 of the encoding apparatus can derive the QP offset for deriving the luma QP based on the expected average luma value and the information on the QP.
[0196] The encoding apparatus according to an embodiment can derive the luma QP based on the QP offset (S720). More specifically, the quantization module 123 of the encoding apparatus can derive the luma QP based on the QP offset.
[0197] The encoding apparatus according to an embodiment can perform quantization on a quantization group including the current block based on the derived luma QP (S730). More specifically, the quantization module 123 of the encoding apparatus can perform quantization on the quantization group including the current block based on the derived luma QP.
[0198] The encoding apparatus according to an embodiment can encode image information including the information on the QP (S740). More specifically, the entropy encoding module 130 can encode the image information including the information on the QP.
[0199] According to Figure 7 and Figure 8 , the encoding apparatus according to an embodiment can derive an expected average luma value of a current block from neighboring available samples (S700), derive a QP offset for deriving a luma QP based on the expected average luma value and the information on the QP (S710), perform quantization on a quantization group including the current block based on the derived luma QP (S730), and encode image information including the information on the QP (S740). Accordingly, a quantization parameter can be efficiently derived, and overall encoding efficiency can be enhanced.
[0200] Figure 9 is a flowchart illustrating an operation of a decoding apparatus according to an embodiment, Figure 10 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment.
[0201] Figure 9 Each of the steps disclosed in Figure 2 may be performed by the decoding apparatus 200 disclosed in Figure 2 . More specifically, S900 can be performed by the entropy decoding module 210 illustrated in Figure 2 , S910 to S940 can be performed by the inverse quantization module 222 illustrated in Figure 2 , S950 can be performed by the inverse transform module 223 illustrated in Figure 2 , S960 can be performed by the prediction module 230 illustrated in Figure 2 , and S970 can be performed by the adder 240 illustrated in Figure 6 . In addition, the operations according to S900 to S970 are based on the aboveSome descriptions described above are described in detail in the detailed description of the embodiments described above. Accordingly, descriptions overlapping with Figure 2 and Figure 6 will be omitted or simplified.
[0202] As shown in Figure 10 , the decoding apparatus according to the embodiments can include an entropy decoding module 210, an inverse quantization module 222, an inverse transform module 223, a prediction module 230, and an adder 240. However, in some cases, all the components shown in Figure 10 may not be essential components, and the encoding apparatus can be implemented by more or less components than those shown in Figure 10 .
[0203] The entropy decoding module 210, the inverse quantization module 222, the inverse transform module 223, the prediction module 230, and the adder 240 in the decoding apparatus according to the embodiments can be implemented as separate chips, or at least two or more components can be implemented through one chip.
[0204] The decoding apparatus according to the embodiments can decode image information including information on QP (S900). More specifically, the entropy decoding module 210 in the decoding apparatus can decode the image information including information on QP.
[0205] In one embodiment, the information on QP is signaled at a sequence parameter set (SPS) level.
[0206] In one embodiment, the image information includes information on an effective data range parameter (EDRP), and the information on the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of input values, mapping information for relating the minimum input value to luminance, mapping information for relating the maximum input value to luminance, and identification information of a transfer function. This method can be indicated as range matching.
[0207] More specifically, the present application can be used for efficient coding of image / video content whose range of codewords (input values) is limited. This can often occur in HDR content (due to the use of a transfer function that supports high luminance). It can also occur when SDR data is transformed using a luminance transform function that corresponds to HDR data. In these cases, a video encoder can be configured to signal an effective data range parameter (EDRP). And a decoder can be configured to receive the EDRP associated with the video data, and utilize the EDRP data in the decoding process. For example, the EDRP data can include a minimum input value, a maximum input value, a dynamic range of input values (indicating the difference between the maximum input value and the minimum input value), mapping information between the minimum input value and its corresponding luminance, mapping information between the maximum input value and its corresponding luminance, a transfer function identification (a known transfer function can be identified by an ID number assigned thereto, and detailed mapping information for each transfer function can be available), etc.
[0208] For example, the EDRP data can be signaled in a slice header, a picture parameter set (PPS), or a sequence parameter set (SPS). In this way, the EDRP data can be used to further modify the coded values during the decoding process.
[0209] The present application introduces a quality control parameter (QCP) to specify further adjustment of the quantization parameter. And a decoder can be configured to receive the QCP associated with the video data, and utilize the QCP data in the decoding process.
[0210] The decoding device according to embodiments can derive an expected average luminance value of the current block from neighboring available samples (S910). More specifically, the inverse quantization module 222 in the decoding device can derive an expected average luminance value of the current block from neighboring available samples.
[0211] The decoding device according to embodiments can derive a QP offset for deriving a luminance QP based on the expected average luminance value and information about the QP (S920). More specifically, the inverse quantization module 222 in the decoding device can derive a QP offset for deriving a luminance QP based on the expected average luminance value and information about the QP.
[0212] In one embodiment, the QP offset is derived based on the following equation.
[0213] [mathematical expression 17]
[0214] Luma_avg_qp = A * (avg_luma - M) + B
[0215] wherein Luma_avg_qp represents the QP offset, avg_luma represents the expected average luma value, A represents a scaling factor for mapping the pixel value difference to the QP difference, M represents a predefined value related to the bit depth, B represents an offset value, and wherein A and B are predetermined values or values included in the picture information.
[0216] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, and the mapping table is determined using the expected average luma value as input.
[0217] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, and the mapping table is determined using a value obtained by dividing the expected average luma value by a predefined constant value.
[0218] In one embodiment, the neighboring available samples include at least one of at least one luma sample adjacent to a left boundary of the quantization group and at least one luma sample adjacent to an upper boundary of the quantization group.
[0219] In one embodiment, the at least one luma sample adjacent to the left boundary of the quantization group is included in a column of luma samples directly adjacent to the left boundary of the quantization group, and the at least one luma sample adjacent to the upper boundary of the quantization group is included in a row of luma samples directly adjacent to the upper boundary of the quantization group.
[0220] In one embodiment, the neighboring available samples include a luma sample adjacent to a left side of a top-left sample of the quantization group, and the neighboring available samples include a luma sample adjacent to an upper side of the top-left sample of the quantization group.
[0221] In one embodiment, the neighboring available samples include at least one of a reconstructed neighboring sample, a sample included in at least one reconstructed neighboring block, a predicted neighboring sample, and a sample included in at least one predicted neighboring block.
[0222] In one embodiment, the avg_luma can be derived from neighboring pixel values (blocks). The avg_luma indicates an expected average luma value obtained from available (already decoded) neighboring pixels (or blocks).
[0223] i) the available neighboring pixels can include:
[0224] - a pixel located in (xQg-1, yQg+K). Here, the luma position (xQg, yQg) specifies the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. (left one row of the current block)
[0225] - the pixel located in (xQg+K, yQg-1). Here, the luma position (xQg, yQg) specifies the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. (previous line of the current block)
[0226] - Instead of one line, multiple lines can be used.
[0227] ii) The available neighboring blocks can be used to compute avg_luma:
[0228] - The block including the pixel located in (xQg-1, yQg) can be used
[0229] - The block including the pixel located in (xQg, yQg-1) can be used
[0230] iii) The Avg_luma value can be computed based on reconstructed neighboring pixels / blocks.
[0231] iv) The Avg_luma value can be computed based on predicted path pixels / blocks.
[0232] In one embodiment, the information on QP includes at least one syntax element related to a QP offset, and deriving the QP offset based on at least one of the expected average luma value and the information on QP includes deriving the QP offset based on the at least one syntax element related to the QP offset.
[0233] A new syntax element can be introduced to take into account Luma_avg_qp. For example, the Luma_avg_qp value can be sent via the bitstream. It can also be specified by two syntax elements such as Luma_avg_qp_abs and Luma_avg_qp_flag. Luma_avg_qp can be indicated by two syntax elements (luma_avg_qp_delta_abs and luma_avg_qp_delta_sign_flag).
[0234] luma_avg_qp_delta_abs specifies the absolute value of the difference CuQpDeltaLumaVal between the luma quantization parameter of the current coding unit and the luma quantization parameter derived without taking luma into account.
[0235] luma_avg_qp_delta_sign_flag specifies the sign of CuQpDeltaLumaVal as follows:
[0236] If luma_avg_qp_delta_sign_flag is equal to 0, the corresponding CuQpDeltaLumaVal has a positive value.
[0237] Otherwise (luma_avg_qp_delta_sign_flag is equal to 1), the corresponding CuQpDeltaLumaVal has a negative value.
[0238] When luma_avg_qp_delta_sign_flag is not present, it is inferred to be equal to 0.
[0239] When luma_avg_qp_delta_sign_flag is present, the variables IsCuQpDeltaLumaCoded and CuQpDeltaLumaVal are derived as follows:
[0240] [mathematical expression 18]
[0241] IsCuQpDeltaLumaCoded = 1
[0242] CuQpDeltaLumaVal = cu_qp_delta_abs * (1 - 2 * luma_avg_qp_delta_sign_flag)
[0243] CuQpDeltaLumaVal specifies the difference between the luma quantization parameter of the coding unit containing Luma_avg_qp and the luma quantization parameter of the coding unit without Luma_avg_qp.
[0244] In one embodiment, the above syntax elements can be transmitted at a quantization group level (or quantization unit level) (e.g., a CU, a CTU, or a predefined block unit).
[0245] The decoding device according to the embodiments can derive a luma QP based on the QP offset (S930). More specifically, the inverse quantization module 222 in the decoding device can derive the luma QP based on the QP offset.
[0246] The decoding device according to the embodiments can perform inverse quantization on a quantization group including a current block based on the derived luma QP (S940). More specifically, the inverse quantization module 222 in the decoding device can perform inverse quantization on the quantization group including the current block based on the derived luma QP.
[0247] In one embodiment, the decoding device can derive a chroma QP from the derived luma QP based on at least one chroma QP mapping table, and perform inverse quantization on the quantization group based on the derived luma QP and the derived chroma QP, wherein the at least one chroma QP mapping table is based on a dynamic range of chroma and the QP offset.
[0248] In one implementation, instead of one chroma QP mapping table, there can be multiple chroma QP derivation tables. Additional information can be needed to specify which QP mapping table to use. Here is an example of another chroma QP derivation table.
[0249] [Table 5]
[0250] qPi Cb ]]> <26 26 27 28 29 30 31 32 33 34 35 36 37 38 39 >39 qP Cb ]]> = qPi Cb ]] 26 27 27 28 29 29 30 30 31 31 32 32 33 33 = qPi Cb -6]]>
[0251] In one implementation, the at least one chroma QP mapping table includes at least one Cb QP mapping table and at least one Cr QP mapping table.
[0252] The decoding device according to the implementation can generate the residual samples of the current block based on the inverse quantization (S950). More specifically, the inverse transform module 223 in the decoding device can generate the residual samples of the current block based on the inverse quantization.
[0253] The decoding device according to the implementation can generate the prediction samples of the current block based on the picture information (S960). More specifically, the prediction module 230 in the decoding device can generate the prediction samples of the current block based on the picture information.
[0254] The decoding device according to the implementation can generate the reconstructed samples of the current block based on the residual samples of the current block and the prediction samples of the current block (S970). More specifically, the adder in the decoding device can generate the reconstructed samples of the current block based on the residual samples of the current block and the prediction samples of the current block.
[0255] According to Figure 9 and Figure 10 , the decoding device according to the implementation can decode the picture information including information about QP (S900), derive an expected average luminance value of the current block from neighboring available samples (S910), derive a QP offset for deriving a luma QP based on the expected average luminance value and the information about QP (S920), derive the luma QP based on the QP offset (S930), perform inverse quantization on a quantization group including the current block based on the derived luma QP (S940), generate the residual samples of the current block based on the inverse quantization (S950), generate the prediction samples of the current block based on the picture information (S960), and generate the reconstructed samples of the current block based on the residual samples of the current block and the prediction samples of the current block (S970). Thus, the quantization parameter can be derived efficiently, and the overall coding efficiency can be enhanced.
[0256] The above-described method according to the present application can be implemented in software, and the encoding device and / or the decoding device according to the present application can be included in an image processing device such as a TV, a computer, a smart phone, a display device, etc.
[0257] When the embodiments of the present application are implemented in software, the above-described method can be implemented by a module (process, function etc.) that performs the above-described functions. The module is stored in a memory and can be executed by a processor. The memory can be internal or external to the processor, and can be coupled to the processor by any of a variety of well-known means. The processor can include an application specific integrated circuit (ASIC), other chip sets, logic circuit, and / or a data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices.
Claims
1. A picture decoding method by a decoding device, the picture decoding method comprising the steps of: obtaining, from a bitstream, picture information including information related to a quantization parameter (QP), wherein the information related to the QP includes information related to a luma QP and information related to a chroma QP, and the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value; deriving the luma QP based on the information related to the luma QP; deriving the chroma QP based on the information related to the chroma QP, wherein the chroma QP includes a first chroma QP for a Cb component and a second chroma QP for a Cr component; and generating a reconstructed picture including reconstructed samples based on the luma QP, the first chroma QP for the Cb component, and the second chroma QP for the Cr component, wherein the step of deriving the luma QP comprises: deriving a predicted luma QP based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) represents a top-left sample position of a current quantization group; deriving the QP delta value based on the information on the absolute value of the QP delta value and the information on the sign of the QP delta value; and deriving the luma QP based on the predicted luma QP and the QP delta value, wherein the step of deriving the chroma QP comprises: deriving the first chroma QP for the Cb component based on a first chroma QP mapping table for the Cb component; and deriving the second chroma QP for the Cr component based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
2. The picture decoding method of claim 1, wherein, deriving the chroma QP based on the luma QP.
3. The picture decoding method of claim 2, wherein, deriving the first chroma QP based on first chroma QP offset information for the Cb component, and wherein the first chroma QP offset information includes a first PPS-level offset and a first slice-level offset.
4. The picture decoding method of claim 2, wherein, deriving the second chroma QP based on second chroma QP offset information for the Cr component, and wherein the second chroma QP offset information includes a second PPS-level offset and a second slice-level offset. 5.A picture encoding method by an encoding device, the picture encoding method comprising the steps of: deriving a luma quantization parameter (QP); deriving a chroma QP, wherein the chroma QP includes a first chroma QP for a Cb component and a second chroma QP for a Cr component; and generating information related to QPs, the information related to the QPs including information related to the luma QP and information related to the chroma QPs; and encoding image information including the information related to the QPs, wherein the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, wherein the luma QP is expressed based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the step of deriving the chroma QPs includes the steps of: deriving the first chroma QP for the Cb component based on a first chroma QP mapping table for the Cb component; and deriving the second chroma QP for the Cr component based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
6. A non-transitory computer-readable storage medium storing a bitstream generated by a method performed by an encoding device, the method comprising the steps of: deriving a luma quantization parameter (QP); deriving chroma QPs, wherein the chroma QPs include a first chroma QP for a Cb component and a second chroma QP for a Cr component; generating information related to QPs, the information related to the QPs including information related to the luma QP and information related to the chroma QPs; and encoding image information including the information related to the QPs to output the bitstream, wherein the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, wherein the luma QP is expressed based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the step of deriving the chroma QPs includes the steps of: deriving the first chroma QP for the Cb component based on a first chroma QP mapping table for the Cb component; and deriving the second chroma QP for the Cr component based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
7. A method for transmitting data for image information, the method comprising the steps of: obtaining a bitstream of the image information, the image information including information related to quantization parameters, QPs, the information related to the QPs including information related to luma QPs and information related to chroma QPs, the chroma QPs including a first chroma QP for a Cb component and a second chroma QP for a Cr component; and transmitting the data including the bitstream of the image information, wherein the information related to the luma QPs includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, wherein the luma QPs are represented based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the first chroma QP for the Cb component is derived based on a first chroma QP mapping table for the Cb component, wherein the second chroma QP for the Cr component is derived based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
Citation Information
Patent Citations
Method For Decoding Chroma Image
CN107277504A
Systems and methods for optimizing video coding based on luminance transfer function or video color component values
CN107852512A