Video coding apparatus and method based on quantization parameter derivation
By deriving quantization parameters using a predicted average luma value and offset, the method improves video coding efficiency, addressing the high data volume challenge in high-resolution images and reducing transmission and storage costs.
Patent Information
- Application Number
- JP2022078055
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-01
- Filing Date
- 2022-05-11
- Publication Date
- 2026-01-22
- Estimated Expiration
- 2039-03-05
AI Technical Summary
The increasing demand for high-resolution and high-quality images leads to higher transmission and storage costs due to increased data volume, necessitating a more efficient video compression technique.
A method and apparatus for improving video coding efficiency by deriving quantization parameters using a predicted average luma value and quantization parameter offset, followed by inverse quantization and reconstruction processes.
Enhances overall image/video compression efficiency and quantization efficiency through efficient derivation of quantization parameters.
Smart Images

Figure 0007804527000006 
Figure 0007804527000007 
Figure 0007804527000008
Abstract
Description
[Technical Field]
[0001] The present invention relates to video coding technology, and more particularly to a video coding apparatus and method based on quantization parameter derivation in a video coding system. [Background technology]
[0002] Recently, demands for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, have been increasing in various fields. Because such image data has high resolution and quality, the amount of information or bits to be transmitted increases compared to existing image data. Therefore, when transmitting or storing image data using a medium such as a wired / wireless broadband line, transmission costs and storage costs may increase.
[0003] Therefore, there is a demand for a highly efficient video compression technique for efficiently transmitting, storing, and reproducing high-resolution and high-quality video information. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention provides a method and apparatus for improving video coding efficiency.
[0005] The present invention also provides a method and apparatus for increasing quantization efficiency.
[0006] The present invention also provides a method and apparatus for efficiently deriving quantization parameters. [Means for solving the problem]
[0007] According to an embodiment of the present invention, there is provided a picture decoding method performed by a decoding device, the picture decoding method including: decoding video information including information on a quantization parameter (QP), deriving a predicted average luma value of a current block from available neighboring samples, deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the predicted average luma value and the information on the QP, deriving the luma QP based on the QP offset, performing inverse quantization on a quantization group including the current block based on the derived luma QP, generating residual samples for the current block based on the inverse quantization, generating predicted samples for the current block based on the video information, and generating reconstructed samples for the current block based on the residual samples for the current block and the predicted samples for the current block.
[0008] According to an embodiment of the present invention, there is provided a decoding device for decoding a picture, the decoding device including: an entropy decoding module configured to decode video information including information on a quantization parameter (QP), an inverse quantization module configured to derive a predicted average luma value of a current block from available neighboring samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the predicted average luma value and the information on the QP, derive the luma QP based on the QP offset, and perform inverse quantization on a quantization group including the current block based on the derived luma QP, an inverse transform module configured to generate residual samples for the current block based on the inverse quantization, a prediction module configured to generate predicted samples for the current block based on the video information, and a reconstruction module configured to generate reconstructed samples for the current block based on the residual samples for the current block and the predicted samples for the current block.
[0009] According to an embodiment of the present invention, there is provided a picture encoding method to be performed by an encoding apparatus, the picture encoding method including: deriving a predicted average luma value of a current block from available neighboring samples, deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the predicted average luma value and information on a QP, deriving the luma QP based on the QP offset, performing quantization on a quantization group including the current block based on the derived luma QP, and encoding video information including information on the QP.
[0010] According to an embodiment of the present invention, there is provided an encoding apparatus for encoding a picture, the encoding apparatus including: a quantization module configured to derive a predicted average luma value of a current block from available neighboring samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the predicted average luma value and information on a QP, derive the luma QP based on the QP offset, and perform quantization for a quantization group including the current block based on the derived luma QP, and an entropy encoding module configured to encode video information including the information on the QP. [Effects of the Invention]
[0011] According to the present invention, the overall image / video compression efficiency can be increased.
[0012] According to the present invention, the quantization efficiency can be increased.
[0013] According to the present invention, the quantization parameters can be derived efficiently. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a schematic diagram illustrating a configuration of an encoding device according to an embodiment. [Figure 2] 1 is a schematic diagram illustrating the configuration of a decoding device according to an embodiment. [Figure 3] An example of a chromaticity diagram is shown below. [Figure 4] 1 shows an example of a mapping of linear light values for SDR and HDR representations. [Figure 5] 1 is a flowchart illustrating a picture restoration process according to one embodiment. [Figure 6] 10 is a flowchart illustrating a picture restoration process according to another embodiment. [Figure 7]10 is a flowchart illustrating an operation of an encoding device according to an embodiment. [Figure 8] 1 is a block diagram showing a configuration of an encoding device according to an embodiment; [Figure 9] 10 is a flowchart illustrating an operation of a decoding device according to an embodiment. [Figure 10] 1 is a block diagram showing the configuration of a decoding device according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0015] According to an embodiment of the present invention, there is provided a picture decoding method performed by a decoding device, the picture decoding method including: decoding video information including information on a quantization parameter (QP), deriving a predicted average luma value of a current block from available neighboring samples, deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the predicted average luma value and the information on the QP, deriving the luma QP based on the QP offset, performing inverse quantization on a quantization group including the current block based on the derived luma QP, generating residual samples for the current block based on the inverse quantization, generating predicted samples for the current block based on the video information, and generating reconstructed samples for the current block based on the residual samples for the current block and the predicted samples for the current block.
[0016] The present invention can be modified in various forms, and specific embodiments will be illustrated in the drawings and described in detail. However, these embodiments are not intended to limit the present invention. Terms used in this specification are used merely to describe specific embodiments and are not intended to limit the present invention. The singular includes the plural unless clearly indicated otherwise. In this specification, the terms "comprise" or "have" mean the presence of any feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0017] Meanwhile, each component in the drawings described in the present invention is illustrated independently for the convenience of explaining the different characteristic functions of the video coding / decoding device, and does not mean that each component is implemented as separate hardware or software. For example, two or more of the components may be integrated into one component, or one component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0018] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In addition, the same reference numerals are used throughout the drawings to designate the same components, and redundant description of the same components will be omitted.
[0019] This specification may be applied to technical fields related to video or image processing. For example, the methods or embodiments disclosed in the following description may be applied to various video coding standards, such as the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), next-generation video / image coding standards after VVC, or video / image coding standards before VVC, such as the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265).
[0020] In this specification, "video" may refer to a set of images over time. A "picture" generally refers to a unit representing one image at a specific time period, and a "slice" is a unit constituting a portion of a picture in coding. One picture may be composed of multiple slices, and pictures and slices may be used interchangeably as needed. Furthermore, the term "image" may refer to a concept including still images and video, which is a set of still images over time. Furthermore, "video" does not necessarily refer to a set of still images over time, but may be interpreted as a concept including still images in some embodiments.
[0021] A pixel or pel can refer to the smallest unit of a picture (or image). The term 'sample' can also be used as a term corresponding to a pixel. A sample can generally refer to a pixel or a pixel value, and can refer to only a pixel / pixel value of a luminance (luma) component, or only a pixel / pixel value of a chroma component.
[0022] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. The unit may be used interchangeably with terms such as block or area. Generally, an M×N block may refer to a set of samples or transform coefficients consisting of M columns and N rows.
[0023] FIG. 1 is a schematic diagram illustrating the configuration of an encoding device according to an embodiment.
[0024] Hereinafter, the encoding / decoding device may include a video encoding / decoding device and / or a video encoding / decoding device. The video encoding / decoding device may be used as a concept including a video encoding / decoding device, and the video encoding / decoding device may be used as a concept including a video encoding / decoding device.
[0025] 1, the encoding apparatus 100 may include a picture partition module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filter module 150, and a memory 160. The residual processing module 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, an inverse quantization module 125, and an inverse transform module 126.
[0026] The picture division module 105 can divide an input picture into at least one processing unit.
[0027] For example, the processing unit is called a coding unit (CU). In this case, the coding units may be recursively divided from the largest coding unit (LCU) using a quad-tree binary-tree (QTBT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure.
[0028] In this case, for example, a quad tree structure may be applied first, followed by a binary tree structure and a ternary tree structure. Alternatively, the binary tree structure / ternary tree structure may be applied first. The coding procedure according to the present invention may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or a coding unit may be recursively divided into coding units of deeper depths and used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described below.
[0029] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding units may be divided into deeper depth coding units from a largest coding unit (LCU) using a quadtree structure. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or a coding unit may be recursively divided into deeper depth coding units and used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit cannot be divided into coding units smaller than the smallest coding unit.
[0030] Here, the term "final coding unit" refers to a coding unit that is partitioned or divided into the prediction unit or the transform unit. The prediction unit is a unit that is partitioned from the coding unit and is a unit of sample prediction. In this case, the prediction unit may be divided into sub-blocks. The transform unit may be divided from the coding unit according to a quadtree structure and is a unit that derives transform coefficients and / or a unit that derives a residual signal from the transform coefficients.
[0031] Hereinafter, the coding unit is referred to as a coding block (CB), the prediction unit is referred to as a prediction block (PB), and the transform unit is referred to as a transform block (TB). The prediction block or the prediction unit may refer to a specific region in a block format within a picture and may include an array of prediction samples. Furthermore, the transform block or the transform unit may refer to a specific region in a block format within a picture and may include an array of transform coefficients or residual samples.
[0032] The prediction module 110 predicts a current block or a residual block to generate a prediction block including prediction samples of the current block. The prediction module 110 performs prediction on a coding block, a transform block, or a prediction block.
[0033] The prediction module 110 predicts a current block or a residual block to generate a prediction block including prediction samples of the current block. The prediction module 110 performs prediction on a coding block, a transform block, or a prediction block.
[0034] The prediction module 110 may determine whether intra prediction or inter prediction is applied to the current block. For example, the prediction module 110 may determine whether intra prediction or inter prediction is applied on a CU basis.
[0035] In the case of intra prediction, the prediction module 110 can derive a prediction sample for a current block based on a reference sample outside the current block within a picture to which the current block belongs (hereinafter, the current picture).
[0036] In this case, the prediction module 110 can (i) derive a prediction sample based on the average or interpolation of neighboring reference samples of the current block, or (ii) derive a prediction sample based on a reference sample that exists in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block.
[0037] Case (i) is called a non-directional mode or a non-angular mode, and case (ii) is called a directional mode or an angular mode. In intra prediction, prediction modes may include, for example, 65 directional prediction modes and at least two or more non-directional modes. The non-directional modes may include a DC prediction mode and a planar mode. The prediction module 110 may determine a prediction mode to be applied to a current block using prediction modes applied to neighboring blocks.
[0038] In the case of inter prediction, the prediction module 110 may derive a prediction sample for a current block based on a sample identified by a motion vector on a reference picture. The prediction module 110 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the prediction module 110 may use motion information of a neighboring block as motion information of the current block.
[0039] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In MVP mode, the motion vector of the current block can be derived by using the motion vector of the neighboring block as the motion vector predictor of the current block.
[0040] In the case of inter-prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in a reference picture. The reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). The motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and output in the form of a bitstream.
[0041] When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on a reference picture list can be used as a reference picture. The reference pictures included in a Picture Order Count (POC) can be sorted based on the POC difference between the current picture and the corresponding reference picture. The POC corresponds to the display order of pictures and can be distinguished from the coding order.
[0042] The subtractor 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When the skip mode is applied, the residual sample is not generated as described above.
[0043] The transform module 122 transforms the residual samples in units of transform blocks to generate transform coefficients. The transform module 122 may perform the transform according to the size of the corresponding transform block and a prediction mode applied to a coding block or a prediction block spatially overlapping with the corresponding transform block.
[0044] For example, if intra prediction is applied to the coding block or the prediction block that overlaps with the transform block and the transform block is a 4x4 residual array, the residual samples may be transformed using a Discrete Sine Transform (DST); otherwise, the residual samples may be transformed using a Discrete Cosine Transform (DCT).
[0045] The quantization module 123 may quantize the transform coefficients to generate quantized transform coefficients.
[0046] The rearrangement module 124 rearranges the quantized transform coefficients. The rearrangement module 124 can rearrange the quantized transform coefficients in a block format into a one-dimensional vector format through a coefficient scanning method. Although the rearrangement module 124 has been described as a separate configuration, the rearrangement module 124 may be a part of the quantization module 123.
[0047] The entropy encoding module 130 may perform entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding module 130 may encode the quantized transform coefficients and information required for video restoration (e.g., syntax element values, etc.) together or separately using entropy encoding or a predetermined method.
[0048] The information encoded by entropy encoding can be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The bitstream can be transmitted over a network or stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0049] The inverse quantization module 125 inversely quantizes the values (quantized transform coefficients) quantized by the quantization module 123, and the inverse transform module 126 inversely transforms the values inversely quantized by the inverse quantization module 125 to generate residual samples.
[0050] The adder 140 reconstructs a picture by combining the residual samples and the prediction samples. The residual samples and the prediction samples may be added in block units to generate a reconstructed block. Here, the adder 140 may be a part of the prediction module 110. Meanwhile, the adder 140 is referred to as a reconstruction module or a reconstructed block generator.
[0051] The filter module 150 may apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through the deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries within the reconstructed picture or distortion in the quantization process may be corrected. The sample adaptive offset may be applied on a sample-by-sample basis and may be applied after the deblocking filtering process is completed. The filter module 150 may also apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset have been applied.
[0052] The memory 160 may store the reconstructed picture (decoded picture) or information necessary for encoding / decoding. Here, the reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter module 150. The stored reconstructed picture may be used as a reference picture for (inter) prediction of another picture. For example, the memory 160 may store (reference) pictures used for inter prediction. In this case, the pictures used for inter prediction may be specified by a reference picture set or a reference picture list.
[0053] FIG. 2 is a schematic diagram illustrating the configuration of a decoding device according to an embodiment.
[0054] 2, the decoding apparatus 200 includes an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filter module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, an inverse quantization module 222, and an inverse transform module 223. Although not shown, the video decoding apparatus 200 may also include a receiver for receiving a bitstream including video information. The receiver may be a separate module or may be included in the entropy decoding module 210.
[0055] When a bitstream containing video / image information is input, the (video) decoding device 200 can restore the video / image / picture in response to the process in which the video / image information is processed in the (video) encoding device 100.
[0056] For example, the video decoding apparatus 200 may perform video decoding using a processing unit applied to the video encoding apparatus. Accordingly, a processing unit block for video decoding may be, for example, a coding unit, a prediction unit, or a transform unit. The coding unit may be divided from a maximum coding unit into a quad tree structure, a binary tree structure, and / or a ternary tree structure.
[0057] A prediction unit and a transform unit may also be used in some cases. The prediction unit block is a block derived or partitioned from the coding unit. The transform unit may be divided from the coding unit using a quadtree structure and is a unit for deriving transform coefficients or a unit for deriving a residual signal from the transform coefficients.
[0058] The entropy decoding module 210 may parse a bitstream and output information necessary for video or picture reconstruction. For example, the entropy decoding module 210 may decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and calculate values of syntax elements necessary for video reconstruction and values of quantized transform coefficients associated with residuals.
[0059] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model based on the syntax element to be decoded and decoding information to be decoded, predicts the occurrence probability of the bin according to the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. After determining the context model, the CABAC entropy decoding method can update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.
[0060] In the entropy decoding module 210, information regarding prediction among the decoded information is provided to the prediction module 230, and residual values on which entropy decoding is performed in the entropy decoding module 210 can be input to the realignment module 221.
[0061] The reordering module 221 may reorder the quantized transform coefficients in a two-dimensional block format. The reordering module 221 may perform reordering in response to coefficient scanning performed by the encoding apparatus. Although the reordering module 221 has been described as a separate configuration, the reordering module 221 may be a part of the dequantization module 222.
[0062] The inverse quantization module 222 may output the inverse quantized transform coefficients by inverse quantizing the quantized transform coefficients based on a (inverse) quantization parameter, where information for deriving the quantization parameter may be signaled from the encoding apparatus.
[0063] The inverse transform module 223 may inverse transform the transform coefficients to derive the residual samples.
[0064] The prediction module 230 may predict a current block and generate a prediction block including prediction samples of the current block. The prediction module 230 may perform prediction on a coding block, a transform block, or a prediction block.
[0065] The prediction module 230 may determine whether to apply intra prediction or inter prediction based on the prediction information. In this case, the unit for determining whether to apply intra prediction or inter prediction differs from the unit for generating prediction samples. Furthermore, the unit for generating prediction samples differs between inter prediction and intra prediction. For example, whether to apply inter prediction or intra prediction may be determined on a CU basis. Furthermore, for example, in inter prediction, a prediction mode may be determined on a PU basis to generate prediction samples. In intra prediction, a prediction mode may be determined on a PU basis to generate prediction samples on a TU basis.
[0066] In the case of intra prediction, the prediction module 230 may derive prediction samples for the current block based on neighboring reference samples in the current picture. The prediction module 230 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined using the intra prediction mode of the neighboring block.
[0067] In the case of inter prediction, the prediction module 230 may derive a prediction sample for the current block based on a sample identified on a reference picture by a motion vector on the reference picture. The prediction module 230 may derive a prediction sample for the current block by applying a skip mode, a merge mode, or an MVP mode. In this case, motion information required for inter prediction of the current block provided to the encoding apparatus, such as information on a motion vector and a reference picture index, may be obtained or derived based on the prediction information.
[0068] In the skip mode and merge mode, motion information of neighboring blocks can be used as motion information of the current block, where the neighboring blocks can include spatial neighboring blocks and temporal neighboring blocks.
[0069] The prediction module 230 constructs a merge candidate list using motion information of available neighboring blocks and can use information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding apparatus. The motion information can include a motion vector and a reference picture. When motion information of temporally neighboring blocks is used in skip mode and merge mode, the top picture on the reference picture list can be used as the reference picture.
[0070] In skip mode, unlike merge mode, the residual between the predicted sample and the original sample is not transmitted.
[0071] In the MVP mode, the motion vector of the current block can be derived using the motion vector of a neighboring block as a motion vector predictor, where the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0072] For example, when a merge mode is applied, a merge candidate list may be generated by using the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vector corresponding to the Col block, which is a temporally neighboring block. In the merge mode, the motion vector of a candidate block selected from the merge candidate list is used as the motion vector of the current block. The prediction information may include a merge index indicating a candidate block having an optimal motion vector selected from the candidate blocks included in the merge candidate list. In this case, the prediction module 230 may derive the motion vector of the current block by using the merge index.
[0073] As another example, when a Motion Vector Prediction (MVP) mode is applied, a motion vector predictor candidate list is generated by using the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col block, which is a temporally neighboring block. That is, the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col block, which is a temporally neighboring block, can be used as motion vector candidates. The prediction information can include a predicted motion vector index indicating an optimal motion vector selected from the motion vector candidates included in the list.
[0074] In this case, the prediction module 230 may select a predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list by using the motion vector index. The prediction module of the encoding apparatus may obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor and output the MVD in the form of a bitstream. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction module 230 may obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction module may obtain or derive a reference picture index, etc., indicating a reference picture from the prediction information.
[0075] The adder 240 may reconstruct a current block or a current picture by adding residual samples and predicted samples. The adder 240 may reconstruct a current picture by adding residual samples and predicted samples in block units. When a skip mode is applied, residuals are not transmitted, and therefore predicted samples may become reconstructed samples. Although the adder 240 has been described as a separate component, it may be part of the prediction module 230. Meanwhile, the adder 240 may be referred to as a reconstruction module or a reconstruction block generator.
[0076] The filter module 250 may apply a deblocking filter, a sample adaptive offset, and / or an ALF to the reconstructed picture. In this case, the sample adaptive offset may be applied in sample units and may be applied after deblocking filtering. The ALF may be applied after deblocking filtering and / or the sample adaptive offset.
[0077] The memory 260 may store a reconstructed picture (a decoded picture) or information necessary for decoding. Here, the reconstructed picture is a reconstructed picture that has undergone a filtering procedure by the filter module 250. For example, the memory 260 may store a picture used for inter prediction. In this case, the picture used for inter prediction may be specified by a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for another picture. In addition, the memory 260 may output the reconstructed picture in an output order.
[0078] FIG. 3 shows an example of a chromaticity diagram.
[0079] The present embodiment relates to video coding, and more particularly to a technique for optimizing video coding according to given conditions such as a specified or predicted luminance transfer function, a dynamic range of video, and luminance values of coding blocks.
[0080] As used herein, the "luminance transfer function" is referred to as the optical-electro transfer function (OETF) or the electro-optical transfer function (EOTF). Even though the optical-electrical transfer function and the electro-optical transfer function are not exact inverses of each other, the optical-electrical transfer function is referred to as the inverse electro-optical transfer function, and the electro-optical transfer function is referred to as the inverse optical-electrical transfer function.
[0081] The techniques described herein can be used to compensate for non-optimal video coding performance that occurs when the mapping of luminance values to digital codewords is not considered equally important. For example, OETF actually allows more bits for dark areas than for bright areas (or vice versa). In this case, video encoders / decoders designed based on the assumption that all digital codewords are coded with equal importance generally do not perform video coding in an optimal manner.
[0082] It should be noted that although the techniques of this disclosure are described in conjunction with the ITU-T H.264 and ITU-T H.265 standards, the techniques of this disclosure are generally applicable to any video coding standard.
[0083] Video compression technology is used in a wide range of devices, including digital televisions, desktop computers, portable computers, tablet computers, digital recording devices, digital media players, video game devices, smartphones, etc. Digital video can be coded according to video coding standards such as ITU-T H.264 and High Efficiency Video Coding (HEVC), also known as ISO / IEC MPEG-4 AVC. Video coding standards allow a specific format (e.g., YUV420) to be coded.
[0084] Conventional digital video cameras initially generate raw data corresponding to the signals generated by each video sensor. For example, digital video capture devices record video as a set of linearly related luminance values. However, human vision cannot perceive linear changes in luminance values. For example, a luminance of 100 cd / m 2 The area of the image associated with the luminance value of 200 cd / m 2The luminance value does not necessarily need to be perceived as twice as bright as the area of the image associated with that luminance value. Thus, a luminance transfer function (e.g., an optical-to-electrical transfer function (OETF) or an electrical-to-optical transfer function (EOTF)) can be used to convert linear luminance data into data that can be perceived in a meaningful manner. The optical-to-electrical transfer function (OETF) can map absolute linear luminance values to digital codewords in a nonlinear manner. As a result, the digital codewords can be converted into a video format supported by a video coding standard.
[0085] Conventional video coding / display systems, such as conventional television video distribution environments, provide a Standard Dynamic Range (SDR), typically between about 0.1 and 100 cd / m 2 (often called "nits"), which is significantly smaller than what you'll encounter in real life. For example, a light bulb has a brightness range of 10,000 cd / m 2 The surface of sunlight can have a brightness of hundreds of thousands of cd / m 2 or brighter than the night sky, which is 0.005 cd / m 2 It can have the following brightness:
[0086] Recently, LCD and OLED displays have become widely used, and the technology of these devices allows for higher brightness and wider color space reproduction. The realization and desired brightness and dynamic range of various displays differs significantly from the brightness and dynamic range of conventional (SDR) capturing and production devices. For example, content production systems can generate or capture content with a contrast of 1,000,000:1. Televisions and other video distribution environments are expected to provide visual effects closer to the real experience, providing users with a strong sense of "realism." The existing SDR brightness range (0.1 to 100 cd / m 2 ) instead of a higher brightness range (0.005 to 10,000 cd / m 2) can be considered. For example, 0.01 cd / m 2 Minimum brightness of 2,000 cd / m 2 When HDR content is displayed on a display that supports a maximum brightness of 10 ...
[0087] Furthermore, Ultra High Definition Television (UHDTV) aims to provide users with a sense of realism. Increasing resolution alone is not sufficient to achieve this goal unless content with higher peak brightness and greater contrast values than current TVs is generated, captured, and displayed. Furthermore, to enhance realism, a richer range of hues than currently available in commonly used color gamuts, such as BT.709, must be rendered. Therefore, new content will not only have several tens of times greater brightness and contrast, but also a significantly wider color gamut (e.g., BT.2020 or a future wider color gamut). The various color gamut ranges are shown in Figure 3.
[0088] Beyond this latest technological development, HDR (High Dynamic Range) image / video playback can now be achieved using suitable transfer functions (OETF / EOTF) at both the production and consumer end.
[0089] FIG. 4 shows an example of a mapping of linear light values for SDR and HDR representations.
[0090] A transfer function can be described as a mapping between input and output in the real (floating point) range [0.0, 1.0]. One example of a luminance transfer function that supports HDR data includes the so-called SMPTE (Society of Movie Picture and Television) HDR (High Dynamic Range) transfer function, which is referred to as SMPTE ST 2084. Another example of a luminance transfer function that supports HDR data includes a hybrid log-gamma transfer function for HDR signals (also referred to as ITU-R BT.2100). Specifically, SMPTE HDR transfer functions include an EOTF and an inverse EOTF. The SMPTE ST 2084 inverse EOTF is described by the following equation:
[0091] [Number 1] L c =R / 10,000
[0092] [Number 2] V=((c1+c2*L c n ) / (1+c3*L c n )) m
[0093] In Equation 1 and Equation 2, c1=c3-c2+1=3424 / 4096=0.8359375, c2=32*2413 / 4096=18.8515625, c3=32*2392 / 4096=18.6875, m=128*2523 / 4096=78.84375, and n=0.25*2610 / 4096=0.1593017578125.
[0094] The SMPTE ST 2084 EOTF can be described by the following mathematical formula:
[0095] [Number 3] L c =((max[(V 1 / m --c1), 0]) / (c2-c3*V 1 / m ))1 / n
[0096] [Number 4] R=10,000*L c
[0097] In the above formula, R is 0 to 10,000 cd / m 2 For example, the same L as 1 c is 10,000 cd / m 2 R can represent an absolute linear luminance value. In the above formula, V is called a luminance value (or a perceptual curve value). OETF can map the perceptual curve value to a digital codeword. Therefore, V is 2 N An example of a function that can be used when mapping V to an N-bit codeword can be defined as follows:
[0098] [Number 5] Digital value = INT((2 N -1)*V)
[0099] In Equation 5, INT(x) rounds down any decimal value less than 0.5 and rounds up any decimal value greater than or equal to 0.5 to generate a constant value.
[0100] As an illustrative example, FIG. 4 shows a BT.709 style transfer function (green curve) ranging from 0.1 to 100 cd / m 2 8-bit SDR systems capable of displaying 0.005 to 10,000 cd / m and other transfer functions (SMPTE ST 2084). 2This diagram is a schematic representation. It does not capture the exact shape of the curve and is presented for illustrative purposes only. In Figure 4, the horizontal axis represents constant code levels, and the vertical axis represents linear light values (scaled to log10). This exemplary mapping includes a conventional code level range ratio to accommodate all foot-room (negative samples below the real range of [0.0, 1.0]) and head-room (samples above real 1.0). By design, the 10-bit HDR transfer function displayed here allocates approximately twice as many code levels [119 to 509] as the existing 8-bit SDR transfer function allocates [16 to 235] code levels in the SDR range, and also provides a similar number of new code levels [510 to 940] to extend brightness. 2 New code levels [64-118] are assigned for dark light intensities below 100.
[0101] In a sense, the 10-bit HDR system described here allocates approximately one additional bit of precision within the existing SDR intensity range, and applies another additional bit to refine the curve to 100 cd / m 2 By extending the range to greater brightness intensities, an additional 2 bits are allocated to existing consumer 8-bit "SDR" video. For comparison, the 10-bit SDR transfer function is also shown (red dash curve).
[0102] Although current video coding standards can code video data without considering the luma transfer function, the performance of the video coding standard can be affected by the luma transfer function because the distribution of codewords can depend on the luma transfer function. For example, video coding standards are based on the assumption that each codeword is generally mapped with the same importance in terms of human visual sensitivity (HVS). However, each codeword does not actually map with the same importance. There are many transfer functions available, and each transfer function has its own mapping rule. Therefore, for these reasons, the performance of video coders such as HEVC is not optimized. For example, as will be described later, HEVC and existing video compression system technologies based on quantization parameter values quantize the entire range of codewords with the same importance regardless of the luma value, resulting in suboptimal coding.
[0103] Meanwhile, some examples for standards to support HDR video processing / coding are described in Table 1.
[0104] [Table 1]
[0105] FIG. 5 is a flowchart illustrating a picture reconstruction process according to one embodiment.
[0106] Video content typically includes video sequences organized into pictures / groups of frames (GOPs). Each video frame or picture can include multiple slices, each of which includes multiple video blocks. A video block can be defined as a maximal array of pixel values (also called samples) that can be predictively coded. Video encoders / decoders apply predictive coding to video blocks and subdivisions of video blocks. ITU-T H.264 specifies macroblocks containing 16x16 luma samples. ITU-T H.265 (or commonly referred to as HEVC) specifies a similar coding tree unit (CTU) structure in which a picture can be divided into equal-sized CTUs, each of which can include coding blocks (CBs) with 16x16, 32x32, or 64x64 luma samples. In JEM, a search model other than HEVC, a CTU can include coding blocks with 128x128, 128x64, 128x32, 64x64, or 16x16 luma samples. Here, the coding block, the prediction block, and the transformation block are the same as each other. Specifically, the coding block (the prediction block and the transformation block) may be a square or non-square block.
[0107] According to one embodiment, a decoding device receives a bitstream (S500), performs entropy decoding (S510), performs inverse quantization (S520), determines whether to perform an inverse transform (S530), performs the inverse transform (S540), performs prediction (S550), and generates reconstruction samples (S560). A more detailed description of the embodiment is provided below.
[0108] As described above, a prediction syntax element can associate the coding block of the prediction syntax element with a corresponding reference sample. For example, for intra-predictive coding, an intra-prediction mode can specify the location of the reference sample. In ITU-T H.265, possible intra-prediction modes for the luma component include a planar prediction mode (predMode: 0), a DC prediction mode (predMode: 1), and a multiple angle prediction mode (predMode: 2-N, where N is 34 or 65 or greater). One or more syntax elements can identify one of the intra-prediction modes. For inter-predictive coding, a motion vector (MV) exploits temporal redundancy in video by identifying reference samples in pictures other than the picture of the coding block being coded. For example, a current coding block can be predicted from a reference block located in a previously coded frame, and the motion vector can be used to indicate the location of the reference block. The motion vector and associated data may describe, for example, the horizontal component of the motion vector, the vertical component of the motion vector, the resolution for the motion vector (e.g., quarter-pixel accuracy), the prediction direction, and / or the reference picture index value. Coding standards such as HEVC may also support motion vector prediction, which allows a motion vector to be determined using the motion vectors of neighboring blocks.
[0109] A video encoder can generate residual data by subtracting a prediction video block from a source video block. The prediction video block can be intra-predicted or inter- (motion vector) predicted. The residual data is obtained in the pixel domain. Transform coefficients are obtained by applying a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block, thereby calculating a set of residual transform coefficients. The transform coefficient generator can output the residual transform coefficients to a coefficient quantizer.
[0110] The Quantization Parameter (QP) derivation process is summarized as follows:
[0111] The first step is to derive the luma QP, which consists of a) deriving a predicted luma quantization parameter (qP) based on previously coded (available) quantization parameters. Y_PRED ), b) obtaining a cu_delta_QP offset indicating the difference between the predicted QP (obtained in a) above) and the actual QP; and c) determining a luma QP value based on the bit depth, the predicted QP, and the cu_delta_QP.
[0112] The second step is to derive the chroma QP, which includes a) deriving the chroma QP from the luma QP, and b) finding the chroma QP offset from the PPS level offset (i.e., pps_cb_qp_offset, pps_cr_qp_offset) and slice level chroma QP offset (i.e., slice_cb_qp_offset, slice_cr_qp_offset).
[0113] Next, details of the process mentioned above.
[0114] the predicted luma quantization parameter qP Y_PRED is derived as follows:
[0115] [Number 6] qP Y_PRED =(qP Y_A +qP Y_B +1)>>1
[0116] The variable qP Y_A and the variable qP Y_B denotes the quantization parameter from the previous quantization group, where qP Y_A is set equal to the luma quantization parameter of the coding unit including the luma coding block covering (xQg-1, yQg), where the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. Y_A and the variable qP Y_B denotes the quantization parameter from the previous quantization group. Y_A is set equal to the luma quantization parameter of the coding unit including the luma coding block covering (xQg-1, yQg), where the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. Y_B is the luma quantization parameter Qp of the coding unit including the luma coding block covering (xQg, yQg-1). Y The variable qP is set to the same value as Y_A or the variable qP Y_B If qP is not available, Y_PREV where qP Y_PREV is the luma quantization parameter Qp of the last coding unit in decoding order in the previous quantization group. Y is set to the same as
[0117] qP Y_PRED Once determined, the luma quantization parameters are updated by adding CuQpDeltaVal as follows:
[0118] [Number 7] Qp Y =((qP Y_PRED +CuQpDeltaVal+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0119] Here, the CuQpDeltaVal value is transmitted in the bitstream via two syntax elements: cu_qp_delta_abs and cu_qp_delta_sign_flag. Y denotes the value of the luma quantization parameter range offset, which depends on bit_depth_luma_minus8 (i.e., luma bit depth - 8) as follows:
[0120] [Number 8] QpBdOffset Y =6*bit_depth_luma_minus8
[0121] Finally, the luma quantization parameter Qp′ Y is derived as follows:
[0122] [Number 9] Luma quantization parameter Qp′ Y =Qp Y +QpBdOffset Y
[0123] The chroma QP is derived from the luma QP taking into account the PPS level offsets (pps_cb_qp_offset, pps_cr_qp_offset) and slice level offsets (slice_cb_qp_offset, slice_cr_qp_offset) as follows:
[0124] [Number 10] qPiC b =Clip3(-QpBdOffset C , 57, Qp Y+pps_cb_qp_offset+slice_cb_qp_offset) qPiC r =Clip3(-QpBdOffset C , 57, Qp Y +pps_cr_qp_offset+slice_cr_qp_offset)
[0125] qPiC b and the qPiC r are calculated based on Tables 2 and 3 below. Cb and qP Cr will be updated to.
[0126] [Table 2]
[0127] Table 2 shows the qPi Cb From qP Cb Here is a mapping to
[0128] [Table 3]
[0129] Table 3 shows the qPi Cr From qP Cr Here is a mapping to
[0130] Finally, the Cb and Cr components, Qp′ Cb and Qp′ Cr The chroma quantization parameters for are derived as follows:
[0131] [Number 11] Qp′ Cb =qP Cb +QpBdOffset C Qp′ Cr =qP Cr +QpBdOffset C
[0132] QpBdOffset C indicates the value of the chroma quantization parameter range offset, which depends on bit_depth_chroma_minus8 (i.e., chroma bit depth - 8) as follows:
[0133] [Number 12] QpBdOffset C =6*bit_depth_chroma_minus8
[0134] Meanwhile, Table 4 shows definitions of syntax elements used in this specification.
[0135] [Table 4]
[0136] FIG. 6 is a flowchart illustrating a picture reconstruction process according to another embodiment.
[0137] S600, S610, and S630 to S670 correspond to S500 to S560 in FIG. 5, and therefore detailed explanations that overlap with the above explanations will be omitted or simplified.
[0138] According to one embodiment, a decoding device receives a bitstream (S600), performs entropy decoding (S610), performs inverse quantization (S630), determines whether to perform an inverse transform (S640), performs the inverse transform (S650), performs prediction (S660), and generates reconstructed samples (S670). The decoding device may also derive a QP offset based on the entropy decoding (S620), and perform inverse quantization based on the derived QP offset.
[0139] S620 can be expressed as follows: Hereinafter, the QP offset can be expressed as "Luma_avg_qp".
[0140] qP Y_PREDがOnce determined, the luma quantization parameter can be updated by adding Luma_avg_qp as follows:
[0141] [Number 13] Qp Y =((qP Y_PRED +CuQpDeltaVal+Luma_avg_qp+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0142] For example, Luma_avg_qp may be derived (or estimated) from the luma values of neighboring pixels (or blocks) that have already been decoded and are available. Luma_avg_qp may be determined from the neighboring pixel values based on a predefined derivation rule. For example, Luma_avg_qp may be derived as follows:
[0143] [Number 14] Luma_avg_qp=A*(avg_luma-M)+B
[0144] In Equation 14, avg_luma is the expected average luma value obtained from the available (decoded) neighboring pixels (or blocks).
[0145] M is a value predefined by the bit depth.
[0146] A is a scaling factor for mapping pixel value differences to qp differences (which can be predefined or transmitted in the bitstream) and indicates the slope of the qp mapping.
[0147] B is an offset value that can be predefined or transmitted in the bitstream.
[0148] The derivation of Luma_avg_qp from the avg_luma value is not limited to one of a number of formulas. As another example, Luma_avg_qp can be obtained from a table mapping as follows:
[0149] [Number 15] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma]
[0150] Here, avg_luma is input to the table and the output of the table is Luma_avg_qp. To reduce the size of the table, the input value (avg_luma) range can be further reduced as follows:
[0151] [Number 16] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma / D]
[0152] where D is a predefined constant value to reduce the range of input values.
[0153] In one embodiment, Luma_avg_qp may be derived based on information about the QP. The decoding device may obtain information about the QP from the bitstream. For example, the information about the QP may include init_qp_minus26, slice_qp_delta, slice_cb_qp_offset, slice_cr_qp_offset, cu_qp_delta_abs, and cu_qp_delta_sign_flag. The information about the QP is not limited to the above examples.
[0154] FIG. 7 is a flowchart showing the operation of the encoding device according to an embodiment, and FIG. 8 is a block diagram showing the configuration of the encoding device according to an embodiment.
[0155] Each step shown in Fig. 7 may be performed by the encoding apparatus 100 shown in Fig. 1. More specifically, steps S700 to S730 may be performed by the quantization module 123 shown in Fig. 1, and step S740 may be performed by the entropy encoding module 130 shown in Fig. 1. Furthermore, the operations of steps S700 to S740 are partially based on the detailed description of Fig. 6. Therefore, detailed descriptions that overlap with those described in Figs. 1 and 6 will be omitted or simplified.
[0156] As shown in Fig. 8, the encoding apparatus according to one embodiment may include a quantization module 123 and an entropy encoding module 130. However, in some cases, not all of the components shown in Fig. 8 may be required, and the encoding apparatus may be implemented with more or fewer components than those shown in Fig. 8.
[0157] In one embodiment, the quantization module 123 and the entropy encoding module 130 in the decoding device may be implemented in separate chips, or at least two or more components may be implemented in one chip.
[0158] According to an embodiment, the encoding apparatus may derive a predicted average luma value of the current block from available neighboring samples (S700). More specifically, the quantization module 123 of the encoding apparatus may derive a predicted average luma value of the current block from available neighboring samples.
[0159] According to an embodiment, the encoding apparatus may derive a QP offset for deriving a luma QP based on the predicted average luma value and information on the QP (S710). More specifically, the quantization module 123 of the encoding apparatus may derive a QP offset for deriving a luma QP based on the predicted average luma value and information on the QP.
[0160] According to an embodiment, the encoding apparatus may derive a luma QP based on the QP offset (S720). More specifically, the quantization module 123 of the encoding apparatus may derive a luma QP based on the QP offset.
[0161] According to an embodiment, the encoding apparatus may perform quantization on a quantization group including the current block based on the derived luma QP (S730). More specifically, the quantization module 123 of the encoding apparatus may perform quantization on a quantization group including the current block based on the derived luma QP.
[0162] According to an embodiment, the encoding apparatus may encode video information including information on a QP (S740). More specifically, the entropy encoding module 130 may encode video information including information on a QP.
[0163] 7 and 8, the encoding apparatus according to one embodiment derives a predicted average luma value of a current block from available neighboring samples (S700), derives a QP offset for deriving a luma QP based on the predicted average luma value and information on QP (S710), derives a luma QP based on the QP offset (S720), performs quantization on a quantization group including the current block based on the derived luma QP (S730), and encodes video information including information on QP (S740). As a result, quantization parameters can be derived efficiently, and overall coding efficiency can be improved.
[0164] FIG. 9 is a flowchart showing the operation of a decoding device according to an embodiment, and FIG. 10 is a block diagram showing the configuration of a decoding device according to an embodiment.
[0165] Each step shown in Figure 9 may be performed by the decoding apparatus 200 shown in Figure 2. More specifically, S900 may be performed by the entropy decoding module 210 shown in Figure 2, S910 to S940 may be performed by the inverse quantization module 222 shown in Figure 2, S950 may be performed by the inverse transform module 223 shown in Figure 2, S960 may be performed by the prediction module 230 shown in Figure 2, and S970 may be performed by the adder 240 shown in Figure 2. Furthermore, the operations of S900 to S970 are partially based on the detailed description of Figure 6. Therefore, detailed description that overlaps with the contents described in Figures 2 and 6 will be omitted or simplified.
[0166] 10, the decoding apparatus according to one embodiment may include an entropy decoding module 210, an inverse quantization module 222, an inverse transform module 223, a prediction module 230, and an adder 240. However, in some cases, not all of the components shown in FIG. 10 may be required, and the encoding apparatus may be implemented with more or fewer components than those shown in FIG.
[0167] In one embodiment of the decoding device, the entropy decoding module 210, the inverse quantization module 222, the inverse transform module 223, the prediction module 230, and the adder 240 may be implemented on separate chips, or at least two or more components may be implemented on a single chip.
[0168] According to an embodiment, the decoding device may decode video information including information on a QP (S900). More specifically, the entropy decoding module 210 of the decoding device may decode video information including information on a QP.
[0169] In one embodiment, information on the QP is signaled at the Sequence Parameter Set (SPS) level.
[0170] In one embodiment, the image information includes information on Effective Data Range Parameters (EDRP), and the information on the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of input values, mapping information for relating a minimum input value to brightness, mapping information for relating a maximum input value to brightness, and identification information of a transfer function. This method can be represented by range matching.
[0171] More specifically, the present disclosure can be used to efficiently code image / video content with a limited codeword (input value) range. This often occurs with HDR content because HDR content uses a transfer function that supports high luminance. This can also occur when converting SDR data using a luminance conversion function corresponding to HDR data. In such cases, a video encoder can be configured to signal an Valid Data Range Parameter (EDRP). A decoder can be configured to receive the EDRP associated with video data and use the EDRP data in the decoding process. The EDRP data can include, for example, a minimum input value, a maximum input value, a dynamic range of input values (indicating the difference between the maximum and minimum input values), mapping information between the minimum input value and its corresponding brightness, mapping information between the maximum input value and its corresponding brightness, transfer function identification (known transfer functions can be identified by assigned ID numbers, and detailed mapping information for each transfer function can be used), etc.
[0172] For example, the EDRP data can be signaled in a slice header, a picture parameter set (PPS), or a sequence parameter set (SPS). In this manner, the EDRP data can be used to further modify coded values during the decoding process.
[0173] The present invention introduces a quality control parameter (QCP) to specify additional adjustments to the quantization parameters. A decoder can be configured to receive the QCP associated with the video data and utilize the QCP data in the decoding process.
[0174] According to an embodiment, the decoding device may derive a predicted average luma value of the current block from available neighboring samples (S910). More specifically, the inverse quantization module 222 of the decoding device may derive a predicted average luma value of the current block from available neighboring samples.
[0175] According to an embodiment, the decoding device may derive a QP offset for deriving a luma QP based on the predicted average luma value and information on the QP (S920). More specifically, the inverse quantization module 222 of the decoding device may derive a QP offset for deriving a luma QP based on the predicted average luma value and information on the QP.
[0176] In one embodiment, the QP offset is derived based on the following formula:
[0177] [Number 17] Luma_avg_qp=A*(avg_luma-M)+B
[0178] In Equation 17, Luma_avg_qp represents a QP offset, avg_luma represents an expected average luma value, A represents a scaling coefficient for mapping a pixel value difference to a QP difference, M represents a predefined value associated with bit depth, B represents an offset value, and A and B are predetermined values or values included in the image information.
[0179] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, the mapping table being determined using the expected average luma value as an input.
[0180] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, and the mapping table is determined using a value obtained by dividing the expected average luma value by a predefined constant value.
[0181] In one embodiment, the available neighboring samples include at least one of at least one luma sample adjacent to the left boundary of the quantization group and at least one luma sample adjacent to the top boundary of the quantization group.
[0182] In one embodiment, at least one luma sample adjacent to the left boundary of the quantization group is included in a luma sample column immediately adjacent to the left boundary of the quantization group, and at least one luma sample adjacent to the top boundary of the quantization group is included in a luma sample row immediately adjacent to the top boundary of the quantization group.
[0183] In one embodiment, the available neighboring samples include the luma sample adjacent to the left of the top-left sample of the quantization group, and the available neighboring samples include the luma sample adjacent to the top-left sample of the quantization group.
[0184] In one embodiment, the available neighboring samples include at least one of the reconstructed neighboring samples, samples included in at least one of the reconstructed neighboring blocks, predicted neighboring samples, and samples included in at least one of the predicted neighboring blocks.
[0185] In one embodiment, avg_luma can be derived from neighboring pixel values (blocks), where avg_luma indicates the expected average luma value obtained from available (already decoded) neighboring pixels (or blocks).
[0186] i) The available neighboring pixels may include:
[0187] The available neighboring pixels include the pixel located at (xQg-1, yQg+K), where the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture (the leftmost line of the current block).
[0188] The available neighboring pixels include the pixel located at (xQg+K, yQg-1), where the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture (the top-most line of the current block).
[0189] Instead of one line, multiple lines can be used.
[0190] ii) avg_luma can be calculated using available neighboring blocks.
[0191] A block containing a pixel located at -(xQg-1, yQg) can be used.
[0192] A block containing a pixel located at -(xQg, yQg-1) can be used.
[0193] iii) The Avg_luma value can be calculated based on the reconstructed neighboring pixels / blocks.
[0194] iv) The Avg_luma value can be calculated based on the predicted neighboring pixels / blocks.
[0195] In one embodiment, the information on the QP includes at least one syntax element associated with a QP offset, and the step of deriving the QP offset based on at least one of the predicted average luma value and the information on the QP includes the step of deriving the QP offset based on the at least one syntax element associated with the QP offset.
[0196] A new syntax element can be introduced to take the luma_avg_qp into account. For example, the luma_avg_qp value can be transmitted via the bitstream. The luma_avg_qp value can be indicated via two syntax elements, such as luma_avg_qp_abs and luma_avg_qp_flag.
[0197] Luma_avg_qp can be indicated by two syntax elements: luma_avg_qp_delta_abs and luma_avg_qp_delta_sign_flag.
[0198] luma_avg_qp_delta_abs indicates the absolute value of CuQpDeltaLumaVal, which is the difference between the luma quantization parameter of the current coding unit and the luma quantization parameter derived without considering luma.
[0199] luma_avg_qp_delta_sign_flag indicates the sign of CuQpDeltaLumaVal as follows:
[0200] If luma_avg_qp_delta_sign_flag is 0, the corresponding CuQpDeltaLumaVal is a positive value.
[0201] Otherwise (luma_avg_qp_delta_sign_flag is 1), the corresponding CuQpDeltaLumaVal is a negative value.
[0202] If luma_avg_qp_delta_sign_flag is not present, it is inferred to be equal to 0.
[0203] If luma_avg_qp_delta_sign_flag is present, the variables IsCuQpDeltaLumaCoded and CuQpDeltaLumaVal are derived as follows:
[0204] [Number 18] IsCuQpDeltaLumaCoded=1 CuQpDeltaLumaVal=cu_qp_delta_abs*(1-2*luma_avg_qp_delta_sign_flag)
[0205] CuQpDeltaLumaVal indicates the difference between the luma quantization parameter for a coding unit that includes Luma_avg_qp and the luma quantization parameter for a coding unit that does not include Luma_avg_qp.
[0206] In one embodiment, the above-described syntax elements may be transmitted at a quantization group level (or a quantization unit level) (eg, CU, CTU, or a predefined block unit).
[0207] According to an embodiment, the decoding device may derive a luma QP based on the QP offset (S930). More specifically, the inverse quantization module 222 of the decoding device may derive a luma QP based on the QP offset.
[0208] According to an embodiment, the decoding device may perform inverse quantization on a quantization group including the current block based on the derived luma QP (S940). More specifically, the inverse quantization module 222 of the decoding device may perform inverse quantization on a quantization group including the current block based on the derived luma QP.
[0209] In one embodiment, the decoding device may derive a chroma QP from the derived luma QP based on at least one chroma QP mapping table, and perform inverse quantization for a quantization group based on the derived luma QP and the derived chroma QP, where the at least one chroma QP mapping table is based on a chroma dynamic range and the QP offset.
[0210] In one embodiment, instead of one chroma QP mapping table, there can be multiple chroma QP derivation tables. Additional information is required to identify which QP mapping table to use. The following is another example of a chroma QP derivation table:
[0211] [Table 5]
[0212] In one embodiment, the at least one chroma QP mapping table includes at least one Cb-QP mapping table and at least one Cr-QP mapping table. According to one embodiment, the decoding device may generate residual samples for the current block based on inverse quantization (S950). More specifically, the inverse transform module 223 of the decoding device may generate residual samples for the current block based on inverse quantization.
[0213] According to an embodiment, the decoding device may generate a prediction sample for the current block based on the image information (S960). More specifically, the prediction module 230 of the decoding device may generate a prediction sample for the current block based on the image information.
[0214] According to an embodiment, the decoding device may generate reconstructed samples for the current block based on residual samples for the current block and predicted samples for the current block (S970). More specifically, the adder of the decoding device may generate reconstructed samples for the current block based on residual samples for the current block and predicted samples for the current block.
[0215] 9 and 10, the decoding apparatus according to one embodiment decodes video information including information on QP (S900), derives a predicted average luma value of a current block from available neighboring samples (S910), derives a QP offset for deriving a luma QP based on the predicted average luma value and information on QP (S920), derives a luma QP based on the QP offset (S930), performs inverse quantization on a quantization group including the current block based on the derived luma QP (S940), generates residual samples for the current block based on the inverse quantization (S950), generates predicted samples for the current block based on the video information (S960), and generates reconstructed samples for the current block based on the residual samples for the current block and the predicted samples for the current block (S970). As a result, quantization parameters can be derived efficiently, and overall coding efficiency can be improved.
[0216] The above-described method according to the present invention can be implemented in software, and the encoding device and / or the decoding device according to the present invention can be included in a video processing device such as a TV, a computer, a smartphone, or a display device.
[0217] When an embodiment of the present invention is implemented in software, the above-described methods may be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor by various known means. The processor may include an application specific integrated circuit (ASIC), other chipset, logic circuit, and / or data processing device. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory card, storage medium, and / or other storage device.
Claims
1. 1. A picture decoding method performed by a decoding device, comprising: obtaining video information including information related to a quantization parameter (QP) from a bitstream, the information related to the QP includes information related to a luma QP; the information related to the luma QP includes information about an absolute value of a QP delta value and information about a sign of the QP delta value; deriving the luma QP based on the information related to the luma QP; deriving chroma QPs including a first chroma QP for a Cb component and a second chroma QP for a Cr component; generating a reconstructed picture including reconstructed samples based on the luma QP, the first chroma QP for the Cb component, and the second chroma QP for the Cr component; deriving the luma QP based on the information related to the luma QP, deriving a predicted luma QP based on a first luma QP of a coding unit that includes a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit that includes a luma coding block covering (xQg, yQg-1), where (xQg, yQg) indicates a top-left sample position of a current quantization group; deriving the QP delta value based on the information regarding the absolute value of the QP delta value and the information regarding the sign of the QP delta value; deriving the luma QP based on the predicted luma QP and the QP delta value; The step of deriving the chroma QP comprises: deriving the first chroma QP for the Cb component based on the luma QP and a first chroma QP mapping table for the Cb component; deriving the second chroma QP for the Cr component based on the luma QP and a second chroma QP mapping table for the Cr component; both the entries of the first chroma QP mapping table and the entries of the second chroma QP mapping table are based on the luma QP; the first chroma QP mapping table for the Cb component and the second chroma QP mapping table for the Cr component exist separately, A picture decoding method, wherein the first chroma QP mapping table for the Cb component has different table values than the second chroma QP mapping table for the Cr component.
2. The picture decoding method of claim 1 , wherein the information related to the QP is signaled at a Sequence Parameter Set (SPS) level.
3. the video information includes information about EDRP (Effective Data Range Parameters); 2. The picture decoding method of claim 1, wherein the information about the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of input values, mapping information for relating the minimum input value to brightness, mapping information for relating the maximum input value to brightness, and identification information of a transfer function.
4. 1. A picture encoding method performed by an encoding device, comprising: deriving a luma quantization parameter (QP) for the current block; deriving chroma QPs, the chroma QPs including a first chroma QP for a Cb component and a second chroma QP for a Cr component for the current block; generating information related to a QP; encoding video information including the information related to the QP; the information related to the QP includes information related to the luma QP and information related to the chroma QP; the information related to the luma QP includes information about an absolute value of a QP delta value and information about a sign of the QP delta value; the luma QP is indicated based on a predicted luma QP and the QP delta value; the predicted luma QP is derived based on a first luma QP of a coding unit including a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit including a luma coding block covering (xQg, yQg-1); (xQg, yQg) indicates the upper left sample position of the current quantization group, The step of deriving the chroma QP comprises: deriving the first chroma QP for the Cb component based on the luma QP and a first chroma QP mapping table for the Cb component; deriving the second chroma QP for the Cr component based on the luma QP and a second chroma QP mapping table for the Cr component; both the entries of the first chroma QP mapping table and the entries of the second chroma QP mapping table are based on the luma QP; the first chroma QP mapping table for the Cb component and the second chroma QP mapping table for the Cr component exist separately, A picture encoding method, wherein the first chroma QP mapping table for the Cb component has different table values than the second chroma QP mapping table for the Cr component.
5. The picture encoding method of claim 4 , wherein the information related to the QP is signaled at a Sequence Parameter Set (SPS) level.
6. the video information includes information about EDRP (Effective Data Range Parameters); 5. The picture encoding method of claim 4, wherein the information about the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of input values, mapping information for relating the minimum input value to brightness, mapping information for relating the maximum input value to brightness, and identification information of a transfer function.
7. 1. A transmission method for data including a bitstream relating to a picture, comprising: generating the bitstream for the picture, the bitstream comprising: deriving a luma quantization parameter (QP) for the current block; deriving chroma QPs, the chroma QPs including a first chroma QP for a Cb component and a second chroma QP for a Cr component for the current block; generating information related to a QP; encoding video information including the information related to the QP; transmitting the data including the bitstream; the information related to the QP includes information related to the luma QP and information related to the chroma QP; the information related to the luma QP includes information about an absolute value of a QP delta value and information about a sign of the QP delta value; the luma QP is indicated based on a predicted luma QP and the QP delta value; the predicted luma QP is derived based on a first luma QP of a coding unit including a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit including a luma coding block covering (xQg, yQg-1); (xQg, yQg) indicates the upper left sample position of the current quantization group, The step of deriving the chroma QP comprises: deriving the first chroma QP for the Cb component based on the luma QP and a first chroma QP mapping table for the Cb component; deriving the second chroma QP for the Cr component based on the luma QP and a second chroma QP mapping table for the Cr component; both the entries of the first chroma QP mapping table and the entries of the second chroma QP mapping table are based on the luma QP; the first chroma QP mapping table for the Cb component and the second chroma QP mapping table for the Cr component exist separately, The transmission method, wherein the first chroma QP mapping table for the Cb component has different table values than the second chroma QP mapping table for the Cr component.
8. The transmission method of claim 7 , wherein the information related to the QP is signaled at a Sequence Parameter Set (SPS) level.
9. the video information includes information about EDRP (Effective Data Range Parameters); 8. The transmission method of claim 7, wherein the information about the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of input values, mapping information for relating the minimum input value to brightness, mapping information for relating the maximum input value to brightness, and identification information of a transfer function.
Citation Information
Patent Citations
Method for determining color difference component quantization parameter and device using the method
US20130329785A1
Systems and methods for optimizing video coding based on a luminance transfer function or video color component values
WO2016199409A1
Signaling of quantization information in non-quadtree-only partitioned video coding
WO2018013706A1