Video coding apparatus and method based on quantization parameter derivation
By deriving quantization parameters based on expected luma values and offsets, the method improves video coding efficiency, addressing the inefficiencies in handling high-resolution videos and reducing transmission and storage costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-08-07
- Publication Date
- 2026-05-26
Smart Images

Figure 0007866126000006 
Figure 0007866126000007 
Figure 0007866126000008
Abstract
Description
Technical Field
[0001] The present invention relates to video coding (video encoding) technology. More specifically, the present invention relates to a video coding apparatus and method based on quantization parameter derivation in a video coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as high-definition (HD) videos and ultra-high-definition (UHD) videos has been increasing in various fields. Since such video data has high resolution and high quality, the amount of information or bits to be transmitted increases compared to existing image data. Therefore, when transmitting or storing video data using a medium such as a wired / wireless broadband line, the transmission cost and storage cost can be increased.
[0003] Therefore, there is a need for a highly efficient video compression technology for efficiently transmitting, storing, and playing back information of high-resolution and high-quality videos.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present invention provides a method and apparatus for improving video coding efficiency.
[0005] The present invention also provides a method and apparatus for increasing quantization efficiency.
[0006] The present invention also provides a method and apparatus for efficiently deriving quantization parameters.
Means for Solving the Problems
[0007] According to one embodiment of the present invention, a picture decoding method is provided that is performed by a decoding device. The picture decoding method includes the steps of: decoding image information including information for quantization parameters (QP); deriving an expected average luma value for the current block from available neighboring samples; deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and the information for QP; deriving a luma QP based on the QP offset; performing inverse quantization on a quantization group including the current block based on the derived luma QP; generating a residual sample for the current block based on the inverse quantization; generating a predicted sample for the current block based on the image information; and generating a restored sample for the current block based on the residual sample for the current block and the predicted sample for the current block.
[0008] According to one embodiment of the present invention, a decoding device for decoding a picture is provided. The decoding device includes: an entropy decoding module configured to decode image information including information for quantization parameters (QP); an inverse quantization module configured to derive an expected mean luma value for the current block from available neighboring samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected mean luma value and the information for QP, derive a luma QP based on the QP offset, and perform inverse quantization on a quantization group including the current block based on the derived luma QP; an inverse transform module configured to generate a residual sample for the current block based on the inverse quantization; a prediction module configured to generate a predicted sample for the current block based on the image information; and a restoration module configured to generate a restored sample for the current block based on the residual sample for the current block and the predicted sample for the current block.
[0009] According to one embodiment of the present invention, a picture encoding method is provided that is performed by an encoding device. The picture encoding method includes the steps of: deriving an expected average luma value for the current block from available neighboring samples; deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and information for QP; deriving a luma QP based on the QP offset; performing quantization on a quantization group including the current block based on the derived luma QP; and encoding video information including information for QP.
[0010] According to one embodiment of the present invention, an encoding device for encoding a picture is provided. The encoding device includes a quantization module configured to derive an expected average luma value of a current block from available neighboring samples, derive a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and information for QP, derive a luma QP based on the QP offset, and perform quantization on a quantization group including the current block based on the derived luma QP; and an entropy encoding module configured to encode image information including information for QP. [Effects of the Invention]
[0011] According to the present invention, the overall video compression efficiency can be increased.
[0012] According to the present invention, the quantization efficiency can be increased.
[0013] According to the present invention, quantization parameters can be efficiently derived. [Brief explanation of the drawing]
[0014] [Figure 1] This is a schematic diagram showing the configuration of an encoding device according to one embodiment. [Figure 2] This is a schematic diagram showing the configuration of a decoding device according to one embodiment. [Figure 3] An example of a chromaticity diagram is shown. [Figure 4] Examples of linear light value mapping for SDR and HDR representation are shown. [Figure 5] This is a flowchart showing the picture restoration process according to one embodiment. [Figure 6] This flowchart shows the picture restoration process according to another embodiment. [Figure 7]This is a flowchart showing the operation of an encoding device according to one embodiment. [Figure 8] Block diagram showing the configuration of an encoding device according to one embodiment. [Figure 9] This is a flowchart showing the operation of a decoding device according to one embodiment. [Figure 10] This is a block diagram showing the configuration of a decoding device according to one embodiment. [Modes for carrying out the invention]
[0015] According to one embodiment of the present invention, a picture decoding method is provided that is performed by a decoding device. The picture decoding method includes the steps of: decoding video information including information for quantization parameters (QP); deriving an expected average luma value for the current block from available neighboring samples; deriving a quantization parameter offset (QP offset) for deriving a luma quantization parameter (luma QP) based on the expected average luma value and the information for QP; deriving a luma QP based on the QP offset; performing inverse quantization on a quantization group including the current block based on the derived luma QP; generating a residual sample for the current block based on the inverse quantization; generating a predicted sample for the current block based on the video information; and generating a restored sample for the current block based on the residual sample for the current block and the predicted sample for the current block.
[0016] The present invention can be modified in various forms, and specific embodiments are illustrated in the drawings and described in detail. However, these embodiments are not intended to limit the present invention. Terms used herein are used solely to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless otherwise clearly defined. In this specification, the terms “includes” or “have” mean that there are features, figures, steps, actions, components, parts or combinations thereof described in the specification, and should be understood not to presuppose the presence or addition of one or more different features, figures, steps, actions, components, parts or combinations thereof.
[0017] On the other hand, each component shown in the drawings described in this invention is illustrated independently for the convenience of illustrating the distinct characteristic functions of the video coding / decoding device, and does not imply that each component is implemented in separate hardware or separate software. For example, two or more components may be integrated into a single component, and a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the invention as long as they do not deviate from the essence of the invention.
[0018] Preferred embodiments of the present invention will be described in detail below with reference to the attached drawings. The same reference numerals are used for the same components throughout the drawings, and redundant descriptions of the same components are omitted.
[0019] This specification can be applied to the technical field of video or image processing. For example, the methods or embodiments disclosed in the following description can be applied to a variety of video coding standards, such as the Versatile Video Coding (VVC) standard (ITU-T Rec.H.266), post-VVC next-generation video / image coding standards, or pre-VVC video / image coding standards such as the High Efficiency Video Coding (HEVC) standard (ITU-T Rec.H.265).
[0020] In this specification, "video" can mean a set of images over time. "Picture" generally refers to a unit representing a single image at a specific point in time, and "slice" is a unit that constitutes a part of a picture in coding. A single picture can consist of multiple slices, and pictures and slices can be used together as needed. Furthermore, the term "image" can encompass the concept of still images and video, which is a set of still images over time. Also, "video" does not necessarily mean a set of still images over time, and in some embodiments, it can be interpreted as a concept that includes the meaning of still images.
[0021] A pixel or pel can refer to the smallest unit of a picture (or image). The term 'sample' can also be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and can represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value.
[0022] A unit represents a basic unit of image processing. A unit can contain at least one of the following: a specific area of a picture and information associated with that area. The unit may be used interchangeably with terms such as block or area. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows.
[0023] Figure 1 schematically illustrates the configuration of an encoding device according to one embodiment.
[0024] Hereinafter, encoding / decoding devices may include video encoding / decoding devices and / or image encoding / decoding devices. The term "video encoding / decoding device" can be used as a concept that includes image encoding / decoding devices, and the term "image encoding / decoding device" can be used as a concept that includes video encoding / decoding devices.
[0025] Referring to Figure 1, the encoding device 100 may include a picture splitting module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filter module 150, and a memory 160. The residual processing module 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, an inverse quantization module 125, and an inverse transform module 126.
[0026] The picture splitting module 105 can split an input picture into at least one processing unit.
[0027] For example, the processing unit is called a coding unit (CU). In this case, the coding unit can be recursively divided from the largest coding unit (LCU) by a quad-tree binary-tree (QTBT) structure. For example, a single coding unit can be divided into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure.
[0028] In this case, for example, a quad-tree structure may be applied first, followed by a binary-tree structure and a ternary-tree structure. Alternatively, a binary-tree structure / ternary-tree structure may be applied first. The coding procedure according to the present invention can be performed based on a final coding unit that is not further subdivided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency due to image characteristics, or a coding unit may be recursively subdivided into coding units of deeper depth and used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and reconstruction, which will be described later.
[0029] As another example, the processing unit may include a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit can be divided into deeper coding units by a quad-tree structure from the largest coding unit (LCU). In this case, the largest coding unit can be used as the final coding unit based on coding efficiency due to video characteristics, or the coding unit can be recursively divided into deeper coding units and used as the final coding unit. If a smallest coding unit (SCU) is set, the coding unit cannot be divided into coding units smaller than the smallest coding unit.
[0030] Here, the term “final coding unit” means a coding unit that is partitioned or divided into the prediction unit or the conversion unit. The prediction unit is a unit partitioned from the coding unit and is a sample prediction unit. In this case, the prediction unit may also be divided into sub-blocks. The conversion unit may be divided from the coding unit by a quad-tree structure and is a unit that derives conversion coefficients and / or a unit that derives a residual signal from the conversion coefficients.
[0031] Hereinafter, the coding unit will be referred to as a coding block (CB), the prediction unit as a prediction block (PB), and the transformation unit as a transform block (TB). The prediction block or prediction unit may represent a specific region in block form within a picture and may include an array of prediction samples. The transform block or transform unit may also represent a specific region in block form within a picture and may include an array of transformation coefficients or residual samples.
[0032] The prediction module 110 predicts the current block or residual block and generates a prediction block containing prediction samples of the current block. The unit of prediction performed by the prediction module 110 is a coding block, or a transformation block, or a prediction block.
[0033] The prediction module 110 predicts the current block or residual block and generates a prediction block containing prediction samples of the current block. The unit of prediction performed by the prediction module 110 is a coding block, or a transformation block, or a prediction block.
[0034] The prediction module 110 can determine whether intra-prediction or inter-prediction is applied to the current block. For example, the prediction module 110 can determine whether intra-prediction or inter-prediction is applied on a CU basis.
[0035] In the case of intra-prediction, the prediction module 110 can derive prediction samples for the current block based on reference samples outside the current block within the picture to which the current block belongs (hereinafter referred to as the current picture).
[0036] In this case, the prediction module 110 can (i) derive a prediction sample based on the average or interpolation of neighboring reference samples in the current block, and (ii) derive a prediction sample based on reference samples in the neighboring reference samples of the current block that exist in a specific (prediction) direction relative to the prediction sample.
[0037] (i) is called a non-directional mode or non-angular mode, and (ii) is called a directional mode or angular mode. In intraprediction, the prediction modes may have, for example, 65 directional prediction modes and at least 2 or more non-directional modes. Non-directional modes may include DC prediction modes and planar modes. The prediction module 110 can determine the prediction mode to be applied to the current block by utilizing the prediction modes applied to adjacent blocks.
[0038] In the case of interpretation, the prediction module 110 can derive prediction samples for the current block based on samples identified by motion vectors on a reference picture. The prediction module 110 can derive prediction samples for the current block by applying one of the following modes: skip mode, merge mode, and MVP (motion vector prediction). In skip mode and merge mode, the prediction module 110 can use motion information of adjacent blocks as motion information for the current block.
[0039] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted. In MVP mode, the motion vector of the current block can be derived by using the motion vector of the adjacent block as the motion vector predictor for the current block.
[0040] In interpretation, adjacent blocks can include spatially adjacent blocks currently present in the picture and temporally adjacent blocks present in the reference picture. The reference picture containing the temporally adjacent blocks can also be called a collocated picture (colPic). The motion information can include motion vectors and reference picture indices. Information such as prediction mode information and motion information can be (entropy) encoded and output in bitstream form.
[0041] When motion information of temporally adjacent blocks is used in skip mode and merge mode, the topmost picture on the reference picture list can be used as the reference picture. The reference picture included in the Picture Order Count (POC) can be sorted based on the POC difference between the current picture and the corresponding reference picture. The POC corresponds to the display order of the pictures and can be distinguished from the coding order.
[0042] The subtractor 121 generates a residual sample, which is the difference between the original sample and the predicted sample. When skip mode is applied, it does not generate a residual sample as described above.
[0043] The transformation module 122 transforms the residual samples in units of transformation blocks to generate transformation coefficients. The transformation module 122 can perform the transformation based on the size of the transformation block and the prediction mode applied to the coding block or prediction block that spatially overlaps with the transformation block.
[0044] For example, if an intra prediction is applied to a coding block or prediction block that overlaps with the transformation block, and the transformation block is a 4x4 resistive array, the resistive samples can be transformed using a discrete sine transform (DST); otherwise, the resistive samples can be transformed using a discrete cosine transform (DCT).
[0045] The quantization module 123 can quantize the conversion coefficients and generate quantized conversion coefficients.
[0046] The realignment module 124 realigns the quantized transformation coefficients. The realignment module 124 can realign the block-type quantized transformation coefficients into a one-dimensional vector form via a coefficient scanning method. Although the realignment module 124 has been described in a separate configuration, the realignment module 124 may be part of the quantization module 123.
[0047] The entropy encoding module 130 can perform entropy encoding on quantized transformation coefficients. Entropy encoding can include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding module 130 can encode the quantized transformation coefficients and information necessary for video reconstruction (e.g., the values of syntax elements) together or separately by entropy encoding or a predetermined method.
[0048] The information encoded by entropy encoding can be transmitted or stored in bitstream form in network abstraction layer (NAL) units. The bitstream can be transmitted over a network or stored on a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include a variety of storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0049] The inverse quantization module 125 inversely quantizes the values (quantized conversion coefficients) quantized by the quantization module 123, and the inverse conversion module 126 inversely converts the values inversely quantized by the inverse quantization module 125 to generate residual samples.
[0050] The adder 140 combines the residual sample and the predicted sample to reconstruct the picture. The residual sample and the predicted sample can be added in block units to generate a reconstructed block. Here, the adder 140 may be part of the prediction module 110. On the other hand, the adder 140 is also called the reconstructed module or the reconstructed block generation unit.
[0051] The filter module 150 can apply a deblocking filter and / or a sample adaptive offset to the reconstructed picture. Through deblocking filtering and / or the sample adaptive offset, artifacts at block boundaries within the reconstructed picture or distortions during the quantization process can be corrected. The sample adaptive offset can be applied on a sample-by-sample basis and can be applied after the deblocking filtering process is complete. The filter module 150 can also apply an Adaptive Loop Filter (ALF) to the reconstructed picture. The ALF can be applied to the reconstructed picture after the deblocking filter and / or the sample adaptive offset has been applied.
[0052] The memory 160 can store the restored picture (decoded picture) or information necessary for encoding / decoding. Here, the restored picture is a restored picture that has undergone a filtering procedure by the filter module 150. The stored restored picture can be used as a reference picture for (inter)prediction of other pictures. For example, the memory 160 can store a (reference) picture used for interprediction. In this case, the picture used for interprediction can be specified by a reference picture set or a reference picture list.
[0053] Figure 2 schematically illustrates the configuration of the decoding apparatus according to the embodiment.
[0054] Referring to Figure 2, the decoding device 200 includes an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filter module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, an inverse quantization module 222, and an inverse transform module 223. Although not shown, the video decoding device 200 may also include a receiver for receiving a bitstream containing video information. The receiver may be a separate module or may be included in the entropy decoding module 210.
[0055] When a bitstream containing video / image information is input, the (video) decoding device 200 can restore the video / image / picture in response to the process by which the video / image information is processed in the (video) encoding device 100.
[0056] For example, the video decoding device 200 can perform video decoding using processing units applied to the video encoding device. Therefore, a video decoding processing unit block is, for example, a coding unit, and other examples include a coding unit, a prediction unit, or a transformation unit. The coding unit can be divided from a maximum coding unit into a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure.
[0057] Prediction units and conversion units may also be used as appropriate. The prediction unit block is a block derived from or partitioned from the coding unit. The conversion unit can be partitioned from the coding unit by a quad-tree structure and is a unit that derives conversion coefficients or a unit that derives a residual signal from conversion coefficients.
[0058] The entropy decoding module 210 can parse the bitstream and output information necessary for video or picture restoration. For example, the entropy decoding module 210 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and calculates the values of syntax elements and quantized transformation coefficients associated with residuals necessary for video restoration.
[0059] More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model based on the syntax element to be decoded and the decoding information to be decoded, predicts the probability of a bin occurring based on the determined context model, and generates a symbol corresponding to the value of each syntax element by performing arithmetic decoding of the bins. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.
[0060] In the entropy decoding module 210, the information for prediction from the decoded information is provided to the prediction module 230, and the residual values for which entropy decoding has been performed in the entropy decoding module 210 can be input to the realignment module 221.
[0061] The realignment module 221 can realign the quantized conversion coefficients in a two-dimensional block form. The realignment module 221 can perform realignment in response to coefficient scanning performed by the encoding device. Although the realignment module 221 has been described in a separate configuration, the realignment module 221 may be part of the inverse quantization module 222.
[0062] The inverse quantization module 222 can output the inversely quantized conversion coefficients by inversely quantizing the quantized conversion coefficients based on the (inverse) quantization parameters. At this time, information for deriving the quantization parameters can be signaled from the encoding device.
[0063] The inverse transformation module 223 can derive the residual sample by inversely transforming the transformation coefficients.
[0064] The prediction module 230 can predict the current block and generate a prediction block containing prediction samples of the current block. The unit of prediction performed by the prediction module 230 is a coding block, a transformation block, or a prediction block.
[0065] The prediction module 230 can decide whether to apply intra-prediction or inter-prediction based on the prediction information. In this case, the unit for deciding whether to apply intra-prediction or inter-prediction is different from the unit for generating prediction samples. Also, the unit for generating prediction samples is different for inter-prediction and intra-prediction. For example, the decision on whether to apply inter-prediction or intra-prediction can be made in units of CU (Unit). Alternatively, for example, in inter-prediction, the prediction mode can be determined in units of PU (Phone Unit), and prediction samples can be generated. In intra-prediction, the prediction mode can also be determined in units of PU, and prediction samples can be generated in units of TU (Unit).
[0066] In the case of intra-prediction, the prediction module 230 can derive prediction samples for the current block based on adjacent reference samples in the current picture. The prediction module 230 can derive prediction samples for the current block by applying a directional or non-directional mode based on adjacent reference samples of the current block. In this case, the prediction mode to be applied to the current block can be determined using the intra-prediction mode of the adjacent block.
[0067] In the case of interpretation, the prediction module 230 can derive prediction samples for the current block based on samples identified on the reference picture by motion vectors on the reference picture. The prediction module 230 can derive prediction samples for the current block by applying skip mode, merge mode, or MVP mode. At this time, information on motion information necessary for interpretation of the current block provided to the encoding device, such as motion vectors and reference picture indices, can be obtained or derived based on the prediction information.
[0068] In skip mode and merge mode, the movement information of adjacent blocks can be used as the movement information of the current block. In this case, adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.
[0069] The prediction module 230 constructs a merge candidate list using motion information of available adjacent blocks, and can use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information can include motion vectors and reference pictures. When motion information of temporally adjacent blocks is used in skip mode and merge mode, the top-level picture on the reference picture list can be used as the reference picture.
[0070] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not transmitted.
[0071] In MVP mode, the motion vector of the current block can be derived by using the motion vector of the adjacent block as a motion vector predictor. In this case, the adjacent block can include both spatially adjacent blocks and temporally adjacent blocks.
[0072] For example, when merge mode is applied, a merge candidate list can be generated by using the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent Col block. In merge mode, the motion vectors of the candidate blocks selected from the merge candidate list are used as the motion vectors of the current block. The prediction information may include a merge index that indicates the candidate block having the optimal motion vector selected from among the candidate blocks included in the merge candidate list. In this case, the prediction module 230 can derive the motion vector of the current block by using the merge index.
[0073] As another example, when MVP (Motion Vector Prediction) mode is applied, a list of motion vector predictor candidates is generated by utilizing the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col. That is, the motion vectors of the restored spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent block, Col, can be used as motion vector candidates. The prediction information may include a predicted motion vector index that indicates the optimal motion vector selected from the motion vector candidates included in the list.
[0074] At this time, the prediction module 230 can select the predicted motion vector of the current block from among the motion vector candidates included in the motion vector candidate list by using the motion vector index. The prediction module of the encoding device can calculate the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, and can output the MVD in bitstream format. That is, the MVD is obtained by subtracting the motion vector predictor from the motion vector of the current block. In this case, the prediction module 230 can obtain the motion vector difference included in the prediction information and derive the motion vector of the current block by adding the motion vector difference and the motion vector predictor. The prediction module can obtain or derive a reference picture index that indicates a reference picture from the prediction information.
[0075] The adder 240 can reconstruct the current block or current picture by adding the residual sample and the predicted sample. The adder 240 can reconstruct the current picture by adding the residual sample and the predicted sample in block units. When skip mode is applied, the residual is not transmitted, so the predicted sample can become the reconstructed sample. Here, the adder 240 has been described in a separate configuration, but the adder 240 may also be part of the prediction module 230. On the other hand, the adder 240 is called a reconstruction module or a reconstruction block generation unit.
[0076] The filter module 250 can apply a deblocking filter, sample-adaptive offset, and / or ALF to the restored picture. The sample-adaptive offset can be applied on a sample-by-sample basis and can be applied after deblocking filtering. The ALF can be applied after deblocking filtering and / or sample-adaptive offset.
[0077] The memory 260 can store a restored picture (a decoded picture) or information necessary for decoding. Here, the restored picture is a restored picture after the filtering procedure has been completed by the filter module 250. For example, the memory 260 can store a picture used for interpretation. In this case, the picture used for interpretation can be specified by a reference picture set or a reference picture list. The restored picture can be used as a reference picture for other pictures. The memory 260 can also output the restored pictures in an output order.
[0078] Figure 3 shows an example of a chromaticity diagram.
[0079] This embodiment relates to video coding, and more particularly to a technique for optimizing video coding under given conditions such as a defined or expected luminance transfer function, the dynamic range of the video, and the luminance values of the coding blocks.
[0080] As used herein, the “luminance transfer function” is also called the optical-electro transfer function (OETF) or the electro-optical transfer function (EOTF). Even though the optical-electro transfer function and the electro-optical transfer function are not exact inverse functions of each other, the optical-electro transfer function is called the inverse electro-optical transfer function, and the electro-optical transfer function is called the inverse optical-electron transfer function.
[0081] The techniques described herein can be used to compensate for non-optimal video coding performance that occurs when the mapping of luminance values to digital codewords is not considered with equal importance. For example, OETF actually allows more bits in dark areas (or vice versa) than in bright areas. In this case, a video encoder / decoder designed under the assumption that all digital codewords are coded with equal importance will generally not perform video coding in an optimal manner.
[0082] Although the technology described herein is explained in relation to the ITU-T H.264 and ITU-T H.265 standards, it should be noted that the technology described herein is generally applicable to any video coding standard.
[0083] Video compression technology is used in a wide range of devices, including digital televisions, desktop computers, portable computers, tablet computers, digital recording devices, digital media players, video game consoles, and smartphones. Digital video can be coded by video coding standards such as ITU-T H.264, also known as ISO / IEC MPEG-4 AVC, and High Efficiency Video Coding (HEVC). Video coding standards enable coding of specific formats (i.e., YUV420).
[0084] Conventional digital video cameras initially generate raw data corresponding to the signals produced by each image sensor. For example, a digital video capture device records the image as a linearly related set of luminance values. However, human vision cannot perceive linear changes in luminance values. That is, for example, 100 cd / m² 2 The area of the image associated with the brightness value is 200 cd / m². 2A luminance value does not necessarily have to be perceived as twice as bright as the associated area of the image. Thus, a luminance transfer function (e.g., optical-electrical transfer function (OETF) or electrical-optical transfer function (EOTF)) can be used to convert linear luminance data into data that can be recognized in a meaningful way. The optical-electrical transfer function (OETF) can non-linearly map absolute linear luminance values to digital codewords. As a result, the digital codewords can be converted to video formats supported by video coding standards.
[0085] Conventional video coding / display systems, such as those found in traditional television video distribution environments, provide a standard dynamic range (SDR), typically ranging from approximately 0.1 to 100 cd / m². 2 It supports a brightness range (often called "nits"). This range is considerably smaller than the range we encounter in real life. For example, a light bulb has a brightness of 10,000 cd / m². 2 It can have a brightness of several hundred thousand cd / m², and the surface of sunlight has a brightness of several hundred thousand cd / m². 2 It can have a brightness of 0.005 cd / m², compared to the night sky's brightness of 0.005 cd / m². 2 It can have the following brightness levels.
[0086] Recently, LCD and OLED displays have become widely used, and the technology of these devices enables the reproduction of even higher brightness and a wider color space. The realization of diverse displays and the desired brightness and dynamic range differ significantly from the brightness and dynamic range of conventional (SDR) capturing and generating devices. For example, content generation systems can generate or capture content with a contrast ratio of 1,000,000:1. Televisions and other video distribution environments are expected to provide users with a strong sense of "realism" by offering visual effects that are closer to actual experience. The brightness range of existing SDR is 0.1 to 100 cd / m². 2 Instead of ), a higher brightness range (0.005 to 10,000 cd / m²) 2) can be considered. For example, a minimum brightness of 0.01 cd / m 2 and a maximum brightness of 2,000 cd / m 2 When HDR content is displayed on a display that supports, the display has a dynamic range of 200,000:1. The dynamic range of a scene can be described by the ratio of the maximum light intensity to the minimum light intensity.
[0087] Also, Ultra High Definition Television (UHDTV) aims to provide users with a "sense of reality". Without generating, capturing, and displaying content with a higher peak brightness and a larger contrast value than current TVs, simply increasing the resolution is not sufficient to achieve such a goal. Also, when trying to enhance the sense of reality, it is necessary to render a richer hue than that provided by currently commonly used colour gamuts, for example, BT.709. Therefore, new content not only has brightness and contrast dozens of times larger, but also has a considerably wide colour gamut (for example, BT.2020 or a wider colour gamut in the future). Various colour gamut ranges are shown in Figure 3.
[0088] Beyond this latest technological development, HDR (High Dynamic Range) video / video playback can be achieved using a transfer function (OETF / EOTF) suitable for both the production side and the consumer side currently.
[0089] Figure 4 shows an example of the mapping of linear light values for SDR and HDR representations.
[0090] A transfer function can be described as a mapping between input and output in the real (floating-point) range [0.0, 1.0]. An example of a luminance conversion function for HDR data includes the so-called SMPTE (Society of Movie Picture and Television) HDR (High Dynamic Range) transfer function, also known as SMPTE ST 2084. Other examples of luminance conversion functions for HDR data include the hybrid log-gamma transfer function for HDR signals (also known as ITU-R BT.2100). Specifically, the SMPTE HDR transfer function includes the EOTF and inverse EOTF. The SMPTE ST 2084 inverse EOTF is described by the following formula:
[0091] [Mathematics 1] L c =R / 10,000
[0092] [Math 2] V=((c1+c2*L c n ) / (1+c3*L c n )) m
[0093] In equations 1 and 2, c1 = c3 - c2 + 1 = 3424 / 4096 = 0.8359375, c2 = 32 * 2413 / 4096 = 18.8515625, c3 = 32 * 2392 / 4096 = 18.6875, m = 128 * 2523 / 4096 = 78.84375, and n = 0.25 * 2610 / 4096 = 0.1593017578125.
[0094] The SMPTE ST 2084 EOTF can be explained by the following formula:
[0095] [Math 3] L c =((block[(V 1 / m --c1), 0) / (c2-c3*V 1 / m ))1 / n
[0096] [Math 4] R = 10,000 * L c
[0097] In the above formula, R is between 0 and 10,000 cd / m². 2 This is a luminance value that has a predicted range. For example, the same L as 1. c is 10,000 cd / m² 2 This is to correspond to the brightness level. R can represent the absolute linear brightness value. Also, in the above formula, V is called the brightness value (or perceptual curve value). OETF can map the perceptual curve value to a digital codeword. Therefore, V is 2 N -It can be mapped to a bitcodeword. An example of a function that can be used when mapping V to an N-bit codeword can be defined as follows:
[0098] [Number 5] Digital value = INT((2 N -1)*V)
[0099] In formula 5, INT(x) rounds down decimal values less than 0.5 and rounds up decimal values greater than or equal to 0.5 to produce a constant value.
[0100] As one example, Figure 4 shows a BT.709 style transfer function (green curve) with a range of 0.1 to 100 cd / m². 2 An 8-bit SDR system capable of displaying a transfer function (SMPTE ST 2084) with a range of 0.005 to 10,000 cd / m² is also available. 2This compares 10-bit HDR systems that can demonstrate the following. This figure is a schematic diagram. This figure does not capture the curve in its exact form, but is represented simply for illustrative purposes. In Figure 4, the horizontal axis represents constant code levels, and the vertical axis represents linear illuminance (adjusted to log10). This exemplary mapping includes conventional code level range ratios to accommodate both foot-room ("negative" samples below the real range [0.0, 1.0]) and head-room (real samples greater than or equal to 1.0). Due to design characteristics, the 10-bit HDR transfer function shown here allocates approximately twice as many code levels [119-509] as existing 8-bit SDR transfer functions allocate [16-235] code levels in the SDR range, and also provides a similar number of new code levels [510-940] to extend brightness. 0.01 cd / m 2 A new code level [64-118] is assigned for dim light intensities below this level.
[0101] In a sense, the 10-bit HDR system described here allocates approximately one additional bit of precision within the existing SDR intensity range, and by applying other additional bits, the curve can be raised to 100 cd / m². 2 By extending to higher brightness intensities, it allocates an additional 2 bits to existing consumer-side 8-bit "SDR" video. For comparison, the 10-bit SDR transfer function is also shown (red dashed curve).
[0102] Current video coding standards can code video data without considering the luminance transfer function (LMS), but their performance can be affected by the LMS because the distribution of codewords depends on it. For example, video coding standards are based on the assumption that each codeword is generally mapped with the same importance in terms of human visual sensitivity (HVS). However, each codeword is not actually mapped with the same importance. Many transfer functions are available, and each has its own unique mapping rules. Therefore, for this reason, the performance of video coders like HEVC is not optimized. For example, as will be discussed later, HEVC and existing video compression system technologies based on quantization parameter values quantize the entire range of codewords with the same importance regardless of LMS values, so coding is not performed optimally.
[0103] On the other hand, some examples of standards for supporting HDR video processing / coding are described in Table 1.
[0104] [Table 1]
[0105] Figure 5 is a flowchart showing the picture restoration process according to one embodiment.
[0106] Video content typically consists of video sequences composed of picture / frame groups (GOPs). Each video frame or picture can contain numerous slices, and each slice contains multiple video blocks. A video block can be defined as the largest array (also called samples) of pixel values that can be predictively coded. A video encoder / decoder applies predictive coding to video blocks and subdivisions of video blocks. ITU-T H.264 identifies macroblocks containing 16x16 luma samples. ITU-T H.265 (or commonly referred to as HEVC) identifies a similar coding tree unit (CTU) structure where a picture can be divided into CTUs of the same size, and each CTU can contain coding blocks (CBs) with 16x16, 32x32, or 64x64 luma samples. In JEM, a search model other than HEVC, a CTU can contain coding blocks with 128x128, 128x64, 128x32, 64x64, or 16x16 luma samples. Here, coding blocks, prediction blocks, and transformation blocks are identical to each other. Specifically, coding blocks (prediction blocks and transformation blocks) are either square or non-square blocks.
[0107] In one embodiment, the decoding device receives a bitstream (S500), performs entropy decoding (S510), performs inverse quantization (S520), decides whether to perform an inverse transform (S530), performs the inverse transform (S540), performs prediction (S550), and generates a reconstructed sample (S560). A more specific description of the embodiment is as follows.
[0108] As mentioned above, predictive syntax elements can associate the coding block of the predictive syntax element with a corresponding reference sample. For example, for intra predictive coding, the intra predictive mode can identify the location of the reference sample. In ITU-T H.265, the possible intra predictive modes for a luma component include planar predictive mode (predMode:0), DC predictive mode (predMode:1), and multiple angle predictive modes (predMode:2-N, where N is 34 or 65 or greater). One or more syntax elements can identify one of the intra predictive modes. For intra predictive coding, motion vectors (MVs) leverage temporal overlap in video by identifying the reference sample in a picture other than the picture of the coding block being coded. For example, the currently coding block may be predicted from a reference block located in a previously coded frame, and motion vectors can be used to indicate the location of the reference block. Motion vectors and associated data can describe, for example, the horizontal component of the motion vector, the vertical component of the motion vector, the resolution for the motion vector (e.g., 1 / 4 pixel precision), the prediction direction, and / or the reference picture index value. Furthermore, coding standards such as HEVC can support motion vector prediction. Motion vector prediction allows the motion vector to be identified using the motion vectors of adjacent blocks.
[0109] A video encoder can generate residual data by subtracting a predicted video block from a source video block. The predicted video block can be intra-predicted or inter(motion vector) predicted. The residual data is obtained in the pixel domain. Transformation coefficients are obtained by applying a transformation such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transformation to the residual block, thereby calculating a set of residual transformation coefficients. A transformation coefficient generator can output the residual transformation coefficients to a coefficient quantizer.
[0110] The process for deriving the quantization parameter (QP) can be summarized as follows:
[0111] The first step is to derive the Luma QP, a) predicting the Luma quantization parameter (qP) based on the previously coded (available) quantization parameters. Y_PRED The steps include: a) finding the (a) predictive QP (obtained in a) above) and the actual QP; b) obtaining a cu_delta_QP offset that shows the difference between the predictive QP (obtained in a) above) and the actual QP; and c) determining the Luma QP value based on the bit depth, predictive QP, and cu_delta_QP.
[0112] The second step is to derive the chroma QP, which includes a) deriving the chroma QP from the luma QP, and b) finding the chroma QP offset from the PPS level offsets (i.e., pps_cb_qp_offset, pps_cr_qp_offset) and the slice level chroma QP offsets (i.e., slice_cb_qp_offset, slice_cr_qp_offset).
[0113] Next, here are the details of the process mentioned above.
[0114] The aforementioned predicted luma quantization parameter qP Y_PRED This is derived as follows:
[0115] [Number 6] qP Y_PRED =(qP Y_A +qP Y_B +1)>>1
[0116] The aforementioned Variable qP Y_A and the aforementioned variable qP Y_B Here, qP Y_A This is set to be the same as the luma quantization parameter of the coding unit including luma coding block covering (xQg-1, yQg). Here, the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. Here, the variable qP Y_A and the aforementioned variable qP Y_B This shows the quantization parameters from the previous quantization group. Specifically, the qP Y_A This is set to the same value as the luma quantization parameter of the coding unit including luma coding block covering (xQg-1, yQg). Here, the luma position (xQg, yQg) indicates the top-left luma sample of the current quantization group relative to the top-left luma sample of the current picture. The variable qP Y_B Qp is the luma quantization parameter of the coding unit including luma coding block covering (xQg, yQg-1). Y The same as the aforementioned variable qP. Y_A or the aforementioned variable qP Y_B If it is not possible to use qP Y_PREV It is set to be the same as qP Y_PREV Qp is the Luma quantization parameter of the last coding unit in the decoding order within the quantization group. Y It will be set to the same value as above.
[0117] qP Y_PRED Once determined, the Luma quantization parameters are updated by adding CuQpDeltaVal as shown below.
[0118] [Number 7] Qp Y =((qP Y_PRED +CuQpDeltaVal+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0119] Here, the CuQpDeltaVal value is transmitted as a bitstream via two syntax elements, cu_qp_delta_abs and cu_qp_delta_sign_flag. QpBdOffset Y This indicates the value of the Luma quantization parameter range offset, which depends on bit_depth_luma_minus8 (i.e., Luma's bit depth - 8) as shown below.
[0120] [Number 8] QpBdOffset Y =6*bit_depth_luma_minus8
[0121] Finally, the Luma quantization parameter Qp′ Y This is derived as follows:
[0122] [Number 9] Luma quantization parameter Qp' Y =Qp Y +QpBdOffset Y
[0123] Chroma QP is derived from Luma QP, taking into account the PPS level offset (pps_cb_qp_offset, pps_cr_qp_offset) and slice level offset (slice_cb_qp_offset, slice_cr_qp_offset) as shown below.
[0124] [Number 10] qPiC b =Clip3(-QpBdOffset C , 57, Qp Y+pps_cb_qp_offset+slice_cb_qp_offset) qPiC r =Clip3(-QpBdOffset C , 57, Qp Y +pps_cr_qp_offset+slice_cr_qp_offset)
[0125] The aforementioned qPiC b and the qPiC r Based on Tables 2 and 3 below, qP Cb and qP Cr It will be updated to [the latest version].
[0126] [Table 2]
[0127] Table 2 shows qPi Cb From qP Cb This shows the mapping to [the specified location].
[0128] [Table 3]
[0129] Table 3 shows qPi Cr From qP Cr This shows the mapping to [the specified location].
[0130] Finally, the Cb and Cr components, Qp' Cb and Qp' Cr The chromatic quantization parameters for are derived as follows:
[0131] [Number 11] Qp' Cb =qP Cb +QpBdOffset C Qp' Cr =qP Cr +QpBdOffset C
[0132] QpBdOffset C This indicates the value of the chroma quantization parameter range offset, which depends on bit_depth_chroma_minus8 (i.e., the chroma bit depth -8) as shown below.
[0133] [Number 12] QpBdOffset C =6*bit_depth_chroma_minus8
[0134] Table 4, on the other hand, shows the definitions of syntax elements used in this specification.
[0135] [Table 4]
[0136] Figure 6 is a flowchart showing the picture restoration process according to another embodiment.
[0137] Since S600, S610, and S630 to S670 correspond to S500 to S560 in Figure 5, detailed explanations that overlap with the above explanation are omitted or simplified.
[0138] According to one embodiment, the decoding device receives a bitstream (S600), performs entropy decoding (S610), performs inverse quantization (S630), decides whether to perform an inverse transform (S640), performs the inverse transform (S650), performs prediction (S660), and generates a recovered sample (S670). The decoding device can also derive a QP offset based on the entropy decoding (S620) and perform inverse quantization based on the derived QP offset.
[0139] S620 can be shown as follows. Below, the QP offset can be represented as “Luma_avg_qp”.
[0140] qP Y_PREDがOnce determined, the luma quantization parameter can be updated by adding Luma_avg_qp as follows.
[0141] [Equation 13] Qp Y =((qP Y_PRED +CuQpDeltaVal+Luma_avg_qp+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0142] As an example, Luma_avg_qp can be derived (or estimated) from the luma values of adjacent pixels (or blocks) that have already been decoded and are available. Luma_avg_qp can be determined from the adjacent pixel values based on a predefined derivation rule. For example, Luma_avg_qp can be derived as follows.
[0143] [Equation 14] Luma_avg_qp = A * (avg_luma - M) + B
[0144] In Equation 14, avg_luma is the predicted average luma value obtained from available (decoded) adjacent pixels (or blocks).
[0145] M is a value predefined by the bit depth.
[0146] A is a scaling coefficient for mapping the difference in pixel values to the difference in qp (which can be predefined or transmitted in the bitstream), indicating the slope of the qp mapping.
[0147] B is an offset value that can be predefined or transmitted in the bitstream.
[0148] The derivation of Luma_avg_qp from the avg_luma value is not limited to just one of several formulas. As another example, Luma_avg_qp can be obtained from a table mapping as shown below.
[0149] [Number 15] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma]
[0150] Here, avg_luma is entered into the table, and the output of the table is Luma_avg_qp. To reduce the size of the table, the range of input values (avg_luma) can be further reduced as follows.
[0151] [Number 16] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma / D]
[0152] Here, D is a predefined constant value used to reduce the range of input values.
[0153] In one embodiment, Luma_avg_qp can be derived based on information regarding QP. The decoding device can obtain information regarding QP from the bitstream. For example, the information regarding QP may include init_qp_minus26, slice_qp_delta, slice_cb_qp_offset, slice_cr_qp_offset, cu_qp_delta_abs, and cu_qp_delta_sign_flag. Furthermore, the information regarding QP is not limited to the examples listed above.
[0154] Figure 7 is a flowchart showing the operation of an encoding device according to one embodiment, and Figure 8 is a block diagram showing the configuration of an encoding device according to one embodiment.
[0155] Each step shown in Figure 7 can be performed by the encoding device 100 shown in Figure 1. More specifically, steps S700 to S730 can be performed by the quantization module 123 shown in Figure 1, and step S740 can be performed by the entropy encoding module 130 shown in Figure 1. Furthermore, the operation of steps S700 to S740 is based on the partial explanation detailed in Figure 6. Therefore, detailed explanations that overlap with those explained in Figures 1 and 6 are omitted or simplified.
[0156] As shown in Figure 8, the encoding device according to one embodiment may include a quantization module 123 and an entropy encoding module 130. However, in some cases, not all of the components shown in Figure 8 may be necessary, and the encoding device may be implemented with more or fewer components than those shown in Figure 8.
[0157] In the decoding apparatus according to one embodiment, the quantization module 123 and the entropy encoding module 130 may be mounted on separate chips, or at least two or more components may be mounted via a single chip.
[0158] The encoding device according to one embodiment can derive the expected average luma value of the current block from available adjacent samples (S700). More specifically, the quantization module 123 of the encoding device can derive the expected average luma value of the current block from available adjacent samples.
[0159] The encoding device according to one embodiment can derive a QP offset for deriving Luma QP based on the expected average Luma value and information on QP (S710). More specifically, the quantization module 123 of the encoding device can derive a QP offset for deriving Luma QP based on the expected average Luma value and information on QP.
[0160] The encoding device according to one embodiment can derive Luma QP based on the QP offset (S720). More specifically, the quantization module 123 of the encoding device can derive Luma QP based on the QP offset.
[0161] The encoding device according to one embodiment can perform quantization on a quantization group including the current block based on the derived Luma QP (S730). More specifically, the quantization module 123 of the encoding device can perform quantization on a quantization group including the current block based on the derived Luma QP.
[0162] The encoding device according to one embodiment can encode video information including information for QP (S740). More specifically, the entropy encoding module 130 can encode video information including information for QP.
[0163] Referring to Figures 7 and 8, the encoding device according to one embodiment derives the expected average lumens value of the current block from available adjacent samples (S700), derives a QP offset for deriving lumens QP based on the expected average lumens value and information on QP (S710), derives lumens QP based on the QP offset (S720), performs quantization on the quantization group including the current block based on the derived lumens QP (S730), and encodes video information including information on QP (S740). As a result, quantization parameters can be efficiently derived, and the overall coding efficiency can be improved.
[0164] Figure 9 is a flowchart showing the operation of a decoding device according to one embodiment, and Figure 10 is a block diagram showing the configuration of a decoding device according to one embodiment.
[0165] Each step shown in Figure 9 can be performed by the decoding device 200 shown in Figure 2. More specifically, S900 can be performed by the entropy decoding module 210 shown in Figure 2, S910 to S940 can be performed by the inverse quantization module 222 shown in Figure 2, S950 can be performed by the inverse transform module 223 shown in Figure 2, S960 can be performed by the prediction module 230 shown in Figure 2, and S970 can be performed by the adder 240 shown in Figure 2. Furthermore, the operation of S900 to S970 is based on the partial explanation detailed in Figure 6. Therefore, detailed explanations that overlap with the content explained in Figures 2 and 6 are omitted or simplified.
[0166] As shown in Figure 10, a decoding device according to one embodiment may include an entropy decoding module 210, an inverse quantization module 222, an inverse transform module 223, a prediction module 230, and an adder 240. However, in some cases, not all of the components shown in Figure 10 are necessarily provided, and the encoding device may be implemented with more or fewer components than those shown in Figure 10.
[0167] In the decoding apparatus according to one embodiment, the entropy decoding module 210, the inverse quantization module 222, the inverse transformation module 223, the prediction module 230, and the adder 240 may be implemented on separate chips, or at least two or more components may be implemented via a single chip.
[0168] The decoding device according to one embodiment can decode video information including information on QP (S900). More specifically, the entropy decoding module 210 of the decoding device can decode video information including information on QP.
[0169] In one embodiment, information for QP is signaled at the Sequence Parameter Set (SPS) level.
[0170] In one embodiment, the video information includes information for Effective Data Range Parameters (EDRP), and the information for the EDRP includes at least one of the following: minimum input value, maximum input value, dynamic range of input values, mapping information for associating the minimum input value with brightness, mapping information for associating the maximum input value with brightness, and transfer function identification information. This method can be demonstrated by range matching.
[0171] More specifically, this disclosure can be used to efficiently code video content with a limited range of codewords (input values). This frequently occurs with HDR content because HDR content uses transfer functions that support high brightness. It can also occur when converting SDR data using a brightness conversion function corresponding to HDR data. In such cases, the video encoder can be configured to signal an Effective Data Range Parameter (EDRP). The decoder can be configured to receive the EDRP associated with the video data and to utilize the EDRP data in the decoding process. The EDRP data may include, for example, a minimum input value, a maximum input value, the dynamic range of the input values (indicating the difference between the maximum and minimum input values), mapping information between the minimum input value and its corresponding brightness, mapping information between the maximum input value and its corresponding brightness, and a transfer function identifier (known transfer functions can be identified by an assigned ID number, and detailed mapping information for each transfer function can be utilized).
[0172] For example, EDRP data can be signaled using a slice header, picture parameter set (PPS), or sequence parameter set (SPS). In this manner, the EDRP data can be used to make additional modifications to the coded values while the decoding process is being performed.
[0173] The present invention introduces quality control parameters (QCPs) to identify additional adjustments to quantization parameters. A decoder can be configured to receive QCPs associated with video data and to utilize the QCP data in the decoding process.
[0174] The decoding apparatus according to one embodiment can derive the expected average luma value of the current block from available neighboring samples (S910). More specifically, the inverse quantization module 222 of the decoding apparatus can derive the expected average luma value of the current block from available neighboring samples.
[0175] The decoding apparatus according to one embodiment can derive a QP offset for deriving Luma QP based on the expected average Luma value and information on QP (S920). More specifically, the inverse quantization module 222 of the decoding apparatus can derive a QP offset for deriving Luma QP based on the expected average Luma value and information on QP.
[0176] In one embodiment, the QP offset is derived based on the following formula.
[0177] [Number 17] Luma_avg_qp=A*(avg_luma-M)+B
[0178] In the above formula 17, Luma_avg_qp represents the QP offset, avg_luma represents the expected average Luma value, A represents a scaling factor for mapping the difference in pixel values to the difference in QP, M represents a predefined value related to the bit depth, B represents the offset value, and A and B are predetermined values or values included in the video information.
[0179] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, and the mapping table is determined using the expected average luma value as input.
[0180] In one embodiment, the QP offset is derived from a mapping table based on the expected average luma value, and the mapping table is determined using a value obtained by dividing the expected average luma value by a predetermined constant value.
[0181] In one embodiment, the available adjacent samples include at least one of the following: at least one Luma sample adjacent to the left boundary of the quantization group and at least one Luma sample adjacent to the upper boundary of the quantization group.
[0182] In one embodiment, at least one luma sample adjacent to the left boundary of a quantization group is included in the luma sample column directly adjacent to the left boundary of the quantization group, and at least one luma sample adjacent to the upper boundary of a quantization group is included in the luma sample row directly adjacent to the upper boundary of the quantization group.
[0183] In one embodiment, available adjacent samples include a Luma sample adjacent to the left of the upper left sample of the quantization group, and available adjacent samples include a Luma sample adjacent to the upper side of the upper left sample of the quantization group.
[0184] In one embodiment, the available neighbor samples include at least one of the restored neighbor samples, samples contained in at least one of the restored neighbor blocks, predicted neighbor samples, and samples contained in at least one of the predicted neighbor blocks.
[0185] In one embodiment, avg_luma can be derived from adjacent pixel values (blocks), and avg_luma represents the expected average luma value obtained from available (already decoded) adjacent pixels (or blocks).
[0186] i) Available adjacent pixels may include the following:
[0187] -Available adjacent pixels include the pixel located at (xQg-1, yQg+K), where the luma position (xQg, yQg) represents the upper-left luma sample of the current quantization group relative to the upper-left luma sample of the current picture (the leftmost line of the current block).
[0188] -Available adjacent pixels include the pixel located at (xQg+K, yQg-1), where the luma position (xQg, yQg) represents the upper left luma sample of the current quantization group relative to the upper left luma sample of the current picture (the uppermost line of the current block).
[0189] - Multiple lines can be used instead of a single line.
[0190] ii) The avg_luma can be calculated using the available adjacent blocks.
[0191] A block containing the pixel located at -(xQg-1, yQg) can be used.
[0192] A block containing the pixel located at -(xQg, yQg-1) can be used.
[0193] iii) The Avg_luma value can be calculated based on the restored adjacent pixels / blocks.
[0194] iv) The Avg_luma value can be calculated based on predicted adjacent pixels / block.
[0195] In one embodiment, the information for QP includes at least one syntax element associated with the QP offset, and the step of deriving the QP offset based on at least one of the expected mean luma value and the information for QP includes the step of deriving the QP offset based on at least one syntax element associated with the QP offset.
[0196] New syntax elements can be introduced to account for Luma_avg_qp. For example, a Luma_avg_qp value can be transmitted via a bitstream. A Luma_avg_qp value can be represented through two syntax elements, such as Luma_avg_qp_abs and Luma_avg_qp_flag.
[0197] Luma_avg_qp can be represented by two syntax elements (luma_avg_qp_delta_abs and luma_avg_qp_delta_sign_flag).
[0198] luma_avg_qp_delta_abs represents the absolute value of CuQpDeltaLumaVal, which is the difference between the current coding unit's luma quantization parameter and the luma quantization parameter derived without considering luminance.
[0199] luma_avg_qp_delta_sign_flag indicates the sign of CuQpDeltaLumaVal as shown below.
[0200] If luma_avg_qp_delta_sign_flag is 0, the corresponding CuQpDeltaLumaVal is a positive value.
[0201] Otherwise (when luma_avg_qp_delta_sign_flag is 1), the corresponding CuQpDeltaLumaVal is a negative value.
[0202] If luma_avg_qp_delta_sign_flag does not exist, it is assumed to be the same as 0.
[0203] If luma_avg_qp_delta_sign_flag exists, the variables IsCuQpDeltaLumaCoded and CuQpDeltaLumaVal are derived as follows:
[0204] [Number 18] IsCuQpDeltaLumaCoded=1 CuQpDeltaLumaVal=cu_qp_delta_abs*(1-2*luma_avg_qp_delta_sign_flag)
[0205] CuQpDeltaLumaVal represents the difference between the Luma quantization parameters for coding units that include Luma_avg_qp and the Luma quantization parameters for coding units that do not include Luma_avg_qp.
[0206] In one embodiment, the syntax elements described above can be transmitted at the quantization group level (or quantization unit level) (e.g., CU, CTU, or a predefined block unit).
[0207] The decoding apparatus according to one embodiment can derive Luma QP based on the QP offset (S930). More specifically, the inverse quantization module 222 of the decoding apparatus can derive Luma QP based on the QP offset.
[0208] The decoding apparatus according to one embodiment can perform inverse quantization on a quantization group including the current block based on the derived Luma QP (S940). More specifically, the inverse quantization module 222 of the decoding apparatus can perform inverse quantization on a quantization group including the current block based on the derived Luma QP.
[0209] In one embodiment, the decoding device can derive a chroma QP from the derived luma QP based on at least one chroma QP mapping table, and perform inverse quantization on the quantization group based on the derived luma QP and the derived chroma QP. Here, at least one chroma QP mapping table is based on the dynamic range of the chroma and the QP offset.
[0210] In one embodiment, instead of a single chroma-QP mapping table, there can be multiple chroma-QP derivation tables. Additional information is needed to identify the QP mapping table to use. The following is another example of a chroma-QP derivation table.
[0211] [Table 5]
[0212] In one embodiment, at least one chroma-QP mapping table includes at least one Cb-QP mapping table and at least one Cr-QP mapping table. The decoding apparatus according to one embodiment can generate a residual sample for the current block based on inverse quantization (S950). More specifically, the inverse transform module 223 of the decoding apparatus can generate a residual sample for the current block based on inverse quantization.
[0213] The decoding device according to one embodiment can generate a predictive sample for the current block based on the video information (S960). More specifically, the prediction module 230 of the decoding device can generate a predictive sample for the current block based on the video information.
[0214] The decoding apparatus according to one embodiment can generate a restored sample for the current block based on a residual sample for the current block and a predicted sample for the current block (S970). More specifically, the adder of the decoding apparatus can generate a restored sample for the current block based on a residual sample for the current block and a predicted sample for the current block.
[0215] Referring to Figures 9 and 10, the decoding apparatus according to one embodiment decodes video information including information for QP (S900), derives the expected average lumens value of the current block from available adjacent samples (S910), derives a QP offset for deriving lumens QP based on the expected average lumens value and information for QP (S920), derives lumens QP based on the QP offset (S930), performs inverse quantization on the quantization group including the current block based on the derived lumens QP (S940), generates a residual sample for the current block based on the inverse quantization (S950), generates a predicted sample for the current block based on the video information (S960), and generates a restored sample for the current block based on the residual sample for the current block and the predicted sample for the current block (S970). As a result, quantization parameters can be efficiently derived, and the overall coding efficiency can be improved.
[0216] The method according to the present invention described above can be implemented in software, and the encoding device and / or decoding device according to the present invention can be included in an image processing device such as a TV, computer, smartphone, or display device.
[0217] When embodiments of the present invention are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. The modules can be stored in memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by a variety of known means. The processor can include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory can include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices.
Claims
1. In a picture decoding method performed by a decoding device, A step of acquiring video information from a bitstream that includes information related to QP (quantization parameter), wherein the information related to QP includes information related to luminous QP and information related to chroma QP. A step of deriving the Luma QP based on the information related to the Luma QP, A step of deriving the chromat QP based on the information related to the chromat QP, wherein the chromat QP includes a first chromat QP for the Cb component and a second chromat QP for the Cr component. The step of generating a reconstructed picture including the reconstructed sample based on the chroma QP, the first chroma QP for the Cb component, and the second chroma QP for the Cr component, The aforementioned Luma QP is derived based on the predicted Luma QP and the QP delta value. The step of deriving the aforementioned chromatic QP is: The steps include: deriving the first chromatic QP for the Cb component based on the first chromatic QP mapping table for the Cb component; The step of deriving the second chromatic QP for the Cr component based on the second chromatic QP mapping table for the Cr component, The first chromatic QP mapping table for the Cb component is a separate mapping table from the second chromatic QP mapping table for the Cr component. The first chroma-QP mapping table for the Cb component is a picture decoding method different from the second chroma-QP mapping table for the Cr component.
2. The predicted Luma QP is derived based on the first Luma QP of the upper coding unit adjacent to the current quantization group and the second Luma QP of the left coding unit adjacent to the current quantization group, The picture decoding method according to claim 1, wherein the QP delta value is derived based on first information relating to the absolute value of the QP delta value and second information relating to the sign of the QP delta value.
3. The picture decoding method according to claim 1, wherein the chroma QP is derived based on the luma QP.
4. The picture decoding method according to claim 3, wherein the first chroma QP for the Cb component is derived based on the first chroma QP offset information for the Cb component.
5. The picture decoding method according to claim 4, wherein the first chroma QP offset information includes a first picture level offset and a first slice level offset.
6. The picture decoding method according to claim 3, wherein the second chroma QP for the Cr component is derived based on the second chroma QP offset information for the Cr component.
7. The picture decoding method according to claim 6, wherein the second chroma QP offset information includes a second picture level offset and a second slice level offset.
8. A picture encoding method performed by an encoding device, The steps include: deriving the Luma QP (quantization parameter), A step of deriving a chromatic QP, wherein the chromatic QP includes a first chromatic QP for the Cb component and a second chromatic QP for the Cr component. A step of generating QP-related information concerning the luma QP and the chroma QP, The step of encoding video information including the information related to the QP, The aforementioned Luma QP is shown based on the predicted Luma QP and the QP delta value. The step of deriving the aforementioned chromatic QP is: The steps include: deriving the first chromatic QP for the Cb component based on the first chromatic QP mapping table for the Cb component; The step of deriving the second chromatic QP for the Cr component based on the second chromatic QP mapping table for the Cr component, The first chromatic QP mapping table for the Cb component is a separate mapping table from the second chromatic QP mapping table for the Cr component. The first chroma-QP mapping table for the Cb component is different from the second chroma-QP mapping table for the Cr component, in a picture encoding method.
9. A method for transmitting data including a bitstream relating to video information, A step of generating the bitstream relating to the video information, wherein the bitstream is: The steps include: deriving the Luma QP (quantization parameter), A step of deriving a chromatic QP, wherein the chromatic QP includes a first chromatic QP for the Cb component and a second chromatic QP for the Cr component. A step of generating QP-related information concerning the luma QP and the chroma QP, A step of encoding the video information including the information related to the QP, and a step of generating the resulting The step of transmitting the data, which includes the bitstream relating to the video information, The aforementioned Luma QP is shown based on the predicted Luma QP and the QP delta value. The step of deriving the aforementioned chromatic QP is: The steps include: deriving the first chromatic QP for the Cb component based on the first chromatic QP mapping table for the Cb component; The step of deriving the second chromatic QP for the Cr component based on the second chromatic QP mapping table for the Cr component, The first chromatic QP mapping table for the Cb component is a separate mapping table from the second chromatic QP mapping table for the Cr component. The transmission method for the first chroma-QP mapping table for the Cb component differs from that for the second chroma-QP mapping table for the Cr component.