Encoding apparatus, decoding apparatus, and data transmitting apparatus
By deriving brightness quantization parameters using quantization parameter offset in an image encoding device, and performing inverse quantization and reconstructed sample generation, the low efficiency problem in the transmission and storage of high-resolution and high-quality images is solved, and the image compression efficiency and quantization efficiency are improved.
Patent Information
- Application Number
- CN202310308659.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-01
- Filing Date
- 2019-03-05
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2039-03-05
AI Technical Summary
Existing technologies are inefficient in the transmission and storage of high-resolution and high-quality images, leading to increased costs. There is a need to improve image compression efficiency and quantization efficiency.
By utilizing the entropy decoding module, inverse quantization module, inverse transform module, prediction module, and reconstruction module in the decoding and encoding devices, the brightness quantization parameters are derived using the quantization parameter offset, and inverse quantization and reconstruction sample generation are performed to improve quantization efficiency.
It improves image compression and quantization efficiency, enabling efficient image transmission and storage.
Smart Images

Figure CN116320455B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 201980006064.3 (International Application No.: PCT / KR2019 / 002520, Application Date: March 5, 2019, Invention Title: Image Coding Device and Method Based on Quantization Parameter Derivation). Technical Field
[0002] This invention relates to image coding technology. More specifically, this invention relates to an image coding device and method based on quantization parameter derivation in an image coding system. Background Technology
[0003] Recently, there has been an increased demand for high-resolution, high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. With higher resolution and quality image data, the amount of information or bits to be transmitted increases relative to existing image data. Therefore, transmission and storage costs can increase when transmitting image data using media such as wired / wireless broadband lines, or when storing it.
[0004] Therefore, an efficient image compression technology is needed for the efficient transmission, storage, and reproduction of high-resolution and high-quality image information. Summary of the Invention
[0005] Technical issues
[0006] This invention provides a method and apparatus for enhancing video encoding efficiency.
[0007] The present invention also provides a method and apparatus for increasing quantization efficiency.
[0008] The present invention also provides a method and apparatus for efficiently deriving quantization parameters.
[0009] Technical solution
[0010] According to an embodiment of the present invention, a picture decoding method performed by a decoding device is provided. The method includes: decoding image information including information about quantization parameters (QP); deriving a expected average brightness value of the current block from neighboring available samples; deriving a quantization parameter offset (QP offset) for deriving the brightness quantization parameter (brightness QP) based on the expected average brightness value and the information about QP; deriving the brightness QP based on the QP offset; performing inverse quantization on a quantization group including the current block based on the derived brightness QP; generating a residual sample of the current block based on the inverse quantization; generating a prediction sample of the current block based on the image information; and generating a reconstructed sample of the current block based on the residual sample and the prediction sample of the current block.
[0011] According to an embodiment of the present invention, a decoding device for decoding an image is provided. The decoding device includes: an entropy decoding module configured to decode image information including information about quantization parameters (QP); an inverse quantization module configured to derive a expected average brightness value of the current block from neighboring available samples, derive a quantization parameter offset (QP offset) for deriving the brightness quantization parameter (brightness QP) based on the expected average brightness value and information about QP, derive the brightness QP based on the QP offset, and perform inverse quantization on a quantization group including the current block based on the derived brightness QP; an inverse transform module configured to generate residual samples of the current block based on inverse quantization; a prediction module configured to generate prediction samples of the current block based on image information; and a reconstruction module configured to generate reconstructed samples of the current block based on the residual samples and prediction samples of the current block.
[0012] According to an embodiment of the present invention, a picture encoding method performed by an encoding device is provided. The method includes: deriving a expected average luminance value of the current block from neighboring available samples; deriving a quantization parameter offset (QP offset) for deriving a luminance quantization parameter (luminance QP) based on the expected average luminance value and information about QP; deriving the luminance QP based on the QP offset; performing quantization on a quantization group including the current block based on the derived luminance QP; and encoding image information including information about QP.
[0013] According to an embodiment of the present invention, an encoding apparatus for encoding an image is provided. The encoding apparatus includes: a quantization module configured to derive a expected average luminance value of a current block from neighboring available samples, derive a quantization parameter offset (QP offset) for deriving a luminance quantization parameter (luminance QP) based on the expected average luminance value and information about QP, derive the luminance QP based on the QP offset, and perform quantization on a quantization group including the current block based on the derived luminance QP; and an entropy encoding module configured to encode image information including information about QP.
[0014] Beneficial effects
[0015] According to the present invention, the overall image / video compression efficiency can be increased.
[0016] According to the present invention, quantization efficiency can be increased.
[0017] According to the present invention, quantization parameters can be derived efficiently. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the configuration of an encoding device according to an embodiment.
[0019] Figure 2This is a schematic diagram illustrating the configuration of a decoding device according to an embodiment.
[0020] Figure 3 An example of a chromaticity diagram is shown.
[0021] Figure 4 An example of a mapping for linear light values used in SDR and HDR representations is shown.
[0022] Figure 5 This is a flowchart illustrating the processing of the reconstructed screen according to an implementation method.
[0023] Figure 6 This is a flowchart illustrating the processing of a reconstructed screen according to another embodiment.
[0024] Figure 7 This is a flowchart illustrating the operation of an encoding device according to an embodiment.
[0025] Figure 8 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.
[0026] Figure 9 This is a flowchart illustrating the operation of a decoding device according to an embodiment.
[0027] Figure 10 This is a block diagram illustrating the configuration of a decoding device according to an embodiment. Detailed Implementation
[0028] According to an embodiment of the present invention, a picture decoding method performed by a decoding device is provided. The method includes: decoding image information including information about quantization parameters (QP); deriving a expected average brightness value of the current block from neighboring available samples; deriving a quantization parameter offset (QP offset) for deriving the brightness quantization parameter (brightness QP) based on the expected average brightness value and the information about QP; deriving the brightness QP based on the QP offset; performing inverse quantization on a quantization group including the current block based on the derived brightness QP; generating a residual sample of the current block based on the inverse quantization; generating a prediction sample of the current block based on the image information; and generating a reconstructed sample of the current block based on the residual sample and the prediction sample of the current block.
[0029] The mode of the present invention
[0030] This invention can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the invention. The terminology used in the following description is for describing particular embodiments only and is not intended to limit the invention. Singular expressions include plural expressions, provided they are interpreted differently. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0031] On the other hand, the elements in the accompanying drawings described in this invention are drawn independently for the convenience of illustrating different specific functions in the image encoding / decoding apparatus, and are not intended to imply that the elements are specifically implemented by independent hardware or independent software. For example, two or more elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or divided are part of this invention without departing from the concept of the invention.
[0032] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, the same reference numerals are used to indicate the same elements, and the same descriptions of the same elements will be omitted.
[0033] The following description can be applied to the technical field of video or image processing. For example, the methods or implementations disclosed in the following description can be applied to various video coding standards, such as the Universal Video Coding (VVC) standard (ITU-T Rec.H.266), the next generation video / image coding standard after VVC, or the previous generation video / image coding standard before VVC (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec.H.265), etc.).
[0034] In this specification, video may refer to a collection of images over time. A frame generally refers to a unit representing an image within a specific time period, and a slice is a unit that constitutes part of a frame in the encoding process. A frame may consist of multiple slices, and frames and slices may be combined if necessary. Additionally, in some cases, the term "image" may refer to a concept that includes both still images and video (a collection of still images over time). Furthermore, "video" does not necessarily mean only a collection of still images over time, but in some embodiments it can be interpreted as a concept that includes the meaning of still images.
[0035] A pixel, or image unit, can refer to the smallest unit of a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent pixel / pixel values for either the luminance component or the chrominance component alone.
[0036] A unit represents a basic unit of image processing. A unit may include at least one of the following: a specific region of the image and information associated with that region. A unit may be used in combination with terms such as a block or region. Typically, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows.
[0037] Figure 1 The configuration of the encoding device according to the embodiment is illustrated schematically.
[0038] Hereinafter, encoding / decoding devices may include video encoding / decoding devices and / or image encoding / decoding devices. The term "video encoding / decoding device" can be used as a concept that includes image encoding / decoding devices, and vice versa.
[0039] Reference Figure 1 The encoding device 100 may include a screen segmentation module 105, a prediction module 110, a residual processing module 120, an entropy encoding module 130, an adder 140, a filtering module 150, and a memory 160. The residual processing module (residual processing unit) 120 may include a subtractor 121, a transform module 122, a quantization module 123, a rearrangement module 124, an inverse quantization module 125, and an inverse transform module 126.
[0040] The screen segmentation module 105 can divide the input screen into at least one processing unit.
[0041] In one example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from the largest coding unit (LCU) according to a quadtree-binary tree (QTBT) structure. For example, a coding unit may be divided into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure.
[0042] In this case, for example, a quadtree structure can be applied first, followed by binary and ternary tree structures. Alternatively, a binary / ternary tree structure can be applied first. The encoding process according to the invention can be performed based on final coding units that are not further subdivided. In this case, the largest coding unit can be directly used as the final coding unit based on image characteristics, coding efficiency, etc., or the coding unit can be recursively subdivided into lower-depth coding units that can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and recovery (described later).
[0043] As another example, the processing unit may include a coding unit (CU), a prediction module (PU), or a transform unit (TU). The coding unit may be split along a quadtree structure from the maximum coding unit (LCU) into deeper coding units. In this case, the maximum coding unit may be directly used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit may be recursively divided into lower-depth coding units that can also be used as the final coding unit. When a minimum coding unit (SCU) is set, the coding unit may not be divided into coding units smaller than the minimum coding unit.
[0044] In this paper, the term "final coding unit" refers to the coding unit that segments or partitions the prediction module or conversion unit. A prediction module is a unit segmented from the coding unit and can be a unit for sample prediction. In this case, the prediction module can be divided into sub-blocks. A conversion unit can be segmented from the coding unit along a quadtree structure and can be a unit used to derive conversion coefficients and / or a unit used to derive the residual signal from the conversion factor.
[0045] Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction module may be referred to as a prediction block (PB), and the transformation unit may be referred to as a transform block (TB). A prediction block or prediction module may refer to a specific region in the form of a block in the image, and may include an array of prediction samples. In addition, a transform block or transform unit may refer to a specific region in the form of a block in the image, and may include an array of transform coefficients or residual samples.
[0046] The prediction module 110 predicts the current block or residual block and generates a block that can be predicted using the prediction samples of the current block. The unit of prediction performed in the prediction module 110 can be a coded block, a transform block, or a prediction block.
[0047] The prediction module 110 predicts the current block or the residual block and generates a predicted block that includes the prediction samples of the current block. The unit of prediction performed in the prediction module 110 can be a coded block, a transform block, or a prediction block.
[0048] The prediction module 110 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block. For example, the prediction module 110 can determine whether to apply intra-frame prediction or inter-frame prediction on a CU-by-CU basis.
[0049] In the case of intra-frame prediction, prediction module 110 can deduce the prediction sample of the current block based on the reference sample outside the current block in the frame to which the current block belongs (hereinafter referred to as the current frame).
[0050] In this case, the prediction module 110 may derive the prediction sample based on the average or interpolation of the neighboring reference samples of the current block (case (i)), or based on the reference sample in the sample that exists in a specific (prediction) direction relative to the prediction sample (case (ii)).
[0051] Case (i) may be referred to as non-directional mode or non-angular mode, and case (ii) may be referred to as directional mode or angular mode. In intra-frame prediction, the prediction mode may have, for example, 65 directional prediction modes and at least two non-directional modes. Non-directional modes may include DC prediction modes and planar modes. Prediction module 110 may use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0052] In the case of inter-frame prediction, prediction module 110 can derive the prediction sample of the current block based on samples specified by motion vectors on a reference frame. Prediction module 110 can derive the prediction sample of the current block by applying one of skip mode, merge mode, and motion vector prediction (MVP) mode. In skip mode and merge mode, prediction module 110 can use the motion information of neighboring blocks as the motion information of the current block.
[0053] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not sent. In MVP mode, the motion vector of the current block can be derived by using the motion vectors of neighboring blocks as motion vector predictors to be used as the motion vector predictors of the current block.
[0054] In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in a reference frame. The reference frame including the temporally neighboring block can be referred to as a colPic. Motion information can include motion vectors and reference frame indices. Information such as prediction mode information and motion information can be (entropy-encoded) and output as a bitstream.
[0055] When using motion information from time-proximity blocks in skip and merge modes, the highest-ranking frame in the reference frame list can be used as the reference frame. Reference frames included in the frame order count can be sorted based on the frame order count (POC) difference between the current frame and its corresponding reference frame. The POC corresponds to the display order of the frames and is distinguishable from the encoding order.
[0056] Subtractor 121 generates residual samples as the difference between the original samples and the predicted samples. When the skip mode is applied, residual samples may not be generated as described above.
[0057] Transform module 122 transforms residual samples based on transform blocks to generate transform coefficients. Transform module (transform unit) 122 can perform transformation according to the size of the transform block and the prediction mode applied to the coding block or prediction block that overlaps with the transform block in space.
[0058] For example, if intra-frame prediction is applied to a coding block or prediction block that overlaps with the transform block and the transform block is a 4×4 residual array, the residual samples are transformed into a Discrete Sine Transform (DST). In other cases, a Discrete Cosine Transform (DCT) transform kernel can be used to transform the residual samples.
[0059] The quantization module (quantization unit) 123 can quantize the transformation coefficients to generate quantized transformation coefficients.
[0060] The rearrangement module 124 rearranges the quantized transform coefficients. The rearrangement module 124 can rearrange the block-quantized transform coefficients into a one-dimensional vector form using a coefficient scanning method. The rearrangement module 124 may be part of the quantization module 123, but is described as an alternative configuration.
[0061] The entropy coding module 130 can perform entropy coding on the quantized transform coefficients. For example, entropy coding can include coding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding module 130 can encode the information required for video recovery (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients, according to entropy coding or a predetermined method.
[0062] Encoded information can be transmitted or stored in NAL (Network Abstraction Layer) units as bitstreams. The bitstreams can be transmitted via a network or stored in digital storage media. Networks may include broadcast networks and / or communication networks, and digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0063] The inverse quantization module 125 performs inverse quantization on the quantized values (quantized transformation coefficients) obtained from the quantization module 123, and the inverse transformation module 126 performs inverse quantization on the inverse quantized values obtained from the inverse quantization module 125 to generate residual samples.
[0064] Adder 140 combines the residual samples and the predicted samples to reconstruct the image. The residual samples and the predicted samples are added in blocks to generate reconstructed blocks. Here, adder 140 may be part of prediction module 110, and may also be referred to as a reconstruction module or reconstruction block generation unit.
[0065] For the reconstructed image, the filtering module 150 can apply a deblocking filter and / or a sample adaptive offset. Deblocking filtering and / or sample adaptive offset can correct artifacts at block boundaries or distortions in quantization processing in the reconstructed image. The sample adaptive offset can be applied sample-by-sample and can be applied after the deblocking filtering process is complete. The filtering module 150 can apply an ALF (Adaptive Loop Filter) to the recovered image. The ALF can be applied to the reconstructed image after applying the deblocking filter and / or sample adaptive offset.
[0066] The memory 160 can store the reconstructed frame (decoded frame) or information required for encoding / decoding. Here, the reconstructed frame can be a frame reconstructed after the filtering process has been completed by the filtering module 150. The stored reconstructed frame can be used as a reference frame for (inter-frame) prediction of another frame. For example, the memory 160 can store (reference) frames for inter-frame prediction. In this case, the frames used for inter-frame prediction can be specified by a set of reference frames or a list of reference frames.
[0067] Figure 2 The configuration of the decoding device according to the embodiment is illustrated schematically.
[0068] Reference Figure 2 The decoding device 200 may include an entropy decoding module 210, a residual processing module 220, a prediction module 230, an adder 240, a filtering module 250, and a memory 260. Here, the residual processing module 220 may include a rearrangement module 221, an inverse quantization module 222, and an inverse transform module 223. Additionally, although not shown, the video decoding device 200 may include a receiver for receiving a bitstream including video information. The receiving unit may be a separate module or may be included in the entropy decoding module 210.
[0069] When an input bitstream includes video / image information, the (video) decoding device 200 can recover the video / image / picture in response to the processing of the video / image information in the (video) encoding device 100.
[0070] For example, video decoding device 200 can use processing units applied in video encoding devices to perform video decoding. Therefore, the processing unit block for video decoding can be, for example, an encoding unit, and in another example, an encoding unit, a prediction module, or a conversion unit. Encoding units can be partitioned from the largest encoding unit along a quadtree structure, a binary tree structure, and / or a ternary tree structure.
[0071] Further prediction modules and conversion units can be used as needed. Prediction blocks can be derived from or segmented from coding units. Conversion units can be partitioned from coding units along a quadtree structure and can be units that derive conversion factors or units that derive residual signals from conversion factors.
[0072] The entropy decoding module 210 can parse the bitstream and output the information required for video restoration or image restoration. For example, the entropy decoding module 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and calculates the values of the syntax elements required for video restoration as well as the coefficient values of the quantization of the residual.
[0073] More specifically, the CABAC entropy decoding method includes: receiving binary bits (bins) corresponding to each syntax element in the bitstream; determining a context model based on the decoding target syntax element and the decoding information of the decoding target; predicting the occurrence probability of the binary bits according to the determined context model; and performing arithmetic decoding of the binary bits to generate symbols corresponding to the values of each syntax element. At this point, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the decoded symbols / bins for the context model of the next symbol / bin.
[0074] Predicted information about the decoding information in the entropy decoding module 210 is provided to the prediction module 230, and the residual value of the entropy decoding performed in the entropy decoding module 210 can be input to the rearrangement module 221.
[0075] Rearrangement module 221 can rearrange the quantized transform coefficients into a two-dimensional block form. Rearrangement module 221 can perform the rearrangement in response to a coefficient scan performed in the encoding device. Rearrangement module 221 may be part of inverse quantization module 222, but is described as an alternative configuration.
[0076] The inverse quantization module 222 can perform inverse quantization on the quantized transform coefficients based on the (inverse)quantization parameters and output the transform coefficients. At this time, the encoding device can be notified by a signal of the information used to derive the quantization parameters.
[0077] The inverse transformation module 223 can reverse the transformation coefficients to derive the residual samples.
[0078] The prediction module 230 can predict the current block and generate a predicted block that includes the prediction samples of the current block. The unit of prediction performed by the prediction module 230 can be a coded block, a transform block, or a prediction block.
[0079] The prediction module 230 can determine whether to apply intra-frame prediction or inter-frame prediction based on the prediction information. In this case, the unit that determines whether to apply intra-frame prediction or inter-frame prediction may be different from the unit that generates the prediction samples. Furthermore, the units that generate prediction samples for inter-frame prediction and intra-frame prediction may also be different. For example, the decision to apply inter-frame prediction or intra-frame prediction may be made at the CU (unit of measurement). Alternatively, for example, in inter-frame prediction, the prediction mode may be determined at the PU (unit of measurement) to generate prediction samples. In intra-frame prediction, the prediction mode may be determined at the PU (unit of measurement), and prediction samples may be generated at the TU (unit of measurement).
[0080] In the case of intra-frame prediction, prediction module 230 can derive the prediction sample of the current block based on neighboring reference samples in the current frame. Prediction module 230 can apply a directional or non-directional mode to derive the prediction sample of the current block based on the neighboring reference samples of the current block. In this case, the intra-frame prediction mode of neighboring blocks can be used to determine the prediction mode to be applied to the current block.
[0081] In the case of inter-frame prediction, prediction module 230 can derive the prediction sample for the current block based on the sample specified by the motion vector on the reference frame. Prediction module 230 can derive the prediction sample for the current block by applying skip mode, merge mode, or MVP mode. At this time, motion information (e.g., information about motion vectors, reference frame index, etc.) required for inter-frame prediction of the current block provided in the encoding device can be obtained or derived based on the prediction information.
[0082] In skip and merge modes, motion information from neighboring blocks can be used as motion information for the current block. In this case, neighboring blocks can include spatially neighboring blocks and temporally neighboring blocks.
[0083] The prediction module 230 can construct a merge candidate list using motion information of available neighboring blocks, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled from the encoding device. The motion information may include motion vectors and reference frames. When using motion information of temporally neighboring blocks in skip mode and merge mode, the highest frame in the reference frame list can be used as the reference frame.
[0084] In skip mode, unlike merge mode, the difference (residual) between the predicted sample and the original sample is not sent.
[0085] In MVP mode, the motion vectors of neighboring blocks can be used as motion vector predictions to derive the motion vector of the current block. In this case, neighboring blocks can include spatially neighboring blocks and temporally neighboring blocks.
[0086] For example, when applying the merge mode, a merge candidate list can be generated using the motion vectors of reconstructed spatially neighboring blocks and / or the motion vectors corresponding to Col blocks (temporally neighboring blocks). In merge mode, the motion vectors of candidate blocks selected from the merge candidate list are used as the motion vector of the current block. Prediction information may include a merge index, which indicates the candidate block with the optimal motion vector selected from the candidate blocks included in the merge candidate list. In this case, the prediction module 230 can use the merge index to derive the motion vector of the current block.
[0087] As another example, when applying the Motion Vector Prediction (MVP) mode, a candidate list of motion vector predictions is generated using the motion vectors of the reconstructed spatial neighbor blocks and / or the motion vectors corresponding to the Col block (temporal neighbor block). That is, the motion vectors of the reconstructed spatial neighbor blocks and / or the motion vectors corresponding to the neighboring Col block can be used as motion vector candidates. Information about the prediction may include the predicted motion vector index, which indicates the optimal motion vector selected from the motion vector candidates included in the list.
[0088] At this time, the prediction module 230 can use the motion vector index to select the predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list. The prediction unit of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector prediction, and can output it as a bit stream. That is, the MVD can be obtained by subtracting the motion vector prediction from the motion vector of the current block. In this case, the prediction module 230 can obtain the motion vector difference included in the information about the prediction, and derive the motion vector of the current block by adding the motion vector difference and the motion vector prediction. The prediction module can also obtain or derive the reference screen index, etc., indicating the reference screen from the information about the prediction.
[0089] Adder 240 can add the residual sample and the predicted sample to reconstruct the current block or the current frame. Adder 240 can reconstruct the current frame by adding the residual sample and the predicted sample block by block. When the skip mode is applied, since no residual is sent, the predicted sample can be the recovered sample. Here, adder 240 is described as an alternative configuration, but adder 240 can be part of prediction module 230. Furthermore, adder 240 can be referred to as a reconstruction module or reconstruction block generation unit.
[0090] The filtering module 250 can apply sample adaptive offset and / or ALF to the reconstructed image using deblocking filtering. The sample adaptive offset can be applied sample-by-sample and can be applied after deblocking filtering. The ALF can be applied after deblocking filtering and / or sample adaptive offset.
[0091] The memory 260 can store reconstructed frames (decoded frames) or information required for decoding. Here, the reconstructed frames can be frames reconstructed after the filtering process has been completed by the filtering module 250. For example, the memory 260 can store frames used for inter-frame prediction. In this case, the frames used for inter-frame prediction can be specified by a set of reference frames or a list of reference frames. The reconstructed frames can be used as reference frames for another frame. In addition, the memory 260 can output the reconstructed frames according to the output order.
[0092] Figure 3 An example of a chromaticity diagram is shown.
[0093] This implementation relates to video coding, and more specifically, to techniques for optimizing video coding based on given conditions such as a defined or anticipated luminance transfer function, the dynamic range of the video, or the luminance values of the coding blocks.
[0094] As used herein, the term luminance transfer function may refer to the photoelectric transfer function (OETF) or the electro-optical transfer function (EOTF). It should be noted that the photoelectric transfer function can be called the inverse electro-optical transfer function, and the electro-optical transfer function can be called the inverse photoelectric transfer function, even though the two transfer functions are not exactly inverses of each other.
[0095] The techniques described in this paper can be used to compensate for suboptimal video coding performance when the mapping from luminance values to digital codewords is not considered with equal importance. For example, in practice, OETF allows for more bits in dark areas compared to bright areas (and vice versa). In such cases, video encoders / decoders designed based on the assumption that all digital codewords are encoded with equal importance often fail to perform video coding optimally.
[0096] It should be noted that although the techniques of this disclosure are described in the ITU-T H.264 and ITU-T H.265 standards, the techniques of this disclosure are generally applicable to any video coding standard.
[0097] Video compression technology has been deployed in many devices, including digital televisions, desktop or laptop computers, tablet computers, digital recording devices, digital media players, video game consoles, and smartphones. Digital video can be encoded according to video coding standards such as ITU-T H.264 (ISO / IEC MPEG-4 AVC) and High Efficiency Video Coding (HEVC). Video coding standards allow encoding in specific formats (i.e., YUV420).
[0098] Traditional digital video cameras initially generate raw data corresponding to the signals generated by their individual image sensors. For example, digital image capture devices record images as a linearly correlated set of brightness values. However, human vision cannot perceive changes in brightness values linearly. That is, for example, an image area associated with a brightness value of 100 cd / m² may not be perceived as twice as bright as an image area associated with a brightness value of 200 cd / m². Therefore, a brightness transfer function (e.g., the photoelectric transfer function (OETF) or the electro-optical transfer function (EOTF)) can be used to convert linear brightness data into data that can be perceived in a meaningful way. The photoelectric transfer function (OETF) maps absolutely linear brightness values to digital codewords in a non-linear manner. The resulting digital codewords can then be converted into video formats supported by video coding standards.
[0099] Traditional video encoding / display systems, such as those used in traditional television video distribution environments, typically support frequencies from approximately 0.1 to 100 cd / m². 2 The standard dynamic range (SDR) is the range of brightness (often referred to as "nits"). This range is significantly smaller than the range encountered in real-world situations. For example, a light bulb can have a luminance range exceeding 10,000 cd / m². 2 Surfaces illuminated by sunlight can have hundreds of thousands of cd / m². 2 The brightness is above that of the night sky, which is 0.005 cd / m². 2 Or lower.
[0100] In recent years, LCD and OLED displays have been widely used, and the technology of these devices allows for higher brightness and wider color space reproduction. The achievable and desired brightness and dynamic range of various displays can differ significantly from those of traditional (SDR) capture and creation devices. For example, content creation systems may be able to create or capture content with a contrast ratio of 1,000,000:1. Television and other video distribution environments are expected to provide a viewing experience closer to real life, offering users a stronger sense of "being there." This contrasts with existing SDR brightness ranges (from 0.1 to 100 cd / m²). 2 A higher brightness range (from 0.005 to 10000 cd / m²) could be considered. 2 For example, when supporting 0.01 cd / m 2 Minimum brightness and 2000 cd / m 2 When displaying HDR content on a monitor at its maximum brightness, it can have a dynamic range of 200,000:1. It's important to note that the dynamic range of a scene can be described as the ratio of maximum light intensity to minimum light intensity.
[0101] Furthermore, the goal of Ultra High Definition Television (UHDTV) is to provide users with a sense of "realism." Simply increasing resolution may not be sufficient to fully achieve this goal without creating, capturing, and displaying content with higher peak brightness and greater contrast values than today's TVs. Additionally, greater realism requires rendering a richer palette of colors than those offered by today's commonly used color gamuts (e.g., BT.709). Therefore, new content will not only have brightness and contrast values several orders of magnitude higher, but also a significantly wider color gamut (e.g., BT.2020, or even wider in the future). Figure 1 It describes various color gamut ranges.
[0102] Beyond this recent technological evolution, HDR (High Dynamic Range) image / video reproduction can now be achieved for both producers and consumers using appropriate transfer functions (OETF / EOTF).
[0103] Figure 4 An example of a mapping for linear light values used in SDR and HDR representations is shown.
[0104] A transfer function can be described as a mapping between inputs and outputs in the real-valued (floating-point) range [0.0, 1.0]. An example of a luminance transformation function corresponding to HDR data includes the so-called SMPTE (Motion Picture and Television Association) High Dynamic Range (HDR) transfer function, which may be referred to as SMPTE ST 2084. Another example of a luminance transformation function corresponding to HDR data includes the mixed log-gamma transfer function (also known as ITU-R BT.2100) for HDR signals. Specifically, the SMPTE HDR transfer function includes an EOTF and an inverse EOTF. The SMPTE ST 2084 inverse EOTF is described by the following set of mathematical expressions.
[0105] [Mathematical Expression 1]
[0106] L c =R / 10000
[0107] [Mathematical Expression 2]
[0108] V=((c1+c 2* L c n ) / (1+c 3* L c n )) m
[0109] In mathematical expressions 1 and 2, c1 = c3 - c2 + 1 = 3424 / 4096 = 0.8359375, c2 = 32 * 2413 / 4096 = 18.8515625, c3 = 32 * 2392 / 4096 = 18.6875, m = 128 * 2523 / 4096 = 78.84375, and n = 0.25 * 2610 / 4096 = 0.1593017578125.
[0110] The SMPTE ST 2084EOTF can be described using the following set of mathematical expressions.
[0111] [Mathematical Expression 3]
[0112] L c =((max[(V) 1 / m -c1),0]) / (c2-c3*V 1 / m ))1 / n
[0113] [Mathematical Expression 4]
[0114] R = 10000 * Lc
[0115] In the above formula, R is a value ranging from 0 to 10000 cd / m². 2 The expected range of brightness values. For example, L equal to 1. c Designed with 10000 cd / m 2 The brightness level corresponds to this. R can indicate the absolute linear brightness value. Furthermore, in the above formula, V can be referred to as the brightness value (or perceptual curve value). Since the OETF can map perceptual curve values to digital codewords, V can therefore be mapped to 2. N Bit codewords. An example of a function that can be used to map V to an N-bit codeword can be defined as:
[0116] [Mathematical Expression 5]
[0117] numeric value = INT((2 N -1)*V)
[0118] In mathematical expression 5, INT(x) generates an integer value by rounding down decimals less than 0.5 and rounding up decimals greater than or equal to 0.5.
[0119] As an illustrative example, Figure 2 This will allow the use of BT.709 style transfer functions (green curve) to represent 0.1 to 100 cd / m³. 2 The 8-bit SDR system can represent 0.005 to 10000 cd / m² using another transfer function (SMPTE ST 2084). 2A comparison is made with a 10-bit HDR system. The graphs in this figure are approximate. They do not capture the exact form of the curves and are shown for illustrative purposes only. In the figure, integer code levels are along the horizontal axis, and linear light values (scaled to log10) are along the vertical axis. This illustrative mapping includes traditional code level range scaling to accommodate both lower margin ("negative" samples below the real value range of [0.0, 1.0]) and upper margin (samples above the real value of 1.0). Due to the nature of the design, the 10-bit HDR transfer function shown here assigns approximately twice as many code levels [119 to 509] compared to the traditional 8-bit SDR transfer function which assigns [16 to 235] in the SDR range, while providing a similar number of new code levels [510 to 940] to extend the luminance. For values below 0.01 cd / m² 2 The darker intensity is assigned to new code levels [64 to 118].
[0120] In a sense, by assigning approximately one additional bit of precision within the traditional SDR intensity range, while simultaneously applying another additional bit, the curve can be extended to greater than 100 cd / m². 2 The intensity of the 10-bit HDR system shown here is such that it distributes 2 extra bits compared to a traditional consumer 8-bit "SDR" video. For comparison, the 10-bit SDR transfer function is also plotted (red dashed curve).
[0121] While current video coding standards can encode video data without considering the luma transfer function, the performance of these standards can be affected by it because codeword distribution can depend on it. For example, a video coding standard might be based on the assumption that individual codewords are generally mapped to equal importance in terms of human visual sensitivity (HVS). However, this may not always be the case in reality. Many transfer functions are available, each with its own mapping rules, which are not universal. Therefore, this can lead to suboptimal performance for video encoders such as HEVC. For example, and as described in more detail below, HEVC and existing video compression systems based on quantization parameter values may not perform optimally because they quantize the entire range of codewords with equal importance regardless of how luma values are quantized.
[0122] In addition, Table 1 below describes some examples of standards that support HDR video processing / coding.
[0123] [Table 1]
[0124]
[0125] Figure 5 This is a flowchart illustrating the processing of the reconstructed screen according to an implementation method.
[0126] Video content typically comprises a sequence of video frames (GOPs). Each video frame or GOP can include multiple slices, where each slice contains multiple video blocks. A video block can be defined as the largest array of predictably coded pixel values (also called samples). The video encoder / decoder applies predictive coding to the video blocks and their subdivisions. ITU-T H.264 specifies macroblocks comprising 16×16 luma samples. ITU-T H.265 (or commonly known as HEVC) specifies a similar Code Tree Unit (CTU) structure, where a picture can be divided into CTUs of equal size, and each CTU can include coded blocks (CBs) with 16×16, 32×32, or 64×64 luma samples. In JEM, an exploratory model beyond HEVC, CTUs can include coded blocks with 128×128, 128×64, 128×32, 64×64, or 16×16, etc. Here, coded blocks, predictive blocks, and transform blocks can be identical to each other. Specifically, the coding blocks (prediction blocks and transform blocks) can be square or non-square blocks.
[0127] According to the implementation method, the decoding device may receive a bitstream (S500), perform entropy decoding (S510), perform inverse quantization (S520), determine whether to perform inverse transform (S530), perform inverse transform (S540), perform prediction (S550), and generate reconstructed samples (S560). A more detailed description of the implementation method is shown below.
[0128] As described above, prediction syntax elements associate their coded blocks with corresponding reference samples. For example, for intra-frame predictive coding, the intra-frame prediction mode can specify the location of the reference sample. In ITU-T H.265, possible intra-frame prediction modes for the luma component include planar prediction mode (predMode:0), DC prediction (predMode:1), and multiple corner prediction modes (predMode:2-N, where N can be 34, 65, or greater). One or more syntax elements can identify one of the pre-intra-frame prediction modes. For inter-frame predictive coding, motion vectors (MVs) identify reference samples in frames outside the frame to be coded, thereby utilizing temporal redundancy in the video. For example, the current coded block can be predicted from a reference block located in a previously coded frame, and motion vectors can be used to indicate the location of the reference block. For example, motion vectors and associated data can describe the horizontal component of the motion vector, the vertical component of the motion vector, the resolution of the motion vector (e.g., quarter-pixel precision), the prediction direction, and / or the reference frame index value. Furthermore, coding standards such as HEVC, for example, support motion vector prediction. Motion vector prediction allows the use of motion vectors from neighboring blocks to specify motion vectors.
[0129] A video encoder generates residual data by subtracting a predicted video block from a source video block. The predicted video block can be predicted intra-frame or inter-frame (motion vector) predictive. The residual data is obtained in the pixel domain. Transform coefficients are obtained by applying a transform (e.g., Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or a conceptually similar transform) to the residual block to generate a set of residual transform coefficients. The transform coefficient generator outputs the residual transform coefficients to the coefficient quantization unit.
[0130] The derivation and processing of the quantization parameter (QP) are summarized as follows.
[0131] The first step is to derive the luminance QP. a) Find the predicted luminance quantization parameter (qP) based on the previously encoded (available) quantization parameters. Y_PRED b) Obtain the cu_delta_QP offset indicating the difference between the predicted QP (obtained in a) and the actual QP, and c) Determine the luminance QP value based on the bit depth, the predicted QP, and cu_delta_QP.
[0132] The second step is to derive the chroma QP. a) Derive the chroma QP from the luminance QP. b) Find the chroma QP offset from the PPS-level offset (i.e., pps_cb_qp_offset, pps_cr_qp_offset) and the slice-level chroma QP offset (i.e., slice_cb_qp_offset, slice_cr_qp_offset).
[0133] The following are the details of the above processing.
[0134] Predicted brightness quantization parameter qP Y_PRED The derivation is as follows.
[0135] [Mathematical Expression 6]
[0136] qP Y_PRED =(qP Y_A +qP Y_B +1)>>1
[0137] Among them, variable qP Y_A and qP Y_B Indicates the quantization parameters from the previous quantization group. Here, qP Y_A The luminance quantization parameter is set to be equal to the luminance quantization parameter of the coding unit containing the luminance coding block covering (xQg-1, yQg). Here, the luminance position (xQg, yQg) specifies the top-left luminance sample of the current quantization group relative to the top-left luminance sample of the current frame. The variable qP... Y_A and qP Y_B Indicates the quantization parameters from the previous quantization group. Specifically, qP Y_AThe luminance quantization parameter is set to be equal to the luminance coded unit containing the luminance coded block covering (xQg-1, yQg). Here, the luminance position (xQg, yQg) specifies the top-left luminance sample of the current quantization group relative to the top-left luminance sample of the current frame. Variable qP Y_B The luminance quantization parameter Qp is set to be equal to the luminance coded unit containing the luminance coded block covering (xQg, yQg-1). Y When qP Y_A or qP Y_B When unavailable, it is set to equal to qP. Y_PREV Here, qP Y_PREV The luminance quantization parameter Qp is set to be equal to the last encoded unit in the previous quantization group in decoding order. Y .
[0138] Once qP is determined Y_PRED The brightness quantization parameter is then updated by adding it to CuQpDeltaVal.
[0139] [Mathematical Expression 7]
[0140] Qp Y =((qP) Y_PRED +CuQpDeltaVal+52+2*QpBdOffset Y )%(52+QpBdOffset Y ))-QpBdOffset Y
[0141] The CuQpDeltaVal value is sent via a bitstream using two syntax elements, such as cu_qp_delta_abs and cu_qp_delta_sign_flag. QpBdOffset Y Specifies the value of the range offset for the luminance quantization parameter, and it depends on bit_depth_luma_minus8 (i.e., the bit depth of luminance - 8).
[0142] [Mathematical Expression 8]
[0143] QpBdOffset Y = 6 * bit_depth_luma_minus8
[0144] Finally, the brightness quantization parameter Qp′ is derived as follows. Y .
[0145] [Mathematical Expression 9]
[0146] Brightness quantization parameter Qp′ Y =Qp Y +QpBdOffsetY
[0147] The following considers the PPS-level offset (pps_cb_qp_offset, pps_cr_qp_offset) and slice-level offset (slice_cb_qp_offset, slice_cr_qp_offset) to derive the chrominance QP from the luminance QP.
[0148] [Mathematical Expression 10]
[0149] qPi Cb =Clip3(-QpBdOffset) C ,57,Qp Y +pps_cb_qp_offset+slice_cb_qp_offset)
[0150] qPi Cr =Clip3(-QpBdOffset) C ,57,Qp Y +pps_cr_qp_offset+slice_cr_qp_offset)
[0151] Based on Tables 2 and 3 below, the above qPi Cb and qPi Cr Further updates to qP Cb and qP Cr .
[0152] [Table 2]
[0153] <![CDATA[qPi Cb ]]> <30 30 31 32 33 34 35 36 37 38 39 40 41 42 43 >43 <![CDATA[qP Cb ]]> <![CDATA[=qPi Cb ]]> 29 30 31 32 33 33 34 34 35 35 36 36 37 37 <![CDATA[=qPi Cb -6]]>
[0154] Table 2 shows the data from qPi Cb to qP Cb The mapping.
[0155] [Table 3]
[0156] <![CDATA[qPi Cr ]]> <30 30 31 32 33 34 35 36 37 38 39 40 41 42 43 >43 <![CDATA[qP Cr ]]> <![CDATA[=qPi Cr ]]> 29 30 31 32 33 33 34 34 35 35 36 36 37 37 <![CDATA[=qPi Cr -6]]>
[0157] Table 3 shows the data from qPi Cr to qP Cr The mapping.
[0158] Finally, the colorimetric parameters Qp′ for the Cb and Cr components are derived as follows. Cb and Qp′ Cr .
[0159] [Mathematical Expression 11]
[0160] Qp′ Cb =qP Cb +QpBdOffsetC
[0161] Qp′ Cr =qP Cr +QpBdOffset C
[0162] QpBdOffset C Specifies the value of the chroma quantization parameter range offset, and it depends on bit_depth_chroma_minus8 (i.e., the bit depth of chroma - 8).
[0163] [Mathematical Expression 12]
[0164] QpBdOffset C = 6 * bit_depth_chroma_minus8
[0165] In addition, Table 4 shows the definitions of the syntax elements used in this specification.
[0166] [Table 4]
[0167]
[0168] Figure 6 This is a flowchart illustrating the processing of a reconstructed screen according to another embodiment.
[0169] Due to S600, S610 and S630 to S670 and Figure 5 The S500 to S560 correspond to the above descriptions, so detailed descriptions that are repeated above will be omitted or simplified.
[0170] According to the implementation method, the decoding device can receive a bitstream (S600), perform entropy decoding (S610), perform inverse quantization (S60), determine whether to perform inverse transform (S640), perform inverse transform (S650), perform prediction (S660), and generate reconstructed samples (S670). Furthermore, the decoding device can derive the QP offset based on entropy decoding (S620) and perform inverse quantization based on the derivation of the QP offset.
[0171] S620 can be specified as shown below. In the following text, the QP offset can be represented by "Luma_avg_qp".
[0172] Once qPY_PRED is determined, the luminance quantization parameter can be updated by adding it to Luma_avg_qp as follows.
[0173] [Mathematical Expression 13]
[0174] QpY=((qPY_PRED+CuQpDeltaVal+Luma_avg_qp+52+2*QpBdOffsetY)%(52+QpBdOffsetY))-QpBdOffsetY
[0175] In one example, Luma_avg_qp can be derived (or inferred) from the luminance values of neighboring pixels (or blocks) that have already been decoded and are available. Luma_avg_qp can be determined from neighboring pixel values based on predefined derivation rules. For example, Luma_avg_qp can be derived as follows.
[0176] [Mathematical Expression 14]
[0177] Luma_avg_qp=A*(avg_luma-M)+B.
[0178] In mathematical expression 14, avg_luma: the expected average luminance value obtained from the available (decoded) neighboring pixels (or blocks).
[0179] M: can depend on a predefined value for the bit depth.
[0180] A: A scaling factor that maps pixel value differences to QP differences (can be predefined or sent in the bitstream). It indicates the slope of the QP mapping, and
[0181] B: Offset value, which can be predefined or sent in the bitstream.
[0182] The derivation of Luma_avg_qp from the value of avg_luma is not restricted by one of the above formulas. In another example, Luma_avg_qp can be obtained from the table mapping as follows.
[0183] [Mathematical Expression 15]
[0184] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma]
[0185] Here, avg_luma is the input to the table, and the output of the table is Luma_avg_qp. To reduce the table size, the range of input values (avg_luma) can be further reduced as follows.
[0186] [Mathematical Expression 16]
[0187] Luma_avg_qp=Mapping_Table_from_luma_to_QP[avg_luma / D]
[0188] Here, D is a predefined constant value to reduce the range of input values.
[0189] In implementations, Luma_avg_qp can be derived based on information about QP. The decoding device can obtain information about QP from the bitstream. In one example, information about QP may include init_qp_minus26, slice_qp_delta, slice_cb_qp_offset, slice_cr_qp_offset, cu_qp_delta_abs, and cu_qp_delta_sign_flag. Furthermore, information about QP is not limited to the examples listed above.
[0190] Figure 7 This is a flowchart illustrating the operation of the encoding device according to an embodiment. Figure 8 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.
[0191] Figure 7 The steps disclosed in the document can be obtained by Figure 1 The encoding device 100 disclosed in the document executes this. More specifically, S700 to S730 can be performed by... Figure 1 The quantization module 123 shown is executed, and S740 can be controlled by... Figure 1 The entropy encoding module 130 shown is executed. Furthermore, the operations of S700 to S740 are based on the above. Figure 6 Some descriptions described in [the text]. Therefore, with [the text] Figure 1 and Figure 6 The detailed descriptions of the above content will be omitted or simplified.
[0192] like Figure 8 As shown, the encoding device according to the embodiment may include a quantization module 123 and an entropy encoding module 130. However, in some cases, Figure 8 All components shown may not be necessary; the encoding device may be derived from... Figure 8 The implementation of more or fewer components is shown.
[0193] The quantization module 123 and entropy encoding module 130 in the decoding device according to the embodiment can be implemented as separate chips, or at least two or more components can be implemented through a single chip.
[0194] The encoding device according to the embodiment can derive the expected average brightness value of the current block from nearby available samples (S700). More specifically, the quantization module 123 of the encoding device can derive the expected average brightness value of the current block from nearby available samples.
[0195] The encoding device according to the embodiment can derive a QP offset for deriving the luminance QP based on the expected average luminance value and information about QP (S710). More specifically, the quantization module 123 of the encoding device can derive a QP offset for deriving the luminance QP based on the expected average luminance value and information about QP.
[0196] The encoding device according to the embodiment can derive the luminance QP based on the QP offset (S720). More specifically, the quantization module 123 of the encoding device can derive the luminance QP based on the QP offset.
[0197] According to the embodiment, the encoding device can perform quantization on the quantization group including the current block based on the derived luminance QP (S730). More specifically, the quantization module 123 of the encoding device can perform quantization on the quantization group including the current block based on the derived luminance QP.
[0198] The encoding device according to the embodiment can encode image information including information about QP (S740). More specifically, the entropy encoding module 130 can encode image information including information about QP.
[0199] according to Figure 7 and Figure 8 According to the embodiment, the encoding device can derive the expected average brightness value of the current block from nearby available samples (S700), derive the QP offset for deriving the brightness QP based on the expected average brightness value and information about QP (S710), perform quantization on the quantization group including the current block based on the derived brightness QP (S730), and encode image information including information about QP (S740). Therefore, quantization parameters can be derived efficiently, and the overall encoding efficiency can be enhanced.
[0200] Figure 9 This is a flowchart illustrating the operation of the decoding device according to an embodiment. Figure 10 This is a block diagram illustrating the configuration of a decoding device according to an embodiment.
[0201] Figure 9 The steps disclosed in the document can be obtained by Figure 2 The decoding device 200 disclosed in the document performs the operation. More specifically, the S900 can be executed by... Figure 2 The entropy decoding module 210 shown is executed, and S910 to S940 can be performed by... Figure 2 The inverse quantization module 222 shown is executed, and S950 can be controlled by... Figure 2 The inverse transformation module 223 shown is executed, and S960 can be performed by... Figure 2 The prediction module 230 shown is executed, and S970 can be performed by... Figure 2 The adder 240 shown is executed. Furthermore, the operations according to S900 to S970 are based on the above. Figure 6Some descriptions described in [the text]. Therefore, with [the text] Figure 2 and Figure 6 The detailed descriptions of the above content will be omitted or simplified.
[0202] like Figure 10 As shown, the decoding device according to the embodiment may include an entropy decoding module 210, an inverse quantization module 222, an inverse transform module 223, a prediction module 230, and an adder 240. However, in some cases, Figure 10 All components shown may not be necessary; the encoding device may be derived from... Figure 10 The implementations shown include more or fewer components.
[0203] The entropy decoding module 210, inverse quantization module 222, inverse transform module 223, prediction module 230, and adder 240 in the decoding device according to the embodiment can be implemented as separate chips, or at least two or more components can be implemented through a single chip.
[0204] The decoding device according to the embodiment can decode image information including information about QP (S900). More specifically, the entropy decoding module 210 in the decoding device can decode image information including information about QP.
[0205] In one implementation, information about the QP is signaled at the Sequence Parameter Set (SPS) level.
[0206] In one implementation, the image information includes information about the Effective Data Range Parameter (EDRP), and the information about the EDRP includes at least one of a minimum input value, a maximum input value, a dynamic range of the input values, mapping information for relating the minimum input value to brightness, mapping information for relating the maximum input value to brightness, and identification information of the transfer function. This method can be instructed as range matching.
[0207] More specifically, this invention can be used for efficient encoding of image / video content where the range of codewords (input values) is limited. This often occurs in HDR content (due to the use of transfer functions that support high brightness). It can also occur when SDR data is transformed using a brightness transformation function corresponding to HDR data. In these cases, the video encoder can be configured to signal the Valid Data Range Parameter (EDRP). And the decoder can be configured to receive the EDRP associated with the video data and utilize the EDRP data in the decoding process. For example, the EDRP data may include a minimum input value, a maximum input value, the dynamic range of the input values (indicating the difference between the maximum and minimum input values), mapping information between the minimum input value and its corresponding brightness, mapping information between the maximum input value and its corresponding brightness, transfer function identifiers (known transfer functions can be identified by their assigned ID numbers, and detailed mapping information for each transfer function may be available), etc.
[0208] For example, EDRP data can be signaled in the slice header, Picture Parameter Set (PPS), or Sequence Parameter Set (SPS). This allows the EDRP data to be used to further modify the encoded values during the decoding process.
[0209] This invention introduces a Quality Control Parameter (QCP) to specify further adjustments to the quantization parameters. Furthermore, the decoder can be configured to receive the QCP associated with the video data and utilize the QCP data during decoding processing.
[0210] According to the embodiment, the decoding device can derive the expected average brightness value of the current block from the neighboring available samples (S910). More specifically, the inverse quantization module 222 in the decoding device can derive the expected average brightness value of the current block from the neighboring available samples.
[0211] The decoding device according to the embodiment can derive a QP offset for deriving the brightness QP based on the expected average brightness value and information about QP (S920). More specifically, the inverse quantization module 222 in the decoding device can derive a QP offset for deriving the brightness QP based on the expected average brightness value and information about QP.
[0212] In one implementation, the QP offset is derived based on the following formula.
[0213] [Mathematical Expression 17]
[0214] Luma_avg_qp=A*(avg_luma-M)+B
[0215] In the formula, Luma_avg_qp represents the QP offset, avg_luma represents the expected average brightness value, A represents the scaling factor used to map the pixel value difference to the QP difference, M represents a predefined value related to the bit depth, B represents the offset value, and A and B are predetermined values or values included in the image information.
[0216] In one implementation, the QP offset is derived from a mapping table based on the expected average brightness value, and the mapping table is determined using the expected average brightness value as input.
[0217] In one implementation, the QP offset is derived from a mapping table based on the expected average brightness value, and the mapping table is determined using a value obtained by dividing the expected average brightness value by a predefined constant value.
[0218] In one implementation, the nearest available samples include at least one of at least one luminance sample adjacent to the left boundary of the quantization group and at least one luminance sample adjacent to the upper boundary of the quantization group.
[0219] In one embodiment, at least one luminance sample adjacent to the left boundary of the quantization group is included in the luminance sample column directly adjacent to the left boundary of the quantization group, and at least one luminance sample adjacent to the upper boundary of the quantization group is included in the luminance sample row directly adjacent to the upper boundary of the quantization group.
[0220] In one implementation, the nearest available samples include the luminance samples that are adjacent to the left side of the top-left sample of the quantization group, and the nearest available samples include the luminance samples that are adjacent to the top side of the top-left sample of the quantization group.
[0221] In one implementation, the available neighbor samples include at least one of reconstructed neighbor samples, samples included in at least one reconstructed neighbor block, predicted neighbor samples, and samples included in at least one predicted neighbor block.
[0222] In one implementation, avg_luma can be derived from neighboring pixel values (blocks). avg_luma indicates the expected average brightness value obtained from available (already decoded) neighboring pixels (or blocks).
[0223] i) Available neighboring pixels may include:
[0224] - Pixels located in (xQg-1, yQg+K). Here, the luminance position (xQg, yQg) specifies the top-left luminance sample of the current quantization group relative to the top-left luminance sample of the current frame. (Leftmost line of the current block)
[0225] - Pixels located in (xQg+K, yQg-1). Here, the luminance position (xQg, yQg) specifies the top-left luminance sample of the current quantization group relative to the top-left luminance sample of the current frame. (Previous line of the current block)
[0226] - Instead of one line, multiple lines can be used.
[0227] ii) Available neighboring blocks can be used to compute avg_luma:
[0228] - Blocks containing pixels located at (xQg-1, yQg) can be used.
[0229] - Blocks containing pixels located at (xQg, yQg-1) can be used.
[0230] iii) The Avg_luma value can be calculated based on the reconstructed neighboring pixels / blocks.
[0231] iv) The Avg_luma value can be calculated based on the predicted path pixels / blocks.
[0232] In one implementation, the information about QP includes at least one syntax element related to QP offset, and deriving the QP offset based on at least one of the expected average luminance value and the information about QP includes deriving the QP offset based on at least one syntax element related to QP offset.
[0233] New syntax elements can be introduced to account for Luma_avg_qp. For example, the Luma_avg_qp value can be sent via a bitstream. It can also be specified by two syntax elements such as Luma_avg_qp_abs and Luma_avg_qp_flag. Luma_avg_qp can be indicated by two syntax elements (luma_avg_qp_delta_abs and luma_avg_qp_delta_sign_flag).
[0234] luma_avg_qp_delta_abs specifies the absolute value of the difference CuQpDeltaLumaVal between the current luminance quantization parameter and the luminance quantization parameter derived without considering luminance.
[0235] The following specifies the symbol for CuQpDeltaLumaVal:
[0236] If luma_avg_qp_delta_sign_flag equals 0, then the corresponding CuQpDeltaLumaVal has a positive value.
[0237] Otherwise (luma_avg_qp_delta_sign_flag equals 1), the corresponding CuQpDeltaLumaVal has a negative value.
[0238] When luma_avg_qp_delta_sign_flag does not exist, it is inferred to be equal to 0.
[0239] When luma_avg_qp_delta_sign_flag exists, the variables IsCuQpDeltaLumaCoded and CuQpDeltaLumaVal are deduced as follows:
[0240] [Mathematical Expression 18]
[0241] IsCuQpDeltaLumaCoded=1
[0242] CuQpDeltaLumaVal=cu_qp_delta_abs*(1-2*luma_avg_qp_delta_sign_flag)
[0243] CuQpDeltaLumaVal specifies the difference between the luminance quantization parameter of a coding unit that includes Luma_avg_qp and the luminance quantization parameter of a coding unit that does not have Luma_avg_qp.
[0244] In one implementation, the above syntax elements can be sent at the quantization group level (or quantization unit level) (e.g., CU, CTU, or predefined block unit).
[0245] The decoding device according to the embodiment can derive the luminance QP based on the QP offset (S930). More specifically, the inverse quantization module 222 in the decoding device can derive the luminance QP based on the QP offset.
[0246] According to the embodiment, the decoding device can perform inverse quantization on the quantization group including the current block based on the derived luminance QP (S940). More specifically, the inverse quantization module 222 in the decoding device can perform inverse quantization on the quantization group including the current block based on the derived luminance QP.
[0247] In one embodiment, the decoding device may derive a chroma QP from a derived luminance QP based on at least one chroma QP mapping table, and perform inverse quantization on a quantization group based on the derived luminance QP and the derived chroma QP, wherein the at least one chroma QP mapping table is based on the dynamic range of chroma and QP offset.
[0248] In one implementation, instead of a single chroma QP mapping table, multiple chroma QP derivation tables may exist. Additional information may be required to specify the QP mapping table to use. Below is an example of another chroma QP derivation table.
[0249] [Table 5]
[0250] <![CDATA[qPi Cb ]]> <26 26 27 28 29 30 31 32 33 34 35 36 37 38 39 >39 <![CDATA[qP Cb ]]> <![CDATA[=qPi Cb ]]> 26 27 27 28 29 29 30 30 31 31 32 32 33 33 <![CDATA[=qPi Cb -6]]>
[0251] In one embodiment, at least one chromaticity QP mapping table includes at least one Cb QP mapping table and at least one Cr QP mapping table.
[0252] According to the implementation, the decoding device can generate a residual sample of the current block based on inverse quantization (S950). More specifically, the inverse transform module 223 in the decoding device can generate a residual sample of the current block based on inverse quantization.
[0253] The decoding device according to the embodiment can generate a prediction sample of the current block based on image information (S960). More specifically, the prediction module 230 in the decoding device can generate a prediction sample of the current block based on image information.
[0254] According to the implementation method, the decoding device can generate a reconstructed sample of the current block based on the residual sample of the current block and the predicted sample of the current block (S970). More specifically, the adder in the decoding device can generate a reconstructed sample of the current block based on the residual sample of the current block and the predicted sample of the current block.
[0255] according to Figure 9 and Figure 10 According to the embodiment, the decoding device can decode image information including information about QP (S900), derive the expected average brightness value of the current block from neighboring available samples (S910), derive the QP offset for deriving the brightness QP based on the expected average brightness value and the information about QP (S920), derive the brightness QP based on the QP offset (S930), perform inverse quantization on the quantization group including the current block based on the derived brightness QP (S940), generate residual samples of the current block based on the inverse quantization (S950), generate prediction samples of the current block based on the image information (S960), and generate reconstructed samples of the current block based on the residual samples and prediction samples of the current block (S970). Therefore, quantization parameters can be derived efficiently, and the overall coding efficiency can be enhanced.
[0256] The method described above according to the present invention can be implemented in software, and the encoding and / or decoding devices according to the present invention can be included in image processing devices such as TVs, computers, smartphones, display devices, etc.
[0257] When embodiments of the present invention are implemented in software, the above-described methods can be implemented by modules (processes, functions, etc.) that perform the above-described functions. These modules are stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by any of a variety of well-known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices.
Claims
1. A decoding device that decodes an image, the decoding device comprising: a memory; and at least one processor connected to the memory, the at least one processor configured to: obtain, from a bitstream, image information including information related to a quantization parameter (QP), wherein the information related to the QP includes information related to a luma QP and information related to a chroma QP, and the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, derive the luma QP based on the information related to the luma QP, derive the chroma QP based on the information related to the chroma QP, wherein the chroma QP includes a first chroma QP for a Cb component and a second chroma QP for a Cr component, and generate a reconstructed picture including reconstructed samples based on the luma QP, the first chroma QP for the Cb component, and the second chroma QP for the Cr component, wherein the luma QP is derived based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the QP delta value is derived based on the information on the absolute value of the QP delta value and the information on the sign of the QP delta value, wherein the first chroma QP for the Cb component is derived based on a first chroma QP mapping table for the Cb component, wherein the second chroma QP for the Cr component is derived based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component. the chroma QP is derived based on the luma QP.
2. The decoding device of claim 1, wherein, the first chroma QP is derived based on first chroma QP offset information for the Cb component, and 3. The decoding device of claim 2, wherein, wherein the first chroma QP offset information includes a first PPS-level offset and a first slice-level offset. the second chroma QP is derived based on second chroma QP offset information for the Cr component, and 4. The decoding device of claim 2, wherein, wherein the second chroma QP offset information includes a second PPS-level offset and a second slice-level offset.
5. An encoding device that encodes an image, the encoding device comprising: a memory; and at least one processor connected to the memory, the at least one processor configured to: derive a luma quantization parameter (QP), derive a chroma QP, wherein the chroma QP includes a first chroma QP for a Cb component and a second chroma QP for a Cr component, generating information related to QPs, the information related to the QPs including information related to a luma QP and information related to a chroma QP, and encoding image information including the information related to the QPs, wherein the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, wherein the luma QP is represented based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the first chroma QP for the Cb component is derived based on a first chroma QP mapping table for the Cb component, wherein the second chroma QP for the Cr component is derived based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
6. A transmitting apparatus for data of image information, the transmitting apparatus comprising: at least one processor configured to obtain a bitstream of the image information, the image information including information related to quantization parameters (QPs), the information related to the QPs including information related to a luma QP and information related to a chroma QP, the chroma QP including a first chroma QP for a Cb component and a second chroma QP for a Cr component; and a transmitter configured to transmit the data including the bitstream of the image information, wherein the information related to the luma QP includes information on an absolute value of a QP delta value and information on a sign of the QP delta value, wherein the luma QP is represented based on a predicted luma QP and the QP delta value, wherein the predicted luma QP is derived based on a first luma QP of a coding unit containing a luma coding block covering (xQg-1, yQg) and a second luma QP of a coding unit containing a luma coding block covering (xQg, yQg-1), wherein (xQg, yQg) denotes a top-left sample position of a current quantization group, wherein the first chroma QP for the Cb component is derived based on a first chroma QP mapping table for the Cb component, wherein the second chroma QP for the Cr component is derived based on a second chroma QP mapping table for the Cr component, wherein the first chroma QP mapping table for the Cb component is a separate mapping table from the second chroma QP mapping table for the Cr component, and wherein the first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component. The first chroma QP mapping table for the Cb component is different from the second chroma QP mapping table for the Cr component.
Citation Information
Patent Citations
Derivation of color gamut scalability parameters and tables in scalable video coding
CN107690808A
Methods of determination for chroma quantization parameter and apparatuses for using the same
KR1020120100837A