Cross-Component Quantization in Video Coding

By using the default chromaticity delta QP and flag bits, the calculation burden problem of the encoder when determining chromaticity QP is solved, the encoding efficiency and quality are improved, and the bit allocation balance between brightness and chroma components is achieved.

CN113454998BActive Publication Date: 2025-07-11ZTE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980092594.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-05
Publication Date
2025-07-11
Estimated Expiration
2039-03-05

AI Technical Summary

Technical Problem

In the current H.266/VVC development, the encoder needs to determine the chromaticity delta QP by itself, resulting in excessive computational burden and affecting the encoding efficiency.

Method used

By calculating the quantization parameters of the chroma component using the default chroma delta QP, the encoder's calculation burden when determining the chroma QP is reduced, the flag bit is used to indicate whether the default chroma delta QP is used, and the flag bit and chroma delta QP are encoded in the bitstream.

Benefits of technology

It reduces the computational burden of the encoder, improves coding efficiency, realizes the bit distribution balance between brightness and chrominance components, and improves the encoding quality and compression effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113454998B_ABST
    Figure CN113454998B_ABST
Patent Text Reader

Abstract

A video decoding technique includes parsing a bitstream to determine a quantization parameter (QP) for a luminance component of a region from data units of a parameter set included in the bitstream, determining a flag from the data units of the parameter set, in a case where the flag is equal to a first value, determining a QP for a chrominance component of the region as a first function of the QP for the luminance component and a default chrominance delta QP indicated by the flag, in a case where the flag is equal to a second value, determining the QP for the chrominance component of the region by: (a) obtaining a delta chrominance QP from the data units, and (b) determining the QP for the chrominance component as a second function of the QP for the luminance component and the delta chrominance QP, and in a case where the parameter set is activated for decoding the region, using the QP for the chrominance component to decode the chrominance component.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent document generally relates to the encoding and decoding of video and images. Background Art

[0002] Video encoding uses compression tools to encode two-dimensional video frames into a compressed bitstream representation, which is more efficient for storage or transmission over a network. As the progress of devices capable of receiving, displaying digital video and images is increasing, the demand for digital video content is also increasing. The industry is currently working on defining next-generation video encoding technologies. Summary of the Invention

[0003] Among other things, this patent document describes techniques for encoding and decoding digital video using cross-component quantization parameter encoding embodiments.

[0004] The embodiments described in this application provide video or picture encoding and decoding methods, encoding and decoding devices to at least solve the problem of signaling default chrominance delta quantization parameter (QP), and to reduce the computational burden on the encoder when determining chrominance delta QP.

[0005] In one exemplary aspect, a method of processing visual information is disclosed. The method includes parsing a bitstream to determine a quantization parameter (QP) of a luminance component of a region of visual information from a data unit of a parameter set included in the bitstream, determining a flag from the data unit of the parameter set, in a case where the flag is equal to a first value, determining a QP for a chrominance component of the region as a first function of the QP for the luminance component and a default chrominance delta QP indicated by the flag, in a case where the flag is equal to a second value, determining a QP for the chrominance component of the region by: (a) obtaining a QP of a delta chrominance from the data unit, and (b) determining the QP for the chrominance component as a second function of the QP for the luminance component and the delta chrominance QP, and in a case where the parameter set is activated for decoding the region, using the QP for the chrominance component to decode the chrominance component.

[0006] In another example aspect, a method for generating an encoded representation of visual information is disclosed. The method includes determining a QP for a luminance component of a region for the visual information and signaling a QP for a chrominance component in the encoded representation of the visual information as follows: in a case where the QP for the chrominance component is equal to a value of a first function of a default QP for the chrominance component and the QP of the luminance component, indicating in the encoded representation by a first value of a flag bit that the default chrominance delta QP will be used for decoding and encoding the flag in a data unit of a parameter set of the encoded representation, otherwise, indicating in the encoded representation by a second value of the flag bit and indicating a value in a data unit of the delta chrominance QP such that the QP for the chrominance component is a second function of the luminance component and the delta chrominance QP.

[0007] In another example aspect, an apparatus for processing one or more bitstreams of video or pictures is disclosed.

[0008] In yet another example, a computer program storage medium is disclosed. The computer program storage medium includes code stored thereon. When the code is executed by a processor, it causes the processor to implement the described method.

[0009] These and other aspects will be described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1A is a flowchart of an example method for visual information processing.

[0011] Figure 1B is a flowchart of an example method for visual information encoding.

[0012] Figure 2 is a diagram showing an example video or picture encoder implementing the method in the present disclosure.

[0013] Figure 3 is a diagram showing an example of dividing a picture into tile groups.

[0014] Figures 4A - 4C is a diagram showing an example of a syntax structure of parameters for representing a chrominance QP in a bitstream.

[0015] Figure 5 is a diagram showing an example video or picture decoder implementing the method in the present disclosure.

[0016] Figure 6 is a diagram showing a first example device including at least the example encoder described in the present disclosure.

[0017] Figure 7 is a diagram showing a second example device including at least the example decoder described in the present disclosure.

[0018] Figure 8 FIG. is a diagram showing an electronic system including a first example device and a second example device. Specific embodiments

[0019] The section headings used in this document are for readability only and do not limit the scope of the embodiments and techniques disclosed in each section. Examples using the H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding) standards describe certain functions. However, the applicability of the disclosed techniques is not limited to H.264 / AVC or H.265 / HEVC systems.

[0020] This disclosure relates to video processing and communication, and in particular, to methods and apparatuses for encoding digital video or pictures to generate a bitstream, and methods and devices for decoding the bitstream to reconstruct digital video or pictures.

[0021] Brief discussion

[0022] In the current H.266 / VVC (Versatile Video Coding) WD (Working Draft), the picture-level chroma delta QP can be signaled in the PPS (Picture Parameter Set) (i.e., pps_cb_qp_offset and pps_cr_qp_offset) and the tile group header (i.e., tile_group_cb_qp_offset and tile_group_cr_qp_offset). The value of the tile group-level QP is set to be equal to the sum of the luma QP, the chroma delta QP signaled in the PPS, and the chroma delta QP signaled in the tile group header.

[0023] In the current H.266 / VVC development, advanced coding tools for chroma, such as dual tree (i.e., separate block partitioning for the luma and chroma components of the coding block), CCLM (Cross-Component Linear Model), etc., have brought great benefits to the coding efficiency of the chroma component. From the perspective of bit allocation, by shifting multiple bits from chroma to luma, the coding efficiency of the luma component will be further improved. This is demonstrated by setting an exemplary chroma delta QP (e.g., equal to 1) at the sequence (or at least picture) level in the common test conditions (CTC) for the H.266 / VVC software (i.e., VTM). As a result, in the bitstream, pps_cb_qp_offset and pps_cr_qp_offset are equal to 1, and tile_group_cb_qp_offset and tile_group_cr_qp_offset are equal to 0.

[0024] Under the advanced coding tools for chrominance, by shifting bits between the chrominance and luminance components to balance the compression gain of chrominance and luminance, the perceived quality is actually improved. In the current H.266 / VVC WD, there is no default (or recommended) value for such chrominance delta QP, and the encoder always needs to determine the chrominance delta QP by itself, which results in a heavy computational burden on the encoder.

[0025] Techniques for compressing digital videos and pictures utilize the correlation characteristics between pixel samples to remove redundancy in videos and pictures. By allocating bits between the luminance and chrominance components, the encoder achieves an operating point for the optimal trade-off between quality and compression when encoding videos. Generally, bit allocation can be performed by determining quantization parameters (QPs) for video sequences, pictures, different color components, and picture regions (where a picture region contains one or more coding blocks).

[0026] In the current H.266 / VVC (Versatile Video Coding) WD (Working Draft), picture-level chrominance delta QP can be signaled in the PPS (Picture Parameter Set) (i.e., pps_cb_qp_offset and pps_cr_qp_offset) and the tile group header (i.e., tile_group_cb_qp_offset and tile_group_cr_qp_offset). The value of the tile group-level QP is set to be equal to the sum of the luminance QP, the chrominance delta QP signaled in the PPS, and the chrominance delta QP signaled in the tile group header.

[0027] In the current H.266 / VVC development, advanced coding tools for chrominance, such as dual-tree (i.e., separate block partitioning for the luminance and chrominance components of coding blocks), CCLM (Cross-Component Linear Mode), etc., have brought great benefits to the coding efficiency of the chrominance component. From the perspective of bit allocation, by shifting multiple bits from chrominance to luminance, the coding efficiency of the luminance component can be further improved. This can be demonstrated by setting an exemplary chrominance delta QP (e.g., equal to 1) at the sequence (or at least picture) level under the common test conditions (CTC) for the H.266 / VVC software (i.e., VTM). As a result, in the bitstream, pps_cb_qp_offset and pps_cr_qp_offset are equal to 1, and tile_group_cb_qp_offset and tile_group_cr_qp_offset are equal to 0.

[0028] According to some embodiments described in this document, there is provided a coding method for processing videos or pictures, including:

[0029] Determine a quantization parameter (QP) for a luminance component of a picture or a picture region;

[0030] Calculate a first chrominance QP for a chrominance component of a picture or a picture region using a default chrominance delta QP;

[0031] If it is determined that the first chrominance QP is used to encode the chrominance component, set a flag to a first value indicating the use of the default chrominance delta QP, and encode the flag in a data unit of a parameter set in a bitstream;

[0032] Otherwise, determine a second chrominance QP for the chrominance component, set the flag to a second value indicating the non - use of the default chrominance delta QP, set the chrominance delta QP equal to the difference between the QP for the luminance component and the second chrominance QP, and encode the flag and the chrominance delta QP in a data unit of a parameter set in a bitstream.

[0033] According to some embodiments of this document, a decoding method for processing a bitstream to reconstruct a video or a picture is provided, including:

[0034] Parse the bitstream to determine a quantization parameter (QP) for a luminance component of a picture or a picture region from a data unit of a parameter set;

[0035] Parse the bitstream to obtain a flag from a data unit of a parameter set;

[0036] If the flag is equal to the first value, set the QP for the chrominance component of the picture or the picture region to be equal to the sum of the QP for the luminance component and the default chrominance delta QP indicated by the flag;

[0037] Otherwise, if the flag is equal to the second value, parse the bitstream to obtain a delta chrominance QP from a data unit of a parameter set, and set the QP for the chrominance component to be equal to the sum of the QP for the luminance component and the delta chrominance QP;

[0038] When the parameter set is activated for decoding a picture or a picture region, use the QP for the chrominance component to decode the chrominance component.

[0039] By the above method, the computational burden on the encoder for determining the chrominance QP is reduced because once the QP for the luminance component is determined, the chrominance QP can be obtained using the default chrominance delta QP, at least as a candidate chrominance QP.

[0040] In some embodiments, a video consists of a sequence of one or more pictures. A bitstream, sometimes also referred to as a video elementary stream, is generated by an encoder processing the video or pictures. The bitstream can also be a transport stream or a media file, which is the output of performing system layer processing on the video elementary stream generated by a video or picture encoder. Decoding the bitstream produces the video or pictures. System layer processing encapsulates the video elementary stream. For example, the video elementary stream is packed as a payload into a transport stream or a media file. System layer processing also includes operations of encapsulating the transport stream or the media file as a stream for transmission or a file for storage as a payload. A data unit generated in system layer processing is called a system layer data unit. Information appended to the system layer data unit during payload encapsulation in system layer processing is called system layer information, e.g., the header of the system layer data unit. Extracting the bitstream obtains a sub-bitstream, which contains a portion of the bits of the bitstream and one or more necessary modifications to syntax elements by the extraction process. Decoding the sub-bitstream produces a video or pictures, which may have a lower resolution and / or a lower frame rate compared to the video or pictures obtained by decoding the bitstream. The video or pictures obtained from the sub-bitstream can also be a region of the video or pictures obtained from the bitstream.

[0041] Embodiment 1

[0042] Figure 2 FIG. shows an encoder that encodes a video or pictures using the method of the present disclosure. The input of the encoder is a video, and the output is a bitstream. When the video includes a sequence of pictures, the encoder processes the images one by one in a preset order (i.e., the encoding order). The encoder order is determined according to the prediction structure specified in the configuration file of the encoder. Note that the encoding order of the pictures in the video (corresponding to the decoding order of the pictures at the decoder side) can be the same as or different from the display order of the pictures.

[0043] The partitioning unit 201 partitions pictures in the input video according to the configuration of the encoder. Generally, a picture can be partitioned into one or more largest coding blocks. The largest coding block is the largest allowed or configured block in the coding process and is typically a square region. A picture can be partitioned into one or more tiles, and a tile can contain an integer number of largest coding blocks, or a non-integer number of largest coding blocks. One option is that a tile can contain one or more stripes. That is, a tile can be further partitioned into one or more stripes, and each stripe can contain an integer number of largest coding blocks, or a non-integer number of largest coding blocks. Another option is that a stripe contains one or more tiles, or equivalently, a tile group contains one or more tiles. That is, one or more tiles in a specific order (e.g., raster scan order) in the picture form a tile group (or equivalently a stripe). In the following description, "tile group" is used as an example. The partitioning unit 201 can be configured to partition a picture using a fixed pattern. For example, the partitioning unit 201 partitions the picture into tile groups, and each tile group has a single tile that contains one row of largest coding blocks. Another example is that the partitioning unit 201 partitions the picture into multiple tiles and forms the tiles into tile groups in raster scan order in the picture. Alternatively, the partitioning unit 201 can also use a dynamic pattern to partition the picture into tile groups, tiles, and blocks. For example, to adapt to the limitations of the maximum transmission unit (MTU) size, the partitioning unit 201 adopts a dynamic tile group partitioning method to ensure that the number of coding bits of each tile group does not exceed the MTU limit.

[0044] Figure 3 FIG. is a diagram showing an example of partitioning a picture into tile groups. The partitioning unit 201 partitions a picture 30 (depicted by a dashed line) having 16 by 8 largest coding blocks into eight tiles 300, 310, 320, 330, 340, 350, 360, and 370. The partitioning unit 201 partitions the picture 30 into three tile groups. The tile group 3000 contains the tile 300, the tile group 3100 contains the tiles 310, 320, 330, 340, and 350, and the tile group 3200 contains the tiles 360 and 370. One or more tile groups or tiles can be referred to as picture regions. Generally, partitioning a picture into one or more tiles is performed according to the encoder configuration file. The partitioning unit 201 sets partitioning parameters to indicate the partitioning manner of partitioning the picture into tiles. For example, the partitioning manner can be to partition the picture into (nearly) equal-sized tiles. Another example is that the partitioning manner can indicate the positions of the tile boundaries in the rows and / or columns to facilitate flexible partitioning.

[0045] The output parameters of the partitioning unit 201 indicate the partitioning manner of the picture.

[0046] The prediction unit 202 determines the prediction samples of the coding block. The prediction unit 202 includes a block partitioning unit 203, a motion estimation (ME) unit 204, a motion compensation (MC) unit 205, and an intra prediction unit 206. The inputs to the prediction unit 202 are the largest coding block output by the partitioning unit 201 and the attribute parameters associated with the largest coding block, such as the position of the largest coding block in the picture, in the tile group, and / or in the tile. The prediction unit 202 uses the largest coding block divided into one or more coding blocks, and can further divide it into smaller coding blocks. One or more partitioning methods can be applied, including quadtree, binary split, and ternary split. The prediction unit 202 determines the prediction samples of the coding blocks obtained in the partitioning. Optionally, the prediction unit 202 can further divide the coding block into one or more prediction blocks to determine the prediction samples. The prediction unit 202 uses one or more pictures in the decoded picture buffer (DPB) unit 214 as references to determine the inter prediction samples of the coding block. The prediction unit 202 can also use the reconstructed part of the picture output by the adder 212 as a reference to derive the prediction samples of the coding block. The prediction unit 202 determines the prediction samples of the coding block and the related parameters for deriving the prediction samples, for example, by using a general rate distortion optimization (RDO) method, which are also the output parameters of the prediction unit 202.

[0047] Inside the prediction unit 202, the block partitioning unit 203 determines the partitioning of the coding block. The block partitioning unit 203 divides the largest coding block into one or more coding blocks, and can further divide it into smaller coding blocks. One or more partitioning methods can be applied, including quadtree, binary split, and ternary split. Optionally, the block partitioning unit 203 can further divide the coding block into one or more prediction blocks to determine the prediction samples. The block partitioning unit 203 can adopt the RDO method in determining the partitioning of the coding block. The output parameters of the block partitioning unit 203 include one or more parameters indicating the partitioning of the coding block.

[0048] The ME unit 204 and the MC unit 205 use one or more decoded pictures from the DPB 214 as reference pictures to determine the inter - frame prediction samples of the coded block. The ME unit 204 constructs one or more reference lists containing one or more reference pictures, and determines one or more matching blocks in the reference pictures for the coded block. The MC unit 205 uses the samples in the matching blocks to derive the prediction samples, and calculates the difference (i.e., the residual) between the original samples and the prediction samples in the coded block. The output parameters of the ME unit 204 indicate the positions of the matching blocks, including the reference list index, the reference index (refIdx), and the motion vector (MV), etc. Among them, the reference list index indicates the reference list containing the reference picture where the matching block is located, the reference index indicates the reference picture in the reference list containing the matching block, and the MV indicates the relative offset between the positions of the coded block and the matching block in the same coordinates to represent the positions of the pixels in the picture. The output parameters of the MC unit 205 are the inter - frame prediction samples of the coded block, and the parameters used to construct the inter - frame prediction samples, for example, the weighting parameter for the samples in the matching block, the filter type, and the parameters for filtering the samples in the matching block. Generally, the RDO method can be jointly applied to the ME unit 204 and the MC unit 205 to obtain the best matching block in terms of rate - distortion (RD) and the corresponding output parameters of the two units.

[0049] Specifically and optionally, the ME unit 204 and the MC unit 205 can use the current picture containing the coded block as a reference to obtain the intra - frame prediction samples of the coded block. Here, intra - frame prediction means using only the data in the picture containing the coded block as a reference for deriving the prediction samples of the coded block. In this case, the ME unit 204 and the MC unit 205 use the reconstructed part in the current picture, where the reconstructed part comes from the output of the adder 212. An example is that the encoder allocates a picture buffer to (temporarily) store the output data of the adder 212. Another method of the encoder is to reserve a special picture buffer in the DPB unit 214 to hold the data from the adder 212.

[0050] The intra prediction unit 206 obtains the intra prediction samples of the coded block by using the reconstructed part of the current picture including the coded block as a reference. The intra prediction unit 206 uses the reconstructed neighboring samples of the coded block as the input to the filter for deriving the intra prediction samples of the coded block, where the filter can be an interpolation filter (e.g., for calculating prediction samples when using angular intra prediction) or a low-pass filter (e.g., for calculating the DC value), or a cross-component filter to derive the prediction value of a (color) component by using the already coded (color) component. Specifically, the intra prediction unit 206 can perform a search operation to obtain a matching block of the coded block within the range of the reconstructed part in the current picture, and set the samples in the matching block as the intra prediction samples of the coded block. The intra prediction unit 206 calls the RDO method to determine the intra prediction mode (i.e., the method for calculating the intra prediction samples of the coded block) and the corresponding prediction samples. In addition to the intra prediction samples, the output of the intra prediction unit 206 further includes one or more parameters indicating the intra prediction mode in use.

[0051] The adder 207 is configured to calculate the difference between the original samples and the prediction samples of the coded block. The output of the adder 207 is the residual of the coded block. The residual can be represented as a two-dimensional matrix of N x M, where N and M are two positive integers, and N and M can be the same or different values.

[0052] The transform unit 208 takes the residual as its input. The transform unit 208 can apply one or more transform methods to the residual. From the perspective of signal processing, the transform method can be represented by a transform matrix. Optionally, the transform unit 208 can determine to use a rectangular block having the same shape and size as the shape and size of the coded block (in the text, a square block is a special case of a rectangular block) as the transform block for the residual. Optionally, the transform unit 208 can determine to divide the residual into several rectangular blocks (which can also include the special case where the width or height of the rectangular block is one sample), and sequentially perform transform operations on these several rectangles, for example, according to the default order (e.g., raster scan order), predefined order (e.g., the order corresponding to the prediction mode or transform method), the selected order for several candidate orders. The transform unit 208 can determine to perform multiple transforms on the residual. For example, the transform unit 208 first performs a core transform on the residual, and then performs a secondary transform on the coefficients obtained after the core transform is completed. The transform unit 208 can use the RDO method to determine the transform parameters, which indicate the execution manner used in the transform process applied to the residual block, such as dividing the residual block into transform blocks, the transform matrix, multiple transforms, etc. The transform parameters are included in the output parameters of the transform unit 208. The output parameters of the transform unit 208 include: the parameters and data obtained after the residual (e.g., transform coefficients) represented by a two-dimensional matrix is transformed.

[0053] After the transform unit 208 transforms the residual, the quantization unit 209 quantizes the data output by the transform unit 208. The quantizer used in the quantization unit 209 can be one or both of a scalar quantizer and a vector quantizer. In most video encoders, the quantization unit 209 employs a scalar quantizer. The quantization step of the scalar quantizer is represented by the quantization parameter (QP) in the video encoder. Generally, the same mapping between the QP and the quantization step is preset or predefined in the encoder and the corresponding decoder.

[0054] The value of the QP (e.g., picture-level QP and / or block-level QP) can be set according to the profile applied to the encoder or can be determined by the encoder control unit in the encoder. For example, the encoder control unit uses a rate control (RC) method to determine the quantization step for a picture and / or a block, and then converts the quantization step to a QP according to the mapping between the QP and the quantization step. Generally, color components (i.e., RGB components, luminance, and chrominance components) contribute and affect the perceived quality differently. For example, the human visual system (HVS) is more sensitive to the luminance component than the chrominance component. Therefore, from the perspective of bit allocation between color components, the encoder needs to achieve a reasonable trade-off between encoding quality and compression. To achieve the best trade-off, the quantization unit 209 uses different QPs to quantize color components.

[0055] An example implementation of the quantization unit 209 is to determine the QP for a picture or a picture region by calling a general RC method. The quantization unit 209 uses this QP as the QP for the luminance component of this picture or this picture region (i.e., luminance QP). In an example implementation of the encoder, a default chroma delta QP can be used to quickly calculate the QP for the chrominance component of this picture or this picture region (i.e., chroma QP), so as to save the computational burden of determining the chroma QP in the RC method. The chroma delta QP is the difference between the luminance QP and the chroma QP, which represents the relative trade-off of bit allocation between the luminance and chrominance components. The default chroma delta QP can be obtained through a large number of tests and experiments on video chips with various characteristics, and the most commonly used value of the chroma delta QP can be burned into the encoder (and the corresponding decoder).

[0056] Alternatively, since the profile / tier / level specifies the maximum size of the pictures supported in the encoder (and decoder) conforming to the profile / tier / level, various default chroma delta QPs can be set for various combinations of the profile / tier / level. For example, the default chroma delta QP for the RGB profile may be different from another default chroma delta QP for the YUV profile. The default chroma delta QP for the first level may be different from another default chroma delta QP for the second level.

[0057] When using the default chrominance QP, after obtaining the luma QP, the encoder calculates a first chrominance QP for the chrominance component of this picture or this picture region using the default chrominance QP. The encoder sets the first chrominance QP to be equal to the sum of the luma QP and the default chrominance QP.

[0058] The encoder determines whether the first chrominance QP is used to encode the chrominance component. For example, the encoder follows a quality standard designed to represent objective or subjective quality to calculate an indication value. If the indication value is within the allowable range, the encoder determines that the first chrominance QP is used as the chrominance QP for the picture or picture region. Otherwise, the encoder derives a second chrominance QP for the chrominance component. One method is to evaluate one or several adjacent values around the first chrominance QP and then select the first QP that results in an indication value within the allowable range, rather than invoking a computationally intensive algorithm, such as a general rate control method that first determines the bits allocated to the chrominance component.

[0059] The control parameter for quantization unit 209 is QP, including the parameter for chrominance QP. In addition to directly using the value of the chrominance QP as the parameter for chrominance QP, the parameter for chrominance QP can be a flag indicating whether to use the default chrominance QP, and if not, the value of the chrominance delta QP is equal to the difference between the QP for the luma component and the chrominance QP. The output of quantization unit 209 is one or more quantized transform coefficients (also called "Levels") represented in a two-dimensional matrix form.

[0060] Inverse quantization 210 performs a scaling operation on the output of quantization 209 to obtain reconstructed coefficients. Inverse transform unit 211 performs an inverse transform on the reconstructed coefficients from inverse quantization unit 210 according to the transform parameters from transform unit 208. The output of inverse transform unit 211 is the reconstructed residual. In particular, when the encoder determines to skip quantization in block coding (e.g., the encoder implements the RDO method to determine whether to apply quantization to the coding block), the encoder bypasses quantization unit 209 and inverse quantization unit 210 and directs the output data of transform unit 208 to inverse transform unit 211.

[0061] Adder 212 takes the predicted samples of the coded block from prediction unit 202 and the reconstructed residual as inputs, calculates the reconstructed samples of the coded block, and places the reconstructed samples in a buffer (e.g., a picture buffer). For example, the encoder allocates a picture buffer to (temporarily) store the output data of adder 212. Another method of the encoder is to reserve a special picture buffer in DPB 214 to hold the data from adder 212.

[0062] The filtering unit 213 performs a filtering operation on the reconstructed picture samples in the decoded picture buffer and outputs a decoded picture. The filtering unit 213 may include one filter or multiple cascaded filters. For example, according to the H.265 / HEVC standard, the filtering unit consists of two cascaded filters: namely, the deblocking filter and the sample adaptive offset (SAO) filter. The filtering unit 213 also includes an adaptive loop filter (ALF). The filtering unit 213 may also include a neural network filter. When the reconstructed samples of all the coded blocks in a picture have been stored in the decoded picture buffer, the filtering unit 213 may start filtering the reconstructed samples of the picture, which may be referred to as "picture-level filtering". Optionally, another implementation of the picture-level filtering of the filtering unit 213 (referred to as "block-level filtering") is: if the reconstructed samples are not used as references in all consecutive coded blocks in the coded picture, then start filtering the reconstructed samples of the coded blocks in the picture. Block-level filtering does not require the filtering unit 213 to maintain the filtering operation until all the reconstructed samples of the picture are available, thereby saving the time delay between threads in the encoder. The filtering unit 213 determines the filtering parameters by invoking the RDO method. The output of the filtering unit 213 is the decoded samples of the picture and the filtering parameters, which include the indication information of the filter, the filter coefficients, and the filter control parameters, etc.

[0063] The encoder stores the decoded picture from the filtering unit 213 in the DPB 214. The encoder may determine one or more instructions applied to the DPB 214, which are used to control the operations on the pictures in the DPB 214. For example, storing the temporal length of the picture in the DPB 214, and outputting the picture from the DPB 214, etc. In some embodiments, these instructions are taken as the output parameters of the DPB 214.

[0064] The entropy coding unit 215 performs binarization and entropy coding on one or more coding parameters of the picture to convert the values of the coding parameters into codewords consisting of binary symbols "0" and "1", and writes the codewords into a bitstream according to the specification or standard. The coding parameters can be classified into texture data and non-texture. The texture data is the transform coefficients of the coded blocks, and the non-texture data is other data in the coding parameters except the texture data, including the output parameters of the units in the encoder, the parameter sets, the headers, and the supplementary information, etc. The output of the entropy coding unit 215 is a bitstream that conforms to the specification or standard.

[0065] The entropy coding unit 215 receives output parameters from the quantization unit 209, in particular the parameters for the chroma QP. In particular, if the parameter for the chroma QP is the chroma QP for the chroma component of a picture or picture region, the entropy coding unit 215 converts the chroma QP into a flag indicating whether to use the default chroma delta QP, and, if not, the chroma delta QP is set to be equal to the difference between the luma QP and the chroma QP. If the parameter for the chroma QP is the above flag and (if necessary) the chroma delta QP, the entropy coding unit 215 does not call such a conversion operation on the parameter for the chroma QP.

[0066] Figures 4A - 4C FIG. is an example of a syntax structure for representing the parameters of the chroma QP in the bitstream, where Figures 4A - 4C the syntax shown in bold is a syntax element represented by a string of one or more bits present in the bitstream, and u(1) and se(v) are two coding methods having the same functions as those in the publicly disclosed standards such as H.264 / AVC and H.265 / HEVC.

[0067] Figure 4A The semantics of the syntax elements in are presented as follows.

[0068] default_chroma_qp_offset_flag being equal to 1 specifies that cb_qp_offset and cr_qp_offset do not exist in the bitstream. default_chroma_qp_offset_flag being equal to 0 specifies that cb_qp_offset and cr_qp_offset exist in the bitstream.

[0069] cb_qp_offset and cr_qp_offset respectively specify the offsets of the luma QP used to derive the chroma QP for the Cb and Cr components. The values of cb_qp_offset and cr_qp_offset will be in the range from -12 to +12 (including the end values). When not presented, the value of cb_qp_offset is inferred to be the first default value, and the value of cr_qp_offset is inferred to be the second default value.

[0070] The first default value and the second default value can be the same value or two different values.

[0071] Figure 4B The semantics of the syntax elements in are presented as follows. Using this syntax structure, the encoder has the ability to separately set the default chroma QP for different chroma components.

[0072] When default_cb_qp_offset_flag equals 1, it specifies that cb_qp_offset does not exist in the bitstream. When default_chroma_qp_offset_flag equals 0, it specifies that cb_qp_offset exists in the bitstream.

[0073] cb_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cb_qp_offset should be in the range of -12 to +12 (including the end values). When not presented, the value of cb_qp_offset is inferred as the first default value.

[0074] When default_cr_qp_offset_flag equals 1, it specifies that cr_qp_offset does not exist in the bitstream. When default_chroma_qp_offset_flag equals 0, it specifies that cr_qp_offset exists in the bitstream.

[0075] cr_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cr_qp_offset should be in the range of -12 to +12 (including the end values). When not presented, the value of cr_qp_offset is inferred as the first default value.

[0076] The first default value and the second default value can be the same value or two different values.

[0077] Figure 4C The semantics of the syntax elements in [ ] are presented as follows. Using this syntax structure, the encoder has the ability to set the default chroma QP separately for different chroma components while maintaining the ability to turn on the default QP for different chroma components simultaneously.

[0078] When default_chroma_qp_offset_flag equals 1, it specifies that default_cb_qp_offset_flag, cb_qp_offset, default_cr_qp_offset_flag, and cr_qp_offset do not exist in the bitstream. When default_chroma_qp_offset_flag equals 0, it specifies that default_cb_qp_offset_flag, cb_qp_offset, default_cr_qp_offset_flag, and cr_qp_offset exist in the bitstream.

[0079] When default_cb_qp_offset_flag equals 1, it specifies that cb_qp_offset does not exist in the bitstream. When default_chroma_qp_offset_flag equals 0, it specifies that cb_qp_offset exists in the bitstream. When not presented, the value of default_cb_qp_offset_flag is inferred to be 1.

[0080] cb_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cb_qp_offset can range from -12 to +12 (including the end values). When not present, the value of cb_qp_offset is inferred to be the first default value.

[0081] When default_cr_qp_offset_flag equals 1, it specifies that cr_qp_offset does not exist in the bitstream. When default_chroma_qp_offset_flag equals 0, it specifies that cr_qp_offset exists in the bitstream. When not presented, the value of default_cr_qp_offset_flag is inferred to be 1.

[0082] cr_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cr_qp_offset should range from -12 to +12 (including the end values). When not presented, the value of cr_qp_offset is inferred to be the first default value.

[0083] The first default value and the second default value can be the same value or two different values.

[0084] In one example, the entropy coding unit 215 encodes the parameters for the chroma QP in the data unit of the parameter set. The parameter set is applied in the encoding of a picture or a picture region (e.g., a tile group), such as a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS). The entropy coding unit 215 sets the value of the parameter set identifier in the header of the tile group to be equal to the identifier of the parameter set. In this way, the parameter set will be activated when decoding the tile group.

[0085] In another example, the entropy coding unit 215 encodes the parameters for the chroma QP in the data unit containing the header of the tile group. In this way, if another parameter set for the chroma QP is available in the parameter set activated in the decoding of the tile group, the parameters for the chroma QP in the header of the tile group override the parameters for the chroma QP in the parameter set in the decoding of the tile group.

[0086] Embodiment 2

[0087] Figure 5 It is a diagram showing that a decoder decodes a bitstream generated by the encoder in Embodiment 1 above using the method of this document. The input of the decoder is the bitstream, and the output of the decoder is the decoded video or picture obtained by decoding the bitstream.

[0088] The parsing unit 501 in the decoder parses the input bitstream. The parsing unit 501 uses the entropy decoding method and the binarization method specified in the standard to convert each codeword in the bitstream composed of one or more binary symbols (i.e., "0" and "1") into a numerical value of the corresponding parameter. The parsing unit 501 also derives parameter values based on one or more available parameters. For example, when there is a flag in the bitstream indicating that the decoded block is the first decoded block in the picture, the parsing unit 501 sets the address parameter (which indicates the address of the first decoded block in the picture area) to 0.

[0089] The parsing unit 501 parses one or more data units in the bitstream, such as parameter set data units and tile group data units, to obtain parameters for the chrominance QP, including a flag indicating whether to use the default chrominance delta QP, and, if not, the chrominance delta QP is the difference between the luma QP and the chrominance QP.

[0090] Figures 4A - 4C It is a diagram showing an example of the syntax structure for representing the parameters for the chrominance QP, where Figures 4A - 4C the bold syntax is a syntax element represented by a string of one or more bits present in the bitstream, and u(1) and se(v) are two decoding methods, whose functions are the same as those in the publicly available standards such as H.264 / AVC and H.265 / HEVC.

[0091] Figure 4A 、 4B and 4C are diagrams showing examples of the syntax structure for the parameters for the chrominance QP parsed from the bitstream by the parsing unit 501. Figure 4A 、 4BThe syntax structures of 4C can be parsed separately and obtained from one or more data units in the bitstream. Optionally, the data unit can be a parameter set data unit. The parameter set is applied to decode a picture or a picture region (such as a tile group), such as a sequence parameter set (SPS), a picture parameter set (PPS), and an adaptive parameter set (APS). The parsing unit 501 obtains the identifier value of the parameter set identifier from the header of the tile group, activates the parameter set whose identifier is equal to the identifier value, and obtains the chroma QP from the activated parameter set. Optionally, the data unit can be a data unit containing the header of the tile group. In addition, if the parsing unit 501 first obtains the parameter for the chroma QP from the parameter set and then from the header of the tile group, the parsing unit 501 overrides the chroma QP obtained from the parameter set with the chroma QP obtained from the tile group. The decoder uses this chroma QP when decoding the chroma component.

[0092] Figure 4A The semantics of the syntax elements in

[0093] When default_chroma_qp_offset_flag is equal to 1, it specifies that cb_qp_offset and cr_qp_offset do not exist in the bitstream. When default_chroma_qp_offset_flag is equal to 0, it specifies that cb_qp_offset and cr_qp_offset exist in the bitstream.

[0094] cb_qp_offset and cr_qp_offset respectively specify the offsets of the luma QP for deriving the chroma QP of the Cb and Cr components. The values of cb_qp_offset and cr_qp_offset can be in the range of -12 to +12 (including the end values). When not presented, the value of cb_qp_offset is inferred as the first default value, and the value of cr_qp_offset is inferred as the second default value.

[0095] The first default value and the second default value can be the same value or two different values.

[0096] Figure 4B The semantics of the syntax elements in

[0097] When default_cb_qp_offset_flag is equal to 1, it specifies that cb_qp_offset does not exist in the bitstream. When default_chroma_qp_offset_flag is equal to 0, it specifies that cb_qp_offset exists in the bitstream.

[0098] cb_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cb_qp_offset should be in the range of -12 to +12 (including the end values). When not present, the value of cb_qp_offset is inferred as the first default value.

[0099] default_cr_qp_offset_flag being equal to 1 specifies that cr_qp_offset is not present in the bitstream. default_chroma_qp_offset_flag being equal to 0 specifies that cr_qp_offset is present in the bitstream.

[0100] cr_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cr_qp_offset should be in the range of -12 to +12 (including the end values). When not present, the value of cr_qp_offset is inferred as the first default value.

[0101] The first default value and the second default value can be the same value or two different values.

[0102] Figure 4C The semantics of the syntax elements in [ ] are presented as follows. Using this syntax structure, the encoder has the ability to set the default chroma QP separately for different chroma components while maintaining the ability to turn on the default QP for different chroma components simultaneously.

[0103] default_chroma_qp_offset_flag being equal to 1 specifies that default_cb_qp_offset_flag, cb_qp_offset, default_cr_qp_offset_flag, and cr_qp_offset are not present in the bitstream. default_chroma_qp_offset_flag being equal to 0 specifies that default_cb_qp_offset_flag, cb_qp_offset, default_cr_qp_offset_flag, and cr_qp_offset are present in the bitstream.

[0104] default_cb_qp_offset_flag being equal to 1 specifies that cb_qp_offset is not present in the bitstream. default_chroma_qp_offset_flag being equal to 0 specifies that cb_qp_offset is present in the bitstream. When not present, the value of default_cb_qp_offset_flag is inferred as 1.

[0105] cb_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cb_qp_offset shall be in the range of -12 to +12 (including the end values). When not present, the value of cb_qp_offset is inferred as the first default value.

[0106] default_cr_qp_offset_flag being equal to 1 specifies that cr_qp_offset does not exist in the bitstream. default_chroma_qp_offset_flag being equal to 0 specifies that cr_qp_offset exists in the bitstream. When not present, the value of default_cr_qp_offset_flag is inferred as 1.

[0107] cr_qp_offset specifies the offset of the luma QP used to derive the chroma QP for the Cb component. The value of cr_qp_offset shall be in the range of -12 to +12 (including the end values). When not present, the value of cr_qp_offset is inferred as the first default value.

[0108] The first default value and the second default value can be the same value or two different values.

[0109] Parse unit 501 obtains the chroma QP through the following steps.

[0110] Step 1: Parse unit 501 parses the input bitstream to determine the QP of the luma component for the picture or picture region from the data units in the bitstream. The data units can be the headers of tile groups and / or one or more parameter sets. Parse unit 501 calculates the luma QP (denoted as QpY) using the same method as H.266 / VVC WD.

[0111] Step 2: Parse unit 501 obtains a flag from the data units. This flag indicates whether to use the default chroma delta QP to derive the chroma QP of the chroma component for the picture or picture region.

[0112] Step 3: Parse unit 501 checks whether the flag is equal to the first value (e.g., "1"). If so, parse unit 501 sets the chroma QP to be equal to the sum of QpY and the default chroma delta QP (denoted as "defaultChromaDeltaQP") indicated by the flag. Otherwise, if the flag is equal to the second value (e.g., "0"), then parse unit 501 parses the bitstream to obtain the delta chroma QP from the data units and sets the QP for the chroma component to be equal to the sum of the QP for the luma component and the delta chroma QP.

[0113] If used Figure 4AFor the syntax elements in, the parsing unit 501 sets the value of the flag to be equal to the value of default_chroma_qp_offset_flag for the Cb and Cr components. When default_chroma_qp_offset_flag is equal to 1, for the Cb component, the parsing unit 501 sets defaultChromaDeltaQP to be equal to the first default value specified in the semantics of cb_qp_offset, and sets the chroma QP for the Cb component to be equal to QpY + defaultChromaDelta QP; for the Cr component, the parsing unit 501 sets defaultChromaDelta QP to be equal to the second default value specified in the semantics of cr_qp_offset, and sets the chroma QP for the Cb component to be equal to QpY + defaultChromaDelta QP. When default_chroma_qp_offset_flag is equal to 0, the parsing unit 501 obtains the value of cb_qp_offset from the data unit, and sets the chroma delta QP for the Cb component to be equal to QpY + cb_qp_offset; the parsing unit 501 obtains the value of cr_qp_offset from the data unit and sets the chroma delta QP for the Cr component to be equal to QpY + cr_qp_offset.

[0114] If used Figure 4BFor the syntax elements in, for the Cb component, the parsing unit 501 sets the flag equal to the value of default_cb_qp_offset_flag. When default_cb_qp_offset_flag is equal to 1, the parsing unit 501 sets defaultChromaDeltaQP equal to the first default value specified in the semantics of cb_qp_offset, and sets the chroma QP for the Cb component equal to QpY + defaultChromaDeltaQP. When default_cb_qp_offset_flag is equal to 0, the parsing unit 501 obtains the value of cb_qp_offset from the data unit, and sets the chroma delta QP for the Cb component equal to QpY + cb_qp_offset. For the Cr component, the parsing unit 501 sets the flag equal to the value of default_cr_qp_offset_flag. When default_cr_qp_offset_flag is equal to 1, the parsing unit 501 sets defaultChromaDeltaQP equal to the second default value specified in the semantics of cr_qp_offset, and sets the chroma QP for the Cb component equal to QpY + defaultChromaDeltaQP. When default_cr_qp_offset_flag is equal to 0, the parsing unit 501 obtains the value of cr_qp_offset from the data unit, and sets the chroma delta QP for the Cr component equal to QpY + cr_qp_offset.

[0115] If used Figure 4CFor the syntax element in it, for the Cr component, the parsing unit 501 sets the flag to be equal to the value of (default_chroma_qp_offset_flag && default_cb_qp_offset_flag). When the flag is equal to 1, the parsing unit 501 sets defaultChromaDeltaQP to be equal to the first default value specified in the cb_qp_offset semantics, and sets the chroma QP for the Cb component to be equal to QpY + defaultChromaDeltaQP. When the flag is equal to 0, the parsing unit 501 obtains the value of cb_qp_offset from the data unit, and sets the chroma delta QP for the Cb component to be equal to QpY + cb_qp_offset. For the Cr component, the parsing unit 501 sets the flag to be equal to the value of (default_chroma_qp_offset_flag && default_cr_qp_offset_flag). When the flag is equal to 1, the parsing unit 501 sets defaultChromaDeltaQP to be equal to the second default value specified in the cr_qp_offset semantics, and sets the chroma QP for the Cb component to be equal to QpY + defaultChromaDeltaQP. When the flag is equal to 0, the parsing unit 501 obtains the value of cr_qp_offset from the data unit, and sets the chroma delta QP for the Cr component to be equal to QpY + cr_qp_offset.

[0116] Corresponding to the implementation of the encoder, the default chroma delta QP can be obtained through a large number of tests and experiments with video chips of various characteristics, and the most commonly used values of the chroma delta QP can be burned into the decoder (and the corresponding encoder). That is, the default chroma delta QP, namely the first default value of cb_qp_offset and the second default value of cr_qp_offset, can be fixed values preset in the codec.

[0117] Alternatively, since the profile / tier / level specifies the maximum size of the pictures supported in a decoder conforming to that profile / tier / level, various default chroma delta QPs can be set for various combinations of that profile / tier / level. For example, the default chroma delta QP for an RGB profile may be different from another default chroma delta QP for a YUV profile. The default chroma delta QP for a first level can be different from another default chroma delta QP for a second level. The parsing unit 501 obtains a first default value of cb_qp_offset and a second default value of cr_qp_offset according to the profile / tier / level information specified in the input bitstream.

[0118] Taking the Cb component as an example. Another implementation of step 3 is that the parsing unit 501 determines the value of cb_qp_offset according to the flag. When the flag is equal to the first value, the parsing unit 501 sets cb_qp_offset to the first default value; otherwise, the parsing unit 501 obtains the value of cb_qp_offset from the data unit. Finally, the parsing unit 501 sets the chroma QP for the Cb component to be equal to QpY+cb_qp_offset. The parsing unit 501 also applies the same steps as above to determine the chroma QP for the Cr component.

[0119] The parsing unit 501 passes the above chroma QP values to other units in the decoder and uses them to decode the corresponding chroma components.

[0120] The parsing unit 501 can pass one or more prediction parameters for deriving the predicted samples of the decoded block to the prediction unit 502. Herein, the prediction parameters can include the output parameters of the partitioning unit 201 and the prediction unit 202 in the above encoder.

[0121] The parsing unit 501 can pass one or more residual parameters for reconstructing the residual of the decoded block to the scaling unit 505 and the transform unit 506. Herein, the residual parameters can include the output parameters of the transform unit 208 and the quantization unit 209 and one or more quantized parameters (e.g., "levels") output by the quantization unit 209 in the above encoder.

[0122] The parsing unit 501 passes the filtering parameters to the filtering unit 508 to filter the reconstructed samples in the picture (e.g., loop filtering).

[0123] The prediction unit 502 may derive the prediction samples of the decoded block according to the prediction parameters. The prediction unit 502 is composed of the MC unit 503 and the intra prediction unit 504. The input of the prediction unit 502 may further include the reconstructed part of the current decoded picture (not processed by the filter unit 508) output from the adder 507 and one or more decoded pictures in the DPB 509.

[0124] When the prediction parameters indicate that the inter prediction mode is used to derive the prediction samples of the decoded block, the prediction unit 502 constructs one or more reference picture lists in the same way as the ME unit 204 in the foregoing encoder. The reference list contains one or more reference pictures from the DPB 509. The MC unit 503 determines one or more matching blocks for the decoded block according to the indication of the reference list, the reference index, and the MV in the prediction parameters, and obtains the inter prediction samples of the decoded block in the same way as the MC unit 205 in the foregoing encoder. The prediction unit 502 outputs the inter prediction samples as the prediction samples of the decoded block.

[0125] Specifically and optionally, the MC unit 503 may use the current decoded picture containing the decoded block as a reference to obtain the intra prediction samples of the decoded block. In this document, intra prediction means using only the data in the picture containing the encoded block as a reference for deriving the prediction samples of the encoded block. In this case, the MC unit 503 uses the reconstructed part in the current picture, where the reconstructed part comes from the output of the adder 507 and is not processed by the filter unit 508. For example, the decoder allocates a picture buffer to (temporarily) store the output data of the adder 507. Another method of the decoder is to reserve a special picture buffer in the DPB 509 to hold the data from the adder 507.

[0126] When the prediction parameter indicates that the intra prediction mode is used to derive the prediction samples of the decoded block, the prediction unit 502 determines the reference samples of the intra prediction unit 504 from the reconstructed neighboring samples of the decoded block in the same way as the intra prediction unit 206 in the aforementioned encoder. The intra prediction unit 504 obtains the intra prediction mode (i.e., the DC mode, the planar mode, or the angular prediction mode), and after the specified processing of the intra prediction mode, uses the reference samples to derive the intra prediction samples of the decoded block. Note that the same derivation process of the intra prediction mode is implemented in the aforementioned encoder (e.g., the intra prediction unit 206) and the decoder (e.g., the intra prediction unit 504). Specifically, if the prediction parameter indicates the matching block (including its position) for the decoded block in the current decoded picture (including the decoded block), the intra prediction unit 504 uses the samples in the matching block to derive the intra prediction samples of the decoded block. For example, the intra prediction unit 504 sets the intra prediction samples to be equal to the samples in the matching block. The prediction unit 502 sets the prediction samples of the decoded block to be equal to the intra prediction samples output by the intra prediction unit 504.

[0127] The decoder passes the QP (including the luminance QP and the chrominance QP) and the quantized coefficients to the scaling unit 505 for inverse quantization processing to obtain the reconstructed coefficients as the output. The decoder feeds the reconstructed coefficients from the scaling unit 505 and the transform parameters in the residual parameters (i.e., the transform parameters in the output of the transform unit 208 in the aforementioned encoder) into the transform unit 506. Specifically, if the residual parameter indicates skipping the scaling during decoding of the block, the decoder bypasses the scaling unit 505 and directs the coefficients in the residual parameters to the transform unit 506.

[0128] The transform unit 506 performs a transform operation on the input coefficients after the transform processing specified in the standard. The transform matrix used in the transform unit 506 is the same as the transform matrix used in the inverse transform unit 211 in the aforementioned encoder. The output of the transform unit 506 is the reconstructed residual of the decoded block.

[0129] Generally, since only the decoding process is specified in the standard, from the perspective of the video coding standard, the processing and related matrices in the decoding process are specified as "transform processing" and "transform matrix" in the standard text. Therefore, in this document, the unit that implements the transform processing specified in the standard text for the decoder is named "transform unit" to conform to the standard. However, considering that the decoding process is regarded as the inverse process of encoding, this unit can always be named "inverse transform unit".

[0130] The adder 507 takes the reconstructed residual in the output of the transform unit 506 and the predicted samples in the output of the prediction unit 502 as input data, and calculates the reconstructed samples of the decoded block. The adder 507 stores the reconstructed samples in the picture buffer. For example, the decoder allocates a picture buffer to (temporarily) store the output data of the adder 507. Another method of the decoder is to reserve a special picture buffer in the DPB 509 to hold the data from the adder 507.

[0131] The decoder passes the filtering parameters from the parsing unit 501 to the filtering unit 508. The filtering parameters for the filtering unit 508 are the same as the filtering parameters in the output of the filtering unit 213 in the aforementioned encoder. The filtering parameters include indication information of one or more filters to be used, filter coefficients, and filtering control parameters. The filtering unit 508 performs a filtering process on the reconstructed samples of the picture stored in the decoded picture buffer using the filtering parameters, and outputs the decoded picture. The filtering unit 508 can be composed of one filter or multiple cascaded filters. For example, according to the H.265 / HEVC standard, the filtering unit is composed of two cascaded filters: namely, the deblocking filter and the sample adaptive offset (SAO) filter. The filtering unit 508 can include an adaptive loop filter (ALF). The filtering unit 508 can also include a neural network filter. When the reconstructed samples of all the coded blocks in the picture have been stored in the decoded picture buffer, the filtering unit 508 can start filtering the reconstructed samples of the picture, which can be referred to as "picture layer filtering". Optionally, an alternative implementation (referred to as "block layer filtering") for the picture layer filtering of the filtering unit 508 is: if the reconstructed samples are not used as references when decoding all the consecutive coded blocks in the picture, start filtering the reconstructed samples of the coded blocks in the picture. Block layer filtering does not require the filtering unit 508 to maintain the filtering operation until all the reconstructed samples of the picture are available, thus saving the time delay between threads in the decoder.

[0132] The decoder stores the decoded picture output by the filtering unit 508 in the DPB 509. In addition, the decoder can perform one or more control operations on the pictures in the DPB 509 according to one or more instructions output by the parsing unit 501. For example, the temporal length of the picture is stored in the DPB 509, and the picture is output from the DPB 509, and so on.

[0133] Embodiment 3

[0134] Figure 6 is a diagram showing a first example device that includes at least an example video encoder or picture encoder as Figure 2 shown.

[0135] The acquisition unit 601 captures videos and pictures. The acquisition unit 601 can be equipped with one or more cameras to shoot videos or pictures of natural scenes. Optionally, the acquisition unit 601 can be implemented with cameras for obtaining depth videos or depth pictures. Optionally, the acquisition unit 601 can include components of an infrared camera. Optionally, the acquisition unit 601 can be configured with a remote sensing camera. The acquisition unit 601 can also be a device or apparatus that generates videos or pictures by using radiation to scan an object.

[0136] Optionally, the acquisition unit 601 can perform preprocessing on the videos or pictures, such as automatic white balance, autofocus, automatic exposure, backlight compensation, sharpening, denoising, stitching, upsampling / downsampling, frame rate conversion, virtual view synthesis, etc.).

[0137] The acquisition unit 601 can also receive videos or pictures from another device or processing unit. For example, the acquisition unit 601 can be a component unit in a code converter. The code converter feeds one or more decoded (or partially decoded) pictures to the acquisition unit 601. Another example is that the acquisition unit 601 obtains videos or pictures from another device via a data link with that device.

[0138] Note that in addition to videos and pictures, the acquisition unit 601 can also be used to capture other media information, such as audio signals. The acquisition unit 601 can also receive artificial information, such as characters, text, and computer-generated videos or pictures, etc.

[0139] The encoder 602 is Figure 2 the illustrated example implementation of an encoder. The input of the encoder 602 is the video or picture output by the acquisition unit 601. The encoder 602 encodes the video or picture and outputs the generated video or picture bitstream.

[0140] The storage / transmission unit 603 receives the video or picture bitstream from the encoder 602 and performs system layer processing on the bitstream. For example, the storage / transmission unit 603 encapsulates the bitstream according to transmission standards and media file formats (such as MPEG-2TS, ISOBMFF, DASH, MMT, etc.). The storage / transmission unit 603 stores the obtained transport stream or media file after encapsulation in the memory or disk of the first example device, or transmits the transport stream or media file via a wired or wireless network.

[0141] Note that in addition to the video or picture bitstream from the encoder 602, the input of the storage / transmission unit 603 can also include audio, text, images, and graphics, etc. The storage / transmission unit 603 generates a transport or media file by encapsulating such different types of media bitstreams.

[0142] The first example device described in this embodiment can be a device capable of generating or processing video (or picture) bitstreams in applications of video communication (e.g., mobile phones, computers, media servers, portable mobile terminals, digital cameras, broadcast devices, content delivery network devices (CDN), surveillance cameras, video conferencing devices, etc.).

[0143] Embodiment 4

[0144] Figure 7 is a diagram showing a second example device that includes at least an example video decoder or picture decoder as shown in Figure 5 the figure.

[0145] The receiving unit 701 receives video or picture bitstreams by obtaining bitstreams from wired or wireless networks, by reading memories or disks in electronic devices, or by obtaining data from other devices via data links.

[0146] The input of the receiving unit 701 may also include transport streams or media files containing video or picture bitstreams. The receiving unit 701 extracts video or picture bitstreams from transport streams or media files according to the specifications of transport or media file formats.

[0147] The receiving unit 701 outputs and delivers the video or picture bitstream to the decoder 702. Note that in addition to video or picture bitstreams, the output of the receiving unit 701 may also include audio bitstreams, characters, texts, images, graphics, etc. The receiving unit 701 outputs to the corresponding processing unit in the second example device. For example, the receiving unit 701 delivers the output audio bitstream to the audio decoder in the device.

[0148] The decoder 702 is Figure 5 an implementation of the example decoder shown in the figure. The input of the decoder 702 is the video or picture bitstream output by the receiving unit 701. The decoder 702 decodes the video or picture bitstream and outputs the decoded video or picture.

[0149] The presentation unit 703 receives the decoded video or picture from the decoder 702. The presentation unit 703 presents the decoded video or picture to the viewer. The presentation unit 703 can be a component of the second example device, e.g., a screen. The presentation unit 703 can also be a device separated from the second example device and having a data link with the second example device, e.g., a projector, a monitor, a television, etc. Optionally, the presentation 703 performs post-processing on the decoded video or picture before presenting it to the viewer, e.g., automatic white balance, autofocus, autoexposure, backlight compensation, sharpening, denoising, stitching, upsampling / downsampling, frame rate conversion, virtual view synthesis, etc.

[0150] Note that, in addition to decoded video or pictures, the input to the rendering unit 703 can be other media data from one or more units of the second example device, such as audio, characters, text, images, graphics, etc. The input to the rendering unit 703 can also include artificial data, such as lines and marks drawn by a local teacher on a slide to draw attention in a distance education application. The rendering unit 703 combines different types of media and then presents the work to the viewer.

[0151] The second example device described in this embodiment can be a device capable of decoding or processing a video (or picture) bitstream in an application of video communication, such as a mobile phone, a computer, a set-top box, a television, an HMD, a monitor, a media server, a portable mobile terminal, a digital camera, a broadcast device, a content delivery network device (CDN), a surveillance device, a video conferencing device, etc.

[0152] Embodiment 5

[0153] Figure 8 is a diagram of an electronic system that includes Figure 6 the first example device shown and Figure 7 the second example device shown.

[0154] The service device 801 is Figure 6 the first example device in

[0155] The storage medium / transmission network 802 can include internal storage resources of a device or an electronic system, external storage resources accessible via a data link, and / or a data transmission network composed of wired and / or wireless networks. The storage medium / transmission network 802 provides storage resources or a data transmission network for the storage / sending unit 603 in the service device 801.

[0156] The destination device 803 is Figure 7 the second example device in . The receiving unit 701 in the destination device 803 receives a video or picture bitstream, a transport stream containing the video or picture bitstream, or a media file containing the video or picture bitstream from the storage medium / transmission network 802.

[0157] The electronic system described in this embodiment can be a device or system capable of generating, storing or transmitting, and decoding a video (or picture) bitstream in an application of video communication (such as a mobile phone, a computer, an IPTV system, an OTT system, a multimedia system on the Internet, a digital television broadcast system, a video surveillance system, a portable mobile terminal, a digital camera, and a video conferencing system, etc.).

[0158] In one embodiment, specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated herein again.

[0159] Obviously, those skilled in the art should know that each module or each behavior in this document can be implemented by a general-purpose computing device, and these modules or behaviors can be concentrated on a single computing device, or can be distributed on a network formed by multiple computing devices, and can optionally be implemented by program code for the computing device to execute, so that these modules or behaviors can be stored in a storage device for the computing device to execute. The behaviors shown or described may, in some cases, be executed in a different order than that shown or described here, or may separately form a single integrated circuit module, or multiple of these modules or behaviors may form a single integrated circuit module for implementation. Therefore, the disclosed technology is not limited to any specific combination of hardware and software.

[0160] Figure 1A A flowchart of an example method 100 for processing visual information is shown. The method may be implemented by a decoder device that generates video or picture frames from a compressed bitstream representation using method 100.

[0161] Method 100 includes parsing (102) the bitstream to determine a quantization parameter (QP) for the luminance component of a region for visual information from a data unit of a parameter set included in the bitstream, determining (104) a flag from the data unit of the parameter set, and in the case where the flag is equal to a first value, determining (106) the QP for the chrominance component of the region as a first function of the QP for the luminance component and a default chrominance delta QP indicated by the flag. In the case where the flag is equal to a second value, determining (108) the QP for the chrominance component of the region by: (a) obtaining (110) a delta chrominance QP from the data unit, and (b) determining (112) the QP for the chrominance component as a second function of the QP for the luminance component and the delta chrominance QP, and decoding (114) the chrominance component using the QP for the chrominance component in the case where the parameter set is activated for decoding the region.

[0162] As indicated by the dashed lines of the boxes corresponding to 106 and 108, it should be understood that during the operation of parsing (102), for a given video region, the video decoding device, when parsing to obtain a flag, the value of the flag can trigger step 106 (the flag is equal to the first value) or step 108 (the flag is equal to a second value different from the first value), but cannot trigger both at the same time. In some cases, there may not be a flag for the region at all (e.g., encoding it without using the advantages of cross-component QP encoding described in this document). Therefore, method 100 can be performed by implementing step 106 or step 108, but not both at the same time for the same region of the video being decoded.

[0163] Figure 1B A flowchart of an example method 150 is shown. Method 150 includes determining (152) the QP of the luminance component of a region for visual information and signaling (154) the QP of the chrominance component in the encoded representation of the visual information as follows: in the case where the QP for the chrominance component is equal to the value of a first function of the default QP for the chrominance component and the QP of the luminance component, indicate in the encoded representation with (156) a first value of a flag bit that the default chroma delta QP is to be used for decoding and encode the flag in a data unit of the parameter set of the encoded representation, otherwise in the encoded representation, indicate with (158) a second value of the flag bit and indicate a value in the data unit to indicate the delta chroma QP such that the QP for the chrominance component is a second function of the luminance component and the delta chroma QP.

[0164] It should be understood that a video encoding device implementing method 150 can implement step 156 or step 158 for encoding a video region, but not both for the same video region. In some cases, a video encoder that performs region encoding without using the cross-component QP signaling technique described in this document may not be able to implement both step 156 and 158. Therefore, as Figure 1B indicated by the dashed borders in, steps 156 and 158 can be implemented as two different and mutually exclusive options for encoding a video region.

[0165] In some embodiments, in methods 100, 150, the first function is an addition function. Therefore, the delta chroma QP can be added to the luminance QP to obtain the chroma QP value. Other possibilities for the first function include context encoding the delta chroma QP and using context information (e.g., adjacent blocks) to obtain the chroma QP. In some embodiments, the first function can be based on a look-up table or a mapping from the delta chroma QP value to the chroma QP value. Figures 4A - 4C An example embodiment of the techniques described in methods 100, 150 is shown.

[0166] Similar to the first function, in various embodiments, the second function can also be a summation or addition function, or can be based on the context or a representation of the values of the luminance QP and the delta chrominance QP based on a table.

[0167] As previously described, the regions used in methods 100, 150 can correspond to the size of a picture or video picture, or can be a smaller portion of the picture.

[0168] In some embodiments, in methods 100, 150, the default chrominance delta QP indicated by a flag is determined based on a mapping, and the mapping remains unchanged in the bitstream. Alternatively, at the picture, tile, or sequence level, the mapping can be changed by signaling in the bitstream.

[0169] Methods 100, 150 can be implemented (e.g., using predictive coding) to encode (or decode) a single picture or image or a sequence of pictures.

[0170] Regarding Figure 2 as described, the encoder can implement method 150 during video encoding. For example, Figure 3 the techniques described in can be used to divide a video into different regions (e.g., tile groups). The quantization unit 209 can determine the luminance and chrominance components and can include corresponding indications in the bitstream, as described regarding method 150. Some example syntax structures are shown in Figures 4A - 4C as follows.

[0171] Regarding Figure 5 as described, the decoder can implement method 100. For example, the parsing unit can parse the flag indicating the chrominance QP. Figures 4A - 4C Various examples of the bitstream syntax that can be parsed by the decoder are shown.

[0172] In some embodiments, a video encoder device can include a processor configured to implement the bitstream generation or encoding techniques described in this document.

[0173] In some embodiments, a video decoder device can include a processor configured to implement the bitstream decoding described in this document. The decoder device can further process the decoded visual information to generate a displayable image (or sequence of images).

[0174] The various device embodiments described herein are for illustrative purposes only and are not intended to limit the implementation of the encoding and decoding operations described in this document.

[0175] Industrial Applicability

[0176] As can be seen from the above description, once the QP for the luminance component is determined, the chrominance QP can be obtained using the default chrominance deltaQP, at least as a candidate chrominance QP. By the above steps, the computational burden on the encoder to determine the chrominance QP is alleviated. Thus, the disadvantages of the existing method are solved by using the aforementioned encoder to generate a bitstream and using the aforementioned decoder to decode the bitstream.

[0177] Figure 8 An example apparatus that can be used to implement the encoder-side or decoder-side techniques described in this document is shown. The apparatus includes a processor that can be configured to perform the encoder-side or decoder-side techniques or both. The apparatus may also include a memory (not shown) for storing processor-executable instructions and for storing video bitstreams and / or display data. The apparatus may include video processing circuitry (not shown), such as transform circuitry, arithmetic coding / decoding circuitry, data coding techniques based on look-up tables, etc. The video processing circuitry may be partially included in the processor and / or partially included in other dedicated circuits such as a graphics processor, a field programmable gate array (FPGA), etc.

[0178] It should be understood that this document provides various techniques that can be implemented by video or image encoder and decoder embodiments to differentially encode chrominance quantization parameters based on information from luminance quantization parameters. In one example aspect, the disclosed techniques can be used to trade off between the visual quality of encoded visual information and the number of bits used for chrominance component encoding. In one advantageous aspect, the decision to trade off bit precision between chrominance and luminance components can be made at sub-picture boundaries and signaled according to the chroma QP value.

[0179] The disclosed and other embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, such as including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver apparatus.

[0180] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0181] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0182] For example, processors suitable for executing computer programs include general and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer will also include or be operatively coupled to receive data from one or more mass storage devices for storing data (e.g., magnetic, magneto-optical or optical disks) or to transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and storage devices, including for example semiconductor storage devices such as EPROM, EEPROM and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special logic circuitry.

[0183] Although the patent document contains many details, these details should not be construed as limitations on any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above may describe features as acting in certain combinations and even initially claim them as such, in some cases one or more features of a claimed combination may be excised from the combination, and the claimed combination may be directed to a sub-combination or variations of a sub-combination.

[0184] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve the desired effect. Additionally, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0185] Only some embodiments and examples have been described, and other embodiments, enhancements and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing visual information, comprising: Parsing a bitstream to determine a quantization parameter QP for a luminance component of a region for the visual information from data units of a parameter set included in the bitstream; Determining a value of a flag from the data units of the parameter set, the value of the flag being a first value indicating the absence of an offset for the QP of the luminance component in the bitstream or a second value indicating the presence of an offset for the QP of the luminance component in the bitstream; In the case where the flag is equal to the first value, determining the QP for the chrominance component of the region as a first function of the QP for the luminance component and a default chrominance delta QP indicated by the flag; In the case where the flag is equal to the second value different from the first value, determining the QP for the chrominance component of the region by the following steps: a. Obtaining a delta chrominance QP from the data unit; And b. Determining the QP for the chrominance component as a second function of the QP for the luminance component and the delta chrominance QP; And In the case where the parameter set is activated for decoding the region, decoding the chrominance component using the QP for the chrominance component, and wherein, depending on whether the flag has the first value or the second value, the QP for the chrominance component of the region is determined based on the first function or the second function using the default chrominance delta QP or the delta chrominance QP indicated by the flag.

2. The method according to claim 1, wherein The first function is an addition function.

3. The method according to any one of claims 1-2, wherein The second function is an addition function.

4. The method according to claim 1, wherein The region corresponds to a picture.

5. The method according to claim 1, wherein The region is a partial region in the picture, and the partial region is one of the following in the picture: one or more tiles; one or more stripes.

6. The method according to claim 1, wherein The default chrominance delta QP indicated by the flag is determined based on a mapping, and wherein the mapping remains unchanged in the bitstream.

7. The method according to claim 1, wherein, The default chrominance delta QP indicated by the flag is determined based on a mapping, and wherein the mapping is changed by signaling in the bitstream.

8. The method according to claim 1, wherein The visual information corresponds to a single picture or a time series of pictures.

9. A method for generating an encoded representation of visual information, comprising: Determining a quantization parameter QP for a luminance component of a region for the visual information; Signaling the QP for the chrominance component in the encoded representation of the visual information by: In the case where the QP for the chrominance component is equal to a value of a first function of the default chrominance delta QP for the chrominance component and the QP for the luminance component, including a first value of a flag indicating that the default chrominance delta QP will be used for decoding in the encoded representation, and encoding the flag in a data unit of a parameter set of the encoded representation; And In the case where the QP for the chrominance component is not equal to the value of the first function of the default chrominance delta QP for the chrominance component and the QP for the luminance component, the second value of the flag and the value in the data unit indicating the delta chrominance QP are included in the coded representation such that the QP for the chrominance component is the second function of the QP of the luminance component and the delta chrominance QP; And wherein whether the flag has the first value or the second value depends on whether the QP for the chrominance component of the region is equal to or different from the value of the first function of the default chrominance delta QP for the chrominance component and the QP for the luminance component, and wherein, depending on whether the flag has the first value or the second value, the QP for the chrominance component of the region is determined by using the default chrominance delta QP indicated by the flag or the delta chrominance QP based on the first function or the second function.

10. The method according to claim 9, wherein The first function is an addition function.

11. The method according to claim 9, wherein The second function is an addition function.

12. The method according to claim 9, wherein, The region corresponds to a picture.

13. The method according to claim 9, wherein, The region is a partial region in the picture, and the partial region is one of the following in the picture: one or more tiles; one or more stripes.

14. The method according to claim 9, wherein, The default chrominance delta QP indicated by the flag is determined based on a mapping, where the mapping remains unchanged in the bitstream.

15. The method according to claim 9, wherein, The default chrominance delta QP indicated by the flag is determined based on a mapping, where the mapping is changed by signaling in the bitstream.

16. The method according to claim 9, wherein, The visual information corresponds to a single picture or a temporal sequence of pictures.

17. A video encoder device, comprising a processor and a memory, the processor being configured to read instructions from the memory to implement the method according to any one of claims 9-16.

18. A video decoder device, comprising a processor and a memory, the processor being configured to read instructions from the memory to implement the method according to any one of claims 1-8.

19. A computer program product, having code stored thereon, which when executed by a processor causes the processor to implement the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Chroma quantization in video coding

    CN107948651A