Single-index quantization matrix design for video encoding and decoding
The unified quantization matrix index simplifies quantization matrix signaling in video encoding and decoding by predicting and deriving QM coefficients based on block size and prediction mode, reducing bit cost and complexity in the VVC standard.
Patent Information
- Application Number
- JP2025043021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-24
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-06-16
AI Technical Summary
Existing video encoding and decoding technologies face inefficiencies in quantization matrix signaling and prediction processes, particularly in the Versatile Video Coding (VVC) standard, leading to increased complexity and bit usage.
A unified quantization matrix (QM) index is introduced, allowing prediction and derivation processes that simplify QM signaling by copying or decimating QM coefficients based on block size, color component, and prediction mode, and transmitting all QM coefficients for size-64 blocks, reducing bit cost and complexity.
The proposed method significantly reduces bit usage and simplifies the quantization matrix signaling process, achieving a substantial reduction in bit cost while maintaining effective video encoding and decoding performance.
Smart Images

Figure 2025094064000001_ABST
Abstract
Description
Technical Field
[0001] Technical Field [1] Generally, this embodiment relates to a method and apparatus for quantization matrix design in video encoding and decoding.
Background Art
[0002] Background [2] To achieve high compression efficiency, image and video encoding schemes typically utilize prediction and transformation to exploit spatial and temporal redundancies in video content. Generally, intra and inter prediction are used to exploit intra or inter picture correlation, in which case the difference between the original block, often called the prediction error or prediction residual, and the predicted block is transformed, quantized, and entropy encoded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy encoding, quantization, transformation, and prediction.
Summary of the Invention
[0003] Summary [3] According to one embodiment, a video decoding method is provided. The method includes obtaining a single identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a block to be decoded within a picture; decoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtaining the quantization matrix based on the reference quantization matrix; inverse quantizing transform coefficients of the block in response to the quantization matrix; and decoding the block in response to the inverse quantized transform coefficients.
[0004] [4] According to another embodiment, a video encoding method is provided. The method includes accessing an encoding block in a picture, accessing a quantization matrix of the block, obtaining a single identifier of the quantization matrix based on the block size, color component, and prediction mode of the block, encoding a syntax element indicating a reference quantization matrix, where the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix, quantizing transform coefficients of the block in response to the quantization matrix, and entropy encoding the quantized transform coefficients.
[0005] [5] According to another embodiment, a video decoding apparatus is provided. The apparatus includes one or more processors configured to obtain a single identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a decoding block in a picture, decode a syntax element indicating a reference quantization matrix, where the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix, obtain the quantization matrix based on the reference quantization matrix, inverse quantize transform coefficients of the block in response to the quantization matrix, and decode the block in response to the inverse quantized transform coefficients.
[0006] [6] According to another embodiment, a video encoding device is provided, the device comprising one or more processors, the one or more processors being configured to access an encoding block in a picture, access a quantization matrix of the block, obtain a single identifier of the quantization matrix based on the block size, color component, and prediction mode of the block, encode a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix, quantize transform coefficients of the block in response to the quantization matrix, and entropy encode the quantized transform coefficients.
[0007] [7] According to another embodiment, a video decoding device is provided, the device comprising means for obtaining a single identifier of a quantization matrix based on the block size, color component, and prediction mode of a decoding block in a picture, means for decoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix, means for obtaining the quantization matrix based on the reference quantization matrix, means for inverse quantizing transform coefficients of the block in response to the quantization matrix, and means for decoding the block in response to the inverse quantized transform coefficients.
[0008] [8] According to another embodiment, a video encoding device is provided. The device includes means for accessing a block to be encoded within a picture, means for accessing a quantization matrix of the block, means for obtaining a single identifier of the quantization matrix based on the block size, color component, and prediction mode of the block, means for encoding a syntax element indicating a reference quantization matrix, where the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix, means for quantizing transform coefficients of the block in response to the quantization matrix, and means for entropy encoding the quantized transform coefficients.
[0009] [9] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to execute an encoding method or a decoding method according to any of the above-described embodiments. One or more of the present embodiments also provide a computer-readable storage medium storing instructions for encoding or decoding video data by the above-described method. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated by the above-described method. One or more embodiments also provide a method and an apparatus for transmitting or receiving a bitstream generated by the above-described method.
Brief Description of the Drawings
[0010] Brief Description of the Drawings
Figure 1
[10] A block diagram of a system capable of implementing the aspects of the present embodiment is shown.
Figure 2
[11] A block diagram of an embodiment of a video encoder is shown.
Figure 3
[12] A block diagram of an embodiment of a video decoder is shown.
Figure 4
[13] For block sizes larger than 32 in VVC draft 5, it is shown that the transform coefficients are presumed to be zero.
Figure 5
[14] Shows a fixed prediction tree as described in JCTVC-H0314.
Figure 6
[15] Shows prediction (decimation) from a larger size according to one embodiment.
Figure 7
[16] Shows a combination of prediction from a larger size and decimation of rectangular blocks according to one embodiment.
Figure 8
[17] Shows the QM derivation process for rectangular blocks in chroma according to one embodiment.
Figure 9
[18] Shows the QM derivation process for rectangular blocks in chroma (adaptation to 4:2:2 format) according to one embodiment.
Figure 10
[19] Shows the QM derivation process for rectangular blocks in chroma (adaptation to 4:4:4 format).
Figure 11
[20] Shows a flowchart for parsing the scaling list data syntax structure according to one embodiment.
Figure 12
[21] Shows a flowchart for encoding the scaling list data syntax structure according to one embodiment.
Figure 13
[22] Shows a flowchart of the QM derivation process according to one embodiment. DETAILED DESCRIPTION
[0011] Detailed Description
[23] FIG. 1 shows a block diagram of an example of a system that can implement various aspects and embodiments. System 100 can be implemented as a device that includes various components described hereinafter and is configured to execute one or more of the aspects described in the present application. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and various electronic devices such as servers. The elements of system 100 can be implemented as a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in the present application.
[0012]
[24] System 100 includes at least one processor 110 configured to execute instructions loaded to implement various aspects described in the present application. Processor 110 can include an embedded memory, an input / output interface, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. Storage device 140 can include, as non-limiting examples, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0013]
[25] System 100 includes an encoder / decoder module 130 configured to process data to provide, for example, encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Further, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or, as is known to those skilled in the art, may be incorporated within the processor 110 as a combination of hardware and software.
[0014]
[26] The program code loaded into the processor 110 or the encoder / decoder 130 to execute the various aspects described herein may be stored in the storage device 140 and subsequently loaded into the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described herein. Such stored items may include, without limitation, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and equations, formulas, operations, and intermediate or final results of operation logic.
[0015]
[27] In some embodiments, the memory within the processor 110 and / or the encoder / decoder module 130 stores instructions and is used to provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (for example, the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140 and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used for storing the television operating system. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations such as MPEG-2, HEVC, or VVC.
[0016]
[28] Inputs to the elements of the system 100 can be provided through various input devices shown in block 105. Such input devices include, without limitation, (i) an RF section that receives RF signals transmitted wirelessly, for example, by a broadcaster, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.
[0017]
[29] In various embodiments, each input processing element known in the art is associated with the input device of block 105. For example, the RF section can be associated with elements suitable for (i) selecting a desired frequency (also called signal selection or bandlimiting of the signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band that can be called a channel in a particular embodiment, (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments can include one or more elements that perform these functions, such as a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various of these functions, including downconverting the received signal to a low frequency (e.g., an intermediate frequency or a frequency near the baseband) or to the baseband. In one set-top box embodiment, the RF section and the associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments perform a rearrangement of the order of the above-described (and other) elements, removal of some of these elements, and / or addition of other elements that perform similar or different functions. The addition of elements can include inserting elements between existing elements, for example, inserting an amplifier and an analog / digital converter. In various embodiments, the RF section includes an antenna.
[0018]
[30] Further, the USB and / or HDMI terminals may include respective interface processors that connect the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, may be performed either within a separate input processing IC or within the processor 110 as needed. Similarly, aspects of the USB or HDMI interface processing may be performed within a separate interface IC or within the processor 110 as needed. The demodulated, error-corrected, and de-multiplexed stream is provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130 that operate in combination with the memory and storage elements to process the data stream as needed for presentation to the output device.
[0019]
[31] The various elements of the system 100 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected using suitable connection means 115, such as internal buses known in the art including I2C buses, wiring, and printed circuit boards, and may transmit data therebetween.
[0020]
[32] The system 100 includes a communication interface 150 that enables it to communicate with other devices via a communication channel 190. The communication interface 150 may include, without limitation, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, without limitation, a modem or a network card, and the communication channel 190 may be implemented, for example, within wired and / or wireless media.
[0021]
[33] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received via a communication channel 190 and a communication interface 150 that are adapted for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to an external network including the Internet to enable streaming applications and other over-the-top communications. Other embodiments provide streaming data to system 100 using a set-top box that distributes data via HDMI communication of input block 105. Still other embodiments provide streaming data to system 100 using the RF connection of input block 105.
[0022]
[34] System 100 may provide output signals to various output devices, including display 165, speaker 175, and other peripheral devices 185. Other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are communicated between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling such as an AV link, CEC, or other communication protocol that enables control between devices with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through each interface 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and speaker 175 may be integrated into a single unit with other components of system 100 within an electronic device, such as a television. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0023]
[35] The display 165 and the speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF section of the input 105 is part of a separate set-top box. In various embodiments where the display 165 and the speaker 175 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0024]
[36] FIG. 2 shows an example video encoder 200 such as a high efficiency video coding (HEVC) encoder. FIG. 2 can also show an encoder that uses a technology similar to HEVC, such as an encoder with improvements to the HEVC standard or a Versatile Video Coding (VVC) encoder being developed by the Joint Video Exploration Team (JVET).
[0025]
[37] In the present application, the terms "reconstructed" and "decoded" may be used synonymously, the terms "encoded" or "coded" may be used synonymously, and the terms "image", "picture", and "frame" may be used synonymously. Although not essential, typically the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.
[0026]
[38] Before encoding, the video sequence may undergo pre-encoding processing (201), such as color conversion (e.g., conversion from RGB4:4:4 to YCbCr4:2:0) or remapping of the input picture components, to obtain a signal distribution with higher resilience to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and attached to the bitstream.
[0027]
[39] In the encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is divided (202) into units such as CUs, for example, and processed in those units. Each unit is encoded using, for example, either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction is performed (260). In the inter mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) whether to use the intra mode or the inter mode for encoding the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting a prediction block from the original image block (210).
[0028]
[40] Next, the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, that is, the residual is encoded directly without applying the transformation or quantization process.
[0029]
[41] The encoder decodes the encoded block to provide a basis for further prediction. The quantized transform coefficients are inverse quantized (240), inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (280).
[0030]
[42] FIG. 3 shows a block diagram of a video decoder 300 of an example. In the decoder 300, a bitstream is decoded by decoder elements described below. The video decoder 300 generally executes an encoding path and a reciprocal decoding path described in FIG. 2. The encoder 200 also executes video decoding as part of video data encoding.
[0031]
[43] In particular, the input to the decoder includes a video bitstream that can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coding information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (335). The transform coefficients are inverse quantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block can be obtained from intra prediction (360) or motion compensation prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (380).
[0032]
[44] The decoded picture can further undergo post-decoding processing (385), for example, color inverse transformation (e.g., transformation from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the reverse of the remapping process executed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0033]
[45] In the HEVC specification, a quantization matrix can be used in the inverse quantization process, and the transformed coefficients are scaled by the current quantization step and further scaled by the quantization matrix (QM) as follows: d[x][y] = Clip3(coeffMin, coeffMax, (((TransCoeffLevel[xTbY][yTbY][cIdx][x][y] * m[x][y] * levelScale[qP % 6]) << (qP / 6)) + (1 << (bdShift - 1))) >> bdShift) where · TransCoeffLevel[…] is the absolute value of the transform coefficient of the current block identified by the spatial coordinates xTbY, yTbY, and the component index cIdx. · x and y are the horizontal / vertical frequency indices. · qP is the current quantization parameter. · Multiplication by levelScale[qP % 6] and left shift by (qP / 6) are equivalent to multiplication by the quantization step qStep = (levelScale[qP % 6] << (qP / 6)). · m[…] […] is a two-dimensional quantization matrix. Here, since the quantization matrix is used for scaling, it is sometimes called a scaling matrix. · bdShift is an additional scaling factor that describes the image sample bit depth. The term (1 << (bdShift - 1)) serves to round to the nearest integer. · d[…] is the resulting inverse-quantized absolute value of the transform coefficient.
[0034]
[46] The syntax used by HEVC for transmitting the quantization matrix is described below:
[0035]
Table 1
[0036]
[47] The following can be noted. · Different matrices are specified for each transformation size (sizeId). In the scaling list data syntax structure, the scaling matrices are scanned into a one-dimensional scaling list (e.g., ScalingList). · For a given transformation size, six matrices for intra / inter coding and Y / Cb / Cr components are specified. · The matrix can be any of the following. · If the scaling_list_pred_mode_flag is zero (the reference matrixId is obtained as matrixId - scaling_list_pred_matrix_id_delta), it can be copied from the matrix sent earlier of the same size. · It can be copied from the default value specified in the standard (when both scaling_list_pred_mode_flag and scaling_list_pred_matrix_id_delta are zero). · It can be fully specified in DPCM coding mode using exponential Golomb entropy coding in the upper-right diagonal scan order. · For block sizes larger than 8×8, only the 8×8 coefficients are sent for signaling the quantization matrix to reduce the coded bits. Then, the coefficients are interpolated using zero-hold (i.e., repetition), except for the DC coefficient sent explicitly.
[0037]
[48] The use of quantization matrices similar to HEVC has been adopted in VVC draft 5 based on the submitted JVET-N0847 (O. Chubach, et al., “CE7-related: Support of quantization matrices for VVC,” JVET-N0847, Geneva, CH, March 2019). The scaling_list_data syntax conforms to the VVC codec shown below.
[0038]
Table 2
[0039]
[49] In the design of VVC draft 5 using JVET-N0847, like in HEVC, the QM is identified by two parameters, matrixId and sizeId. This is shown in the following two tables.
[0040]
Table 3
[0041]
Table 4
[0042]
[50] The combinations of both identifiers are shown in the following table.
[0043]
Table 5
[0044]
[51] Like in HEVC, for block sizes larger than 8×8, only the 8×8 coefficients and the DC coefficient are transmitted. The QM of the correct size is reconstructed using zero-hold interpolation. For example, for a 16×16 block, all coefficients are repeated twice in both directions, and the DC coefficient is replaced with the transmitted one.
[0045]
[52] For rectangular blocks, the size held in the QM selection (sizeId) is the larger dimension, i.e., the larger of the width and height. For example, for a 4×16 block, the QM of the 16×16 block size is selected. Then, the reconstructed 16×16 matrix is vertically decimated by a factor of 4 to obtain the final 4×16 quantization matrix (i.e., 3 out of 4 lines are skipped).
[0046]
[53] Hereinafter, in relation to the sizeId and the square block size used, the QM of a given family of block sizes (square or rectangular) is called size-N. For example, for block sizes of 16×16 or 16×4, QM is identified as size-16 (sizeId4 in VVC draft 5). The size-N notation is used to distinguish from the exact block shape and the number of QM coefficients signaled (limited to 8×8 as shown in Table 3).
[0047]
[54] Further, in VVC draft 5, for size-64, the QM coefficients in the lower right quadrant are not transmitted (assumed to be 0, hereinafter referred to as "zeroed out"). This is implemented by the "x>=4&&y>=4" condition in the scaling_list_data syntax. This avoids the transmission of QM coefficients that are never used in the transform / quantization process. In fact, in VVC, when converting block sizes exceeding 32 in any dimension (64×N, N×64, N≦64), any transform coefficient having x / y frequency coordinates of 32 or more is not transmitted and is assumed to be zero, and thus no quantization matrix coefficient is required for its quantization. This is shown in FIG. 4, where the hatched area corresponds to the transform coefficients assumed to be zero.
[0048]
[55] Compared with HEVC, in VVC, more quantization matrices are required due to the larger number of block sizes. However, in VVC draft 5, QM prediction is still limited to copying the same block size matrix, which can lead to waste of bits. Further, in VVC, by using only block size 2×2 for chroma and only block size 64×64 for luma, the syntax related to QM is more complex. Also, JVET-N0847 describes a specific matrix derivation process for each block size, similar to HEVC.
[0049]
[56] During the HEVC standardization, for example, in JCTVC-E073 (see J. Tanaka, et al., “Quantization Matrix for HEVC,” JCTVC-E073, Geneva, CH, March 2011) and JCTVC-H0314 (see Y. Wang, et al., “Layered quantization matrices representation and compression,” JCTVC-H0314, San Jose, CA, USA, February 2012), several QM prediction techniques have been explored.
[0050]
[57] JCTVC-E073: The QM is transmitted with a specific parameter set (QMPS). Within the QMPS, the QM is transmitted in ascending order of size (sizeId / matrixId as in HEVC). Prediction (= copy) from any previously encoded QM coefficients, including the previous QMPS, has been proposed. Up-conversion using linear interpolation is used for adaptation from a smaller reference QM, while simple downsampling is used for adaptation from a larger reference QM. During the HEVC standardization, this was ultimately rejected.
[0051]
[58] JCTVC-H0314: The QM is transmitted in ascending or descending order. As shown in FIG. 5, it is possible to copy the previously transmitted QM instead of transmitting a new QM using a fixed prediction tree (without an explicit reference index). When the reference QM is larger, simple downsampling is used. During the HEVC standardization, this was ultimately rejected.
[0052]
[59] These two proposals are related to HEVC and do not address the complexity introduced by VVC.
[0053]
[60] This application proposes to simplify the quantization matrix signaling and prediction process of VVC Draft 5 (after the adoption of JVET-N0847) while enhancing the quantization matrix signaling and prediction process so that any QM can be predicted from what has been signaled to any destination by incorporating one or more of the following. - Unify the QM index to include both size and type so that the reference index difference can handle what has been sent to any destination. - Transmit the quantization matrix in descending order of block size reduction. - Specify the prediction process as either a copy or a decimation process as needed. - Transmit all QM coefficients at size-64 so that size-64 QM can be used as a predictor.
[0054]
[61] Furthermore, the QM derivation process that includes upsampling of blocks larger than 8×8 and downsampling of rectangular blocks is described as selecting the QM index according to the block parameters and adapting the QM signaling size to the actual block size.
[0055]
[62] To facilitate notation, the process of predicting the quantization matrix from the default value or what has been sent previously is regarded as the QM prediction process, and the process of adapting the transmitted or predicted QM to the size and chroma format of the transform block is regarded as the QM derivation process. The QM prediction process can be, for example, part of the scaling list data parsing process at the picture level. The derivation process is usually at a lower level, for example, at the transform block level. Various aspects are presented in more detail below, followed by draft text examples and performance results. · Derivation and use of one matrix index to identify QM. With one identifier, when using prediction (copy), the matrix sent to any destination can be referenced, and by sending the larger matrix first, interpolation in the prediction process is avoided. ·A QM prediction process that includes copying or decimating the QM (reference QM) signaled to a previous one that is the transmitted, predicted, or default reference QM. ·A QM derivation process for a given transform block that includes selecting a QM index based on the block size, color component, and prediction mode, and then adapting the size of the selected QM to the size of the block. The resizing process is based on bit shifts of the x and y coordinates within the transform block that index the coefficients of the selected QM. ·Transmission of all coefficients of size-64 QM.
[0056]
[63] Compared with VVC draft 5, these aspects simplify the specification (halving the text changes compared with JVET-N0847) and result in a large bit limit (the bit cost of scaling_list_data can be halved).
[0057]
[64] Unified QM index
[65] The QM used for quantization / inverse quantization of the transform block is identified by one parameter matrixId. In one embodiment, the unified matrixId (QM index) is the following composite. - A size identifier related to the CU size (i.e., only square size matrices are transmitted, so the CU enclosing the square), rather than the block size. Here, for either luma or chroma, the size identifier is controlled by the luma block size, e.g., max(luma block width, luma block height). When the luma and chroma trees are separated, for chroma, the "CU size" refers to the size of the block projected onto the luma plane. - The matrix type that lists the luma QM first because the luma QM can be larger than the chroma (e.g., for the 4:2:0 chroma format).
[0058]
[66] According to this embodiment, the QM index derivation is shown in Table 4, Table 5, and Equation (1).
[0059]
Table 6
[0060]
Table 7
[0061]
[67] The unified matrixId is derived as follows: matrixId = N * sizeId + matrixTypeId (1) where N is the number of possible type identifiers, for example, N = 6.
[0062]
[68] In another embodiment, if seven or more QM types are defined, sizeId should be multiplied by the correct number which is the number of quantization matrix types. In other embodiments, other parameters, such as a specific block size, the signaled matrix size (here limited to 8×8), or the presence or absence of DC coefficients, can also be different. Here, QM is listed in the order of decreasing block size and is identified by one index as shown in Table 6.
[0063]
Table 8
[0064]
[69] QM prediction process
[70] Instead of transmitting QM coefficients, it is possible to predict QM from default values or from QM coefficients transmitted to any destination. In one embodiment, if the reference QM is of the same size, QM is copied; in other cases, as shown in an example in Figure 6, it is decimated by a related ratio. In Figure 6, a size-4 luminance QM is predicted from a size-8 one.
[0065]
[71] Decimation is described by the following equation: ScalingMatrix[matrixId][x][y] = refScalingMatrix[i][j] (2) where matrixSize = (matrixId < 20)? 8 : (matrixId < 26)? 4 : 2 x = 0.. matrixSize - 1, y = 0.. matrixSize - 1, i = x << (log2(refMatrixSize) - log2(matrixSize)), and j = y << (log2(refMatrixSize) - log2(matrixSize)). where refMatrixSize matches the size of refScalingMatrix (and thus the range of the i and j variables).
[0066]
[72] In the example shown in FIG. 6, the luminance size - 4QM (4×4 array: matrixSize is 4) is predicted from the luminance size - 8QM which is an 8×8 array (refMatrixSize is 8); one of the two lines and one of the two columns are dropped to produce a 4×4 array (i.e., the element (2x,2y) in the reference QM is copied to the element (x,y) in the current QM).
[0067]
[73] Equation (3) takes the following form: ScalingMatrix[matrixId][x][y] = refScalingMatrix[i][j] (3) x = 0.. 3, y = 0.. 3, i = x << 1, and j = y << 1.
[0068]
[74] When the reference QM has a DC value, if the current QM requires a DC value, the reference QM is copied as the DC value, and if the current QM does not require a DC value, the reference QM is copied to the top - left QM coefficient.
[0069]
[75] This QM prediction process is, in a preferred embodiment, part of the QM decoding process, but in another embodiment, it can be postponed until the QM derivation process, in which case the decimation for prediction purposes is merged with the QM resizing sub-process.
[0070]
[76] QM derivation process
[77] The proposed derivation process of the quantization matrix first selects the right QM index according to the block parameters as described above (the unified QM index), and then unifies the processes of decimation in rectangular blocks, iteration in blocks of a certain size, for example, larger than 8×8, and chroma format adaptation into one process. The proposed process is based on bit shifts of the x and y output coordinates. To select the right line / column of the selected QM, only a left shift followed by a right shift of the x / y output coordinates is required as shown in the following equation: m[x][y] = ScalingMatrix[matrixId][i][j] (4) where i = (x << log2MatrixSize) >> log2(blkWidth), and j = (y << log2MatrixSize) >> log2(blkHeight). where log2MatrixSize is the log2 of the size of ScalingMatrix[matrixId] (a square 2D array), blkWidth and blkHeight are the width and height of the current transform block respectively, x ranges from 0 to blkWidth - 1, and y ranges from 0 to blkHeight - 1.
[0071]
[78] Hereinafter, several examples are provided to illustrate the QM derivation process. In the example shown in FIG. 7, the QM of the luminance 16×8 block is derived from the luminance size - 16 QM which is actually an 8×8 array with the DC coefficient added. In this example, blkWidth is equal to 16, blkHeight is equal to 8, log2MatrixSize is equal to 3, and thus, Equation (5) takes the following form: m[x][y] = ScalingMatrix[matrixId][i][j] (5) where i = (x << 3) >> 4, and j = (y << 3) >> 3, where x = 0..15 and y = 0..7. Here, x is right-shifted by 1 only, and y remains unchanged (i.e., column i in the selected QM is now column 2 * i and 2 * in the current QM is copied to i+1). Further, since the selected QM has the DC coefficient, it is copied to m[0][0].
[0072]
[79] In another example shown in FIG. 8, a QM (4:2:0 format) for a chroma 4×2 block of an 8×4 CU is generated. This matches the 8×4 CU size where the surrounding square is 8×8. Thus, the selected QM is a size-8 QM, where the chroma QM is encoded as a 4×4 array. Here, blkWidth is equal to 4, blkHeight is equal to 2, log2MatrixSize is equal to 2, and thus, Equation (6) takes the following form: m[x][y] = ScalingMatrix[matrixId][i][j] (6) where i = (x << 2) >> 2, and j = (y << 2) >> 1 where x = 0..3 and y = 0..1. Here, x remains unchanged, and y is left-shifted by 1 only (i.e., row 2y in the reference QM is copied to row y in the current QM).
[0073]
[80] In the following example, the proposed adaptation to 4:2:2 and 4:4:4 formats is different from VVC draft 5. Instead of finding a QM that matches the chroma block size (except for the 64×64 case where there is no chroma matrix), the size match is based on the same (luma) CU size (i.e., the size of the block projected onto the luma plane), and if necessary, the coefficients are repeated. This decouples the QM design from the chroma format.
[0074]
[81] In the example shown in FIG. 9, a QM (4:2:2 format) of an 8×4 CU chroma 8×2 block is generated. The selected QM is the same as the above example shown in FIG. 8, but the 4:2:2 chroma format requires twice the number of columns. Here, the columns are repeated, so x is shifted right by only 1 and y is still shifted left by only 1. In particular, blkWidth is equal to 8, blkHeight is equal to 2, log2MatrixSize is equal to 2, so Equation (7) takes the following form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (7) where i = ( x << 2 ) >> 3, and j = ( y << 2 ) >> 1, where x = 0..7 and y = 0..1.
[0075]
[82] In the example shown in FIG. 10, a QM (4:4:4 format) of an 8×4 CU chroma 8×4 block is generated. The selected QM is still the same as the examples shown in FIGS. 8 and 9, but the 4:4:4 chroma format requires twice the number of rows and columns of the 4:2:0 chroma format. Here, the columns must be repeated, so x is shifted right by only 1, but the row decimation (for the rectangle) can be skipped, so y is not shifted. In particular, blkWidth is equal to 8, blkHeight is equal to 4, log2MatrixSize is equal to 2, so Equation (8) takes the following form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (8) where i = ( x << 2 ) >> 3, and j = ( y << 2 ) >> 2, where x = 0..7 and y = 0..3.
[0076]
[83] Number of coefficients transmitted for size - 64
[84] In one embodiment, all coefficients of size-64 are sent in scaling list syntax so that even if the lower right quadrant is never used by the VVC transform and quantization processes, a QM smaller than size-64 can be predicted. Generally, in the present invention, all coefficients of the maximum QM can be sent.
[0077]
[85] However, when not used as a predictor, since the syntax element scaling_list_delta_coef can be set to zero in the lower right quadrant of size-64 QM, the overhead associated with this increase in the number of coefficients sent compared to the previous work (JVET-N0847) is, in the worst case, 2×16 bits: in the lower right quadrant, when 4×4 = 16 delta coefficients are signaled and forced to zero (encoded using exponential Golomb), each taking 1 bit, and there are two size-64 QMs (luma intra / inter), it is worth noting that it can be limited to.
[0078]
[86] The tests are described in Table 8, which shows that this overhead is negligible compared to the gain brought about by the prediction improvement.
[0079]
[87] In another embodiment, the coefficients in the lower right quadrant of size-64 QM are not sent as part of the size-64 QM signaling, but are sent as supplementary parameters when a smaller QM is first predicted from a given size-64 QM.
[0080]
[88] Table 7 provides some comparison between the method described in JVET-N0847 and the proposed method.
[0081]
Table 9
[0082]
[89] Some syntax and semantics according to one embodiment will be described below.
[0083]
[90] PPS Syntax and Semantics (Minor Conformance)
[0084]
Table 10
[0085] The pps_scaling_list_data_present_flag equal to 1 specifies that the scaling list data used for the picture called PPS is derived based on the scaling list specified by the active SPS and the scaling list specified by the PPS. The pps_scaling_list_data_present_flag equal to 0 specifies that the scaling list data used for the picture called PPS is presumed to be equal to that specified by the active SPS. When the scaling_list_enabled_flag is equal to zero, the value of the pps_scaling_list_data_present_flag shall be a value equal to 0. When the scaling_list_enabled_flag is equal to 1, the sps_scaling_list_data_present_flag is equal to 0, the pps_scaling_list_data_present_flag is equal to 0, and the default scaling matrix is used to derive the array ScalingMatrix as described in the scaling list data semantics as specified in Section 7.4.5.
[0086]
[91] Note that this syntax / semantics is an example intended to be close to the HEVC standard or VVC draft and is not limiting. For example, the conveyance of scaling_list_data is not limited to SPS or PPS and can also be transmitted by other means.
[0087]
[92] Signaling List Data Syntax / Semantics (Simplified)
[0088]
Table 11
[0089] A scaling_list_pred_mode_flag[matrixId] equal to 0 specifies that the scaling matrix is derived from the values of the reference scaling matrix. The reference scaling matrix is specified by scaling_list_pred_matrix_id_delta[matrixId]. A scaling_list_pred_mode_flag[matrixId] equal to 1 specifies that the values of the scaling list are explicitly signaled. scaling_list_pred_matrix_id_delta[matrixId] specifies the reference scaling matrix used for the derivation of the scaling matrix as follows. The value of scaling_list_pred_matrix_id_delta[matrixId] shall be within the range from 0 to matrixId (including the fractional part). When scaling_list_pred_mode_flag[matrixId] is equal to zero: - The variables refMatrixSize and the array refScalingMatrix are first derived as follows: · When scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, the following are applied as default values: · refMatrixSize is set equal to 8, · When matrixId is even, refScalingMatrix = (9) { { 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for intra default value { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, · In other cases refScalingMatrix = (10) { { 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the inter-default value { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, · In other cases (when scaling_list_pred_matrix_id_delta[matrixId] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (11) refMatrixSize = (refMatrixId < 20)? 8 : (refMatrixId < 26)? 4 : 2 ) (12) refScalingMatrix = ScalingMatrix[ refMatrixId ] (13) Next, the array ScalingMatrix[matrixId] is derived as follows: ScalingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (14) where matrixSize = (matrixId < 20)? 8 : (matrixId < 26)? 4 : 2 ) x = 0.. matrixSize - 1, y = 0.. matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ), and j = y << ( log2(refMatrixSize) - log2( matrixSize ) ) scaling_list_dc_coef_minus8[matrixId]+8, if relevant, specifies the first value of the scaling matrix as described in section xxx. Assume that the value of scaling_list_dc_coef_minus8[matrixId] is in the range from -7 to 247 (including fractions). If scaling_list_pred_mode_flag[matrixId] is equal to zero, scaling_list_pred_matrix_id_delta[matrixId] is greater than zero and refMatrixId < 14, and the following applies: - If matrixId < 14, it is assumed that scaling_list_dc_coef_minus8[matrixId] is equal to scaling_list_dc_coef_minus8[refMatrixId], - Otherwise, ScalingMatrix[matrixId][0][0] is set equal to scaling_list_dc_coef_minus8[refMatrixId]+8. When scaling_list_pred_mode_flag[matrixId] is equal to zero, scaling_list_pred_matrix_id_delta[matrixId] is equal to zero (indicating the default value), and matrixId < 14, it is assumed that scaling_list_dc_coef_minus8[matrixId] is equal to 8. scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[matrixId][i] and the previous matrix coefficient ScalingList[matrixId][i - 1] when scaling_list_pred_mode_flag[matrixId] is equal to 1. The value of scaling_list_delta_coef shall be in the range from -128 to 127 (including fractions). The value of ScalingList[matrixId][i] shall be greater than 0. If it exists (i.e., scaling_list_pred_mode_flag[matrixId] is equal to 1), the array ScalingMatrix[matrixId] is derived as follows: ScalingMatrix[ matrixId ][ i ][ j ] = ScalingList[ matrixId ][ k ] (15) where k = 0.. coefNum - 1, i = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0000] , and j = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0001]
[0090]
[93] The main simplifications compared to the JVET-N0847 syntax are the removal of one for() loop and the index simplification from [sizeId][matrixId] to [matrixId].
[0091]
[94] The "xxx section" refers to the uncertain section number introduced in the VVC specification that matches the scaling matrix derivation process of this document.
[0092]
[95] This syntax / semantics is an example, not intended to be limiting, of being close to the HEVC standard or VVC draft 5. For example, the coefficient range is not limited to 1, ···, 255, and may be, for example, 1, ···, 127 (7 bits) or -64, ···, 63. Also, it is not limited to 30 QMs organized as 6 types × 5 sizes (there can be 8 types and fewer or more sizes. See Table 5 for compatibility. In that case, simple compatibility to coefNum and conditions for the existence of the DC coefficient are required). The type of QM prediction (here only copy) is also not limited. For example, a scaling factor or offset can be added, and explicit coding can be added on top of the prediction as a residual. The same can be said for the method used for coefficient transmission (here DPCM), the existence of the DC coefficient, and the number of coefficients to be fixed (only a subset can be transmitted).
[0093]
[96] Regarding the default values, it is not limited to the two default QMs associated with MODE_INTRA and MODE_INTER, and as here, it can be filled with all 16 values until the relevant default values match (for example, the same default QM as HEVC can be selected).
[0094]
[97] Also, the number of coefficients to be signaled, coefNum, can also be mathematically expressed instead of a series of comparisons, and the result is the same: coefNum = Min(64, 4096 >> ((matrixId + 4) / 6) * 2), which introduces a division that is close to the HEVC or current VVC draft style but may not be welcome.
[0095]
[98] Note that here, a larger matrix is sent first and one index is used, whereby the prediction reference (indicated by scaling_list_pred_matrix_id_delta) can be any previously sent matrix or a default value (e.g., when scaling_list_pred_matrix_id_delta is zero), regardless of the intended block size or type.
[0096]
[99] FIG. 11 shows a process (1100) for parsing a scaling list data syntax structure according to an embodiment. In this embodiment, the input is an encoded bitstream and the output is an array of ScalingMatrix. For clarity, details about the DC value are omitted. In particular, in step 1110, the QM prediction mode is decoded from the bitstream. If QM is predicted (1120), the decoder further determines whether QM is inferred (predicted) or signaled in the bitstream according to the above flag. In step 1130, the decoder decodes the QM prediction data from the bitstream, which is necessary for inferring QM, for example, when the QM index difference scaling_list_pred_matrix_id_delta is not signaled. Next, the decoder determines whether QM is predicted from the default value (for example, when scaling_list_pred_matrix_id_delta is zero) or from the previously decoded QM (1140). If the reference QM is the default QM, the decoder selects the default QM as the reference QM (1150). For example, depending on the parity of matrixId, there may be several default QMs as the selection source. In other cases, the decoder selects the previously decoded QM as the reference QM (1155). The index of the reference QM is derived from matrixId and the above index difference. In step 1160, the decoder predicts QM from the reference QM. The prediction consists of a simple copy if the reference QM is the same size as the current QM, or decimation if it is larger than expected. The result is stored in ScalingMatrix[matrixId].
[0097]
[0100] If QM is not predicted (1120), the decoder determines (1170) the number of QM coefficients to be decoded from the bitstream according to matrixId. For example, if matrixId is lower than 20, it is 64; if matrixId is between 20 and 25, it is 16; and in other cases, it is 4. In step 1175, the decoder decodes the relevant number of QM coefficients from the bitstream. In step 1180, the decoder arranges the decoded QM coefficients into a 2D matrix according to the scan order, for example, diagonal scan. The result is stored in ScalingMatrix[matrixId]. Using ScalingMatrix[matrixId], the decoder can obtain the quantization matrix m[][] for inverse quantization of the transform block that can be non-square and / or of different chroma formats using the QM derivation process.
[0098]
[0101] In step 1190, the decoder checks whether the current QM is the last QM being parsed. If it is not the last one, the control returns to step 1110; otherwise, when all QMs have been parsed from the bitstream, the QM parsing process stops.
[0099]
[0102] Figure 12 shows a process (1200) for encoding a scaling list data syntax structure on the encoder side according to one embodiment. On the encoder side, QMs are scanned from the larger block size to the smaller block size (e.g., from matrixId0 to 30) in the order of description. In step 1210, the encoder searches for a prediction preference to determine whether the current QM is a copy (or decimation) of what was previously encoded. The QMs can be designed to optimize the efficiency of QM prediction. For example, if some QMs are initially close enough or close to the default QM, they can be forced to be equal (or decimated if they have different sizes). Further, the coefficients in the lower right quadrant of the size-64 QM can be optimized to better predict subsequent QMs or to reduce the QM bit cost if they are never reused for prediction. Once determined, the QM prediction mode is encoded into the bitstream.
[0100]
[0103] In particular, if the encoder determines to use prediction (1220), in step 1230, the prediction mode is encoded (e.g., scaling_list_pred_mode_flag = 0). In step 1240, the prediction parameters (e.g., QM index difference scaling_list_pred_matrix_id_delta) are encoded: zero index difference for the default QM value, or the relevant index difference if the previous QM was selected as the prediction reference. On the other hand, if explicit signaling is determined, in step 1250, the prediction mode is encoded (e.g., scaling_list_pred_mode_flag = 0). Then a diagonal scan (1260) is performed, and then QM coefficient encoding (1270) is performed.
[0101]
[0104] In step 1280, the encoder checks whether the current QM is the last QM to be encoded. If not, the control returns to step 1210. Otherwise, when all QMs have been encoded into the bitstream, the QM encoding process stops.
[0102]
[0105] Figure 13 shows a QM derivation process 1300 according to an embodiment. The input includes a ScalingMatrix array and can convert block parameters such as size (width / height), prediction mode (intra / inter / IBC, ···), and color component (Y / U / V). The output is a QM having the same size as the conversion block. For clarity, details about the DC value are omitted. In particular, in step 1310, the decoder determines the QM index matrixId according to the current conversion block size (width / height), prediction mode (intra / inter / IBC, ···), and color component (Y / U / V) as described above (uniform QM index). In step 1320, the decoder resizes the QM (ScalingMatrix[matrixId]) selected to match the conversion block size as described above. In one variation, step 1320 can include the decimation required for prediction.
[0103]
[0106] The QM derivation process is the same on the encoder side. Quantization divides the transform coefficient by the QM value, while inverse quantization multiplies. However, the QM is the same. In particular, the QM required for reconstruction in the encoder matches what is signaled in the bitstream.
[0104]
[0107] Conceptually, the transform coefficient d[x][y] can be quantized as follows, where qStep is the quantization step size and m[][] is the quantization matrix: TransCoeffLevel[xTbY][yTbY][cIdx][x][y] = d[x][y] / qStep / m[x][y]
[0105]
[0108] However, in the case of integer calculations, to avoid division, usually, TransCoeffLevel[xTbY][yTbY][cIdx][x][y] = ( ( d[x][y] * im[x][y] * ilevelScale[ qP%6 ] >> ((qP / 6 ) ) + ( 1 << ( bdShift - 1 ) ) ) >> bdShift ) looks like, for example, im[x][y]≒65536 / m[x][y], ilevelScale[0..5]=65536 / levelScale[0..5], and bdShift is an appropriate value. In fact, in the case of software coder, im * ilevelScale is usually pre-calculated and stored in a table.
[0106]
[0109] In the above, the QM prediction process and the QM derivation process are executed separately. In another embodiment, the QM prediction can be postponed until the QM derivation process. This embodiment does not change the QM signaling syntax. This embodiment can be functionally different by continuous resizing, and postpones the prediction part (reference QM acquisition + copy / downscale) and the diagonal scan until the "QM derivation process" described later.
[0107]
[0110] In that case, in one embodiment, the output of the scaling list data parsing process 1100 is, together with the prediction flag and the valid prediction parameters, an array of ScalingList() instead of ScalingMatrix: The ScalingMatrixPredId array always contains the index of the defined ScalingList (default or signaling). This array is recursively constructed during QM decoding by interpreting scaling_list_pred_matrix_id_delta, whereby the QM derivation process can directly use this index to obtain the actual values for constructing the QM used for inverse quantization of the current transform block.
[0108]
[0111] Hereinafter, an example is provided to show the scaling list semantics according to one embodiment.
[0109] Scaling matrix derivation process (new: partially replace the description with scaling list semantics; this is section xxx)
[0112] The input to this process is the prediction mode predMode, the color component variable cIdx, the block width blkWidth, and the block height blkHeight. The output of this process is a (blkWidth)×(blkHeight) array m[x][y] (scaling matrix), where x and y are the horizontal and vertical coefficient positions. Note that SubWidthC and SubHeightC depend on the chroma format and indicate the ratio of the number of samples in the luminance and chroma components. The variable matrixId is derived as follows: matrixId = 6 * sizeId + matrixTypeId (xxx - 1) where subWidth = (cIdx > 0)? SubWidthC : 1, subHeight = (cIdx > 0)? SubHeightC : 1, sizeId = 6 - max(log2(blkWidth * subWidth), log2(blkHeight * subHeight)), and matrixTypeId = (2 * cIdx + (predMode == MODE_INTER? 1 : 0)) The variable log2MatrixSize is derived as follows: log2MatrixSize = (matrixId < 20)? 3 : (matrixId < 26)? 2 : 1 (xxx - 2) The output array m[x][y] is derived by applying the following, where x ranges from 0 to blkWidth - 1 including fractions, and y ranges from 0 to blkHeight - 1 including fractions: m[x][y] = ScalingMatrix[matrixId][i][j] (xxx - 3) where i = (x << log2MatrixSize) >> log2(blkWidth), and j = ( y << log2MatrixSize ) >> log2( blkHeight ) If matrixId is lower than 14, m[0][0] is further modified as follows: m
[0000]
[0000] = scaling_list_dc_coef_minus8[ matrixId ] + 8 (xxx-4)
[0110]
[0113] Similar to the scaling_list_data syntax and semantics, this is an example and it should be noted that it is not limiting. For example, it is not limited to the two default QMs associated with MODE_INTRA and MODE_INTER. Here, as long as the relevant default values match (for example, the same default QM as HEVC can be selected), it is filled with all 16 values. There may be one default QM, or there may be three or more default QMs. MatrixId calculation can be difficult, for example, when there are more or fewer types than 6 for each block size. What is important is that the horizontal and vertical downscaling and upscaling to fit different block sizes from the selected QM are preferably performed in one simple (here, a right shift following a left shift) process.
[0111]
[0114] For rectangular blocks, it is not limited to the selection of the QM identifier of the current block surrounding the square: the derivation of sizeId in formula xxx-1 can follow different rules.
[0112]
[0115] Also, the selected QM size log2MatrixSize can be mathematically expressed instead of a series of comparisons, and the result is the same: log2MatrixSize = min(3, 6 - (matrixId + 4) / 6), which introduces a division that may not be welcome.
[0113]
[0116] Hereinafter, an example is provided to explain the semantics of the scaling process according to an embodiment.
[0114] (Adapted) Scaling Process of Conversion Coefficient
[0117] […] For the derivation of the scaled conversion coefficient d[x][y] where x = 0, ..., nTbW-1 and y = 0, ..., nTbH-1, the following is applied: - An intermediate scaling factor array m of (nTbW) × (nTbH) is derived as follows: - If one or more of the following conditions are true, m[x][y] is set equal to 16: - scaling_list_enabled_flag is equal to 0. - transform_skip_flag[xTbY][yTbY] is equal to 1. - Otherwise, m is the output of the scaling matrix derivation process specified in xxx section, called with the prediction mode CuPredMode[xTbY][yTbY], color component variable cIdx, block width nTbW, and block height nTbH as inputs. - The scaling factor ls[x][y] is derived as follows: […]
[0115]
[0118] The main change compared to VVC draft 5 is to call the xxx section instead of copying the part of the array described in the scaling_list_data semantics.
[0116]
[0119] As described above, it should be noted that this is an example aimed at minimizing changes to the current VVC draft and is not limiting. For example, the color component input of the scaling matrix derivation process may be different from cIdx. Also, QM is not limited to use as a scaling factor and can be used, for example, as a QP offset.
[0117]
[0120] The same test set used in HEVC normalization was augmented and used with QMs derived from the recommendations or default QMs of common standards (JPEG, MPEG2, AVC, HEVC) and QMs seen in actual broadcasts to test QM coding performance. In all tests, some QMs are copied from one type to another (e.g., from luma to chroma or from intra to inter) and / or from one size to another size.
[0118]
[0121] The following table reports the number of bits required to encode scaling_list_data using three different methods: HEVC, JVET-N0847, and the present proposal. In particular, HEVC uses the HEVC test set (24 QMs per test), and the other two use the derived test set (30 QMs per test, with additional sizes: size - 2 for chroma and size - 64 for luma; the size - 2 QMs are downsampled from size - 4, the size - 64 QMs are copied from size - 32, the size - 32 QMs are copied from size - 16, and sizes 16, 8, 4 are maintained as they are).
[0119]
Table 12
[0120]
[0122] In this test, it can be seen that the proposed technique reduces a large amount of bits even when compared with HEVC, while the proposed method encodes more QMs.
[0121]
[0123] Referring again to the technique in JCTVC-E073, the reference indexing (triple: QMPS, size, type) in JCTVC-E073 is more complex than that proposed here because the previous QMPS indexing requires the storage of the previous QMPS. Linear interpolation introduces complexity. Downsampling is the same as that proposed here.
[0122]
[0124] Referring again to the approach in JCTVC-H0314, the transmission from large to small is similar to what is proposed here, but the fixed prediction tree in JCTVC-H0314 is less flexible than the unified indexing and explicit criteria proposed here.
[0123]
[0125] QM in Intra Block Copy Mode
[0126] As described above, different QMs are specified for two block prediction modes, namely Intra and Inter. However, in addition to Intra and Inter, VVC has a new prediction mode: IBC (Intra Block Copy), where a block can be predicted from the reconstructed samples of the same picture using an appropriate displacement vector. In the QM selection for the IBC prediction mode, both JVET-N0847 and the above embodiments use the same QM as the Intra mode.
[0124]
[0127] Since the IBC mode is closer to Inter than Intra, in one embodiment, it is proposed to reuse the QM signaled in the Inter mode (instead of Intra). However, while IBC is close to Inter prediction, it is different: the displacement vectors do not match the movement of the object or camera and are used for texture copy. This can lead to specific artifacts in different embodiments where a particular QM can be useful for copy optimization of IBC blocks. Hereinafter, it is proposed to change the QM selection in the IBC prediction mode. · A preferred embodiment is to select the same QM as the Inter mode (instead of Intra), because IBC is closer to Inter prediction than Intra prediction. · Another option is to have a specific QM for the IBC mode. · These may be signaled explicitly in the syntax or may be inferred (e.g., the average of the Intra QM and the Inter QM).
[0125]
[0128] In a preferred embodiment, the QM selection or derivation process for a particular transform block selects inter-QM if the block has the IBC prediction mode. Referring again to FIG. 12, step 1210 of the QM derivation process needs to be adjusted as described below.
[0126]
[0129] In the previously proposed draft text, the QM selection is described by Equation (xxx-1), which can be changed as follows.
Number
[0127]
[0130] Note that the block can be encoded in intra mode (MODE_INTRA), inter mode (MODE_inter), or intra block copy mode (MODE_IBC). When matrixTypeId is set as in Table 6 or matrixTypeId = (2 * cIdx+(predMode == MODE_INTER? 1:0))), the MODE_IBC block selects matrixTypeId as if it were a MODE_INTRA block. Change in (xxx-1): matrixTypeId = (2 * Using cIdx+(predMode == MODE_INTRA? 0:1), the MODE_IBC block selects matrixTypeId as if it were a MODE_INTER block.
[0128]
[0131] In the draft text proposed by JVET-N0847, the QM selection is described in Table 7-14 and can be changed as follows. In particular, the matrixId of MODE_IBC is not the same as MODE_INTRA as in JVET-N0847, but is assigned in the same way as MODE_INTER.
[0129]
Table 13
[0130]
[0132] Variant 1: Explicit Signaling of QM for IBC
[0133] In this variant, specific QMs (different from intra QM and inter QM) are used for IBC blocks, and these QMs are explicitly signaled in the bitstream. This creates more QMs and requires compliance with the scaling_list_data syntax and matrixId mapping, with an impact on the bit cost. According to this variant, the QM selection described in Equation (xxx - 1) can be changed as follows.
Number
[0131]
[0134] In JVET - N0847, the QM selection table can be changed as follows:
[0132]
Table 14
[0133]
[0135] Variant 2: Inference of QM for IBC Mode
[0136] In this variant, specific QMs (different from intra QM and inter QM) are used for IBC blocks. However, those QMs are not signaled in the bitstream and are inferred: for example, as the average of intra QM and inter QM, a specific default value, or a specific modification to the inter QM, like scaling and offset.
[0134]
[0137] Variant 3: Explicit IBC QM for Luminance Only
[0138] In this variant, the additional QM for IBC is limited to luminance only, and the chroma QM for IBC can either reuse the inter QM as in Variant 1 or infer new QMs as in Variant 2.
[0135]
[0139] Various methods are described herein, and each method includes one or more steps or operations to achieve the methods described above. Except where a particular order of steps or operations is required for proper operation of the method, the order and / or use of particular steps and / or operations may be changed or combined. Further, terms such as "first", "second", etc., e.g., "first decoding" and "second decoding", may be used in various embodiments to modify elements, components, steps, operations, etc. The use of such terms does not imply an order to the modified operations, except where specifically required. Thus, in this example, the first decoding need not be performed before the second decoding and may be performed, for example, before, during, or overlapping with a time period of the second decoding.
[0136]
[0140] The various methods and other aspects described in this application can be used to modify modules of the video encoder 200 and decoder 300 shown in FIGS. 2 and 3, such as quantization and inverse quantization modules (230, 240, 340). Further, each aspect is not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations and extensions of any such standards and recommendations. Except where otherwise noted or technically excluded, the aspects described in this application can be used individually or in combination.
[0137]
[0141] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0138]
[0142] Various implementations include decoding. "Decoding", as used in this application, can include all or part of a process that is performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process typically includes one or more of the processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or, more generally, to a broader decoding process will be apparent based on the context of the specific description and is considered to be well understood by those skilled in the art.
[0139]
[0143] Various implementations include encoding. Similar to the above discussion regarding "decoding", "encoding" as used in this application can include all or part of a process that is performed, for example, on an input video sequence to produce an encoded bitstream.
[0140]
[0144] Note that the syntax elements used in this specification are descriptive terms. Thus, the use of other syntax element names is not excluded. Above, the syntax elements of PPS and scaling lists have been mainly used to describe various embodiments. It should be noted that these syntax elements may be arranged in other syntax structures.
[0141]
[0145] The implementations and aspects described in this specification can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if it is only discussed in the context of one form of implementation (for example, only discussed as a method), the implementation of the features discussed can also be implemented in other forms (for example, an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. The method can be implemented, for example, with a processor generally referred to as a processing device, including an apparatus such as a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, for example, a computer, a mobile phone, a portable / personal information terminal ("PDA"), and other devices that facilitate the communication of information between end users.
[0142]
[0146] References to "one embodiment", "an embodiment", "one implementation", or "an implementation", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the appearance of the phrases "in one embodiment", "in an embodiment", "in one implementation", or "in an implementation", and any other variations that appear in various places throughout this application do not necessarily all refer to the same embodiment.
[0143]
[0147] Furthermore, this application may refer to "determining" various information. The determination of information may include, for example, one or more of the estimation of information, the calculation of information, the prediction of information, or the retrieval of information from memory.
[0144]
[0148] Furthermore, this application may refer to "accessing" various information. Access to information may include, for example, one or more of the reception of information, the retrieval of information (for example, from memory), the storage of information, the transfer of information, the copying of information, the calculation of information, the determination of information, the prediction of information, or the estimation of information.
[0145]
[0149] Furthermore, the present application may refer to the "receiving" of various information. Receiving, like "access", is intended to be a broad term. Receiving of information may include, for example, one or more of accessing the information or searching for the information (e.g., from a memory). Further, "receiving" typically involves, in some form, storing the information during operation, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0146]
[0150] For example, it should be understood that the use of any of the following " / ", "and / or", and "at least one of" in the cases of "A / B", "A and / or B", and "at least one of A and B" is intended to encompass the selection of only the first-listed option (A), the selection of only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to encompass the selection of only the first-listed option (A), the selection of only the second-listed option (B), the selection of only the third-listed option (C), the selection of only the first and second-listed options (A and B), the selection of only the first and third-listed options (A and C), the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This can be extended, as will be apparent to those skilled in the art, to any number of items listed.
[0147]
[0151] Also, as used herein, the term "signaling" specifically refers to something being directed to a corresponding decoder. For example, in certain embodiments, the encoder signals a quantization matrix for inverse quantization. Thus, in embodiments, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has certain parameters etc., signaling is used without transmission (implicit signaling) to enable the decoder to simply know and select the specific parameters. By avoiding the transmission of any actual functionality, bit reduction is achieved in various embodiments. It should be understood that signaling can be accomplished in a variety of ways. In various embodiments, for example, one or more syntax elements, flags, etc. are used for signaling information to the corresponding decoder. The above relates to the verb form of the word "signal", although the word "signal" can also be used as a noun herein.
[0148]
[0152] As will be apparent to those skilled in the art, an implementation can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions to execute a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, a signal can be transmitted via a variety of different wired or wireless links. A signal can be stored on a processor-readable medium.
Claims
1. Obtaining a unique identifier of a quantization block based on a block size, a color component, and a prediction mode of a block to be decoded in a picture; decoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtaining the quantization matrix based on the reference quantization matrix; dequantizing transform coefficients of the block in response to the quantization matrix; decoding the block in response to the dequantized transform coefficients; The method includes:
2. 1. An apparatus comprising one or more processors, the one or more processors comprising: Obtaining a unique identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a block to be decoded in a picture; decoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtaining the quantization matrix based on the reference quantization matrix; dequantizing transform coefficients of the block in response to the quantization matrix; decoding the block in response to the dequantized transform coefficients; An apparatus configured to:
3. The method according to claim 1 or the apparatus according to claim 2, wherein the size of the block is different from the size of the block to which the reference quantization matrix is applied for inverse quantization.
4. The method of claim 1 or 3 or the apparatus of claim 2 or 3, wherein the elements of the quantization matrix are used as scaling factors when dequantizing each transform coefficient of the block.
5. The method according to any one of claims 1 to 4 or the apparatus according to any one of claims 2 to 4, wherein the elements of the quantization matrix are used as offsets when dequantizing each transform coefficient of the block.
6. Accessing a block to code within a picture; accessing a quantization matrix for the block; obtaining a unique identifier for the quantization matrix based on a block size, a color component, and a prediction mode of the block; encoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantizing transform coefficients of the block in response to the quantization block; entropy coding the quantized transform coefficients; The method includes:
7. 1. An apparatus comprising one or more processors, the one or more processors comprising: accessing a block to code in a picture; accessing a quantization matrix for the block; obtaining a unique identifier for the quantization matrix based on a block size, a color component, and a prediction mode of the block; encoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantizing transform coefficients of the block in response to the quantization matrix; entropy coding the quantized transform coefficients; An apparatus configured to:
8. The method according to claim 6 or the apparatus according to claim 7, wherein the size of the block is different from the size of the block to which the reference quantization matrix is applied for quantization.
9. 9. The method of claim 6 or 8 or the apparatus of claim 7 or 8, wherein the elements of the quantization matrix are used as scaling factors when quantizing each transform coefficient of the block.
10. 9. The method of claim 6 or 8 or the apparatus of claim 7 or 8, wherein elements of the quantization matrix are used as offsets when quantizing each transform coefficient of the block.
11. The method according to any one of claims 1 and 3 to 6 or the apparatus according to any one of claims 2 to 5 and 7 to 10, wherein the reference quantization matrix is signaled in advance.
12. The method according to any one of claims 1, 3 to 6 and 8 to 11 or the device according to any one of claims 2 to 5 and 7 to 11, wherein the quantization matrix is obtained from the reference quantization matrix through copying or decimation.
13. The method of claim 12 or the apparatus of claim 12, wherein the quantization matrix is obtained from the reference quantization matrix through copying, in response to the quantization matrix having the same size as the reference quantization matrix.
14. 13. The method of claim 12 or the apparatus of claim 12, wherein the quantization matrix is obtained from the reference quantization matrix through decimation by a corresponding ratio in response to the quantization matrix having a different size from the reference quantization matrix.
15. The method of any one of claims 1, 3 to 6 and 8 to 14 or the apparatus of any one of claims 2 to 5 and 7 to 14, wherein the block size is MxN, where M is the width and N is the height, and the identifier of the block is based on a size max(M,N), where max(M,N) is defined as the greater of M and N.
16. A method according to any one of claims 1, 3 to 6 and 8 to 15 or an apparatus according to any one of claims 2 to 5 and 7 to 15, wherein the set of quantization matrices is signaled in order of increasing identifier, with the quantization matrix of the largest block size being signaled first.
17. 17. The method of claim 16 or the apparatus of claim 16, wherein when signaling the set of quantization matrices, quantization matrices for luma color components are signaled before quantization matrices for chroma color components.
18. A method according to any one of claims 1, 3 to 6 and 8 to 17 or an apparatus according to any one of claims 2 to 5 and 7 to 17, wherein when signaling the set of quantization matrices, quantization matrices of larger block sizes are signaled before quantization matrices of smaller block sizes.
19. The identifier is matrixId=N * The method of any one of claims 1, 3 to 6 and 8 to 18 or the apparatus of any one of claims 2 to 5 and 7 to 18, wherein the matrix type id is derived as sizeId+matrixTypeId, where N is the number of possible type ids, sizeID indicates the block size, and matrixTypeId indicates the color components and the prediction mode.
20. The method of any one of claims 1, 3 to 6 and 8 to 19 or the apparatus of any one of claims 2 to 5 and 7 to 19, wherein the one or more processors are further configured to perform an adaptation of the reference quantization matrix to the block size.
21. The method of any one of claims 1, 3 to 6 and 8 to 20 or the apparatus of any one of claims 2 to 5 and 7 to 20, wherein the one or more processors are further configured to perform an adaptation of the reference quantization matrix to a chroma format of the block that is different from a default chroma format.
22. 21. The method or apparatus of claim 20, wherein the default chroma format is 4:2:
0.
23. A method according to any one of claims 20 to 22 or an apparatus according to any one of claims 20 to 22, wherein the adaptation is based on bit-shifting x and y coordinates in the quantization matrix to index coefficients of the reference quantization matrix.
24. A method according to any one of claims 1, 3 to 6 and 8 to 23 or an apparatus according to any one of claims 2 to 5 and 7 to 23, wherein the identifier is obtained based on whether the prediction mode of the block is an intra prediction mode or an inter prediction mode.
25. The method of any one of claims 1, 3 to 6 and 8 to 24 or the apparatus of any one of claims 2 to 5 and 7 to 24, wherein an intra block copy prediction mode is considered as an inter prediction mode when obtaining the identifier.
26. A method according to any one of claims 1, 3 to 6 and 8 to 25 or an apparatus according to any one of claims 2 to 5 and 7 to 25, wherein the prediction mode is intra block copy, a quantization matrix is signaled for the luminance component of the block, and a quantization matrix for the chroma component is derived by interpreting the prediction mode as an inter prediction mode.
27. A method according to any one of claims 1, 3 to 6 and 8 to 26 or an apparatus according to any one of claims 2 to 5 and 7 to 26, wherein the prediction mode is intra block copy, the reference quantization matrix is obtained by considering the prediction mode as an intra mode, and another reference quantization matrix is obtained by considering the prediction mode as an inter mode, and the quantization matrix is obtained as the average of the reference quantization matrix and the another reference quantization matrix.
28. A signal containing encoded video formed by carrying out a method according to any one of claims 6 and 8 to 27.
29. A computer readable storage medium having stored thereon instructions for encoding or decoding video data according to the method of any one of claims 1, 3-6 and 8-27.
Citation Information
Patent Citations
Image encoder, image encoding method and program, image decoder, image decoding method and program
JP2014011482A
Method and device for encoding / decoding image
US20150078442A1
Method for encoding and decoding quantized matrix and apparatus using same
US20150334396A1
Decoding method, decoder apparatus, encoding method, and encoder apparatus
WO2011052215A1
Image processing device and image processing method
WO2012108237A1