Single-index quantization matrix design for video encoding and decoding
A unified quantization matrix index system addresses inefficiencies in VVC by allowing flexible QM prediction and transmission, reducing bit costs and enhancing encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-25
AI Technical Summary
Existing video encoding and decoding technologies, such as VVC, face complexity and inefficiency in quantization matrix signaling and prediction processes, leading to increased bit costs and unnecessary bit waste due to limited QM prediction techniques.
A unified quantization matrix (QM) index system is introduced, allowing QM prediction from any previously signaled data, with processes for upsampling and downsampling based on block parameters and chroma format, ensuring all QM coefficients are transmitted for larger blocks, and simplifying the QM derivation and prediction processes.
This approach reduces bit costs by half compared to previous methods, simplifies the specification, and enhances the efficiency of QM signaling and prediction, thereby improving video encoding and decoding performance.
Smart Images

Figure 2026053380000017 
Figure 2026053380000018 
Figure 2026053380000019
Abstract
Description
Technical Field
[0001] Technical Field [1] Generally, the present embodiment relates to a method and apparatus for quantization matrix design in video encoding and decoding.
Background Art
[0002] Background [2] To achieve high compression efficiency, image and video encoding schemes typically utilize prediction and transformation to exploit spatial and temporal redundancies in video content. Generally, intra and inter prediction are used to exploit intra or inter picture correlation, in which case the difference between the original block, often called the prediction error or prediction residual, and the predicted block is transformed, quantized, and entropy encoded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy encoding, quantization, transformation, and prediction.
Summary of the Invention
[0003] Summary [3] According to one embodiment, a video decoding method is provided, the method including obtaining a single identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a decoding block in a picture; decoding a syntax element indicating a reference quantization matrix, the syntax element specifying a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtaining the quantization matrix based on the reference quantization matrix; inverse quantizing transform coefficients of the block in response to the quantization matrix; and decoding the block in response to the inverse quantized transform coefficients.
[0004] [4] According to another embodiment, a video coding method is provided, which includes accessing a block to be coded in a picture; accessing a quantization matrix of the block; obtaining a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; coding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantizing the transformation coefficients of the block in response to the quantization matrix; and entropy coding the quantized transformation coefficients.
[0005] [5] According to another embodiment, a video decoding device is provided, comprising one or more processors, the one or more processors configured to: obtain a single identifier of a quantization matrix based on the block size, color components and prediction mode of a block in a picture; decode a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtain the quantization matrix based on the reference quantization matrix; de-quantize the transformation coefficients of the block in response to the quantization matrix; and decode the block in response to the de-quantized transformation coefficients.
[0006] [6] According to another embodiment, a video encoding apparatus is provided, comprising one or more processors, the one or more processors, which access a block to encode in a picture; access a quantization matrix of the block; obtain a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; encode a syntax element indicating a reference quantization matrix, the syntax element being configured to encode a difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantize the transformation coefficients of the block in response to the quantization matrix; and entropy encode the quantized transformation coefficients.
[0007] [7] According to another embodiment, a video decoding apparatus is provided, comprising means for obtaining a single identifier of a quantization matrix based on the block size, color components and prediction mode of a block to be decoded in a picture; decoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; means for obtaining the quantization matrix based on the reference quantization matrix; means for inversely quantizing the transformation coefficients of the block in response to the quantization matrix; and means for decoding the block in response to the inversely quantized transformation coefficients.
[0008] [8] According to another embodiment, a video encoding apparatus is provided, comprising means for accessing a block to be encoded in a picture; means for accessing a quantization matrix of the block; means for obtaining a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; means for encoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; means for quantizing the transformation coefficients of the block in response to the quantization matrix; and means for entropy encoding the quantized transformation coefficients.
[0009] [9] One or more embodiments also provide a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform an encoding or decoding method according to any of the embodiments described above. One or more embodiments also provide a computer-readable storage medium storing instructions for encoding or decoding video data by the methods described above. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated by the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated by the methods described above. [Brief explanation of the drawing]
[0010] Brief explanation of the drawing [Figure 1]
[10] A block diagram of a system that can implement an embodiment of this model is shown. [Figure 2]
[11] A block diagram of one embodiment of a video encoder is shown. [Figure 3]
[12] A block diagram of one embodiment of a video decoder is shown. [Figure 4]
[13] VVC Draft 5 shows that for block sizes greater than 32, the conversion factor is presumed to be zero. [Figure 5]
[14] A fixed prediction tree is shown as described in JCTVC-H0314. [Figure 6]
[15] A prediction (decimation) from a larger size according to one embodiment is shown. [Figure 7]
[16] One embodiment shows a combination of prediction from a larger size and decimation of a rectangular block. [Figure 8]
[17] A QM derivation process for a rectangular block in chroma according to one embodiment is shown. [Figure 9]
[18] A QM derivation process for a rectangular block in chroma according to one embodiment (adapted to 4:2:2 format) is shown. [Figure 10]
[19] The QM derivation process for rectangular blocks in chroma (adaptation to 4:4:4 format) is shown. [Figure 11]
[20] A flowchart for parsing the scaling list data syntax structure according to one embodiment is shown. [Figure 12]
[21] A flowchart illustrating the encoding of a scaling list data syntax structure according to one embodiment is shown. [Figure 13]
[22] A flowchart of the QM derivation process according to one embodiment is shown. [Modes for carrying out the invention]
[0011] Detailed explanation
[23] Figure 1 shows a block diagram of an example of a system that can carry out various embodiments and designs. System 100 may be implemented as a device including various components described later and configured to perform one or more of the embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 100 may be implemented individually or in combination as one integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 100 are distributed across multiple ICs and / or discrete components. In various embodiments, System 100 is communicably coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, System 100 is configured to carry out one or more of the embodiments described herein.
[0012]
[24] System 100 includes, for example, at least one processor 110 configured to execute instructions loaded to carry out various embodiments described herein. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the Art. System 100 includes at least one memory 120 (e.g., volatile memory devices and / or non-volatile memory devices). System 100 includes a storage device 140, which may include, but is not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives, and may include non-volatile memory and / or volatile memory. Storage device 140 may include, as an example not limited to, internal storage, mounted storage, and / or network-accessible storage.
[0013]
[25] System 100 includes an encoder / decoder module 130 configured to process data to provide, for example, encoded video or decoded video. The encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Further, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated within the processor 110 as a combination of hardware and software, as is known to those skilled in the art.
[0014]
[26] The program code loaded into the processor 110 or the encoder / decoder 130 to execute the various aspects described herein may be stored in the storage device 140 and subsequently loaded into the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described herein. Such stored items may include, without limitation, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and equations, formulas, operations, and intermediate or final results of operational logic.
[0015]
[27] In some embodiments, the memory within the processor 110 and / or the encoder / decoder module 130 stores instructions and is used to provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140 and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used for storing the television operating system. In at least one embodiment, high-speed external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations such as MPEG-2, HEVC, or VVC.
[0016]
[28] Inputs to the elements of the system 100 can be provided through various input devices shown in block 105. Such input devices include, without limitation, (i) an RF section that receives RF signals transmitted wirelessly, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0017]
[29] In various embodiments, the input devices of block 105 are associated with various input processing elements known in the Art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also called signal selection or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting it to a narrower frequency band to select a signal frequency band which may be called a channel in a particular embodiment, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, e.g., frequency selectors, signal selectors, band-limiters, channel selectors, filters, down-converters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions, including, for example, down-converting the received signal to a low frequency (e.g., an intermediate frequency or near-baseband frequency) or to the baseband. In one set-top box embodiment, the RF unit and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cord) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments involve rearranging the order of the elements described above (and others), removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF unit includes an antenna.
[0018]
[30] Furthermore, the USB and / or HDMI terminals may include interface processors that connect the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be performed in separate input processing ICs or within processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be performed in separate interface ICs or within processor 110 as needed. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including processor 110 and encoder / decoder 130, which work in conjunction with memory and storage elements to process the data stream as needed for presentation to an output device.
[0019]
[31] Various elements of system 100 may be provided in an integrated enclosure, within which the various elements may be interconnected using suitable connectivity equipment 115, for example, an I2C bus, wiring, and an internal bus known in the art, including a printed circuit board, and data may be transmitted between them.
[0020]
[32] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, transceivers configured to send and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.
[0021]
[33] In various embodiments, the data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. In these embodiments, the communication channel 190 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable the streaming application and other over-the-top communication. In other embodiments, the streaming data is provided to the system 100 using a set-top box that distributes the data via HDMI communication of the input block 105. In yet another embodiment, the streaming data is provided to the system 100 using an RF connection of the input block 105.
[0022]
[34] System 100 may provide output signals to various output devices, including a display 165, a speaker 175, and other peripherals 185. Other peripherals 185 may include, in various embodiments, one or more standalone DVRs, disc players, stereo systems, lighting systems, and other devices that provide functionality based on the output of System 100. In various embodiments, control signals are communicated between System 100 and the display 165, speaker 175, or other peripherals 185 using signaling such as AV Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicably coupled to System 100 via dedicated connections through each of the interfaces 160, 170, and 180. Alternatively, output devices may be connected to System 100 using a communication channel 190 via a communication interface 150. The display 165 and speaker 175 may be integrated into a single unit with other components of System 100, such as an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0023]
[35] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, for example, an HDMI port, a USB port, or a COMP output.
[0024]
[36] Figure 2 shows a video encoder 200, which is an example of a high-efficiency video coding (HEVC) encoder. Figure 2 may also show an encoder that has been improved from the HEVC standard or an encoder that uses technology similar to HEVC, such as a VVC (Versatile Video Coding) encoder currently under development by JVET (Joint Video Exploration Team).
[0025]
[37] In this application, the terms “reconstructed” and “decoded” may be used synonymously, the terms “encoded” or “coded” may be used synonymously, and the terms “image,” “picture,” and “frame” may be used synonymously. Although not required, the term “reconstructed” is usually used on the encoder side, while “decoded” is usually used on the decoder side.
[0026]
[38] Before encoding, the video sequence may undergo preprocessing (201), for example, applying a color conversion (e.g., a conversion from RGB4:4:4 to YCbCr4:2:0) to the input color picture, or performing a remapping of the input picture components, in order to obtain a signal distribution that is more resilient to compression (e.g., by using histogram equalization of one of the color components). Metadata may be associated with the preprocessing and attached to the bitstream.
[0027]
[39] In encoder 200, the picture is encoded by encoder elements as described below. The picture to be encoded is divided into units, for example, CUs (202), and processed by those units. Each unit is encoded using either intra-mode or inter-mode, for example. If a unit is encoded in intra-mode, intra-prediction is performed (260). In inter-mode, motion estimation (275) and compensation (270) are performed. The encoder decides whether to use intra-mode or inter-mode for encoding a unit (205), and indicates the intra / inter decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting the prediction block from the original image block (210).
[0028]
[40] Next, the predicted residual is transformed (225) and quantized (230). The quantized transformed coefficients, as well as the motion vector and other syntax elements, are entropy encoded (245) and output as a bitstream. The encoder may skip the transformation and apply quantization directly to the untransformed residual signal. The encoder may bypass both transformation and quantization, i.e., the residual is coded directly without applying any transformation or quantization process.
[0029]
[41] The encoder decodes the coded blocks to provide a basis for further prediction. The quantized transformation coefficients are inversely quantized (240), inversely transformed (250), and the prediction residuals are decoded. The decoded prediction residuals and the prediction blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0030]
[42] Figure 3 shows a block diagram of an example video decoder 300. In the decoder 300, the bitstream is decoded by the decoder elements described below. The video decoder 300 generally performs the coding path and the corresponding decoding path described in Figure 2. The encoder 200 also performs video decoding as part of the video data coding.
[0031]
[43] In particular, the input to the decoder includes a video bitstream that can be generated by the video encoder 200. The bitstream is first entropy-decoded (330) to obtain transformation coefficients, motion vectors, and other coded information. Picture division information indicates how the picture has been divided. Thus, the decoder can divide the picture according to the decoded picture division information (335). The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. The prediction blocks can be obtained from intra-predictions (360) or motion-compensated predictions (i.e., inter-predictions) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0032]
[44] The decoded picture may undergo further post-decoded processing (385), such as inverse color transformation (e.g., conversion from YCbCr4:2:0 to RGB4:4:4) or reverse remapping, which is the reverse of the remapping process performed in pre-encoding processing (201). The post-decoded processing may use metadata derived in pre-encoding processing and signaled in the bitstream.
[0033]
[45] The HEVC specification allows the use of a quantization matrix in the inverse quantization process, where the transformed coefficients are scaled by the current quantization step and further scaled by the quantization matrix (QM) as follows: d[ x ][ y ] = Clip3( coeffMin, coeffMax, ( ( TransCoeffLevel[ xTbY ][ yTbY ][ cIdx ][ x ][ y ] * m[ x ][ y ] * levelScale[ qP%6 ] << (qP / 6 ) ) + ( 1 << ( bdShift - 1 ) ) ) >> bdShift ) During the ceremony, TransCoeffLevel[...] is the absolute value of the transformation coefficient of the current block, identified by the spatial coordinates xTbY, yTbY and component index cIdx. x and y are horizontal / vertical frequency indices. • qP is currently the quantization parameter. Multiplication by levelScale[qP%6] and left shift by (qP / 6) are equivalent to multiplication by the quantization step qStep = (levelScale[qP%6] << (qP / 6)). m[…][…] is a two-dimensional quantization matrix. Here, the quantization matrix is sometimes called the scaling matrix because it is used for scaling. bdShift is an additional scaling factor that describes the image sample bit depth. The term (1 << (bdShift-1)) is responsible for rounding to the nearest integer. ·d[...] is the absolute value of the resulting inversely quantized transformation coefficient.
[0034]
[46] The syntax used by HEVC for transmitting the quantization matrix is as follows:
[0035] [Table 1]
[0036]
[47] The following points should be noted: Each transformation size (sizeId) is assigned a different matrix. In the scaling list data syntax structure, the scaling matrix is scanned into a one-dimensional scaling list (e.g., ScalingList). Given a given transformation size, six matrices are specified: intra / intercoded and Y / Cb / Cr components. The matrix can be any of the following: If scaling_list_pred_mode_flag is zero (the criterion matrixId is obtained as matrixId - scaling_list_pred_matrix_id_delta), then it can be copied from a previously sent matrix of the same size. - It can be copied from the default value specified in the standard (if both scaling_list_pred_mode_flag and scaling_list_pred_matrix_id_delta are zero). • The scan order can be fully specified using exponential Golomb entropy coding in DPCM coding mode, starting from the upper right diagonal. For block sizes larger than 8x8, only 8x8 coefficients are sent to signal the quantization matrix to conserve coded bits. The coefficients are then interpolated using zero-holds (i.e., iterations), except for the DC coefficients that are explicitly sent.
[0037]
[48] The use of quantization matrices similar to those in HEVC is adopted in VVC Draft 5 based on contribution JVET-N0847 (see O. Chubach, et al., “CE7-related: Support of quantization matrices for VVC,” JVET-N0847, Geneva, CH, March 2019). The scaling_list_data syntax is adapted to the following VVC codecs.
[0038] [Table 2]
[0039]
[49] In the VVC Draft 5 design using JVET-N0847, the QM is identified by two parameters, matrixId and sizeId, as in HEVC. These are shown in the following two tables.
[0040] [Table 3]
[0041] [Table 4]
[0042]
[50] The combinations of both identifiers are shown in the table below.
[0043] [Table 5]
[0044]
[51] For block sizes larger than 8x8, as in HEVC, only the 8x8 coefficients and DC coefficients are sent. The QM of the correct size is reconstructed using zero-hold interpolation. For example, for a 16x16 block, every coefficient is repeated twice in both directions, and the DC coefficient is replaced with the one that was sent.
[0045]
[52] For rectangular blocks, the size held in the QM selection (sizeId) is the larger of the two dimensions, namely the width and the height. For example, for a 4x16 block, a QM of 16x16 block size is selected. The reconstructed 16x16 matrix is then vertically decimated by factor 4 to obtain the final 4x16 quantized matrix (i.e., three of the four lines are skipped).
[0046]
[53] Hereinafter, in relation to sizeId and the square block size used, the QM of a given family block size (square or rectangle) will be referred to as size-N. For example, for a block size of 16×16 or 16×4, the QM is identified as size-16 (sizeId4 in VVC Draft 5). The size-N notation is used to distinguish it from the exact block shape and the number of QM coefficients signaled (limited to 8×8, as shown in Table 3).
[0047]
[54] Furthermore, in VVC Draft 5, for size -64, QM coefficients in the lower right quadrant are not transmitted (presumed to be 0, which will hereafter be referred to as "zero-out"). This is done by the "x>=4&&y>=4" condition in the scaling_list_data syntax. This avoids transmitting QM coefficients that are never used in the transformation / quantization process. In fact, in VVC, when transforming block sizes greater than 32 for any dimension (64×N, N×64, N≦64), any transformation coefficients with x / y frequency coordinates greater than 32 are not transmitted and are presumed to be zero, and therefore quantization matrix coefficients are not needed for their quantization. This is shown in Figure 4, where the shaded area corresponds to the transformation coefficients that are presumed to be zero.
[0048]
[55] Compared to HEVC, VVC requires more quantization matrices due to the larger number of block sizes. However, in VVC Draft 5, QM prediction is still limited to copies of the same block size matrix, which can lead to bit waste. Furthermore, the syntax related to QM is more complex in VVC by using only a 2x2 block size for chroma and only a 64x64 block size for luminance. Also, JVET-N0847 describes a specific matrix derivation process for each block size, similar to HEVC.
[0049]
[56] During the HEVC standardization process, several QM prediction techniques were explored, for example, in JCTVC-E073 (see J. Tanaka, et al., “Quantization Matrix for HEVC,” JCTVC-E073, Geneva, CH, March 2011) and JCTVC-H0314 (see Y. Wang, et al., “Layered quantization matrices representation and compression,” JCTVC-H0314, San Jose, CA, USA, February 2012).
[0050]
[57] JCTVC-E073:QM is transmitted in a specific parameter set (QMPS). Within the QMPS, QM is transmitted in increasing order of size (sizeId / matrixId, as in HEVC). Prediction (=copy) from any previously encoded QM coefficients, including the aforementioned QMPS, has been proposed. Upconversion using linear interpolation is used for fitting from smaller reference QMs, while simple downsampling is used for fitting from larger reference QMs. This was ultimately rejected during the HEVC standardization.
[0051]
[58] JCTVC-H0314: QMs are sent in ascending or ascending order. As shown in Figure 5, it is possible to copy previously sent QMs instead of sending new ones using a fixed prediction tree (without an explicit criterion index). If the criterion QM is larger, simple downsampling is used. This was ultimately rejected during the HEVC standardization.
[0052]
[59] These two proposals relate to HEVC and do not address the complexity introduced by VVC.
[0053]
[60] The present invention proposes to simplify the quantization matrix signaling and prediction process of VVC Draft 5 (after adoption of JVET-N0847) while enhancing the quantization matrix signaling and prediction process so that any QM can be predicted from any previously signaled data by incorporating one or more of the following: - Unify the QM index to include both size and type so that the base index difference can handle data sent to any destination. - The quantization matrices are sent in order of decreasing block size. - Specify the prediction process as either a copy or decimation process, as needed. -Send all QM coefficients with size -64 so that they can be used as predictors.
[0054]
[61] Furthermore, the QM derivation process, which includes upsampling of blocks larger than 8x8 and downsampling of rectangular blocks, is described as selecting a QM index according to the block parameters and fitting the QM signaling size to the actual block size.
[0055]
[62] For ease of notation, the process of predicting the quantization matrix from default values or from previously transmitted data is considered the QM prediction process, and the process of fitting the transmitted or predicted QM to the size of the transformation block and the chroma format is considered the QM derivation process. The QM prediction process can be, for example, part of a scaling list data parsing process at the picture level. The derivation process is usually at a lower level, for example, at the transformation block level. Various embodiments are presented in more detail below, followed by draft text examples and performance results. • Derivation and use of a single matrix index to identify the QM. With a single identifier, when using prediction (copy), it is possible to reference matrices sent to any destination, and by sending the larger matrix first, interpolation in the prediction process is avoided. A QM prediction process that includes copying or decimating a previously signaled QM (reference QM) that is transmitted, predicted, or is the default reference QM. For a given transformation block, the QM derivation process includes selecting a QM index based on the block size, color components, and prediction mode, and then fitting the size of the selected QM to the size of the block. The resizing process is based on bit shifts of the x and y coordinates within the transformation block that index the coefficients of the selected QM. • Send all coefficients for size -64QM.
[0056]
[63] Compared to VVC Draft 5, these embodiments simplify the specification (halving the text changes compared to JVET-N0847) and result in a larger bit limit (halving the bit cost of scaling_list_data).
[0057]
[64] Unified QM Index
[65] The QM used for quantization / dequantization of the transformation block is identified by a single parameter matrixId. In one embodiment, the unified matrixId (QM index) is a composite of the following: - A size identifier related to the CU size (i.e., the CU enclosing a square, since only square-sized matrices are transmitted), rather than the block size. Here, for either luminance or chroma, the size identifier is controlled by the luminance block size, e.g., max(luminance block width, luminance block height). When luminance and chroma trees are separated, for chroma, "CU size" refers to the size of the block projected onto the luminance plane. - Luminance QM can be larger than chroma (for example, in a 4:2:0 chroma format), so the matrix type lists luminance QM first.
[0058]
[66] According to this embodiment, the QM index is derived in Tables 4, 5, and Equation (1).
[0059] [Table 6]
[0060] [Table 7]
[0061]
[67] The unified matrixId is derived as follows: matrixId=N * sizeId+matrixTypeId (1) In the formula, N is the number of possible types of identifiers, for example, N=6.
[0062]
[68] In another embodiment, if seven or more QM types are defined, sizeId should be multiplied by the correct number which is the number of quantization matrix types. In other embodiments, other parameters may also differ, for example, a specific block size, the size of the signaling matrix (limited here to 8x8), or the presence or absence of a DC coefficient. Here, the QMs are listed in order of decreasing block size and identified by a single index, as shown in Table 6.
[0063] [Table 8]
[0064]
[69] QM prediction process
[70] Instead of sending QM coefficients, it is possible to predict QM from default values or from QM coefficients sent to any destination. In one embodiment, if the reference QM is the same size, the QM is copied; otherwise, it is decimated by the relevant ratio, as shown in an example in Figure 6, where a size-4 luminance QM is predicted from a size-8.
[0065]
[71] Decimation is described by the following formula: ScalingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (2) In the formula, matrixSize = (matrixId < 20) ? 8 : (matrixId < 26) ? 4 : 2 ) x = 0 .. matrixSize - 1, y = 0 .. matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ), and j = y << ( log2(refMatrixSize) - log2( matrixSize ) ). In the formula, refMatrixSize corresponds to the size of refScalingMatrix (and consequently the range of the variables i and j).
[0066]
[72] In the example shown in Figure 6, the luminance size -4QM (4x4 array: matrixSize is 4) is predicted from the luminance size -8QM, which is an 8x8 array (refMatrixSize is 8); one of the two lines and one of the two columns are dropped to generate the 4x4 array (i.e., the element (2x,2y) in the reference QM is copied to the element (x,y) in the current QM).
[0067]
[73] Equation (3) takes the following form: ScalingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (3) x = 0 .. 3, y = 0 .. 3, i = x << 1, and j = y << 1.
[0068]
[74] If the reference QM has a DC value, the reference QM is copied as the DC value when the current QM requires a DC value, and the reference QM is copied to the upper left QM coefficient when the current QM does not require a DC value.
[0069]
[75] In a preferred embodiment, this QM prediction process is part of the QM decoding process, but in another embodiment, it can be deferred to the QM derivation process, in which case the decimation for prediction purposes is merged with the QM resize subprocess.
[0070]
[76] QM Derivation Process
[77] The proposed derivation process for the quantization matrix first selects the right QM index according to the block parameters as described above (unified QM index), and then unifies the processes of decimation in rectangular blocks, iteration in blocks of a certain size, e.g., larger than 8x8, and chroma format fitting into a single process. The proposed process is based on bit shifts of the x and y output coordinates. To select the right line / column of the selected QM, only a left shift followed by a right shift of the x / y output coordinates is required, as shown in the following equation: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (4) In the formula, i = ( x << log2MatrixSize ) >> log2( blkWidth ), and j = ( y << log2MatrixSize ) >> log2( blkHeight ). In the formula, log2MatrixSize is the log2 of the size of ScalingMatrix[matrixId] (a 2D square array), blkWidth and blkHeight are the width and height of the current transformation block, respectively, x is in the range of 0 to blkWidth-1, and y is in the range of 0 to blkHeight-1.
[0071]
[78] The following are some examples to illustrate the QM derivation process. In the example shown in Figure 7, the QM of a luminance 16 × 8 block is derived from the luminance size - 16 QM, which is actually an 8 × 8 array plus a DC factor. In this example, blkWidth is equal to 16, blkHeight is equal to 8, log2MatrixSize is equal to 3, and therefore equation (5) takes the following form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (5) In the equation, i = ( x << 3 ) >> 4, and j = ( y << 3 ) >> 3, where x = 0..15 and y = 0..7. Here, x is shifted to the right by 1, and y remains unchanged (i.e., column i in the selected QM is now column 2 in the QM). * i and 2 * (It is copied to i+1). Furthermore, since the selected QM has a DC coefficient, it is copied to m[0][0].
[0072]
[79] In another example shown in Figure 8, a QM (4:2:0 format) of chroma 4x2 blocks of 8x4CU is generated. This matches the 8x4CU size, where the surrounding square is 8x8. Thus the selected QM is a QM of size -8, where the chroma QM is encoded as a 4x4 array. Here, blkWidth is equal to 4, blkHeight is equal to 2, log2MatrixSize is equal to 2, and therefore equation (6) takes the form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (6) In the equation, i = ( x << 2 ) >> 2, and j = ( y << 2 ) >> 1, where x = 0..3 and y = 0..1. Here, x remains unchanged, while y is shifted left by 1 (i.e., row 2y in the reference QM is copied to row y in the current QM).
[0073]
[80] In the following examples, the proposed adaptations to the 4:2:2 and 4:4:4 formats differ from those in VVC Draft 5. Instead of searching for a QM that matches the chroma block size (except for the 64x64 case where no chroma matrix exists), size matching is based on the same (luminance) CU size (i.e., the size of the blocks projected onto the luminance plane), with coefficients repeated where necessary. This makes the QM design independent of the chroma format.
[0074]
[81] In the example shown in Figure 9, a QM (4:2:2 format) of an 8×4CU chroma 8×2 block is generated. The selected QM is the same as in the example shown in Figure 8 above, but the 4:2:2 chroma format requires twice as many columns. Here, the columns are repeated, and therefore x is shifted to the right by 1 and y is shifted to the left by another 1. In particular, blkWidth is equal to 8, blkHeight is equal to 2, log2MatrixSize is equal to 2, and therefore equation (7) takes the following form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (7) In the equation, i = ( x << 2 ) >> 3, and j = ( y << 2 ) >> 1, where x = 0..7 and y = 0..1.
[0075]
[82] In the example shown in Figure 10, a QM (4:4:4 format) of an 8x4 CU chroma 8x4 block is generated. The selected QM is still the same as in the examples shown in Figures 8 and 9, but the 4:4:4 chroma format requires twice the number of rows and columns as the 4:2:0 chroma format. Here, the columns must be repeated, and therefore x is shifted to the right by 1, but row decimation (due to rectangles) can be skipped, and therefore y is not shifted. In particular, blkWidth is equal to 8, blkHeight is equal to 4, log2MatrixSize is equal to 2, and therefore equation (8) takes the form: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (8) In the equation, i = ( x << 2 ) >> 3, and j = ( y << 2 ) >> 2, where x = 0..7 and y = 0..3.
[0076]
[83] Number of transmission coefficients when size is -64
[84] In one embodiment, all coefficients of size -64 are transmitted in scaling list syntax so that smaller QMs can be predicted from size -64, even if the lower right quadrant is never used by the VVC transformation and quantization process. Generally, the present invention may transmit all coefficients of the maximum QM.
[0077]
[85] However, it is worth noting that if not used as a predictor, the syntax element scaling_list_delta_coef can be set to zero in the lower right quadrant of size -64QM, so the overhead associated with this increase in the number of coefficients transmitted compared to the previous work (JVET-N0847) can be limited to 2 × 16 bits in the worst case: in the lower right quadrant, 4 × 4 = 16 delta coefficients are signaled and forced to zero (encoded using exponential golombs), each taking 1 bit, and there are two sizes -64QM (luminance intra / inter):
[0078]
[86] The tests are described in Table 8, which shows that this overhead is negligible compared to the gains brought about by the prediction improvement.
[0079]
[87] In another embodiment, the coefficients of the lower right quadrant of size -64QM are not transmitted as part of size -64QM signaling, but are transmitted as supplemental parameters when a smaller QM is first predicted from a given size -64QM.
[0080]
[88] Table 7 provides some comparison between the method described in JVET-N0847 and the proposed method.
[0081] [Table 9]
[0082]
[89] The following describes some syntax and semantics according to one embodiment.
[0083]
[90] PPS Syntax and Semantics (Minor Conformity)
[0084] [Table 10]
[0085] A pps_scaling_list_data_present_flag equal to 1 specifies that the scaling list data used for the picture called PPS is derived based on the scaling list specified by the active SPS and the scaling list specified by the PPS. A pps_scaling_list_data_present_flag equal to 0 specifies that the scaling list data used for the picture called PPS is assumed to be equal to that specified by the active SPS. If scaling_list_enabled_flag is equal to zero, the value of pps_scaling_list_data_present_flag is assumed to be equal to 0. If scaling_list_enabled_flag is equal to 1, sps_scaling_list_data_present_flag is equal to 0, and pps_scaling_list_data_present_flag is equal to 0, and the default scaling matrix is used to derive the array ScalingMatrix as described in the scaling list data semantics, as specified in Section 7.4.5.
[0086]
[91] Note that this syntax / semantics is an example, but not an limitation, of being intended to be close to the HEVC standard or the VVC draft. For example, the scaling_list_data transport is not limited to SPS or PPS and can be transmitted by other means.
[0087]
[92] Signaling list data syntax / semantics (simplified)
[0088] [Table 11]
[0089] A `scaling_list_pred_mode_flag[matrixId]` equal to 0 specifies that the scaling matrix is derived from the values of the base scaling matrix, which is specified by `scaling_list_pred_matrix_id_delta[matrixId]`. A `scaling_list_pred_mode_flag[matrixId]` equal to 1 specifies that the scaling list values are explicitly signaled. `scaling_list_pred_matrix_id_delta[matrixId]` specifies the base scaling matrix used to derive the scaling matrix, as follows: The value of `scaling_list_pred_matrix_id_delta[matrixId]` must be within the range of 0 to `matrixId` (including fractions). If scaling_list_pred_mode_flag[matrixId] is equal to zero: The variables refMatrixSize and the array refScalingMatrix are first derived as follows: If scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, the following default values will be applied: • refMatrixSize is set to equal 8, If matrixId is even, refScalingMatrix = (9) { { 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for intra-default value { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, Other cases refScalingMatrix = (10) { { 16, 16, 16, 16, 16, 16, 16, 16} / / Placeholder for the default value of the interface { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} { 16, 16, 16, 16, 16, 16, 16, 16} }, • In other cases (where scaling_list_pred_matrix_id_delta[matrixId] is greater than zero), the following applies: refMatrixId = matrixId - scaling_list_pred_matrix_id_delta[ matrixId ] (11) refMatrixSize = (refMatrixId < 20) ? 8 : (refMatrixId < 26) ? 4 : 2 ) (12) refScalingMatrix = ScalingMatrix[ refMatrixId ] (13) Next, the array ScalingMatrix[matrixId] is derived as follows: ScalingMatrix[ matrixId ][ x ][ y ] = refScalingMatrix[ i ][ j ] (14) In the formula, matrixSize = (matrixId < 20) ? 8 : (matrixId < 26) ? 4 : 2 ) x = 0 .. matrixSize - 1, y = 0 .. matrixSize - 1, i = x << ( log2(refMatrixSize) - log2( matrixSize ) ), and j = y << ( log2(refMatrixSize) - log2( matrixSize ) ) `scaling_list_dc_coef_minus8[matrixId]+8` specifies the first value of the scaling matrix, if relevant, as described in section xxx. The value of `scaling_list_dc_coef_minus8[matrixId]` is assumed to be in the range of -7 to 247 (including fractions). If scaling_list_pred_mode_flag[matrixId] is equal to zero, then scaling_list_pred_matrix_id_delta[matrixId] is greater than zero and refMatrixId < 14, and the following applies: -If matrixId<14, then scaling_list_dc_coef_minus8[matrixId] is presumed to be equal to scaling_list_dc_coef_minus8[refMatrixId], -In other cases, ScalingMatrix[matrixId][0][0] is set to equal to scaling_list_dc_coef_minus8[refMatrixId]+8. If scaling_list_pred_mode_flag[matrixId] is equal to zero, then scaling_list_pred_matrix_id_delta[matrixId] is equal to zero (indicating the default value), and matrixId < 14, so it is inferred that scaling_list_dc_coef_minus8[matrixId] is equal to 8. `scaling_list_delta_coef` specifies the difference between the current matrix coefficient `ScalingList[matrixId][i]` and the previous matrix coefficient `ScalingList[matrixId][i-1]`, assuming `scaling_list_pred_mode_flag[matrixId]` is equal to 1. The value of `scaling_list_delta_coef` must be in the range of -128 to 127 (including fractions). The value of `ScalingList[matrixId][i]` must be greater than 0. If it exists (i.e., scaling_list_pred_mode_flag[matrixId] is equal to 1), the array ScalingMatrix[matrixId] is derived as follows: ScalingMatrix[ matrixId ][ i ][ j ] = ScalingList[ matrixId ][ k ] (15) where k = 0 .. coefNum - 1, i = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0000] , and j = diagScanOrder[ log2(coefNum) / 2 ][ log2(coefNum) / 2 ][ k ]
[0001]
[0090]
[93] The main simplifications compared to the JVET-N0847 syntax are the elimination of one for() loop and the simplification of the index from [sizeId][matrixId] to [matrixId].
[0091]
[94] "xxx section" refers to an indeterminate section number introduced in the VVC specification that corresponds to the scaling matrix derivation process in this document.
[0092]
[95] Note that this syntax / semantics is an example, not an limitation, of being intended to be close to the HEVC standard or VVC Draft 5. For example, the coefficient range is not limited to 1, ..., 255, but may be, for example, 1, ..., 127 (7 bits) or -64, ..., 63. Nor is it limited to 30 QMs organized as 6 types × 5 sizes (there may be 8 types and fewer or more sizes; see Table 5 which can be adapted, in which case a simple adaptation to coefNum and conditions on the presence of DC coefficients are required). The type of QM prediction (copy only here) is also not limited. For example, a scaling factor or offset may be added, and explicit coding may be added on top of the prediction as residuals. The same is true for the method used for coefficient transmission (DPCM here), the presence of DC coefficients, and the number of coefficients fixed (only a subset may be transmitted).
[0093]
[96] With respect to default values, we are not limited to the two default QMs associated with MODE_INTRA and MODE_INTER, but can be populated with all 16 values until a corresponding default value matches (for example, the same default QM as HEVC can be selected), as shown here.
[0094]
[97] Also, the number of signaling coefficients coefNum can be expressed mathematically instead of a series of comparisons, and the result is the same: coefNum = Min(64,4096>>((matrixId+4) / 6) * 2) This is similar to the HEVC or current VVC draft format, but introduces division which may not be welcomed.
[0095]
[98] Note that, here, the larger matrix is sent first, one index is used, and so the prediction criterion (indicated by scaling_list_pred_matrix_id_delta) can be any matrix sent first or a default value (for example, if scaling_list_pred_matrix_id_delta is zero), regardless of the intended block size or type.
[0096]
[99] Figure 11 shows the process (1100) of parsing a scaling list data syntax structure according to one embodiment. In this embodiment, the input is a coded bitstream and the output is an array of ScalingMatrix. For clarity, details about DC values are omitted. In particular, in step 1110, the QM prediction mode is decoded from the bitstream. If the QM is predicted (1120), the decoder further determines, depending on the flag above, whether the QM is inferred (predicted) or signaled in the bitstream. In step 1130, the decoder decodes the QM prediction data from the bitstream, which is necessary to infer the QM, for example, if the QM index difference scaling_list_pred_matrix_id_delta is not signaled. The decoder then determines whether the QM is predicted from a default value (for example, if scaling_list_pred_matrix_id_delta is zero) or from the previously decoded QM (1140). If the reference QM is the default QM, the decoder selects the default QM as the reference QM (1150). For example, depending on the parity of matrixId, there may be several default QMs to choose from. Otherwise, the decoder selects the previously decoded QM as the reference QM (1155). The index of the reference QM is derived from matrixId and the above index difference. In step 1160, the decoder predicts the QM from the reference QM. The prediction consists of a simple copy if the reference QM is the same size as the current QM, or a decimation if it is larger than expected. The result is stored in ScalingMatrix[matrixId].
[0097]
[0100] If QM is not predicted (1120), the decoder determines the number of QM coefficients to be decoded from the bitstream according to matrixId (1170). For example, if matrixId is less than 20, it is 64; if matrixId is between 20 and 25, it is 16; and otherwise, it is 4. In step 1175, the decoder decodes the corresponding number of QM coefficients from the bitstream. In step 1180, the decoder organizes the decoded QM coefficients into a 2D matrix according to the scan order, e.g., diagonal scan. The result is stored in ScalingMatrix[matrixId]. Using ScalingMatrix[matrixId], the decoder can use the QM derivation process to obtain a quantization matrix m[][] that dequantizes the transformation block, which may be non-square and / or of a different chroma format.
[0098]
[0101] In step 1190, the decoder checks whether the current QM is the last QM to be parsed. If it is not the last, control returns to step 1110; otherwise, the QM parsing process stops when all QMs have been parsed from the bitstream.
[0099]
[0102] Figure 12 shows the process (1200) of encoding the scaling list data syntax structure on the encoder side according to one embodiment. On the encoder side, the QMs are scanned in the order described, from larger block sizes to smaller block sizes (e.g., from matrixId0 to 30). In step 1210, the encoder searches for prediction preferences to determine whether the current QM is a copy (or decimation) of one previously encoded. The QMs can be designed to optimize the efficiency of QM prediction, for example, if several QMs are initially close enough or close to the default QM, they can be forced to be equal (or forced to be decimated if they are different in size). Furthermore, the coefficients in the lower right quadrant of size -64QM can be optimized to better predict subsequent QMs or to reduce the QM bit cost if they are never reused for prediction. Once determined, the QM prediction modes are encoded into a bitstream.
[0100]
[0103] In particular, if the encoder decides to use prediction (1220), in step 1230 the prediction mode is encoded (e.g., scaling_list_pred_mode_flag=0). In step 1240 the prediction parameters (e.g., QM index difference scaling_list_pred_matrix_id_delta): zero index difference for the default QM value, or the relevant index difference if the previous QM is selected as the prediction criterion. On the other hand, if explicit signaling is decided, in step 1250 the prediction mode is encoded (e.g., scaling_list_pred_mode_flag=0). Next, a diagonal scan (1260) is performed, followed by QM coefficient coding (1270).
[0101]
[0104] In step 1280, the encoder checks whether the current QM is the last QM to be encoded. If it is not the last, control returns to step 1210; otherwise, the QM encoding process stops when all QMs have been encoded into the bitstream.
[0102]
[0105] Figure 13 shows a QM derivation process 1300 according to one embodiment. The input includes a ScalingMatrix array that can transform block parameters such as size (width / height), prediction mode (intra / inter / IBC, ...), and color components (Y / U / V). The output is a QM having the same size as the transformed block. For clarity, details about the DC value are omitted. In particular, in step 1310, the decoder determines the QM index matrixId according to the current transformed block size (width / height), prediction mode (intra / inter / IBC, ...), and color components (Y / U / V), as described above (unified QM index). In step 1320, the decoder resizes the selected QM (ScalingMatrix[matrixId]) to match the transformed block size, as described above. In one variation, step 1320 may include the decimation required for prediction.
[0103]
[0106] The QM derivation process is similar on the encoder side. Quantization divides the transformation coefficients by the QM value, while inverse quantization multiplies them. However, the QM remains the same. In particular, the QM required for reconstruction in the encoder matches that which is signaled in the bitstream.
[0104]
[0107] Conceptually, the transformation coefficients d[x][y] can be quantized as follows, where qStep is the quantization step size and m[][] is the quantization matrix: TransCoeffLevel[xTbY][yTbY][cIdx][x][y] = d[x][y] / qStep / m[x][y]
[0105]
[0108] However, in integer calculations, to avoid division, usually, TransCoeffLevel[xTbY][yTbY][cIdx][x][y] = ( ( d[x][y] * im[x][y] * ilevelScale[ qP%6 ] >> (qP / 6 ) ) + ( 1 << ( bdShift - 1 ) ) ) >> bdShift ) It appears as follows, for example, im[x][y]≈65536 / m[x][y], ilevelScale[0..5]=65536 / levelScale[0..5], and bdShift is an appropriate value. In fact, for a software coder, im * ilevelScale is typically calculated in advance and stored in a table.
[0106]
[0109] In the above, the QM prediction process and the QM derivation process are executed separately. In another embodiment, QM prediction can be postponed until the QM derivation process. This embodiment does not change the QM signaling syntax. This embodiment can be functionally different by sequential resizing, postponing the prediction portion (reference QM acquisition + copy / downscale) and diagonal scanning until the "QM derivation process" described later.
[0107]
[0110] In one embodiment, the output of the scaling list data parsing process 1100 is an array of ScalingList() instead of ScalingMatrix, along with prediction flags and valid prediction parameters: the ScalingMatrixPredId array always contains the index of the defined ScalingList (default or signaling). This array is recursively constructed during QM decoding by interpreting scaling_list_pred_matrix_id_delta, so that the QM derivation process can directly use this index to obtain the actual value for constructing the QM used to dequantize the current transform block.
[0108]
[0111] Below, we provide an example to illustrate scaling list semantics according to one embodiment.
[0109] Scaling matrix derivation process (New: Partially replaces the explanation in Scaling List Semantics; this is in section xxx)
[0112] The inputs to this process are the prediction mode (predMode), the color component variable (cIdx), the block width (blkWidth), and the block height (blkHeight). The output of this process is a (blkWidth) × (blkHeight) array m[x][y] (scaling matrix), where x and y are the horizontal and vertical coefficient positions. SubWidthC and SubHeightC depend on the chroma format and represent the ratio of the number of samples in the luminance component and the chroma component. The variable matrixId is derived as follows: matrixId = 6 * sizeId + matrixTypeId (xxx-1) In the formula, subWidth = (cIdx > 0) ? SubWidthC : 1, subHeight = (cIdx > 0) ? SubHeightC : 1, sizeId = 6 - max( log2( blkWidth * subWidth ), log2( blkHeight * subHeight ) ), and matrixTypeId = ( 2 * cIdx + ( predMode = = MODE_INTER ? 1 : 0 ) ) The variable log2MatrixSize is derived as follows: log2MatrixSize = (matrixId < 20) ? 3 : (matrixId < 26) ? 2 : 1 (xxx-2) The output array m[x][y] is derived by applying the following, where x ranges from 0 to blkWidth-1, including fractions, and y ranges from 0 to blkHeight-1, including fractions: m[ x ][ y ] = ScalingMatrix[ matrixId ][ i ][ j ] (xxx-3) In the formula, i = ( x << log2MatrixSize ) >> log2( blkWidth ), and j = ( y << log2MatrixSize ) >> log2( blkHeight ) If matrixId is lower than 14, m[0][0] is further modified as follows: m
[0000]
[0000] = scaling_list_dc_coef_minus8[ matrixId ] + 8 (xxx-4)
[0110]
[0113] As with the scaling_list_data syntax and semantics, it should be noted that this is an example, not an exhaustive one. For example, it is not limited to the two default QMs associated with MODE_INTRA and MODE_INTER, but can be filled with all 16 values until the associated default values match (for example, the same default QM as HEVC can be selected), as shown here. There may be one default QM, or there may be three or more default QMs. Calculating the matrixId can be difficult, for example, if there are more or fewer types than 6 for each block size. Importantly, horizontal and vertical downscaling and upscaling to fit different block sizes with the selected QM should preferably be done in a single, simple process (here, a left shift followed by a right shift).
[0111]
[0114] In the case of rectangular blocks, the selection of the QM identifier of the current block enclosing the square is not limited to the derivation of sizeId in expression xxx-1; different rules may apply.
[0112]
[0115] Furthermore, the selected QM size log2MatrixSize can also be expressed mathematically instead of a series of comparisons, and the result is the same: log2MatrixSize=min(3,6-(matrixId+4) / 6), but this introduces division, which can be unwelcome.
[0113]
[0116] Below, we provide an example to illustrate the semantics of a scaling process according to one embodiment.
[0114] Scaling process of (adapted) transformation coefficients
[0117] [...] To derive the scaled transformation coefficients d[x][y] where x=0,...,nTbW-1 and y=0,...,nTbH-1, the following applies: The -(nTbW)×(nTbH) intermediate scaling factor array m is derived as follows: -If one or more of the following conditions are true, m[x][y] is set to be equal to 16: -scaling_list_enabled_flag is equal to 0. -transform_skip_flag[xTbY][yTbY] is equal to 1. -Otherwise, m is the output of the scaling matrix derivation process specified in section xxx, which is called using the prediction mode CuPredMode[xTbY][yTbY], the color component variable cIdx, the block width nTbW, and the block height nTbH as inputs. - The scaling factor ls[x][y] is derived as follows: [...]
[0115]
[0118] The main change compared to VVC Draft 5 is that instead of copying the array portion described in the scaling_list_data semantics, the xxx clause is called.
[0116]
[0119] As stated above, please note that this is just one example, and not an exhaustive one, aimed at minimizing changes to the current VVC draft. For example, the color component inputs for the scaling matrix derivation process may differ from cIdx. Also, QM is not limited to use as a scaling factor; it can also be used, for example, as a QP offset.
[0117]
[0120] The same test set used during the HEVC standardization process was used to test QM encoding performance, supplemented with QMs derived from recommended or default QMs of common standards (JPEG, MPEG2, AVC, HEVC) and QMs seen in actual broadcasts. In all tests, several QMs were copied from one type to another (e.g., luminance to chroma or intra to inter), and / or from one size to another.
[0118]
[0121] The table below reports the number of bits required to encode scaling_list_data using three different methods: HEVC, JVET-N0847, and this proposal. Specifically, HEVC uses the HEVC test set (24 QMs per test), while the other two use a derived test set (30 QMs per test, with additional sizes: size -2 for chroma, size -64 for luminance; size -2 QMs are downsampled from size -4, size -64 is copied from size -32, size -32 is copied from size -16, and sizes 16, 8, and 4 are kept unchanged).
[0119] [Table 12]
[0120]
[0122] This test shows that, even when compared to HEVC, the proposed technique saves a large number of bits, while the proposed method encodes more QM.
[0121]
[0123] Referring again to the method in JCTVC-E073, the reference indexing (triple: QMPS, size, type) in JCTVC-E073 is more complex than the one proposed here because the QMPS indexing in the former requires the memory of the QMPS. Linear interpolation introduces complexity. Downsampling is similar to the one proposed here.
[0122]
[0124] Referring again to the method in JCTVC-H0314, the transmission from larger to smaller objects is similar to what is proposed here, but the fixed prediction tree in JCTVC-H0314 is less flexible than the unified indexing and explicit criteria proposed here.
[0123]
[0125] QM in intrablock copy mode
[0126] In the above, different QMs are specified for two block prediction modes, namely intra and inter. However, in addition to intra and inter, VVC has a new prediction mode: IBC (intrablock copy), in which blocks can be predicted from reconstructed samples of the same picture using appropriate displacement vectors. In QM selection in IBC prediction mode, both JVET-N0847 and the above embodiments use the same QM as in intra mode.
[0124]
[0127] Since IBC mode is closer to inter than intra, in one embodiment it is proposed to reuse QMs signaled in inter mode (instead of intra). However, while IBC is closer to inter prediction, it differs in that the displacement vectors do not match the movement of the object or camera and are used for texture copying. This can lead to certain artifacts in different embodiments where certain QMs may be helpful in optimizing the copying of IBC blocks. Below, we propose changing the QM selection in IBC prediction mode. A preferred embodiment is to select the same QM as the intermode (instead of the intra) because the IBC is closer to the intermode prediction than the intramode prediction. Another option is to have a specific QM in IBC mode. These may be explicitly signaled in the syntax or inferred (e.g., the average of intraQM and interQM).
[0125]
[0128] In a preferred embodiment, the QM selection or derivation process for a particular transformation block selects an interQM if the block has an IBC prediction mode. Referring again to Figure 12, step 1210 of the QM derivation process needs to be adjusted as described below.
[0126]
[0129] In the previously proposed draft text, the QM selection is given in equation (xxx-1), which can be changed as follows:
number
[0127]
[0130] Blocks can be encoded in intra-mode (MODE_INTRA), inter-mode (MODE_inter), or intra-block copy mode (MODE_IBC). matrixTypeId is shown in Table 6 or matrixTypeId=(2 * When set as cIdx+(predMode==MODE_INTER?1:0))), the MODE_IBC block is like MODE _ Select matrixTypeId as if it were an INTRA block. Change in (xxx-1): matrixTypeId=(2 * Using cIdx+(predMode==MODE_INTRA?0:1)), the MODE_IBC block selects matrixTypeId as if it were a MODE_INTER block.
[0128]
[0131] In the draft text proposed by JVET-N0847, the QM selection is listed in Table 7-14 and can be modified as follows: In particular, the matrixId of MODE_IBC is assigned the same as MODE_INTER, rather than the same as MODE_INTRA, as in JVET-N0847.
[0129] [Table 13]
[0130]
[0132] Variation 1: Explicitly signaling the QM of the IBC
[0133] In this variation, specific QMs (different from intra-QMs and inter-QMs) are used for IBC blocks, and these QMs are explicitly signaled in the bitstream. This creates more QMs, requires adaptation of the scaling_list_data syntax and matrixId mapping, and has an impact on bit cost. According to this variation, the QM selection described in equation (xxx-1) can be changed as follows:
number
[0131]
[0134] In JVET-N0847, the QM selection table can be modified as follows:
[0132] [Table 14]
[0133]
[0135] Variation 2: Inferring QM in IBC mode
[0136] In this variation, specific QMs (different from intra-QMs and inter-QMs) are used for the IBC block. However, these QMs are not signaled in the bitstream and are inferred: for example, as an average of intra-QMs and inter-QMs, a specific default value, or a specific modification to the inter-QM, such as scaling and offset.
[0134]
[0137] Variation 3: Explicit IBC QM based on luminance only
[0138] In this variation, the additional QM for IBC is limited to luminance only, and the chroma QM for IBC can either reuse the interQM as in variation 1, or infer a new QM as in variation 2.
[0135]
[0139] Various methods are described herein, each method comprising one or more steps or actions to achieve the methods described above. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., e.g., “first decryption” and “second decryption,” may be used in various embodiments to modify elements, components, steps, actions, etc. The use of such terms does not imply an order to the modified actions, unless specifically required. Therefore, in this example, the first decryption does not need to be performed before the second decryption, but may, for example, be performed before the second decryption, during the second decryption, or during a time period overlapping with the second decryption.
[0136]
[0140] The various methods and other embodiments described herein can be used to modify the modules of the video encoder 200 and decoder 300 shown in Figures 2 and 3, for example, the quantization and dequantization modules (230, 240, 340). Furthermore, each embodiment is not limited to VVC or HEVC and can be applied to other standards and recommendations, for example, and any extensions of such standards and recommendations. Unless otherwise specified or technically excluded, the embodiments described herein can be used individually or in combination.
[0137]
[0141] Various numerical values are used in this application. The specific values are for illustrative purposes only, and the embodiments described are not limited to these specific values.
[0138]
[0142] Various implementations include decoding. “Decoding,” as used in this application, may encompass all or part of the processes performed on a received encoded sequence, for example, to produce a final output suitable for display. In various embodiments, such processes typically include one or more processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase “decoding process” is intended to refer to a particular subset of operations or to a broader decoding process in general will become clear from the context of the specific description and will be well understood by those skilled in the art.
[0139]
[0143] Various implementations include encoding. Similar to the above discussion of "decoding," the term "encoding" as used in this application may encompass all or part of the process performed on, for example, an input video sequence to generate an encoded bitstream.
[0140]
[0144] Note that the syntax elements used herein are descriptive terms. Therefore, the use of other syntax element names is not excluded. In the above, the syntax elements of PPS and scaling lists are mainly used to describe various embodiments. Note that these syntax elements may be placed in other syntax structures.
[0141]
[0145] The implementations and embodiments described herein may be implemented, for example, by methods or processes, apparatus, software programs, data streams, or signals. Even if a feature is discussed only in the context of one form of implementation (for example, only as a method), the implementation of the feature discussed may also be implemented in other forms (for example, apparatus or programs). Apparatus may be implemented, for example, by appropriate hardware, software, and firmware. Methods may be implemented by a processor, for example, a processor, including apparatus such as a computer, microprocessor, integrated circuit, or programmable logic device, and commonly referred to as a processing device. Processors also include communication devices, such as computers, mobile phones, portable / personal data entry devices ("PDAs"), and other devices that facilitate the communication of information between end users.
[0142]
[0146] The terms "one embodiment," "one embodiment," "one implementation," or "implementation," and references to other variations thereof, mean that the specific features, structures, characteristics, etc., described in relation to that embodiment are included in at least one embodiment. Therefore, the appearance of the phrases "in one embodiment," "in one embodiment," "in one implementation," or "in implementation," and any other variations appearing in various places throughout this application, do not necessarily all refer to the same embodiment.
[0143]
[0147] Furthermore, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.
[0144]
[0148] Furthermore, this application may refer to “access” to various types of information. Access to information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0145]
[0149] Furthermore, this application may refer to various forms of "reception" of information. Reception, like "access," is intended to be a broad term. Reception of information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "reception" usually involves, in some way, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information during operation.
[0146]
[0150] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the following " / ", "and / or", and "at least one" is intended to include the selection of only the first listed option (A), only the second listed option (B), or both options (A and B). As further examples, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to include the selection of only the first listed option (A), only the second listed option (B), only the third listed option (C), only the first and second listed options (A and B), only the first and third listed options (A and C), only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended to as many items as are listed, as will be obvious to those skilled in the art.
[0147]
[0151] Furthermore, as used herein, the term "signaling" specifically refers to something to a corresponding decoder. For example, in certain embodiments, an encoder signals a quantization matrix for inverse quantization. Thus, in embodiments, the same parameters are used on both the encoder and decoder sides. Therefore, for example, an encoder can transmit certain parameters to a decoder so that the decoder can use the same specific parameters (explicit signaling). Conversely, if the decoder already has certain parameters, signaling is used without transmission (implicit signaling), so that the decoder simply knows and can select certain parameters. Bit saving is achieved in various embodiments by avoiding the transmission of any actual function. It should be understood that signaling can be achieved in a wide variety of ways. In various embodiments, for example, one or more syntax elements, flags, etc., are used to signal information to a corresponding decoder. The above relates to the verb form of the word "signal," but the word "signal" can also be used as a noun herein.
[0148]
[0152] As will be apparent to those skilled in the art, the implementation may generate a wide variety of signals formatted to carry information that can be stored or transmitted. The information may include, for example, data generated by instructions for performing the method or by one of the implementations described. For example, a signal may be formatted to carry a bitstream of the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a wide variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.
Claims
1. Obtaining a single identifier for a quantized block based on the block size, color components, and prediction mode of the blocks to be decoded within the picture, Decoding a syntax element representing a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix. Obtaining the quantization matrix based on the aforementioned reference quantization matrix, The transformation coefficients of the block are inversely quantized in response to the quantization matrix, Decoding the block in response to the inversely quantized transformation coefficients, A method that includes this.
2. A device comprising one or more processors, wherein the one or more processors are Obtaining a single identifier for the quantization matrix based on the block size, color components, and prediction mode of the blocks to be decoded within the picture, Decoding a syntax element representing a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix. Obtaining the quantization matrix based on the aforementioned reference quantization matrix, The transformation coefficients of the block are inversely quantized in response to the quantization matrix, Decoding the block in response to the inversely quantized transformation coefficients, A device configured to perform the following actions.
3. The method according to claim 1 or the apparatus according to claim 2, wherein the size of the block is different from the size of the block to which the reference quantization matrix is applied for inverse quantization.
4. The method according to claim 1 or 3, or the apparatus according to claim 2 or 3, wherein the elements of the quantization matrix are used as scaling factors when inverse quantizing each transformation coefficient of the block.
5. The method according to any one of claims 1 to 4 or the apparatus according to any one of claims 2 to 4, wherein the elements of the quantization matrix are used as offsets when inverse quantizing each transformation coefficient of the block.
6. Accessing the encoded blocks within the picture, Accessing the quantization matrix of the aforementioned block, Obtaining a single identifier for the quantization matrix based on the block size, color components, and prediction mode of the block, Encoding a syntax element representing a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix. quantizing the transformation coefficients of the block in response to the quantization block, Entropy coding the quantized transformation coefficients, A method that includes this.
7. A device comprising one or more processors, wherein the one or more processors access the blocks to be encoded in the picture, Accessing the quantization matrix of the aforementioned block, Obtaining a single identifier for the quantization matrix based on the block size, color components, and prediction mode of the block, Encoding a syntax element representing a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix. Quantizing the transformation coefficients of the block in response to the quantization matrix, Entropy coding the quantized transformation coefficients, A device configured to perform the following actions.
8. The method according to claim 6 or the apparatus according to claim 7, wherein the size of the block is different from the size of the block to which the reference quantization matrix is applied for quantization.
9. The method according to claim 6 or 8, or the apparatus according to claim 7 or 8, wherein the elements of the quantization matrix are used as scaling factors when quantizing each transformation coefficient of the block.
10. The method according to claim 6 or 8, or the apparatus according to claim 7 or 8, wherein the elements of the quantization matrix are used as offsets when quantizing each transformation coefficient of the block.
11. The reference quantization matrix is signaled first, according to the method of any one of claims 1 and 3 to 6, or the apparatus according to any one of claims 2 to 5 and 7 to 10.
12. The quantization matrix is obtained from the reference quantization matrix through copying or decimation, according to the method of any one of claims 1, 3-6, and 8-11, or the apparatus according to any one of claims 2-5 and 7-11.
13. The method or apparatus according to claim 12, wherein the quantization matrix is obtained from the reference quantization matrix through a copy, in response that the quantization matrix has the same size as the reference quantization matrix.
14. The method according to claim 12 or the apparatus according to claim 12, wherein the quantization matrix is obtained from the reference quantization matrix by decimation by a corresponding ratio in response to the quantization matrix having a different size from the reference quantization matrix.
15. The method according to any one of claims 1, 3 to 6, and 8 to 14, where the block size is M × N, where M is the width and N is the height, and the identifier of the block is defined based on the size max(M, N), where max(M, N) is the larger of M and N. The apparatus according to any one of claims 2 to 5 and 7 to 14.
16. The method according to any one of claims 1, 3 to 6, and 8 to 15, wherein a set of quantization matrices are signaled in increasing order of identifiers, and the quantization matrix with the largest block size is signaled first, or the apparatus according to any one of claims 2 to 5 and 7 to 15.
17. The method according to claim 16 or the apparatus according to claim 16, wherein when signaling the set of quantization matrices, the quantization matrix of the luminance color component is signaled before the quantization matrix of the chroma color component.
18. The method according to any one of claims 1, 3 to 6, and 8 to 17, or the apparatus according to any one of claims 2 to 5 and 7 to 17, wherein when signaling the set of quantization matrices, a quantization matrix with a larger block size is signaled before a quantization matrix with a smaller block size.
19. The aforementioned identifier is matrixId=N * The method according to any one of claims 1, 3 to 6, and 8 to 18, or the apparatus according to any one of claims 2 to 5 and 7 to 18, wherein the matrixId is derived as sizeId + matrixTypeId, where N is the number of possible type identifiers, sizeID indicates the block size, and matrixTypeId indicates the color component and the prediction mode.
20. The method according to any one of claims 1, 3 to 6, and 8 to 19, or the apparatus according to any one of claims 2 to 5 and 7 to 19, wherein the one or more processors are further configured to perform fitting the reference quantization matrix to the block size.
21. The method according to any one of claims 1, 3 to 6, and 8 to 20, or the apparatus according to any one of claims 2 to 5 and 7 to 20, wherein the one or more processors are further configured to perform fitting the reference quantization matrix of the block to the chroma format of the block, which is different from the default chroma format.
22. The method according to claim 20 or the apparatus according to claim 20, wherein the default chroma format is 4:2:
0.
23. The conformance is based on a bit shift of the x and y coordinates in the quantization matrix to the index coefficients of the reference quantization matrix, according to the method of any one of claims 20 to 22 or the apparatus according to any one of claims 20 to 22.
24. The method according to any one of claims 1, 3 to 6, and 8 to 23, or the apparatus according to any one of claims 2 to 5 and 7 to 23, wherein the identifier is obtained based on whether the prediction mode of the block is an intra-prediction mode or an inter-prediction mode.
25. The intrablock copy prediction mode is considered as an inter-prediction mode when acquiring the identifier, according to the method of any one of claims 1, 3 to 6, and 8 to 24, or the apparatus according to any one of claims 2 to 5 and 7 to 24.
26. The method according to any one of claims 1, 3 to 6, and 8 to 25, or the apparatus according to any one of claims 2 to 5 and 7 to 25, wherein the prediction mode is an intrablock copy, the quantization matrix is signaled for the luminance component of the block, and the quantization matrix for the chroma component is derived by interpreting the prediction mode as an interprediction mode.
27. The method according to any one of claims 1, 3 to 6, and 8 to 26, or the apparatus according to any one of claims 2 to 5 and 7 to 26, wherein the prediction mode is an intrablock copy, the reference quantization matrix is obtained by considering the prediction mode as an intramode, another reference quantization matrix is obtained by considering the prediction mode as an intermode, and the quantization matrix is obtained as the average of the reference quantization matrix and the other reference quantization matrix.
28. A signal including encoded video formed by performing the method according to any one of claims 6 and 8 to 27.
29. A computer-readable storage medium storing instructions for encoding or decoding video data by any one of claims 1, 3 to 6, and 8 to 27.