Single index quantization matrix design for video encoding and decoding

By designing a single index quantization matrix in video encoding and decoding, the problem that the quantization matrix in the prior art is difficult to adapt to different block sizes, color components and prediction modes, and a more efficient video encoding and decoding effect is achieved.

CN119996695APending Publication Date: 2025-05-13INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171177.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-06-24
Filing Date
2020-06-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technologies, it is difficult to effectively adapt to different block sizes, color components and prediction modes, resulting in poor encoding efficiency and decoding quality.

Method used

A single-index quantization matrix design method is proposed, which obtains a single identifier of the quantization matrix by obtaining the block size, color components and prediction mode of the block, and decodes or encodes the reference quantization matrix.

Benefits of technology

Through this design method, it is possible to more effectively adapt to the characteristics of different video blocks, improve the efficiency and quality of video encoding and decoding, and reduce the redundancy of the bitstream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996695A_ABST
    Figure CN119996695A_ABST
Patent Text Reader

Abstract

The invention relates to a single index quantization matrix design for video encoding and decoding, in particular to different quantization matrices that can be transmitted corresponding to different block sizes, color components, and prediction modes. In order to more effectively signal the coefficients of the quantization matrix, in one implementation, a uniform matrix identifier matrixId is used based on a size identifier (sizeId) related to a first listed larger-sized CU size and a first listed matrix type (matrixTypeId) of the luma QM. For example, a unified identifier is derived as: matrixId = N * sizeId + matrixTypeId, where N is the number of possible type identifiers, e.g., N = 6. This single identifier allows any previously transmitted matrices to be referenced when prediction (replication) is used, and first transmission of the larger matrices avoids interpolation in the prediction process. When a block uses an intra block copy prediction mode, a QM identifier may be obtained as if the block uses an inter prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

This application is a divisional application of Chinese patent application No. 202080045038.4, whose filing date is June 16, 2020 and entitled “Single-index quantization matrix design for video encoding and decoding”, and the contents of the Chinese patent application are cited herein in their entirety. Technical Field

[0001] This embodiment mainly relates to a quantization matrix design method and device in video encoding or decoding. Background Art

[0002] In order to achieve high compression efficiency, image and video coding schemes usually use prediction and transformation to exploit spatial and temporal redundancy in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame picture correlation, and then the difference between the original block and the predicted block, which is usually represented as a prediction error or prediction residual, is transformed, quantized, and entropy decoded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to entropy decoding, quantization, transformation, and prediction. Summary of the invention

[0003] According to one embodiment, a video decoding method is provided, comprising: obtaining a single identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a block to be decoded in a picture; decoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtaining the quantization matrix based on the reference quantization matrix; dequantizing transform coefficients of the block in response to the quantization matrix; and decoding the block in response to the dequantized transform coefficients.

[0004] According to another embodiment, a method for video encoding is provided, comprising: accessing a block to be encoded in a picture; accessing a quantization matrix of the block; obtaining a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; encoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantizing transform coefficients of the block in response to the quantization matrix; and entropy encoding the quantized transform coefficients.

[0005] According to another embodiment, a device for video decoding is provided, which includes one or more processors, wherein the one or more processors are configured to: obtain a single identifier for a quantization matrix based on a block size, color component, and prediction mode of a block to be decoded in a picture; decode a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; obtain the quantization matrix based on the reference quantization matrix; dequantize transform coefficients of the block in response to the quantization matrix; and decode the block in response to the dequantized transform coefficients.

[0006] According to another embodiment, a device for video encoding is provided, which includes one or more processors, wherein the one or more processors are configured to: access a block to be encoded in a picture; access a quantization matrix of the block; obtain a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; encode a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; quantize the transform coefficients of the block in response to the quantization matrix; and entropy encode the quantized transform coefficients.

[0007] According to another embodiment, a video decoding device is provided, comprising: a device for obtaining a single identifier of a quantization matrix based on a block size, a color component, and a prediction mode of a block to be decoded in a picture; a device for decoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; a device for obtaining the quantization matrix based on the reference quantization matrix; a device for dequantizing transform coefficients of the block in response to the quantization matrix; and a device for decoding the block in response to the dequantized transform coefficients.

[0008] According to another embodiment, a video encoding device is provided, comprising: a device for accessing a block to be encoded in a picture; a device for accessing a quantization matrix of the block; a device for obtaining a single identifier of the quantization matrix based on the block size, color components and prediction mode of the block; a device for encoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies the difference between the identifier of the reference quantization matrix and the obtained identifier of the quantization matrix; a device for quantizing the transform coefficients of the block in response to the quantization matrix; and a device for entropy encoding the quantized transform coefficients.

[0009] One or more embodiments also provide a computer program including instructions, which, when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any of the above embodiments. One or more embodiments of the present invention also provide a computer-readable storage medium on which instructions for encoding or decoding video data according to the method described above are stored. One or more embodiments also provide a computer-readable storage medium on which a bit stream generated according to the above method is stored. One or more embodiments also provide a method and apparatus for transmitting or receiving a bit stream generated according to the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A block diagram of a system is shown in which aspects of the present embodiments may be implemented.

[0011] Figure 2 A block diagram of an embodiment of a video encoder is shown.

[0012] Figure 3 A block diagram of an embodiment of a video decoder is shown.

[0013] Figure 4 It is shown that transform coefficients are inferred to zero for block sizes greater than 32 in VVC draft 5.

[0014] Figure 5 The fixed prediction tree described in JCTVC-H0314 is shown.

[0015] Figure 6 Prediction (decimation) from a larger size is shown according to an embodiment.

[0016] Figure 7 Combining prediction from a larger size and decimation of rectangular blocks according to an embodiment is shown.

[0017] Figure 8 The QM derivation process for rectangular blocks of chrominance according to an embodiment is shown.

[0018] Fig. 9 The QM derivation process for rectangular blocks of chroma (adapted to 4:2:2 format) according to an embodiment is shown.

[0019] Fig.10 The QM derivation process for rectangular blocks of chroma (adapted to 4:4:4 format) is shown.

[0020] Fig.11 A flow chart for parsing a zoom list data syntax structure according to an embodiment is shown.

[0021] Fig.12A flow chart for encoding a zoom list data syntax structure according to an embodiment is shown.

[0022] Fig.13 A flow chart for a QM derivation process according to an embodiment is shown. DETAILED DESCRIPTION

[0023] Figure 1 A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be implemented as a device including various components described below, and is configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed on multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or other electronic devices via, for example, a communication bus or by a dedicated input and / or output port communication. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0024] The system 100 includes at least one processor 110, which is configured to execute instructions loaded therein, for implementing various aspects described in the present application, for example. The processor 110 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include a non-volatile memory and / or a volatile memory, including but not limited to an EEPROM, a ROM, a PROM, a RAM, a DRAM, an SRAM, a flash memory, a disk drive, and / or an optical drive. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device.

[0025] The system 100 includes an encoder / decoder module 130, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated into the processor 110 as a combination of hardware and software as known to those skilled in the art.

[0026] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. These stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0027] In several embodiments, memory inside the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be a memory 120 and / or a storage device 140, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations such as MPEG-2, HEVC, or VVC.

[0028] Input to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0029] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band limiting the signal to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its relevant input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the frequency band of expectation again by filtering, down-conversion and perform frequency selection.Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of similar or different functions.Adding element can include inserting element between existing element, for example, inserting amplifier and analog-to-digital converter.In various embodiments, the RF part comprises antenna.

[0030] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, in a separate input processing IC or processor 110. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either in a separate interface IC or in processor 110. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.

[0031] The various components of the system 100 may be disposed within an integrated housing, and the various components may be interconnected and transmit data therebetween using a suitable connection arrangement 115 such as an internal bus as known in the art including an I2C bus, wiring, and printed circuit boards.

[0032] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0033] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 100 using a set-top box that passes data through an HDMI connection of the input box 105. Still other embodiments provide streaming data to the system 100 using an RF connection of the input box 105.

[0034] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripherals 185. In examples of various embodiments, the other peripherals 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 100. In various embodiments, control signals are transmitted between the system 100 and the display 165, the speaker 175, or other peripherals 185 using signaling such as AV. Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 100 via dedicated connections through the respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to the system 100 via the communication interface 150 using the communication channel 190. The display 165 and the speaker 175 can be integrated into a single unit in an electronic device (e.g., a television) along with other components of the system 100. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0035] For example, if the RF portion of input 105 is part of a separate set-top box, the display 165 and speaker 175 may alternatively be separate from one or more of the other components. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0036] Figure 2An example video encoder 200 is shown, such as a High Efficiency Video Coding (HEVC) encoder. Figure 2 Also shown are encoders in which improvements are made to the HEVC standard or encoders that employ techniques similar to HEVC, such as the VVC (Versatile Video Coding) encoder under development by the JVCT (Joint Video Discovery Team).

[0037] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "encoding" or "decoding" can be used interchangeably, and the terms "image", "picture" and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0038] Before being encoded, the video sequence may undergo a pre-encoding process (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.

[0039] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is divided (202) and processed in units such as CUs. Each unit is encoded using, for example, intra or inter mode. When the unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of intra mode or inter mode to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, a prediction residual is calculated by subtracting (210) the predicted block from the original image block.

[0040] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass the transform and quantization, i.e., directly decode the residual without applying a transform or quantization process.

[0041] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0042] Figure 3 A block diagram of an example video decoder 300 is shown. In the decoder 300, a bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs the same Figure 2 The encoder 200 typically also performs video decoding as part of encoding the video data, as well as a decoding pass that is the inverse of the encoding pass described in .

[0043] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other decoding information. The picture segmentation information indicates how the picture is segmented. The decoder can therefore divide (335) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined (355) with the prediction block to reconstruct the image block. The prediction block may be obtained (370) from intra-frame prediction (360) or motion compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0044] The decoded image may be further subjected to post-decoding processing (385), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing an inverse remapping of the remapping process performed in the pre-encoding process (201). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0045] The HEVC specification allows the use of a quantization matrix in the dequantization process, where the transform coefficients are scaled by the current quantization step size and further scaled by the quantization matrix (QM), as follows: d[x][y]=Clip3(coeffMin,coeffMax,((TransCoeffLevel[xTbY][yTbY][cIdx][x][y]*m[x][y]*levelScale[qP%6]<<(qP / 6))+(1<<(bdShift-1)))>>bdShift) in: ●TransCoeffflevel[…] is the absolute value of the transform coefficient of the current block identified by the spatial coordinates xTbY, yTbY and its component index cIdx of the current block. ●x and y are the horizontal / vertical frequency indices. ●qP is the current quantization parameter. ●Multiplying by levelScale[qP%6] and shifting left by (qP / 6) is equivalent to multiplying by the quantization step size qStep= (levelScale[qP%6]<<(qP / 6)). ●m[…][…] is a two-dimensional quantization matrix. Here, since the quantization matrix is ​​used for scaling, it can also be called a scaling matrix. bdShift is an additional scaling factor to take into account the image sampling bit depth. The term 1<<(bdShift-1)) is used for rounding purposes to the nearest integer. ●d[…] is the resulting absolute value of the dequantized transform coefficient.

[0046] The syntax used by HEVC to transmit the quantization matrix is ​​described as follows:

[0047] It can be noticed • Specify a different matrix for each transform size (sizeId). In the scaling list data syntax structure, the scaling matrices are scanned into a 1-D scaling list (eg, ScalingList). • For a given transform size, six matrices are specified for intra / inter coding and Y / Cb / Cr components. ●The matrix can be any of the following o If scaling_list_pred_mode_flag is zero (reference matrixId is obtained as matrixId - scaling_list_pred_matrix_id_delta), the matrix is ​​copied from a previously transmitted matrix of the same size ○ Copy the matrix according to the default values ​​specified in the standard (if scaling_list_pred_mode_flag and scaling_list_pred_matrix_id_delta are both zero) o Use exp-Golomb entropy coding, fully specified in DPCM coding mode in top-right diagonal scanning order. • For block sizes larger than 8x8, only 8x8 coefficients are transmitted for signaling the quantization matrix in order to save decoding bits. The coefficients are then interpolated using zero-holding (ie, repetition), except for the DC coefficient which is explicitly transmitted.

[0048] In VVC draft 5, the use of quantization matrices similar to HEVC has been adopted based on the contribution of JVET-N0847 (see O. Chubach et al., "CE7-related: Support of quantization matrices for VVC", JVET-N0847, Geneva, CH, March 2019). The scaling_list_data syntax has been adapted to the VVC codec as shown below.

[0049] In the design of VVC draft 5 adopting JVET-N0847, as in HEVC, the QM is identified by two parameters matrixId and sizeId. This is shown in the following two tables. Table 1: Block size identifiers (JVET-N0847) brightness Chroma sizeId - - 0 - 2x2 1 4x4 4x4 2 8x8 8x8 3 16x16 16x16 4 32x32 32x32 5 64x64 - 6 Table 2: QM Type Identifiers (JVET-N0847) CuPredMode cIdx (color component) matrixId MODE_INTRA 0(Y) 0 MODE_INTRA 1(Cb) 1 MODE_INTRA 2(Cr) 2 MODE_INTER 0(Y) 3 MODE_INTER 1(Cb) 4 MODE_INTER 2(Cr) 5 Note: MODE_INTRA QM is also used for MODE_IBC (Intra Block Copy).

[0050] The combinations of the two identifiers are shown in the following table: Table 3: (MatrixId, SizeId) combinations (JVET-N0847)

[0051] As in HEVC, for block sizes larger than 8×8, only the 8×8 coefficients and the DC coefficient are transmitted. Zero-preserving interpolation is used to reconstruct the QM of the correct size. For example, for a 16×16 block, each coefficient is repeated twice in both directions, and then the DC coefficient is replaced by the transmitted coefficient.

[0052] For rectangular blocks, the reserved size (sizeId) selected for the QM is the larger size, i.e., the maximum of the width and height. For example, for a 4×16 block, a QM of 16×16 block size is selected. The reconstructed 16×16 matrix is ​​then vertically decimated by a factor of 4 to obtain the final 4×16 quantization matrix (i.e., 3 of the 4 rows are skipped).

[0053] For the following, we refer to the QM for a given family of block sizes (square or rectangular) as size-N, which is related to the sizeId and the square block size it is used for. For example, for block sizes 16x16 or 16x4, the QM is identified as size-16 (sizeId 4 in VVC draft 5). The notation of size-N is used to distinguish the exact block shape, as well as to distinguish the number of QM coefficients that are signaled (restricted to 8×8, as shown in Table 3).

[0054] In addition, in VVC draft 5, for the case of size -64, the QM coefficients of the lower right quadrant are not transmitted (they are inferred to be 0, referred to as "zero output" below). This is achieved by the "x>=4&y>=4" condition in the scaling_list_data syntax. This avoids transmitting QM coefficients that are never used by the transform / quantization process. In fact, in VVC, for transform block sizes greater than 32 in any dimension (64xN, Nx64, where N<=64), any transform coefficient with an x / y frequency coordinate greater than or equal to 32 is not transmitted and is inferred to be zero, so no quantization matrix coefficient is required to quantize it. This is in Figure 4 , where the shaded areas correspond to transform coefficients that are inferred to be zero.

[0055] Compared to HEVC, VVC requires more quantization matrices due to the larger number of block sizes. However, in VVC draft 5, QM prediction is still limited to the replication of matrices of the same block size, which may lead to bit waste. In addition, since only block size 2x2 is used for chroma and only block size 64x64 is used for luma in VVC, the syntax associated with QMs is more complex. In addition, as in HEVC, JVET-N0847 describes a specific matrix derivation process for each block size.

[0056] During HEVC standardization, some QM prediction techniques have been explored, for example, in JCTVC-E073 (see J. Tanaka et al., “Quantization Matrix for HEVC”, JCTVC-E073, Geneva, CH, March 2011) and JCTVC-H0314 (see Y. Wang et al., “Layered quantization matrices representation and compression”, JCTVC-H0314, San José, CA, USA, February 2012).

[0057] JCTVC-E073: QMs are transmitted in a specific parameter set (QMPS). Within the QMPS, QMs are transmitted in increasing size order (sizeId / matrixId similar to HEVC). Prediction (= copying) of QM coefficients from any previously decoded, including the previous QMPS, was proposed. Upconversion with linear interpolation was used to accommodate smaller reference QMs, while simple downsampling was used to accommodate larger reference QMs. This was ultimately rejected during HEVC standardization.

[0058] JCTVC-H0314: QM is transmitted in descending order. Figure 5 The fixed prediction tree shown (without explicit reference indexes) copies the previously sent QM instead of transmitting a new QM, using simple downsampling if the reference QM is larger. This was ultimately rejected during HEVC standardization.

[0059] Both proposals involve HEVC and cannot cope with the complexity introduced by VVC.

[0060] This application proposes to simplify the quantization matrix signaling and prediction process of VVC draft 5 (after adopting JVET-N0847) by combining one or more of the following, while enhancing them so that any QM can be predicted from any previously signaled QM: - Unified QM indexes to include size and type so that reference index differences can resolve any prior The QM index of the previous transmission; -Transmit quantization matrices in descending order of block size; -Specify the prediction process as a replication or extraction process as needed; - All QM coefficients of size-64 are transmitted so that size-64 QM can be used as predictor.

[0061] Furthermore, a QM derivation process involving upsampling of blocks larger than 8x8 and downsampling of rectangular blocks is described to select the QM index according to the block parameters and to adapt the QM signaled size to the actual block size.

[0062] For ease of presentation, we refer to the process of predicting a quantization matrix based on default values ​​or based on previously sent values ​​as the QM prediction process, and the process of adapting the sent or predicted QM to the size and chroma format of the transform block as the QM derivation process. The QM prediction process can be part of the process of parsing scaling list data, for example, at the picture level. The derivation process is typically at a lower level, for example at the transform block level. Each aspect is presented in further detail below, followed by draft text examples and performance results. • Derivation and use of a single matrix index to identify the QM. A single identifier allows referencing any previously transmitted matrix when using prediction (copying), and transmitting larger matrices first to avoid interpolation during prediction. • The QM prediction process consists in copying or extracting the previously signaled QM (reference QM), which is the transmitted, predicted or default reference QM. The QM derivation process, which for a given transform block consists in selecting a QM index based on the block size, color components and prediction mode, and then adapting the size of the selected QM to the size of the block. The resizing process is based on bit shifting of the x and y coordinates within the transform block to index the coefficients of the selected QM. • Transmission of all coefficients for size -64QM.

[0063] These aspects simplify the specification compared to VVC draft 5 (text changes are halved compared to JVET-N0847) and bring significant bit savings (the bit cost of scaling_list_data can be halved).

[0064] Unified QM Index

[0065] The QM used for quantization / dequantization of a transform block is identified by a single parameter matrixId. In one embodiment, the unified matrixId (QM index) is a combination of the following items: - with CU size (i.e. CU encompasses a square shape since only square size matrices are transmitted) Rather than a block size related size identifier. Note that here, for luma or chroma, the size identifier is controlled by the luma block size (e.g., max(luma block width, luma block height)). When luma and chroma trees are separated, for chroma, "CU size" will refer to the size of the block projected on the luma plane. - Matrix types for luma QMs are listed first, since they can be larger than chroma (e.g. in case of 4:2:0 chroma format)

[0066] According to this embodiment, the QM index derivation is shown in Tables 4 and 5 and equation (1). Table 4: Size identifiers (proposed) brightness Chroma sizeId 64x64 32x32 0 32x32 16x16 1 16x16 8x8 2 8x8 4x4 3 4x4 2x2 4 Table 5: Matrix type identifiers (proposed) CuPredMode cIdx (color component) matrixTypeId MODE_INTRA 0(Y) 0 MODE_INTER 0(Y) 1 MODE_INTRA 1(Cb) 2 MODE_INTER 1(Cb) 3 MODE_INTRA 2(Cr) 4 MODE_INTER 2(Cr) 5

[0067] The unified matrixId is derived as follows: MatrixId=N*sizeId+MatrixTypeId (1) Where N is the number of possible type identifiers, for example N=6.

[0068] In another embodiment, if more than six QM types are defined, the sizeId should be multiplied by the correct number, which is the number of quantization matrix types. In other embodiments, other parameters may also be different, such as a specific block size, a signaled matrix size (here limited to 8×8), or the presence of a DC coefficient. Note here that the QMs are listed by decreasing block size and are identified by a single index, as shown in Table 6. Table 6: Unified matrixId (proposed)

[0069] QM Prediction Process

[0070] Instead of transmitting the QM coefficients, the QM may be predicted based on default values ​​or based on any previously transmitted values. In one embodiment, when the reference QM is of the same size, the QM is copied, otherwise it is decimated by a relevant ratio, such as Figure 6 This is shown in the example in which a size-4 brightness QM is predicted from a size-8.

[0071] The extraction is described by the following equation: ScalingMatrix[matrixId][x][y]=refScalingMatrix[i][j] (2) Where matrixSize = (matrixId<20)?8:(matrixId<26)?4:2) x=0…matrixSize-1, y=0…matrixSize-1, i=x<<(log2(refMatrixSize)-log2(matrixSize)), and j=y<<(log2(refMatrixSize)-log2(matrixSize)). where refMatrixSize matches the size of refScalingMatrix (and therefore the range of the i and j variables).

[0072] exist Figure 6In the example shown, a luma size-4QM (which is a 4×4 array: matrixSize is 4) is predicted based on a luma size-8QM, which is an 8×8 array (refMatrixSize is 8); one of the two rows and one of the two columns are discarded to generate a 4×4 array (i.e., the element (2x, 2y) in the reference QM is copied to the element (x, y) in the current QM).

[0073] Equation (3) takes the following form: ScalingMatrix[matrix Id][x][y]=refScalingMatrix[i][j] (3) x=0...3, y=0...3, i=x<<1, j=y<<1.

[0074] When the reference QM has a DC value, it is copied as the DC value if the current QM requires a DC value; otherwise, it is copied to the upper-left QM coefficient.

[0075] In a preferred embodiment, this QM prediction process is part of the QM decoding process, but in another embodiment it may be deferred to the QM derivation process, where decimation for prediction purposes is merged with the QM resizing sub-process.

[0076] QM derivation process

[0077] The proposed derivation process for the quantization matrix first selects the correct QM index according to the block parameters (unified QM index) as described before, and then unifies the processes for decimation of rectangular blocks, repetition for blocks larger than, for example, 8×8 size, and chroma format adaptation into a single process. The proposed process is based on bit shifting of the x and y output coordinates. In order to select the right row / column of the selected QM, it is only necessary to left-shift and then right-shift the x / y output coordinates, as shown in the following equations. m[x][y]=ScalingMatrix[matrixId][i][j] (4) where i=(x<<log2MatrixSize)> >log2(blkwidth), and j=(y<<log2MatrixSize)> >log2(blkHeight). where log2MatrixSize is the log2 of the size of ScalingMatrix[matrixId] (which is a square 2D array), blkWidth and blkHeight are the width and height of the current transform block, respectively, with x ranging from 0 to blkWidth-1 and y ranging from 0 to blkHeight-1.

[0078] The following are several examples to illustrate the QM derivation process. Figure 7 In the example shown in , the QM for the luma 16×8 block is derived from the luma size-16QM, which is actually an 8×8 array plus a DC coefficient. For this example, blkWidth is equal to 16, blkHeight is equal to 8, and log2MatrixSize is equal to 3, so equation (5) takes the following form: m[x][y]=ScalingMatrix[matrixId][i][j] (5) Where i=(x<<3)>>4, and j=(y<<3)>>3, where x=0...15 and y=0...7. Here, x is shifted right by 1, and y is unchanged (i.e., column i in the selected QM is copied to columns 2*i and 2*i+1 in the current QM). In addition, since the selected QM has a DC coefficient, it is copied to m[0][0].

[0079] In such Figure 8 In another example illustrated in , a QM for the chroma 4x2 blocks of an 8x4 CU (4:2:0 format) is generated. This matches the 8x4 CU size, whose enclosing square is 8x8. Therefore, the selected QM is a size-8 QM, where the chroma QM is decoded as a 4x4 array. Here, blkWidth is equal to 4, blkHeight is equal to 2, and log2MatrixSize is equal to 2, so equation (6) takes the following form: m[x][y]=ScalingMatrix[matrixId][i][j] (6) Where i=(x<<2)>>2, and j=(y<<2)>>1, where x=0...3 and y=0...1. Here, x is unchanged and y is shifted left by 1 (ie, row 2y in the reference QM is copied to row y in the current QM).

[0080] In the following example, the proposed adaptation to 4:2:2 and 4:4:4 formats differs from VVC draft 5. Instead of finding a QM that matches the chroma block size (except for 64x64 where there is no chroma matrix), the size matching is based on the same (luminance) CU size (i.e. the size of the block projected on the luminance plane), and coefficients are repeated if necessary. This makes the QM design independent of the chroma format.

[0081] exist Fig. 9 In the example shown in , a QM for a chroma 8x2 block of an 8x4 CU (4:2:2 format) is generated. The selected QM is similar to Figure 8The above example shown in is the same, but the 4:2:2 chroma format requires twice as many columns. Here, the columns are repeated, so x is shifted right by 1 and y is still shifted left by 1. In particular, blkWidth is equal to 8, blkHeight is equal to 2, and log2MatrixSize is equal to 2, so equation (7) takes the following form: m[x][y]=ScalingMatrix[matrix Id][i][j] (7) Where i=(x<<2)>>3, and j=(y<<2)>>1, where x=0...7 and y=0...1.

[0082] exist Fig.10 In the example shown in , a QM for the chroma 8x4 block of an 8x4 CU (4:4:4 format) is generated. The selected QM is still the same as Figure 8 and 9 , but the 4:4:4 chroma format requires twice as many rows and columns as the 4:2:0 chroma format. Here, the columns must be repeated, so x is right-shifted by 1, but the decimation of the rows can be skipped (due to the rectangular shape), so y is not shifted. In particular, blkWidth is equal to 8, blkHeight is equal to 4, and log2MatrixSize is equal to 2, so equation (8) takes the following form: m[x][y]=ScalingMatrix[matrixId][i][j] (8) Where i=(x<<2)>>3, and j=(y<<2)>>2, where x=0...7 and y=0...3.

[0083] The number of coefficients transmitted for size-64

[0084] In one embodiment, to enable prediction from smaller QMs of size-64, all coefficients of size-64 are transmitted in the scaling list syntax, even though the lower right quadrant is never used by the VVC transform and quantization process. In general, we can send all coefficients of the maximum QM.

[0085] However, it is worth noting that compared to previous work (JVET-N0847), the overhead associated with this increased number of transmitted coefficients can be limited to 2×16 bits in the worst case, since the syntax element scaling_list_delta_coef can be set to zero for the lower right quadrant of Size-64QM when it is not used as a predictor: for the lower right quadrant, 4×4=16 delta coefficients are signaled, each costing 1 bit if forced to zero (using exp-Golomb decoding), and there are two Size-64QMs (luminance intra / intermediate).

[0086] The tests described in Table 8 show that this overhead is negligible compared to the gain from improved prediction.

[0087] In another embodiment, the coefficients of the lower right quadrant of the Size-64 QMs are not transmitted as part of the signaling of the Size-64 QM, but are transmitted as supplementary parameters when a smaller QM is first predicted from a given Size-64 QM.

[0088] Table 7 provides some comparisons between the method described in JVET-N0847 and the proposed method: Table 7

[0089] Below, some syntax and semantics according to an embodiment are described.

[0090] PPS syntax and semantics (minor adjustments) pps_scaling_list_data_present_flag equal to 1 specifies that the scaling list data for pictures referencing the PPS are derived based on the scaling lists specified by the active SPS and the scaling lists specified by the PPS. pps_scaling_list_data_present_flag equal to 0 specifies that the scaling list data for pictures referencing the PPS are inferred to be equal to those specified by the active SPS. When scaling_list_enabled_flag is equal to 0, the value of pps_scaling_list_data_present_flag shall be equal to 0. When scaling_list_enabled_flag is equal to 1, sps_scaling_list_data_present_flag is equal to 0 and pps_scaling_list_data_present_flag is equal to 0, the default scaling matrix is ​​used to derive the array ScalingMatrix as described in the Scaling List Data Semantics specified in Section 7.4.5.

[0091] Please note that this syntax / semantics is an example close to the HEVC standard or VVC draft, and is not restrictive. For example, scaling_list_data transport is not limited to SPS or PPS and can be sent by other means.

[0092] Scaling List Data Syntax / Semantics (Simplified) scaling_list_pred_mode_flag[matrixId] equal to 0 specifies that the scaling matrix is ​​derived from the value of the reference scaling matrix. The reference scaling matrix is ​​specified by scaling_list_pred_matrix_id_delta[matrixId]. scaling_list_pred_mode_flag[matrixId] equal to 1 specifies that the value of the scaling list is explicitly signaled. scaling_list_pred_matrix_id_delta[matrixId] specifies the reference scaling matrix used to derive the scaling matrix, as described below. The value of scaling_list_pred_matrix_id_delta[matrixId] should be in the range of 0 to matrixId, inclusive. When scaling_list_pred_mode_flag[matrixId] is equal to zero: - The variable refMatrixSize and the array refScalingMatrix are first derived as follows: ○ If scaling_list_pred_matrix_id_delta[matrixId] is equal to zero, then the following applies to setting the default value: ■ refMatrixSize is set equal to 8, ■If matrixId is an even number, ■Otherwise ○ Otherwise (if scaling_list_pred_matrix_id_delta[matrixId] is greater than zero), the following applies: refMatrixId=matrixId-scaling_list_pred_matrix_id_delta[matrixId](11) refMatrixSize=(refMatrixId<20)? 8:(refMatrixId<26)? 4:2) (12) refScalingMatrix=ScalingMatrix[refMatrixId] (13) -Then derive the array ScalingMatrix[matrixId] as follows: ScalingMatrix[matrixId][x][y]=refScalingMatrix[i][j] (14) Where matrixSize = (matrixId<20)?8:(matrixId<26)?4:2) x=0…matrixSize-1, y=0…matrixSize-1, i=x<<(log2(refMatrixSize)-log2(matrixSize)), and j=y<<(log2(refMatrixSize)-log2(matrixSize)) scaling_list_dc_coef_minus8[matrixId] plus 8 specifies the first value of the scaling matrix when relevant, as described in clause xxx. The value of scaling_list_dc_coef_minus8[matrixId] shall be in the range -7 to 247, inclusive. When scaling_list_pred_mode_flag[matrixId] is equal to zero, scaling_list_pred_matrix_id_delta[matrixId] is greater than zero, and refMatrixId < 14, the following applies: - if matrixId < 14, scaling_list_dc_coef_minus8[matrixId] is inferred to be equal to scaling_list_dc_coef_minus8[refMatrixId], Otherwise, ScalingMatrix[matrixId][0][0] is set equal to scaling_list_dc_coef_minus8[refMatrixId]+8 When scaling_list_pred_mode_flag[matrixId] is equal to zero, scaling_list_pred_matrix_id_delta[matrixId] is equal to zero (indicating a default value), and matrixId < 14, scaling_list_dc_coef_minus8[matrixId] is inferred to be equal to 8 scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[matrixId][i] and the previous matrix coefficient ScalingList[matrixId][i-1] when scaling_list_pred_mode_flag[matrixId] is equal to 1. The value of scaling_list_delta_coef should be in the range of -128 to 127, inclusive. The value of ScalingList[matrixId][i] should be greater than 0. When present (i.e. scaling_list_pred_mode_flag[matrixId] is equal to 1), the array ScalingMatrix[matrix_Id] is derived as follows: ScalingMatrix[matrixId][i][j]=ScalingList[matrixId][k] (15) Where k = 0…coefNum-1, i = diagScanOrder[log2(coefNum) / 2][log2(coefNum) / 2][k][0], and j=diagScanOrder[log2(coefNum) / 2][log2(coefNum) / 2][k][1]

[0093] The main simplifications compared to the JVET-N0847 syntax are the removal of a for() loop and the simplification of indexing from [sizeId][matrixId] to [matrixId].

[0094] “Clause xxx” refers to an undetermined section number to be introduced in the VVC specification, which matches the scaling matrix derivation process in this paper.

[0095] Note that this syntax / semantics is an example close to the HEVC standard or VVC Draft 5, and is not restrictive. For example, the coefficient range is not limited to 1...255, it can be 1...127 (7 bits) or -64...63, for example. Moreover, it is not limited to 30 QMs, organized as 6 types × 5 sizes (there can be 8 types, and fewer or more sizes. See Table 5, which can be modified. Then a simple modification of the coefNum and the conditions for the existence of the DC coefficient will be required). The type of QM prediction (here just a copy) is also not restrictive. For example, scaling factors or offsets can be added, and explicit decoding can be added as a residual on top of the prediction. The same situation represents the method used for coefficient transmission (here DPCM), the presence of a DC coefficient, and a fixed number of coefficients (only a subset can be transmitted).

[0096] Regarding the default values, they are not limited to the two default values ​​QM associated with MODE_INTRA and MODE_INTER, and are similarly filled here with all 16 values ​​until the relevant default values ​​are reached (eg, the same default value QM as HEVC may be selected).

[0097] Likewise, the number of coefficients coefNum to be signaled can be expressed mathematically instead of comparing sequences, with the same result: coefNum=Min(64,4096>>((matrixId+4) / 6)*2), which is closer to HEVC or the current VVC draft approach, but introduces a potentially undesirable split.

[0098] Note here that the larger matrix is ​​transmitted first, and a single index is used so that the prediction reference (indicated by scaling_list_pred_matrix_id_delta) can be any previously transmitted matrix, regardless of its intended block size or type, or the default value (e.g., if scaling_list_pred_matrix_id_delta is zero).

[0099] Fig.11A process (1100) for parsing a scaling list data syntax structure according to an embodiment is shown. For this embodiment, the input is a decoded bitstream and the output is a ScalingMatrix array. For clarity, details about DC values ​​are omitted. Specifically, in step 1110, a QM prediction mode is decoded from the bitstream. If the QM is predicted (1120), the decoder further determines whether to infer (predict) the QM in the bitstream or to signal the QM based on the aforementioned flag. In step 1130, the decoder decodes QM prediction data from the bitstream, which is required to infer the QM when there is no signal notification, such as the QM index difference scaling_list_pred_matrix_id_delta. The decoder then determines (1140) whether to predict the QM based on a default value (e.g., if scaling_list_pred_matrix_id_delta is zero) or based on a previously decoded QM. If the reference QM is a default QM, the decoder selects (1150) the default QM as the reference QM. There may be several default QMs to choose from, for example depending on the parity of matrixId. Otherwise, the decoder selects (1155) a previously decoded QM as reference QM. The index of the reference QM is derived from the matrix and the above index difference. In step 1160, the decoder predicts a QM based on the reference QM. If the reference QM is the same size as the current QM, the prediction consists of a simple copy, or if it is larger than expected, it consists of decimation. The result is stored in ScalingMatrix[matrixId].

[0100] If QM is not predicted (1120), the decoder determines (1170) the number of QM coefficients to be decoded from the bitstream based on matrixId. For example: 64 if matrixId is less than 20, 16 if matrixId is between 20 and 25, otherwise 4. In step 1175, the decoder decodes the relevant number of QM coefficients from the bitstream. In step 1180, the decoder organizes the decoded QM coefficients in a 2D matrix based on a scan order, such as a diagonal scan. The result is stored in ScalingMatrix[matrixId]. Using ScalingMatrix[matrixId], the decoder can use the QM derivation process to obtain the quantization matrix m[][] for dequantizing transform blocks that may be in a non-square shape and / or in a different chroma format.

[0101] At step 1190, the decoder checks whether the current QM is the last QM to be parsed. If not, control returns to step 1110; otherwise, when all QMs are parsed from the bitstream, the QM parsing process stops.

[0102] Fig.12 A process (1200) for encoding a scaling list data syntax structure on the encoder side according to an embodiment is shown. On the encoder side, the QMs are scanned from larger to smaller block sizes (e.g., matrixIds from 0 to 30) in the order described. In step 1210, the encoder searches for prediction preferences to determine whether the current QM is a copy (or extraction) of a previously encoded QM. The QMs can be designed in a way that optimizes QM prediction efficiency, for example, if a QM is initially close enough, or close to a default value QM, they can be forced to be equal (or extracted if the sizes are different). In addition, the coefficients in the lower right quadrant of the size-64 QM can be optimized to better predict subsequent QMs, or to reduce the QM bit cost if they are never reused for prediction. Once decided, the QM prediction mode is encoded in the bitstream.

[0103] Specifically, if the encoder decides (1220) to use prediction, then at step 1230, the prediction mode is encoded (e.g., scaling_list_pred_mode_flag=0). At step 1240, the prediction parameters (e.g., QM index difference scaling_list_pred_matrix_id_delta) are encoded: zero index difference for the default QM value, or the relevant index difference if the previous QM is selected as the prediction reference. On the other hand, if explicit signaling is decided, then at step 1250, the prediction mode is encoded (e.g., scaling_list_pred_mode_flag=0). Then, a diagonal scan is performed (1260), followed by QM coefficient encoding (1270).

[0104] At step 1280, the encoder checks whether the current QM is the last QM to be encoded. If not, control returns to step 1210; otherwise, the QM encoding process stops when all QMs are encoded into the bitstream.

[0105] Fig.13A process 1300 for QM derivation according to an embodiment is shown. The input may include a ScalingMatrix array and transform block parameters, such as size (width / height), prediction mode (intra / inter / IBC, ...), and color components (Y / U / V). The output is a QM with the same size as the transform block. For clarity, details about DC values ​​are omitted. Specifically, in step 1310, the decoder determines the QM index matrixId based on the current transform block size (width / height), prediction mode (intra / inter / IBC, ...), color component (Y / U / V) (unified QM index) as described above. In step 1320, the decoder adjusts the size of the selected QM (ScalingMatrix[matrixId]) to match the transform block size, as described above. In a variant, step 1320 may include the selection required for prediction.

[0106] The QM derivation process is similar on the encoder side. Quantization divides the transform coefficients by the QM value, while dequantization multiplies them by the QM value. But the QMs are the same. In particular, the QMs used for reconstruction in the encoder will match those signaled in the bitstream.

[0107] Conceptually, the transform coefficients d[x][y] can be quantized as follows, where qStep is the quantization step size and m[][] is the quantization matrix: TransCoeffLevel[xTbY][yTbY][cIdx][x][y]=d[x][y] / qStep / m[x][y]

[0108] But for integer calculations and to avoid division, it usually looks like: TransCoeffLevel[xTbY][yTbY][cIdx][x][y]=((d[x][y]*im[x][y]*ilevelScale[qP%6]>>(qP / 6))+(1<<(bdShift-1)))>>bdShift) where, for example, im[x][y]~=65536 / m[x][y] and ilevelScale[0...5]=65536 / levelScale[0...5], and bdShift is an appropriate value. In practice, for software encoders, im*ilevelScale is usually pre-calculated and stored in a table.

[0109] In the above, the QM prediction process and the QM derivation process are performed separately. In another embodiment, the QM prediction can be postponed to the QM derivation process. This embodiment does not change the QM signaling syntax. The present embodiment, which is functionally different due to continuous resizing, postpones the prediction part (obtaining reference QM+copy / reduction) and diagonal scanning to the "QM derivation process" described later.

[0110] In one embodiment, the output of the scaling list data parsing process 1100 is an array of ScalingList() instead of an array of ScalingMatrix, along with prediction flags and valid prediction parameters: The ScalingMatrixPredId array always contains the index of the defined ScalingList (default or signaled). During QM decoding, this array is recursively constructed by interpreting scaling_list_pred_matrix_id_delta, so that the QM derivation process can directly use this index to obtain the actual value to construct the QM for dequantizing the current transform block.

[0111] In the following, an example is provided to illustrate the zoom list semantics according to an embodiment.

[0112] Scaling matrix derivation procedure (new: partially replaces the description in Scaling list semantics; this is section XXX) The inputs to this process are the prediction mode predMode, the color component variable cIdx, the block width blkWidth and the block height blkHeight. The output of this process is a (blkWidth) x (blkHeight) array m[x][y] (scaling matrix), where x and y are the horizontal and vertical coefficient positions. Note that SubWidthC and SubHeightC depend on the chroma format and indicate the ratio of the number of samples in the luma component and the chroma components. The variable matrixId is derived as follows: matrixId=6*sizeId+matrixTypeId(xxx-1)where, subWidth=(cIdx>0)? SubWidthC:1, subHeight=(cIdx>0)? SubHeightC:1, sizeId=6-max(log2(blkWidth*subWidth),log2(blkHeight*subHeight)), and matrixTypeId=(2*cIdx+(predMode==MODE_INTER?1:0)) The variable log2MatrixSize is derived as follows: log2MatrixSize=(matrixId<20)?3:(matrixId<26)?2:1(xxx-2) For an x ​​range from 0 to blkWidth-1, inclusive, and a y range from 0 to blkHeight-1, inclusive, the output array m[x][y] is derived by applying the following formula: m[x][y]=ScalingMatrix[matrixId][i][j](xxx-3) where i=(x<<log2MatrixSize)> >log2(blkWidth), and j=(y<<log2MatrixSize)> >log2(blkHeight) If matrixId is lower than 14, m[0][0] is further modified as follows: m[0][0]=scaling_list_dc_coef_minus8[matrixId]+8 (xxx-4)

[0113] Note that this is an example, not a limitation, for the scaling_list_data syntax and semantics. For example, it is not limited to the two default QMs associated with MODE_INTRA and MODE_INTER, and here it is similarly filled with all 16 values ​​until the relevant default value is reached (for example, the same default QM as HEVC can be selected). There can be a single default QM or more than 2 default QMs. For example, if there are more or less than six types of matrices per block size, the MatrixId calculation can be different. Importantly, horizontal and vertical reduction and enlargement to accommodate block sizes different from the selected QM are completed in a single process, preferably simple (left shift followed by right shift here).

[0114] For rectangular blocks, we are not restricted to choosing a QM identifier for the current block that encloses a square: the derivation of sizeId in equation xxx-1 may follow different rules.

[0115] Likewise, the chosen QM size log2MatrixSize can be expressed mathematically rather than as a series of comparisons, where the result is the same: log2MatrixSize = min(3, 6 - (MatrixId + 4) / 6), but this introduces a division that may be undesirable.

[0116] In the following, examples are provided describing semantics for a zoom process according to an embodiment.

[0117] Scaling of transform coefficients (adaptation) […] For the derived scaled transform coefficients d[x][y], where x=0...nTbW-1, y=0...nTbH-1, the following applies: The intermediate scaling factor array m of -(nTbW)x(nTbH) is derived as follows: - m[x][y] is set equal to 16 if one or more of the following conditions are true: -scaling_list_enabled_flag is equal to 0. -transform_skip_flag[xTbY][yTbY] is equal to 1. Otherwise, m is the output of the scaling matrix derivation process as specified in clause xxx, which is called with the prediction mode CuPredMode[xTbY][yTbY], the color component variables cIdx, the block width nTbW and the block height nTbH as input. The scaling factor ls[x][y] is derived as follows: […]

[0118] The main change compared to VVC draft 5 is the invocation of clause xxx instead of copying part of the array described in the scaling_list_data semantics.

[0119] Note that, as mentioned above, this is an example intended to minimize changes to the current VVC draft, and is not intended to be limiting. For example, the color component inputs used for the scaling matrix derivation process may be different from cIdx. Furthermore, QM is not limited to use as a scaling factor, but may be used as a QP offset, for example.

[0120] The QM decoding performance was tested using the same test set as used during HEVC standardization, augmented with a set of QMs derived from recommended or default QMs of public standards (JPEG, MPEG2, AVC, HEVC) and QMs found in actual broadcasts. In all tests, some QMs were copied from one type to another (e.g., luma to chroma, or intra to inter), and / or one size to another.

[0121] The following table reports the number of bits required to encode scaling_list_data using the following three different methods: HEVC, JVET-N0847, and the proposed. Specifically, HEVC uses the HEVC test set (24 QMs per test), and the other two use the derived test set (30 QMs per test, with additional sizes: size-2 for chroma and size-64 for luma; size-2 QM is downsampled from size-4, size-64 is copied from size-32, size-32 is copied from size-16, and size-16, 8, 4 remain as is). Table 8

[0122] For this test, it can be seen that the proposed technique saves a significant amount of bits, even compared to HEVC, while the proposed method encodes more QMs.

[0123] Returning to the approach in JCTVC-E073, the reference index (triplet: QMPS, size, type) in JCTVC-E073 is more complex than what is proposed here, because the previous QMPS index requires storing the previous QMPS. Linear interpolation introduces complexity. Downsampling is similar to what is proposed here.

[0124] Returning to the approach in JCTVC-H0314, the large-to-small transfer is close to what is proposed here, however the fixed prediction tree in JCTVC-H0314 is less flexible than the unified index and explicit reference proposed here.

[0125] QM for intra block copy mode

[0126] In the above, different QMs are specified for the two block prediction modes, namely intra and inter. However, in addition to intra and inter, there is a new prediction mode in VVC: IBC (Intra Block Copy), where blocks can be predicted from reconstructed samples of the same picture using appropriate displacement vectors. For QM selection in IBC prediction mode, JVET-N0847 and the above embodiments use the same QMs as intra mode.

[0127] Since the IBC mode is closer to inter than intra prediction, in one embodiment, it is proposed to reuse the signaled QM for inter mode (instead of intra). However, although close to inter prediction, IBC is different: the displacement vectors do not match the object or camera motion, but are used for texture replication. This may lead to specific artifacts, where a specific QM can help optimize the decoding of IBC blocks in different embodiments. Below, we propose changing the QM selection for the IBC prediction mode: • The preferred embodiment is to choose the same QM as inter mode (rather than intra), since IBC is closer to inter prediction than intra prediction. • Another option is to have a specific QM for IBC mode. o These can be explicitly signaled in the syntax, or inferred (eg average of intra and inter QMs).

[0128] In a preferred embodiment, the QM selection or derivation process for a particular transform block selects the inter QM if the block has IBC prediction mode. Fig.12 , the QM derivation process of step 1210 needs to be adjusted as described below.

[0129] In the draft text presented above, the QM choice is described in equation (xxx-1), which can be changed as follows: matrixId=6*sizeId+matrixTypeId(xxx-1) Where subWidth=(cIdx>0)? SubWidthC:1, subHeight=(cIdx>0)? SubHeightC:1, sizeId=6-max(log2(blkWidth*subWidth),log2(blkHeight*subHeight)), and matrixTypeId=(2*cIdx+(predMode==MODE_INTRA?0:1))

[0130] Note that a block can be encoded in intra mode (MODE_INTRA), inter mode (MODE_INTER), or intra block copy mode (MODE_IBC). When matrixTypeId is set as in Table 6, or (matrixTypeId = (2*cIdx + (predMode == MODE_INTER? 1:0))), the MODE_IBC block selects matrixTypeId as if it were a MODE_INTRA block. Where the change in (xxx-1): matrixTypeId = (2*cIdx + (predMode == MODE_INTRA? 0:1)), the MODE_IBC block selects matrixTypeId as if it were a MODE_INTER block.

[0131] In the draft text proposed by JVET-N0847, the QM selection is described in Table 7-14, which can be changed as follows: In particular, the matrixId of MODE_IBC is assigned in the same way as MODE_INTER, rather than the same as MODE_INTRA as in JVET-N0847. Table 9: JVET-N0847 specification changes to matrixId according to sizeId, prediction mode and color components

[0132] Variant 1: Explicit signaling of QM for IBC

[0133] In this variant, specific QMs are used for IBC blocks (different from intra and inter QMs), and these QMs are explicitly signaled in the bitstream. This results in more QMs, which requires adaptation of scaling_list_data syntax and matrixId mapping, and has bit cost impact. According to this variant, the QM selection described in equation (xxx-1) can be changed as follows matrixId=6*sizeId+matrixTypeId(xxx-1) Where subWidth=(cIdx>0)? SubWidthC:1, subHeight=(cIdx>0)? SubHeightC:1, sizeId=6-max(log2(blkWidth*subWidth),log2(blkHeight*subHeight)), and matrixTypeId=(3*cIdx+(predMode==MODE_INTRA?0:(predMode==MODE_ INTER? 1:2)))

[0134] In JVET-N0847, the QM selection table can be changed as follows: Table 10: Changes to JVET-N0847 to add more QMs for IBC sizeId CuPredMode cIdx (color component) matrixId 2,3,4,5,6 MODE_INTRA 0(Y) 0 1,2,3,4,5 MODE_INTRA 1(Cb) 1 1,2,3,4,5 MODE_INTRA 2(Cr) 2 2,3,4,5,6 MODE_INTER 0(Y) 3 1,2,3,4,5 MODE_INTER 1(Cb) 4 1,2,3,4,5 MODE_INTER 2(Cr) 5 2,3,4,5,6 MODE_IBC 0(Y) 06 1,2,3,4,5 MODE_IBC 1(Cb) 1 7 1,2,3,4,5 MODE_IBC 2(Cr) 2 8

[0135] Variant 2: Inferred QM for IBC mode

[0136] In this variant, specific QMs are used for IBC blocks (different from intra and inter QMs). However, these QMs are not signaled in the bitstream, but are inferred, for example, as an average of intra and inter QMs, or a specific default value, or a specific variation of the inter QM, such as scaling and offset.

[0137] Variant 3: Explicit IBC QM for Luminance Only

[0138] In this variant, the additional QM for IBC is limited to luma; the chroma QM for IBC can reuse the inter-frame QM as in variant 1, or infer a new chroma QM as in variant 2.

[0139] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. In addition, terms such as "first", "second" and the like may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding" in various embodiments. Unless specifically required, the use of these terms does not mean the sequencing of the modified operation. Therefore, in this example, the first decoding does not need to be performed before the second decoding, and may occur, for example, before, during, or in a time period overlapping with the second decoding.

[0140] Various methods and other aspects described in this application can be used to modify modules, such as Figure 2 and Figure 3 The quantization and inverse quantization modules (230, 240, 340) of the video encoder 200 and decoder 300 shown, in addition, aspects of the present invention are not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, and any extensions of such standards and recommendations. Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.

[0141] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0142] Various implementations involve decoding. "Decoding" as used in this application may include, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be clear based on the context of the specific description and is believed to be fully understood by those skilled in the art.

[0143] Various implementations involve encoding.In a manner similar to the discussion above regarding "decoding," "encoding" as used in this application may include, for example, all or part of a process performed on an input video sequence to produce an encoded bitstream.

[0144] Note that the syntax elements used here are descriptive terms. Therefore, they do not exclude the use of other syntax element names. In the above, the syntax elements for PPS and scaling lists are mainly used to illustrate various embodiments. It should be noted that these syntax elements can be placed in other syntax structures.

[0145] The implementations and aspects described herein can be implemented in, for example, methods or processes, devices, software programs, data streams, or signals. Even if only discussed in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., devices or programs). For example, the device can be implemented with appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0146] References to "one embodiment" or "an embodiment" or "an implementation" or "implementation" and other variations mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations in various places in this application are not necessarily all referring to the same embodiment.

[0147] Additionally, the present application may refer to “determining” various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0148] Furthermore, the present application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0149] Additionally, the present application may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is often involved in one way or another.

[0150] It should be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to include selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to multiple items listed, as will be apparent to one of ordinary skill in this and related arts.

[0151] In addition, as used herein, the word "signal / signaling" refers in particular to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a quantization matrix for dequantization. In this way, in an embodiment, the same parameters are used on the encoder side and the decoder side. Therefore, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various embodiments by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing relates to the verb form of the word "signal / signaling", the word "signal / signaling" can also be used as a noun in this article.

[0152] As will be apparent to one of ordinary skill in the art, implementations may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

Claims

1. A method comprising: obtaining a single identifier for a quantization matrix based on a block size, a color component, and a prediction mode of a block to be decoded in the picture, wherein when the single identifier is obtained, an intra block copy prediction mode is considered to be an inter prediction mode; decoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the single identifier of the quantization matrix; Based on the reference quantization matrix, obtaining the quantization matrix; Dequantizing transform coefficients of the block based on the quantization matrix; and The block is decoded based on the dequantized transform coefficients.

2. The method according to claim 1, wherein: The single identifier is obtained based on whether the prediction mode of the block is an intra prediction mode or an inter prediction mode.

3. The method of claim 1, signaling a quantization matrix for a luma component of the block, and deriving a quantization matrix for chroma components by treating the prediction mode as an inter prediction mode.

4. The method according to claim 1, wherein: Another reference quantization matrix is ​​obtained by considering the prediction mode as an intra mode, and wherein the quantization matrix is ​​obtained as an average of the reference quantization matrix and the another reference quantization matrix.

5. An apparatus comprising one or more processors, wherein the one or more processors are configured to: Based on the block size, color components and prediction mode of the block to be decoded in the picture, a single identifier of the quantization matrix is ​​obtained, wherein When the single identifier is obtained, the intra block copy prediction mode is regarded as the inter prediction mode; decoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the single identifier of the quantization matrix; Based on the reference quantization matrix, obtaining the quantization matrix; Dequantizing transform coefficients of the block based on the quantization matrix; as well as The block is decoded based on the dequantized transform coefficients.

6. The device according to claim 5, wherein: The single identifier is obtained based on whether the prediction mode of the block is an intra prediction mode or an inter prediction mode.

7. The device of claim 5, signaling a quantization matrix for a luma component of the block, and deriving a quantization matrix for chroma components by treating the prediction mode as an inter prediction mode.

8. The device according to claim 5, wherein: Another reference quantization matrix is ​​obtained by considering the prediction mode as an intra mode, and wherein the quantization matrix is ​​obtained as an average of the reference quantization matrix and the another reference quantization matrix.

9. A method comprising: accessing a quantization matrix for a block to be encoded in a picture; obtaining a single identifier for the quantization matrix based on a block size, a color component, and a prediction mode of the block, wherein when the single identifier is obtained, an intra block copy prediction mode is considered to be an inter prediction mode; encoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the single identifier of the quantization matrix; quantizing transform coefficients of the block based on the quantization matrix; and The quantized transform coefficients are entropy encoded.

10. The method according to claim 9, wherein: The single identifier is obtained based on whether the prediction mode of the block is an intra prediction mode or an inter prediction mode.

11. The method of claim 9, signaling a quantization matrix for a luma component of the block, and deriving a quantization matrix for chroma components by treating the prediction mode as an inter prediction mode.

12. The method according to claim 9, wherein: Another reference quantization matrix is ​​obtained by considering the prediction mode as an intra mode, and wherein the quantization matrix is ​​obtained as an average of the reference quantization matrix and the another reference quantization matrix.

13. An apparatus comprising one or more processors, wherein the one or more processors are configured to: accessing a quantization matrix for a block to be encoded in a picture; A single identifier of the quantization matrix is ​​obtained based on a block size, a color component and a prediction mode of the block, wherein: When the single identifier is obtained, the intra block copy prediction mode is regarded as the inter prediction mode; encoding a syntax element indicating a reference quantization matrix, wherein the syntax element specifies a difference between an identifier of the reference quantization matrix and the single identifier of the quantization matrix; quantizing transform coefficients of the block based on the quantization matrix; as well as The quantized transform coefficients are entropy encoded. 14 . The device of claim 13 , wherein the single identifier is obtained based on whether the prediction mode of the block is an intra-prediction mode or an inter-prediction mode.

15. The device of claim 13, signaling a quantization matrix for a luma component of the block, and deriving a quantization matrix for chroma components by treating the prediction mode as an inter prediction mode.

16. The device according to claim 13, wherein: Another reference quantization matrix is ​​obtained by considering the prediction mode as an intra mode, and wherein the quantization matrix is ​​obtained as an average of the reference quantization matrix and the another reference quantization matrix.