Use of Default Scaling Matrix and User-Defined Scaling Matrix
By dynamically selecting and applying the scaling matrix and transformation matrix in the conversion between the video block and the codec representation, the problem of inefficient video encoding and codec in the prior art is solved, and more efficient video compression and codec performance is achieved.
Patent Information
- Application Number
- CN202080059117.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-20
- Filing Date
- 2020-08-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-08-20
AI Technical Summary
When existing video encoding and decoding technologies work hard to effectively utilize the scaling matrix and transformation matrix when processing video blocks, resulting in ineffective encoding and decoding.
A video processing method is proposed, by dynamically selecting and applying the scaling matrix and the transformation matrix based on the codec mode and the transformation mode in the conversion between the video block and the codec representation. Specifically, the method includes: determining the transform skip mode according to the codec mode, selecting an appropriate loop filter or a post-reconstruction filter; and selecting and applying a scaling matrix based on the size and codec mode of the video block.
By dynamically selecting and applying the scaling matrix and transformation matrix, the efficiency of video encoding and decoding is improved, bandwidth usage is reduced, and video compression performance is optimized.
Smart Images

Figure CN114342398B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is the national - phase entry application in China of international patent application PCT / CN2020 / 110275 filed on August 20, 2020, which claims the priority of international patent application No. PCT / CN2019 / 101555 filed on August 20, 2019. The entire disclosure of the above applications is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background art
[0004] Despite the progress in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth required for digital video usage is expected to continue to grow. Summary of the invention
[0005] Disclosed are devices, systems, and methods related to digital video encoding and decoding, particularly related to video encoding and decoding and decoding using a scaling matrix and / or a transformation matrix.
[0006] In one exemplary aspect, a video processing method is disclosed. The method includes: performing a conversion between a video block of a video and an encoded - decoded representation of the video, where the encoded - decoded representation conforms to format rules, where the format rules specify the applicability of a transform skip mode to the video block based on the encoding conditions of the video block, where the format rules specify omitting a syntax element indicating the applicability of the transform skip mode from the encoded - decoded representation, and where the transform skip mode includes skipping applying a forward transform to at least some coefficients before encoding into the encoded - decoded representation, or skipping applying an inverse transform to at least some coefficients before decoding from the encoded - decoded representation during decoding.
[0007] In another exemplary aspect, a video processing method is disclosed. The method includes: determining whether to use a loop filter or a post - reconstruction filter for the conversion between two adjacent video blocks of a video and the encoded - decoded representation of the video according to whether a forward transform or an inverse transform is used for the conversion, where the forward transform includes skipping applying a forward transform to at least some coefficients before encoding into the encoded - decoded representation, or skipping applying an inverse transform to at least some coefficients before decoding from the encoded - decoded representation during decoding; and performing the conversion based on the use of the loop filter or the post - reconstruction filter.
[0008] In another exemplary aspect, a video processing method is disclosed. The method includes: determining a factor of a scaling tool based on a coding / decoding mode of a video block for conversion between the video block of a video and a coded / decoded representation of the video; and performing the conversion using the scaling tool, wherein use of the scaling tool includes: scaling at least some coefficients representing the video block during encoding or descaling at least some coefficients from the coded / decoded representation during decoding.
[0009] In another exemplary aspect, a video processing method is disclosed. The method includes: determining that use of a scaling tool is prohibited for conversion between a video block of a video and a coded / decoded representation of the video due to a block differential pulse coding modulation (BDPCM) coding / decoding tool or a quantization residual BDPCM (QR-BDPCM) coding / decoding tool for the conversion of the video block; and performing the conversion without using the scaling tool, wherein use of the scaling tool includes: scaling at least some coefficients representing the video block during encoding or descaling at least some coefficients from the coded / decoded representation during decoding.
[0010] In another exemplary aspect, a video processing method is disclosed. The method includes: selecting a scaling matrix based on a transform matrix selected for the conversion for conversion between a video block of a video and a coded / decoded representation of the video, wherein the scaling matrix is used to scale at least some coefficients of the video block and wherein the transform matrix is used to transform at least some coefficients of the video block during the conversion; and performing the conversion using the scaling matrix.
[0011] In another exemplary aspect, a video processing method is disclosed. The method includes: determining whether to apply a scaling matrix based on a rule based on whether a quadratic transform matrix is applied to a portion of a video block of a video, wherein the scaling matrix is used to scale at least some coefficients of the video block and wherein the quadratic transform matrix is used to transform at least some residual coefficients of the portion of the video block during the conversion; and performing conversion between the video block of the video and a bitstream representation of the video using the selected scaling matrix.
[0012] In another exemplary aspect, a video processing method is disclosed. The method includes: for a video block having a non-square shape, determining a scaling matrix used in conversion between the video block of the video and a coded / decoded representation of the video, wherein a syntax element in the coded / decoded representation signals the scaling matrix and wherein the scaling matrix is used to scale at least some coefficients of the video block during the conversion; and performing the conversion based on the scaling matrix.
[0013] In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video block of a video and an encoded / decoded representation of the video, wherein, based on a rule, the video block includes a first number of positions at which a scaling matrix is applied during the conversion, and the video block further includes a second number of positions at which the scaling matrix is not applied during the conversion.
[0014] In another exemplary aspect, a video processing method is disclosed. The method includes: determining a scaling matrix to be applied during a conversion between a video block of a video and an encoded / decoded representation of the video; and performing the conversion based on the scaling matrix, wherein the encoded / decoded representation indicates a number of elements in the scaling matrix, and wherein the number depends on whether coefficient zeroing is applied to coefficients of the video block.
[0015] In another exemplary aspect, a video processing method is disclosed. The method includes: performing a conversion between a video block of a video and an encoded / decoded representation of the video according to a rule, wherein, after applying a K×L transform matrix to transform coefficients of the video block and after zeroing all transform coefficients except for upper left M×N transform coefficients, the video block is represented by the encoded / decoded representation, wherein the encoded / decoded representation is configured to exclude signaling of elements of a scaling matrix at positions corresponding to the zeroed ones, wherein the scaling matrix is used to scale the transform coefficients.
[0016] In another exemplary aspect, a video processing method is disclosed. The method includes, during a conversion between a video block of a video and an encoded / decoded representation of the video, based on a rule determining whether a single quantization matrix will be used based on the size of the video block, wherein all video blocks having the size use the single quantization matrix; and performing the conversion using the quantization matrix.
[0017] In another exemplary aspect, a video processing method is disclosed. The method includes: determining, based on encoded / decoding mode information, whether to enable a transform skip mode for a conversion between an encoded / decoded representation of a video block of a video and the video block; and performing the conversion based on the determination, wherein, in the transform skip mode, application of a transform to at least some coefficients representing the video block is skipped during the conversion.
[0018] In another exemplary aspect, another video processing method is disclosed. The method includes: determining to use a scaling matrix for a conversion due to using a block differential pulse codec modulation (BDPCM) or a quantized residual BDPCM (QR-BDPCM) mode for the conversion between an encoded / decoded representation of a video block and the video block; and performing the conversion using the scaling matrix, wherein the scaling matrix is used to scale at least some coefficients representing the video block during the conversion.
[0019] In another exemplary aspect, another video processing method is disclosed. The method includes: prohibiting the use of a scaling matrix for the transformation due to the use of block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR-BDPCM) mode for the transformation between the codec representation of a video block and the video block; and performing the transformation using the scaling matrix, where the scaling matrix is used to scale at least some of the coefficients representing the video block during the transformation.
[0020] In another exemplary aspect, another video processing method is disclosed. The method includes: determining the applicability of a loop filter for the transformation between the codec representation of a video block of a video and the video block according to whether the transform skip mode is enabled for the transformation; and performing the transformation based on the applicability of the loop filter, where in the transform skip mode, the application of the transformation to at least some of the coefficients representing the video block is skipped during the transformation.
[0021] In another exemplary aspect, another video processing method is disclosed. The method includes: selecting a scaling matrix for the transformation between a video block of a video and the codec representation of the video block such that the same scaling matrix is used for the transformations based on inter-frame codec and intra-block copy codec; and performing the transformation using the selected scaling matrix, where the scaling matrix is used to scale at least some of the coefficients of the video block.
[0022] In another exemplary aspect, another video processing method is disclosed. The method includes: selecting a scaling matrix for the transformation based on the transformation matrix selected for the transformation between a video block of a video and the codec representation of the video block; and performing the transformation using the selected scaling matrix, where the scaling matrix is used to scale at least some of the coefficients of the video block, and where the transformation matrix is used to transform at least some of the coefficients of the video block during the transformation.
[0023] In another exemplary aspect, another video processing method is disclosed. The method includes: selecting a scaling matrix for the transformation based on the second-order transformation matrix selected for the transformation between a video block of a video and the codec representation of the video block; and performing the transformation using the selected scaling matrix, where the scaling matrix is used to scale at least some of the coefficients of the video block, and where the second-order transformation matrix is used to transform at least some of the residual coefficients of the video block during the transformation.
[0024] In another exemplary aspect, another video processing method is disclosed. The method includes: for a video block having a non-square shape, determining a scaling matrix used in the conversion between the video block and its codec representation, wherein a syntax element in the codec representation signals the scaling matrix; and performing the conversion based on the scaling matrix, wherein the scaling matrix is used to scale at least some coefficients of the video block during the conversion.
[0025] In another exemplary aspect, another video processing method is disclosed. The method includes: determining a scaling matrix to be partially applied during the conversion between the codec representation of a video block and the video block; and performing the conversion by partially applying the scaling matrix such that the scaling matrix is applied at a first set of positions of the video block and disabled at the remaining positions of the video block.
[0026] In another exemplary aspect, another video processing method is disclosed. The method includes: determining a scaling matrix to be applied during the conversion between the codec representation of a video block and the video block; and performing the conversion based on the scaling matrix, wherein the codec representation signals the number of elements of the scaling matrix, and wherein the number depends on the application of coefficient zeroing in the conversion.
[0027] In another exemplary aspect, another video processing method is disclosed. The method includes: during the conversion between a video block and its codec representation, determining a single quantization matrix to be used based on the size of a specific type of video block; and performing the conversion using the quantization matrix.
[0028] In yet another representative aspect, the above method is embodied in the form of processor-executable code and stored in a computer-readable program medium.
[0029] In yet another representative aspect, a device is disclosed that is configured to or can be used to perform the above method. The device may include a processor programmed to implement such method.
[0030] In yet another exemplary aspect, a video decoder device may implement the methods described herein.
[0031] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a block diagram of an exemplary video encoder implementation.
[0033] Figure 2 shows an example of a secondary transformation.
[0034] Figure 3 Shows an example of a reduced quadratic transform (RST).
[0035] Figure 4 Illustrates two scalar quantizers used in the proposed dependency quantization scheme.
[0036] Figure 5 Shows an example of the state transition and quantizer selection of the proposed dependency quantization.
[0037] Figures 6A - 6B Shows an example of a diagonal scan order.
[0038] Figure 7 Shows an example of the selected positions for QM signaling notification (32x32 transform size).
[0039] Figure 8 Shows an example of the selected positions for QM signaling notification (64x64 transform size).
[0040] Figure 9 Shows an example of applying zeroing to coefficients.
[0041] Figure 10 Shows an example of signaling only the selected elements within a dotted region (e.g., MxN region).
[0042] Figure 11 Is a block diagram of an example of a video processing hardware platform.
[0043] Figure 12 Is a flowchart of an exemplary method of video processing.
[0044] Figure 13 Is a block diagram showing an example of a video decoder.
[0045] Figure 14 Is a block diagram showing an exemplary video processing system in which various techniques disclosed herein can be implemented.
[0046] Figure 15 Is a block diagram showing an exemplary video codec system that can utilize the techniques of the present disclosure.
[0047] Figure 16 Is a block diagram showing an example of a video encoder.
[0048] Figures 17 - 27 Is a flowchart of an exemplary method of video processing. Detailed Description
[0049] Embodiments of the disclosed technology can be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve compression performance. In this document, section headings are used to enhance readability of the description, and do not in any way limit the discussion or embodiments (and / or implementations) to their respective sections.
[0050] 1. Overview
[0051] This document relates to image / video coding technology. Specifically, this document relates to residual coding in image / video coding. It can be applied to existing video coding standards such as HEVC, or standards under consideration (Versatile Video Coding). It can also be applicable to future video coding standards or video codecs.
[0052] 2. Background
[0053] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, in which temporal prediction plus transform coding is employed. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), which is dedicated to researching the VVC standard targeting a 50% bitrate reduction compared to HEVC.
[0054] The latest version of the VVC draft, i.e., Versatile Video Coding (Draft 5), can be found at the following URL:
[0055] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=6640
[0056] The latest reference software for VVC, called VTM, can be found at the following URL:
[0057] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-4.02.1
[0058] 2.1. Color Space and Chrominance Subsampling
[0059] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually three or four values or color components (e.g., RGB). Basically, a color space is an exposition of a coordinate system and a subspace.
[0060] For video compression, the most frequently used color spaces are YCbCr and RGB.
[0061] YCbCr, Y′CbCr or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a family of color spaces that are used as part of the color image pipeline in video and digital photography systems. Y′ is the luminance component, and CB and CR are the blue-difference and red-difference chrominance components. Y′ (preferred) is different from Y as luminance, which means non-linearly encoding the light intensity based on the gamma-corrected RGB primaries.
[0062] Chrominance subsampling is the convention of encoding an image by applying a lower resolution to chrominance information than to luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 2.1.1. 4:4:4
[0064] Each of the three Y'CbCr components has the same sampling rate, so there is no chrominance subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2
[0066] The two chrominance components are sampled at half the sampling rate of luminance: halving the horizontal chrominance resolution. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference. 2.1.3. 4:2:0
[0068] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but since only the Cb and Cr channels are sampled interlaced in this scheme, the vertical resolution is halved. Thus, the data rate is the same. Each of Cb and Cr is subsampled by one-half in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme with different horizontal and vertical siting.
[0069] · In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are addressed between pixels vertically (interpolation addressing).
[0070] · In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are addressed by interpolation halfway between every other luminance sample.
[0071] · In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternate lines.
[0072] 2.2. Coding and Decoding Processes of Typical Video Codecs
[0073] Figure 1 An example of an encoder block diagram of VVC is shown, which contains three loop filter blocks: the deblocking filter (DF), sample adaptive offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets and applying finite impulse response (FIR) filters respectively through signaling the coding side information of the offsets and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts established in the previous stage.
[0074] 2.3. Quantization Matrix
[0075] The well-known spatial frequency sensitivity of the human visual system (HVS) has been a key driving factor behind many aspects of the design of modern image and video coding algorithms and standards, including JPEG, MPEG2, H.264 / AVC High Profile, and HEVC.
[0076] The quantization matrix used in MPEG2 is an 8x8 matrix. In H.264 / AVC, the quantization matrix block sizes include both 4x4 and 8x8. These QMs are coded into the SPS (Sequence Parameter Set) and PPS (Picture Parameter Set). The compression method for QM signaling in H.264 / AVC is differential pulse coding modulation (DPCM).
[0077] In H.264 / AVC High Profile, 4x4 block size and 8x8 block size are used. There are six QMs for the 4x4 block size (i.e., separate matrices for intra / inter coding and Y / Cb / Cr components) and two QMs for the 8x8 block size (i.e., separate matrices for intra / inter Y component), so only eight quantization matrices need to be coded into the bitstream.
[0078] 2.4. Transform and Quantization Design in VVC
[0079] 2.4.1. Transform
[0080] HEVC specifies two-dimensional transforms of various sizes from 4×4 to 32×32, which are finite-precision approximations relative to the discrete cosine transform (DCT). In addition, HEVC also specifies an alternative 4×4 integer transform based on the discrete sine transform (DST) for use in conjunction with 4×4 intra-prediction residual blocks for luminance. In addition, transform skip may be allowed in some block size cases.
[0081] For nS = 4, 8, 16, and 32 and DCT-II, the transform matrix cij (i, j = 0..nS-1) is defined as follows:
[0082] nS = 4
[0083] {64,64,64,64}
[0084] {83,36,-36,-83}
[0085] {64,-64,-64,64}
[0086] {36,-83,83,-36}
[0087] nS = 8
[0088] {64,64,64,64,64,64,64,64}
[0089] {89,75,50,18,-18,-50,-75,-89}
[0090] {83,36,-36,-83,-83,-36,36,83}
[0091] {75,-18,-89,-50,50,89,18,-75}
[0092] {64,-64,-64,64,64,-64,-64,64}
[0093] {50,-89,18,75,-75,-18,89,-50}
[0094] {36,-83,83,-36,-36,83,-83,36}
[0095] {18,-50,75,-89,89,-75,50,-18}
[0096] nS = 16
[0097] {64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64}
[0098] {90 87 80 70 57 43 25 9 -9-25-43-57-70-80-87-90}
[0099] {89 75 50 18-18-50-75-89-89-75-50-18 18 50 75 89}
[0100] {87 57 9-43-80-90-70-25 25 70 90 80 43-9-57-87}
[0101] {83 36-36-83-83-36 36 83 83 36-36-83-83-36 36 83}
[0102] {80 9-70-87-25 57 90 43-43-90-57 25 87 70 -9-80}
[0103] {75-18-89-50 50 89 18-75-75 18 89 50-50-89-18 75}
[0104] {70-43-87 9 90 25-80-57 57 80-25-90 -9 87 43-70}
[0105] {64-64-64 64 64-64-64 64 64-64-64 64 64-64-64 64}
[0106] {57-80-25 90-9-87 43 70-70-43 87 9-90 25 80-57}
[0107] {50-89 18 75-75-18 89-50-50 89-18-75 75 18-89 50}
[0108] {43-90 57 25-87 70 9-80 80-9-70 87-25-57 90-43}
[0109] {36-83 83-36-36 83-83 36 36-83 83-36-36 83-83 36}
[0110] {25 - 70 90 - 80 43 9 - 57 87 - 87 57 - 9 - 43 80 - 90 70 - 25}
[0111] {18 - 50 75 - 89 89 - 75 50 - 18 - 18 50 - 75 89 - 89 75 - 50 18}
[0112] {9 - 25 43 - 57 70 - 80 87 - 90 90 - 87 80 - 70 57 - 43 25 - 9}
[0113] nS = 32
[0114] {64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 64 6464 64 64 64 64 64 64 64 64}
[0115] {90 90 88 85 82 78 73 67 61 54 46 38 31 22 13 4 -4 -13 -22 -31 -38 -46 -54 -61 -67 -73 -78 -82 -85 -88 -90 -90}
[0116] {90 87 80 70 57 43 25 9 -9 -25 -43 -57 -70 -80 -87 -90 -90 -87 -80 -70 -57 -43 -25 -9 9 25 43 57 70 80 87 90}
[0117] {90 82 67 46 22 -4 -31 -54 -73 -85 -90 -88 -78 -61 -38 -13 13 38 61 78 88 90 8573 54 31 4 -22 -46 -67 -82 -90}
[0118] {89 75 50 18 -18 -50 -75 -89 -89 -75 -50 -18 18 50 75 89 89 75 50 18 -18 -50 -75 -89 -89 -75 -50 -18 18 50 75 89}
[0119] {88 67 31 -13 -54 -82 -90 -78 -46 -4 38 73 90 85 61 22 -22 -61 -85 -90 -73 -38 446 78 90 82 54 13 -31 -67 -88}
[0120] {87 57 9-43-80-90-70-25 25 70 90 80 43 -9-57-87-87-57 -9 43 80 90 7025-25-70-90-80-43 9 57 87}
[0121] {85 46-13-67-90-73-22 38 82 88 54 -4-61-90-78-31 31 78 90 61 4-54-88-82-38 22 73 90 67 13-46-85}
[0122] {83 36-36-83-83-36 36 83 83 36-36-83-83-36 36 83 83 36-36-83-83-36 3683 83 36-36-83-83-36 36 83}
[0123] {82 22-54-90-61 13 78 85 31-46-90-67 4 73 88 38-38-88-73 -4 67 90 46-31-85-78-13 61 90 54-22-82}
[0124] {80 9-70-87-25 57 90 43-43-90-57 25 87 70 -9-80-80 -9 70 87 25-57-90-43 43 90 57-25-87-70 9 80}
[0125] {78-4-82-73 13 85 67-22-88-61 31 90 54-38-90-46 46 90 38-54-90-31 6188 22-67-85-13 73 82 4-78}
[0126] {75-18-89-50 50 89 18-75-75 18 89 50-50-89-18 75 75-18-89-50 50 8918-75-75 18 89 50-50-89-18 75}
[0127] {73-31-90-22 78 67-38-90-13 82 61-46-88 -4 85 54-54-85 4 88 46-61-8213 90 38-67-78 22 90 31-73}
[0128] {70-43-87 9 90 25-80-57 57 80-25-90 -9 87 43-70-70 43 87 -9-90-25 8057-57-80 25 90 9-87-43 70}
[0129] {67-54-78 38 85-22-90 4 90 13-88-31 82 46-73-61 61 73-46-82 31 88-13-90 -4 90 22-85-38 78 54-67}
[0130] {64-64-64 64 64-64-64 64 64-64-64 64 64-64-64 64 64-64-64 64 64-64-6464 64-64-64 64 64-64-64 64}
[0131] {61-73-46 82 31-88-13 90 -4-90 22 85-38-78 54 67-67-54 78 38-85-22 904-90 13 88-31-82 46 73-61}
[0132] {57-80-25 90-9-87 43 70-70-43 87 9-90 25 80-57-57 80 25-90 9 87-43-7070 43-87-9 90-25-80 57}
[0133] {54-85-4 88-46-61 82 13-90 38 67-78-22 90-31-73 73 31-90 22 78-67-3890-13-82 61 46-88 4 85-54}
[0134] {50-89 18 75-75-18 89-50-50 89-18-75 75 18-89 50 50-89 18 75-75-1889-50-50 89-18-75 75 18-89 50}
[0135] {46-90 38 54-90 31 61-88 22 67-85 13 73-82 4 78-78-4 82-73-13 85-67-22 88-61-31 90-54-38 90-46}
[0136] {43-90 57 25-87 70 9-80 80-9-70 87-25-57 90-43-43 90-57-25 87-70 -980-80 9 70-87 25 57-90 43}
[0137] {38-88 73 -4-67 90-46-31 85-78 13 61-90 54 22-82 82-22-54 90-61-1378-85 31 46-90 67 4-73 88-38}
[0138] {36-83 83-36-36 83-83 36 36-83 83-36-36 83-83 36 36-83 83-36-36 83-8336 36-83 83-36-36 83-83 36}
[0139] {31-78 90-61 4 54-88 82-38-22 73-90 67-13-46 85-85 46 13-67 90-73 2238-82 88-54-4 61-90 78-31}
[0140] {25-70 90-80 43 9-57 87-87 57 -9-43 80-90 70-25-25 70-90 80-43 -9 57-87 87-57 9 43-80 90-70 25}
[0141] {22-61 85-90 73-38 -4 46-78 90-82 54-13-31 67-88 88-67 31 13-54 82-9078-46 4 38-73 90-85 61-22}
[0142] {18-50 75-89 89-75 50-18-18 50-75 89-89 75-50 18 18-50 75-89 89-7550-18-18 50-75 89-89 75-50 18}
[0143] {13-38 61-78 88-90 85-73 54-31 4 22-46 67-82 90-90 82-67 46-22-4 31-54 73-85 90-88 78-61 38-13}
[0144] {9-25 43-57 70-80 87-90 90-87 80-70 57-43 25 -9 -9 25-43 57-70 80-87 90-90 87-80 70-57 43-25 9}
[0145] {4-13 22-31 38-46 54-61 67-73 78-82 85-88 90-90 90-90 88-85 82-78 73-67 61-54 46-38 31-22 13-4}
[0146] 2.4.2. Quantization
[0147] The HEVC quantizer design is similar to that of H.264 / AVC, where the quantization parameter (QP) in the range of 0 - 51 (for 8-bit video sequences) is mapped to the quantizer step size, which doubles whenever the QP value increases by 6. However, the key difference is that the transform basis norm correction factor that needs to be incorporated into the de-scaling matrix of H.264 / AVC is no longer required in HEVC, thus simplifying the quantizer design. For quantization groups as small as 8×8 samples, the QP value (in the form of ΔQP) can be transmitted to achieve rate control and perceptual quantization purposes. The QP prediction value used to calculate ΔQP uses a combination of the left, upper, and previous QP values. HEVC also supports frequency-dependent quantization by using a quantization matrix for all transform block sizes. Details will be described in Section 2.4.3.
[0148] The quantized transform coefficient qij (i, j = 0..nS - 1) is derived from the transform coefficient dij (i, j = 0..nS - 1) as follows:
[0149] q ij =(d ij *f[QP % 6]+offset)>>(29 + QP / 6 – nS – BitDepth), where i, j = 0,..., nS - 1
[0150] where,
[0151] f[x] = {26214, 23302, 20560, 18396, 16384, 14564}, x = 0,…,5
[0152] 2 28+QP / 6–nS-BitDepth <offset < 2 29+QP / 6–nS-BitDepth
[0153] QP represents the quantization parameter of a transform unit, and BitDepth represents the bit depth associated with the current color component.
[0154] In HEVC, the range of QP is [0, 51].
[0155] 2.4.3. Quantization Matrix
[0156] The quantization matrix (QM) has been adopted in image coding standards such as JPEG and JPEG - 2000 and video standards such as MPEG2, MPEG4, and H.264 / AVC. QM can improve subjective quality through frequency weighting of different frequency coefficients. In the HEVC standard, the quantization block size can reach up to 32×32. QMs with sizes 4×4, 8×8, 16×16, and 32×32 can be coded into the bitstream. For each block size, different quantization matrices are required for intra / inter prediction types and Y / Cb / Cr color components. A total of 24 quantization matrices (separate matrices for the four block sizes of 4×4, 8×8, 16×16, and 32×32) should be coded.
[0157] The parameters for the quantization matrix can be copied directly from a reference quantization matrix or can be signaled explicitly. When signaled explicitly, the first parameter (i.e., the value of the (0,0) component of the matrix) is coded / decoded directly. And the remaining parameters are coded / decoded using predictive coding according to the raster scan of the matrix.
[0158] The coding and signaling of the scaling matrix in HEVC imply three modes: OFF, DEFAULT, and USER_DEFINED. It should be noted that for transform unit sizes larger than 8×8 (i.e., 16×16, 32×32), the scaling matrix is obtained from the 8×8 scaling matrix by upsampling to a larger size (duplication of elements). For scaling matrices of TBs larger than 8×8, additional DC values must be signaled.
[0159] In HEVC, the maximum number of coded / decoded values of a scaling matrix is equal to 64.
[0160] For all TB sizes, the DC value of the default mode is equal to 16.
[0161] 2.4.3.1. Syntax and Semantics
[0162] 7.3.2.2 Sequence Parameter Set RBSP Syntax
[0163] 7.3.2.2.1 General Sequence Parameter Set RBSP Syntax
[0164]
[0165] 7.3.2.3 Picture Parameter Set RBSP Syntax
[0166] 7.3.2.3.1 General picture parameter set RBSP syntax
[0167]
[0168] 7.3.4 Scaling list data syntax
[0169]
[0170] When scaling_list_enabled_flag equals 1, it is specified that the scaling list is used for the scaling process of transform coefficients. When scaling_list_enabled_flag equals 0, it is specified that the scaling list is not used for the scaling process of transform coefficients.
[0171] When sps_scaling_list_data_present_flag equals 1, it is specified that the scaling_list_data() syntax structure exists in the SPS. When sps_scaling_list_data_present_flag equals 0, it is specified that the scaling_list_data() syntax structure does not exist in the SPS. When sps_scaling_list_data_present_flag does not exist, the value of sps_scaling_list_data_present_flag is inferred to be equal to 0.
[0172] When pps_scaling_list_data_present_flag equals 1, it is specified that the scaling list data for the pictures referred to by this PPS is derived based on the scaling list specified by the valid SPS and the scaling list specified by this PPS. When pps_scaling_list_data_present_flag equals 0, it is specified that the scaling list data for the pictures referred to by this PPS is inferred to be equal to those specified by the valid SPS. When scaling_list_enabled_flag equals 0, the value of pps_scaling_list_data_present_flag must equal 0. When scaling_list_enabled_flag equals 1, sps_scaling_list_data_present_flag equals 0, and pps_scaling_list_data_present_flag equals 0, the default scaling list data derivation array ScalingFactor is used as described in the scaling list data semantics specified in Clause 7.4.5.
[0173] 7.4.5 Scaling list data semantics
[0174] scaling_list_pred_mode_flag[sizeId][matrixId] being equal to 0 specifies that the value of the ScalingList is the same as that of the reference ScalingList. The reference ScalingList is specified by scaling_list_pred_matrix_id_delta[sizeId][matrixId]. scaling_list_pred_mode_flag[sizeId][matrixId] being equal to 1 specifies that the value of the ScalingList is signaled explicitly.
[0175] scaling_list_pred_matrix_id_delta[sizeId][matrixId] is used to derive the reference ScalingList for ScalingList[sizeId][matrixId] as follows:
[0176] – If scaling_list_pred_matrix_id_delta[sizeId][matrixId] is equal to 0, then the ScalingList is inferred from the default ScalingList ScalingList[sizeId][matrixId][i] (for i = 0..Min(63, (1<<(4+(sizeId<<1)))-1)) specified in Tables 7-5 and 7-6.
[0177] – Otherwise, the ScalingList is inferred from the reference ScalingList as follows:
[0178] refMatrixId = matrixId -
[0179] scaling_list_pred_matrix_id_delta[sizeId][matrixId] * (sizeId == 3? 3 : 1) (7-42)
[0180] ScalingList[sizeId][matrixId][i] = ScalingList[sizeId][refMatrixId][i]
[0181] where i = 0..Min(63, (1<<(4+(sizeId<<1)))-1) (7-43)
[0182] If sizeId is less than or equal to 2, then the value of scaling_list_pred_matrix_id_delta[sizeId][matrixId] must be in the range from 0 to matrixId (inclusive of the endpoints). Otherwise (sizeId equals 3), then the value of scaling_list_pred_matrix_id_delta[sizeId][matrixId] must be in the range from 1 to matrixId / 3 (inclusive of the endpoints).
[0183] Table 7-3 – Specification of sizeId
[0184] Size of the quantization matrix sizeId 4×4 0 8×8 1 16×16 2 32×32 3
[0185] Table 7-4 – Specification of matrixId according to sizeId, prediction mode, and color component
[0186]
[0187] scaling_list_dc_coef_minus8[sizeId - 2][matrixId] plus 8 specifies the value of the variable ScalingFactor[2][matrixId][0][0] for the scaling list for the 16x16 size when sizeId equals 2 and the value of ScalingFactor[3][matrixId][0][0] for the scaling list for the 32x32 size when sizeId equals 3. The value of scaling_list_dc_coef_minus8[sizeId - 2][matrixId] must be in the range from -7 to 247 (inclusive of the endpoints).
[0188] When scaling_list_pred_mode_flag[sizeId][matrixId] equals 0, scaling_list_pred_matrix_id_delta[sizeId][matrixId] equals 0, and sizeId is greater than 1, the value of scaling_list_dc_coef_minus8[sizeId - 2][matrixId] is inferred to be equal to 8.
[0189] When scaling_list_pred_matrix_id_delta[sizeId][matrixId] is not equal to 0 and sizeId is greater than 1, infer the value of scaling_list_dc_coef_minus8[sizeId - 2][matrixId] to be equal to scaling_list_dc_coef_minus8[sizeId - 2][refMatrixId], where the value of refMatrixId is given by Equation 7-42.
[0190] When scaling_list_pred_mode_flag[sizeId][matrixId] is equal to 1, scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[sizeId][matrixId][i] and the previous matrix coefficient ScalingList[sizeId][matrixId][i - 1]. The value of scaling_list_delta_coef must be in the range from -128 to 127 (inclusive of the endpoints). The value of ScalingList[sizeId][matrixId][i] must be greater than 0.
[0191] Table 7-5 - Specification of the default values of ScalingList[0][matrixId][i], where i = 0..15
[0192]
[0193] Table 7-6 - Specification of the default values of ScalingList[1..3][matrixId][i], where i = 0..63
[0194]
[0195] The four-dimensional array ScalingFactor[sizeId][matrixId][x][y] (where x, y = 0..(1 << (2 + sizeId)) - 1) specifies an array of scaling factors according to the variable sizeId specified in Table 7-3 and the variable matrixId specified in Table 7-4.
[0196] Derive the elements ScalingFactor[0][matrixId][][] of the quantization matrix with size 4x4 as follows:
[0197] ScalingFactor[0][matrixId][x][y] = ScalingList[0][matrixId][i] (7-44)
[0198] where i = 0..15, matrixId = 0..5, x = ScanOrder[2][0][i][0] and
[0199] y = ScanOrder[2][0][i][1]
[0200] The following derives the elements ScalingFactor[1][matrixId][][] of the quantization matrix with size 8x8:
[0201] ScalingFactor[1][matrixId][x][y] = ScalingList[1][matrixId][i] (7-45)
[0202] where i = 0..63, matrixId = 0..5, x = ScanOrder[3][0][i][0] and
[0203] y = ScanOrder[3][0][i][1]
[0204] The following derives the elements ScalingFactor[2][matrixId][][] of the quantization matrix with size 16x16:
[0205] ScalingFactor[2][matrixId][x*2 + k][y*2 + j] = ScalingList[2][matrixId][i] (7-46)
[0206] where i = 0..63, j = 0..1, k = 0..1, matrixId = 0..5,
[0207] x = ScanOrder[3][0][i][0] and y = ScanOrder[3][0][i][1]
[0208] ScalingFactor[2][matrixId][0][0] = scaling_list_dc_coef_minus8[0][matrixId] + 8 (7-47)
[0209] where matrixId = 0..5
[0210] Derive the element ScalingFactor[3][matrixId][][] of the quantization matrix with size 32x32 as follows:
[0211] ScalingFactor[3][matrixId][x*4+k][y*4+j] = ScalingList[3][matrixId][i] (7-48)
[0212] where i = 0..63, j = 0..3, k = 0..3, matrixId = 0,3
[0213] x = ScanOrder[3][0][i][0] and y = ScanOrder[3][0][i][1]
[0214] ScalingFactor[3][matrixId][0][0] = scaling_list_dc_coef_minus8[1][matrixId] + 8 (7-49)
[0215] where matrixId = 0,3
[0216] When ChromaArrayType is equal to 3, derive the element ScalingFactor[3][matrixId][][] of the chroma quantization matrix with size 32x32 (where matrixId = 1, 2, 4, and 5) as follows:
[0217] ScalingFactor[3][matrixId][x*4+k][y*4+j] = ScalingList[2][matrixId][i] (7-50)
[0218] where i = 0..63, j = 0..3, k = 0..3, x = ScanOrder[3][0][i][0] and y = ScanOrder[3][0][i][1]
[0219] ScalingFactor[3][matrixId][0][0] = scaling_list_dc_coef_minus8[0][matrixId] + 8 (7-51)
[0220] 2.5 Transform and Quantization Design in VVC
[0221] 2.5.1 MTS (Multiple Transform Selection)
[0222] The family of discrete sine transforms includes the well-known discrete Fourier transform, discrete cosine transform, discrete sine transform, and discrete Karhunen-Loeve (under the first-order Markov condition) transform. Among all the members, there are 8 types of transforms based on the cosine function and 8 types of transforms based on the sine function, namely DCT-I, II... VIII and DST-I, II... VIII respectively. These variants of the discrete cosine and sine transforms originate from the different symmetries of their corresponding symmetric periodic sequences
[22] . The transform basis functions of the selected types of DCT and DST used in the proposed method are formulated by the following Table 1.
[0223] Table 1 Transform basis functions of DCT-II / V / VIII and DST I / VII for N-point input
[0224]
[0225] For a block, transform skip or DCT2 / DST7 / DCT8 can be selected. Such a method is called multi-transform selection (MTS).
[0226] To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled on the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applied to the luminance. The MTS CU-level flag is signaled when the following conditions are met.
[0227] - Both the width and height are less than or equal to 32
[0228] - The CBF flag is equal to 1
[0229] If the MTS CU flag is equal to 0, then DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, then two other flags are signaled additionally to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in the following table. When it comes to the accuracy of the transform matrix, an 8-bit main transform core is adopted. Therefore, all the transform cores used in HEVC are kept the same, which includes 4-point DCT-2 and DST-7 as well as 8-point, 16-point, and 32-point DCT-2. Moreover, other transform cores include 64-point DCT-2, 4-point DCT-8, and 8-point, 16-point, 32-point DST-7 and DCT-8 main transform cores.
[0230] Table 1 Transform basis functions of DCT-II / V / VIII and DST I / VII for N-point input
[0231]
[0232] Similar to HEVC, the residual of a block can be coded and decoded in transform skip mode. To avoid redundancy in syntax coding, the transform skip flag is not signaled when MTS_CU_flag at the CU level is not equal to zero. The block size limit for transform skip is the same as that for MTS in JEM4, which indicates that transform skip is applicable for a CU when both the block width and height are equal to or less than 32.
[0233] 2.5.1.1. High-frequency zeroing
[0234] In VTM4, large block size transforms up to 64×64 in size are enabled, which are mainly used for higher resolution videos, such as 1080p sequences and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are zeroed, so that only the low-frequency coefficients are retained. For example, for an M×N transform block (where M is the block width and N is the block height), when M is equal to 64, only the left 32 columns of transform coefficients are left. Similarly, when N is equal to 64, only the top 32 rows of transform coefficients are retained. When using the transform skip mode for large blocks, the entire block is used without zeroing any values.
[0235] To reduce the complexity of large-size DST-7 and DCT-8, the high-frequency transform coefficients are zeroed for DST-7 blocks and DCT-8 blocks with a size (width or height, or both width and height) equal to 32. Only the coefficients in the 16×16 lower frequency region are retained.
[0236] 2.5.2. Reduced second transform
[0237] In JEM, a second transform is applied between the forward primary transform and quantization (at the encoder) and between the inverse quantization and inverse primary transform (at the decoder). As Figure 2 shown, the execution of the 4×4 (or 8×8) second transform depends on the block size. For example, a 4×4 second transform is applied to small blocks (i.e., min(width, height) < 8), and an 8×8 second transform is applied to larger blocks (i.e., min(width, height) > 4) for each 8×8 block.
[0238] For the second transform, an inseparable transform is applied, so it is also known as the non-separable second transform (NSST). There are a total of 35 transform sets, and each transform set uses 3 non-separable transform matrices (kernels, each with a 16×16 matrix).
[0239] The reduced secondary transform (RST) was introduced in JVET-K0099 based on the intra prediction direction and four transform sets (instead of 35 transform sets) were introduced in JVET-L0133. In this document, a 16×48 matrix and a 16×16 matrix are used for 8×8 blocks and 4×4 blocks respectively. For convenience of notation, the 16×48 transform is denoted as RST 8×8 and the 16×16 transform is denoted as RST 4×4. Such a method has recently been adopted by VVC.
[0240] Figure 3 An example of the reduced secondary transform (TST) is shown.
[0241] The secondary forward transform and inverse transform are processing steps separate from those of the primary transform.
[0242] For the encoder, the primary forward transform is first performed, followed by the secondary forward transform, quantization, and CABAC bit coding. For the decoder, CABAC bit decoding, inverse quantization, and the subsequent secondary inverse transform are first performed, followed by the primary inverse transform.
[0243] RST is only applicable to intra-coded / decoded TUs.
[0244] 2.5.3. Quantization
[0245] In VTM4, the maximum QP was extended from 51 to 63, and the signaling of the initial QP was changed accordingly. When encoding / decoding non-zero values of slice_qp_delta, the initial value of SliceQpY is modified at the slice segment layer. Specifically, the value of init_qp_minus26 is modified to be in the range of -(26 + QpBdOffsetY) to +37.
[0246] In addition, the same HEVC scalar quantization is combined with a new concept called dependent scalar quantization. Dependent scalar quantization refers to a scheme in which a set of admissible reconstruction values of transform coefficients depends on the values of transform coefficient magnitudes that are earlier in the reconstruction order than the current transform coefficient magnitude. The main effect of this scheme is that, compared with the conventional independent scalar quantization used in HEVC, the admissible reconstruction vectors are more densely packed in the N-dimensional vector space (N represents the number of transform coefficients in the transform block). This means that for a given average number of admissible reconstruction vectors per unit volume of the N-dimensional cell, the average distortion between the input vector and the closest reconstruction vector is reduced. The dependent scalar quantization scheme is achieved by the following operations: (a) defining two scalar quantizers with different reconstruction magnitudes and (b) defining a process for switching between the two scalar quantizers.
[0247] Figure 4It is an illustration of two scalar quantizers used in the proposed dependency quantization scheme.
[0248] In Figure 4 are shown the two scalar quantizers used, denoted as Q0 and Q1. The positions of the available reconstruction amplitudes are uniquely specified by the quantization step size Δ. The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer for the current transform coefficient is determined by the parity of the transform coefficients that are in front of the current transform coefficient in the encoding / decoding / reconstruction order.
[0249] As Figure 5 shown, the switching between the two scalar quantizers (Q0 and Q1) is implemented by a state machine with four states. The state can take four different values: 0, 1, 2, 3. It is uniquely determined by the parity of the amplitudes of the transform coefficients that are in front of the current transform coefficient in the encoding / decoding / reconstruction order. At the start of the inverse quantization for a transform block, the state is set to 0. The transform coefficients are reconstructed in scan order (i.e., in the same order as they are entropy decoded). After reconstructing the current transform coefficient, the state is updated as Figure 5 shown, where k represents the value of the transform coefficient amplitude.
[0250] 2.5.4 User-Defined Quantization Matrix in JVET-N0847
[0251] In this document, it is proposed to add support for signaling default and user-defined scaling matrices based on VTM4.0. This proposal is compliant with a larger block size range (from 4×4 to 64×64 for luma and from 2×2 to 32×32 for chroma), rectangular TBs, dependency quantization, multiple transform selection (MTS), large transforms that zero out high-frequency coefficients (consistent with the one-step definition process of the scaling matrix for TBs), intra-prediction word block partitioning sub-block partitioning (ISP), and intra-block copy (IBC, also known as current picture reference CPR).
[0252] It is proposed to add syntax to VTM4.0 to support signaling default and user-defined scaling matrices, which is compliant with the following:
[0253] – Three modes for the scaling matrix: OFF (off), DEFAULT (default), and USER_DEFINED (user-defined)
[0254] – Larger block size range (from 4×4 to 64×64 for luma and from 2×2 to 32×32 for chroma)
[0255] – Rectangular transform blocks (TBs)
[0256] – Dependency quantization
[0257] – Multiple Transform Selection (MTS)
[0258] – Large transform that zeros high-frequency coefficients
[0259] – Intra Sub-Block Partitioning (ISP)
[0260] – Intra Block Copy (IBC, also known as Current Picture Reference CPR) shares the same QM as Intra Codec Block
[0261] – For all TB sizes, the DEFAULT scaling matrix is flat and has a default value of 16
[0262] – The scaling matrix should “not” be applied to the following
[0263] ο TS for all TB sizes
[0264] ο Second transform (also known as RST)
[0265] 2.5.4.1. QM signaling for square transform sizes
[0266] 2.5.4.1.1. Scanning order of elements in the scaling matrix
[0267] These elements are decoded in the same scanning order as the coefficient decoding (i.e., diagonal scanning order). An example of the diagonal scanning order is depicted in Figures 6A - 6B which shows an example of the diagonal scanning order.
[0268] Figures 6A - 6B shows an example of the diagonal scanning order. Figure 6A shows an example of the scanning direction. Figure 6B shows the coordinates and scanning order index of each element.
[0269] The corresponding specifications for this order are defined as follows:
[0270] 6.5.2 Right-upper diagonal scanning order array initialization process
[0271] The input to this process is the block width blkWidth and the block size height blkHeight.
[0272] The output of this process is the array diagScan[sPos][sComp]. The array index sPos specifies the scanning position within the range from 0 to (blkWidth * blkHeight) - 1. The array index sComp equal to 0 specifies the horizontal component, and the array index sComp equal to 1 specifies the vertical component. Based on the values of blkWidth and blkHeight, the array diagScan is derived as follows:
[0273]
[0274]
[0275] 2.5.4.1.2. Encoding and Decoding of Selected Elements
[0276] The following scaling matrices are encoded and decoded separately for the DC value (i.e., the element at the scan index equal to 0 in the upper left of the matrix). 16×16, 32×32, and 64×64.
[0277] For a TB (N) with a size less than or equal to 8×8 (N <= 8) × N)
[0278] For a TB with a size less than or equal to 8×8, all elements within a scaling matrix are signaled.
[0279] For a TB (N×N) with a size greater than 8×8 (N > 8)
[0280] If the TB has a size greater than 8×8, then only 64 elements in an 8×8 scaling matrix are signaled as the basic scaling matrix. These 64 elements correspond to the coordinates (m*X, m*Y), where m = N / 8 and X and Y are [0…7]. In other words, an NxN block is divided into multiple non-overlapping m*m regions, and for each region, they share the same element, and this shared element is signaled.
[0281] To obtain a square matrix with a size greater than 8×8, the 8×8 basic scaling matrix is upsampled (by copying elements) to the corresponding square size (i.e., 16×16, 32×32, 64×64).
[0282] Taking 32×2 and 64×64 as examples, the selected positions of the elements to be signaled are marked with circles. Each square represents an element.
[0283] Figure 7 An example of the selected positions for QM signaling (32×32 transform size) is shown.
[0284] Figure 8 An example of the selected positions for QM signaling (64×64 transform size) is shown.
[0285] 2.5.4.2. QM Derivation for Non-Square Transform Sizes
[0286] For non-square transform sizes, there is no additional signaling of QM. Instead, the QM for non-square transform sizes is derived from the QM for square transform sizes. An example is shown in Figure 7 .
[0287] More specifically, when generating the scaling matrix for rectangle TB, two cases are considered:
[0288] 1. The height H of the rectangular matrix is greater than the width W. Thus, the scaling matrix ScalingMatrix for rectangle TB with size WxH is defined as follows by a reference scaling matrix with size baseL×baseL, where baseL is equal to min(log2(H), 3):
[0289]
[0290] For i = 0:W-1, j = 0:H-1 and
[0291] 2. The height H of the rectangular matrix is less than the width W. Thus, the scaling matrix ScalingMatrix for rectangle TB with size WxH is defined as follows by a reference scaling matrix with size baseL×baseL, where baseL is equal to min(log2(W), 3):
[0292]
[0293] For i = 0:W-1, j = 0:H-1 and
[0294] Here, int(x) modifies the value of x by truncating the fractional part.
[0295] Figure 8 Examples of QM derivation of non-square blocks by square blocks are shown. (a) QM of a 2×8 block derived from an 8×8 block, (b) QM of an 8×2 block derived from an 8×8 block.
[0296] 2.5.4.3. QM signaling for transform blocks with zeroing
[0297] In addition, when zeroing the high-frequency coefficients for a 64-point transform, the corresponding high frequencies of the scaling matrix are also zeroed. That is, if the width or height of the TB is greater than or equal to 32, then only the left half or the top half of the coefficients are retained, and zeros are assigned to the remaining coefficients, as Figure 9 shown. Checks are performed for this case when obtaining the rectangular matrix according to equations (1) and (2), and 0 is assigned to the corresponding elements in ScalingMatrix(i, j).
[0298] 2.5.4.4. Syntax and semantics for quantization matrix
[0299] The same syntax elements as in HEVC are added to the SPS and PPS. However, the signaling of the scaling list data syntax is changed to:
[0300] 7.3.2.11 Scaling list data syntax
[0301]
[0302]
[0303] 7.4.3.11 Scaling list data semantics
[0304] scaling_list_pred_mode_flag[sizeId][matrixId] being equal to 0 specifies that the value of the scaling list is the same as the value of the reference scaling list. The reference scaling list is specified by scaling_list_pred_matrix_id_delta[sizeId][matrixId]. scaling_list_pred_mode_flag[sizeId][matrixId] being equal to 1 specifies that the value of the scaling list is signaled explicitly.
[0305] scaling_list_pred_matrix_id_delta[sizeId][matrixId] specifies the reference scaling list used to derive ScalingList[sizeId][matrixId]. The derivation of ScalingList[sizeId][matrixId] is as follows based on scaling_list_pred_matrix_id_delta[sizeId][matrixId]:
[0306] – If scaling_list_pred_matrix_id_delta[sizeId][matrixId] is equal to 0, then this scaling list is inferred from the default scaling list ScalingList[sizeId][matrixId][i] (for i = 0..Min(63,(1<<(sizeId<<1))-1)) specified in Tables 7-15, 7-16, 7-17, 7-18.
[0307] – Otherwise, this scaling list is inferred from the reference scaling list as follows:
[0308] For sizeId = 1…6,
[0309] refMatrixId = matrixId -
[0310] scaling_list_pred_matrix_id_delta[sizeId][matrixId] * (sizeId == 6? 3 : 1) * (7 - XX)
[0311] If sizeId is equal to 1, then the value of refMatrixId must not be equal to 0 or 3. Otherwise, if sizeId is less than or equal to 5, then the value of scaling_list_pred_matrix_id_delta[sizeId][matrixId] must be in the range from 0 to matrixId (inclusive of the endpoints). Otherwise (sizeId is equal to 6), then the value of scaling_list_pred_matrix_id_delta[sizeId][matrixId] must be in the range from 0 to matrixId / 3 (inclusive of the endpoints).
[0312] Table 7 - 13 - Specification of sizeId
[0313]
[0314] Table 7 - 14 - Specification of matrixId according to sizeId, prediction mode, and color component
[0315]
[0316]
[0317] scaling_list_dc_coef_minus8[sizeId][matrixId] plus 8 specifies the value of the variable ScalingFactor[4][matrixId][0][0] for the 16x16 size in this scaling list when sizeId is equal to 4, and specifies the value of ScalingFactor[5][matrixId][0][0] for the 32x32 size in this scaling list when sizeId is equal to 5, and specifies the value of ScalingFactor[6][matrixId][0][0] for the 64x64 size in this scaling list when sizeId is equal to 6. The value of scaling_list_dc_coef_minus8[sizeId][matrixId] must be in the range from -7 to 247 (inclusive of the endpoints).
[0318] When scaling_list_pred_mode_flag[sizeId][matrixId] is equal to 0, scaling_list_pred_matrix_id_delta[sizeId][matrixId] is equal to 0, and sizeId is greater than 3, infer the value of scaling_list_dc_coef_minus8[sizeId][matrixId] to be equal to 8.
[0319] When scaling_list_pred_matrix_id_delta[sizeId][matrixId] is not equal to 0 and sizeId is greater than 3, infer the value of scaling_list_dc_coef_minus8[sizeId][matrixId] to be equal to scaling_list_dc_coef_minus8[sizeId][refMatrixId], where the value of refMatrixId is given by Equation 7-XX.
[0320] When scaling_list_pred_mode_flag[sizeId][matrixId] is equal to 1, scaling_list_delta_coef specifies the difference between the current matrix coefficient ScalingList[sizeId][matrixId][i] and the previous matrix coefficient ScalingList[sizeId][matrixId][i - 1]. The value of scaling_list_delta_coef must be in the range of -128 to 127 (including the endpoints). The value of ScalingList[sizeId][matrixId][i] must be greater than 0. When scaling_list_pred_mode_flag[sizeId][matrixId] is equal to 1 and scaling_list_delta_coef does not exist, infer the value of ScalingList[sizeId][matrixId][i] to be 0.
[0321] Table 7-15 - Specification of the default values of ScalingList[1][matrixId][i], where i = 0..3
[0322]
[0323] Table 7-16 - Specification of the default values of ScalingList[2][matrixId][i], where i = 0..15
[0324]
[0325] Table 7-17 – Specification of the default values of ScalingList[3..6][matrixId][i], where i = 0..63
[0326]
[0327]
[0328] Table 7-18 – Specification of the default values of ScalingList[6][matrixId][i], where i = 0..63
[0329]
[0330] The five-dimensional array ScalingFactor[sizeId][sizeId][matrixId][x][y] (where x, y = 0..(1<<sizeId)-1) specifies an array of scaling factors according to the variable sizeId specified in Table 7-13 and the variable matrixId specified in Table 7-14.
[0331] Derivation of the elements ScalingFactor[1][matrixId][][] of the quantization matrix with size 2×2 is as follows:
[0332] ScalingFactor[1][1][matrixId][x][y] = ScalingList[1][matrixId][i] (7-XX)
[0333] where i = 0..3, matrixId = 1, 2, 4, 5, x = DiagScanOrder[1][1][i][0] and y = DiagScanOrder[1][1][i][1]
[0334] Derivation of the elements ScalingFactor[2][matrixId][][] of the quantization matrix with size 4×4 is as follows:
[0335] ScalingFactor[2][2][matrixId][x][y] = ScalingList[2][matrixId][i] (7-XX)
[0336] where \(i = 0..15\), \(matrixId = 0..5\), \(x = DiagScanOrder[2][2][i][0]\) and \(y = DiagScanOrder[2][2][i][1]\)
[0337] Derive the elements \(ScalingFactor[3][matrixId][][]\) of the quantization matrix with size \(8\times8\) as follows:
[0338] \(ScalingFactor[3][3][matrixId][x][y]=ScalingList[3][matrixId][i](7 - XX)\)
[0339] where \(i = 0..63\), \(matrixId = 0..5\), \(x = DiagScanOrder[3][3][i][0]\) and \(y = DiagScanOrder[3][3][i][1]\)
[0340] Derive the elements \(ScalingFactor[4][matrixId][][]\) of the quantization matrix with size \(16\times16\) as follows:
[0341] \(ScalingFactor[4][4][matrixId][x*2 + k][y*2 + j]=ScalingList[4][matrixId][i](7 - XX)\)
[0342] where \(i = 0..63\), \(j = 0..1\), \(k = 0..1\), \(matrixId = 0..5\), \(x = DiagScanOrder[3][3][i][0]\) and \(y = DiagScanOrder[3][3][i][1]\)
[0343] \(ScalingFactor[4][4][matrixId][0][0]=scaling_list_dc_coef_minus8[0][matrixId]+8(7 - XX)\)
[0344] where \(matrixId = 0..5\)
[0345] Derive the elements \(ScalingFactor[5][matrixId][][]\) of the quantization matrix with size \(32\times32\) as follows:
[0346] \(ScalingFactor[5][5][matrixId][x*4 + k][y*4 + j]=ScalingList[5][matrixId][i](7 - XX)\)
[0347] where \(i = 0..63\), \(j = 0..3\), \(k = 0..3\), \(matrixId = 0..5\), \(x = DiagScanOrder[3][3][i][0]\) and \(y = DiagScanOrder[3][3][i][1]\)
[0348] ScalingFactor[5][5][matrixId][0][0]=scaling_list_dc_coef_minus8[1][matrixId]+8 (7 - XX)
[0349] where \(matrixId = 0..5\)
[0350] The elements ScalingFactor[6][matrixId][][] of the quantization matrix with size \(64\times64\) are derived as follows:
[0351] ScalingFactor[6][6][matrixId][x*8 + k][y*8 + j]=ScalingList[6][matrixId][i] (7 - XX)
[0352] where \(i = 0..63\), \(j = 0..7\), \(k = 0..7\), \(matrixId = 0,3\), \(x = DiagScanOrder[3][3][i][0]\) and \(y = DiagScanOrder[3][3][i][1]\)
[0353] ScalingFactor[6][6][matrixId][0][0]=scaling_list_dc_coef_minus8[2][matrixId]+8 (7 - XX)
[0354] where \(matrixId = 0,3\)
[0355] When ChromaArrayType is equal to 3, the elements ScalingFactor[6][6][matrixId][][](where \(matrixId = 1, 2, 4\) and \(5\)) of the chroma quantization matrix with size \(64\times64\) are derived as follows:
[0356] ScalingFactor[6][6][matrixId][x*8 + k][y*8 + j]=ScalingList[5][matrixId][i] (7 - XX)
[0357] where i = 0..63, j = 0..7, k = 0..7, x = DiagScanOrder[3][3][i][0] and y = DiagScanOrder[3][3][i][1]
[0358] ScalingFactor[6][6][matrixId][0][0] = scaling_list_dc_coef_minus8[1][matrixId] + 8 (7 - XX)
[0359] / / Non-square case
[0360] For a quantization matrix with rectangular dimensions,
[0361] The five-dimensional array ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y] (where x = 0..(1<<sizeIdW)-1, y = 0..(1<<sizeIdH)-1, sizeIdW!= sizeIdH) is an array that specifies the scaling factors according to the variables sizeIdW and sizeIdH specified in Table 7-19
[0362] ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y] can be generated from ScalingList[sizeLId][matrixId][i] according to the following rules, where sizeLId = max(sizeIdW, sizeIdH), sizeIdW = 0, 1..6, sizeIdH = 0, 1..6, matrixId = 0..5, x = 0..(1<<sizeIdW)-1, y = 0..(1<<sizeIdH)–1,, x = DiagScanOrder[k][k][i][0] and y = DiagScanOrder[k][k][i][1], k = min(sizeLId, 3) and ratioW = (1<<sizeIdW) / (1<<k), ratioH = (1<<sizeIdH) / (1<<k) and ratioWH = (1<<abs(sizeIdW-sizeIdH))
[0363] - If (sizeIdW > sizeIdH)
[0364] ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y]
[0365] =
[0366] ScalingList[sizeLId][matrixId][Raster2Diag[(1<<k)*((y*ratioWH) / ratioW)+x / ratioW]]
[0367] - else
[0368] ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y]
[0369] = ScalingList[sizeLId][matrixId][Raster2Diag[(1<<k)*(y / ratioH)+(x*ratioWH) / ratioH]],
[0370] where Raster2Diag[] is a function that converts the raster scan position in an 8×8 block to the diagonal scan position
[0371] / / Zeroing case
[0372] For samples that meet the following conditions, the quantization matrix with rectangular dimensions must be zeroed
[0373] - x > 32
[0374] - y > 32
[0375] - the decoded tu is not encoded / decoded by the default transform mode, (1<<sizeIdW) == 32 and x > 16
[0376] - the decoded tu is not encoded / decoded by the default transform mode, (1<<sizeIdH) == 32 and y > 16
[0377] Table 7-19 – Specifications of sizeIdW and sizeIdH
[0378]
[0379] 2.6 Quantized Residual Block Differential Pulse Coding and Modulation
[0380] In JVET-M0413, Quantized Residual Block Differential Pulse Coding and Modulation (QR-BDPCM) was proposed to efficiently encode and decode screen content
[0381] The prediction directions used in QR-BDPCM can be vertical and horizontal prediction modes. Intra-frame prediction for the entire block is performed by replicating samples within a prediction direction (horizontal or vertical prediction) similar to intra-frame prediction. The residual is quantized, and the Δ between the quantized residual and its predictor (horizontal or vertical) quantization value is encoded and decoded. This situation can be described as follows: For a block of size M (rows) × N (columns), let r i,j (0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1) be the prediction residual after performing intra-frame prediction horizontally (copying the left neighbor pixel values line by line across the prediction block) or vertically (copying the top neighbor line to each line in the prediction block) using the unfiltered samples from the upper or left block boundary samples. Let Q(r i,j )(0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1) represent the quantized version of the residual r i,j , where the residual is the difference between the initial block and the prediction block values. Then, block DPCM is applied to the quantized residual samples to obtain a modified M×N array whose elements are When signaling vertical BDPCM:
[0382]
[0383] For horizontal prediction, similar rules apply, and the residual quantized samples are obtained through the following equations
[0384]
[0385] The residual quantized samples are sent to the decoder.
[0386] On the decoder side, the above calculations are reversed to produce Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1. For the vertical prediction case,
[0387]
[0388] For the horizontal case,
[0389]
[0390] The inverse quantized residual Q -1 (Q(r i,j )) is added to the intra-frame block prediction value to produce the reconstructed sample value.
[0391] The main benefit of this scheme is that inverse DPCM can be performed instantaneously during coefficient parsing simply by adding the predictor, or it can be performed after parsing.
[0392] The draft text changes of QR-BDPCM are shown below.
[0393] 7.3.6.5 Coding and Decoding Unit Syntax
[0394]
[0395]
[0396]
[0397] bdpcm_flag[x0][y0] being equal to 1 specifies that there is a bdpcm_dir_flag in the coding and decoding unit of the luminance coding and decoding block including the position (x0, y0).
[0398] bdpcm_dir_flag[x0][y0] being equal to 0 specifies that the prediction direction to be used in the bdpdcm block is the horizontal direction; otherwise, it is the vertical direction.
[0399] 3. Exemplary Technical Problems Solved by the Embodiments and Techniques
[0400] The current VVC design has the following problems with respect to the quantization matrix:
[0401] 1. For QR-BDPCM blocks and BDPCM blocks, no transformation is applied. Therefore, it is not preferred to apply the scaling matrix to such blocks in a manner similar to other blocks where transformation is applied.
[0402] 2. IBC and the intra coding mode share the same scaling matrix. However, IBC is more similar to an inter coding tool. Such a design seems unreasonable.
[0403] 3. For larger blocks with zeroing, several elements are signaled but are reset to zero during the decoding process, which wastes bits.
[0404] 4. In some frameworks, quantization is performed in a way that always applies the quantization matrix. Thus, the quantization matrix for small blocks may cause frequent matrix changes.
[0405] 4. Enumeration of the Embodiments and Techniques
[0406] The inventions detailed below should be regarded as examples for explaining general concepts. These inventions should not be interpreted narrowly. Additionally, these aspects can be combined in any way.
[0407] 1. When the transform skip (TS) flag is not signaled in the bitstream, how to set the value of this flag can depend on the coding and decoding mode information.
[0408] a. In one example, for a BDPCM / QR-BDPCM encoding / decoding block, the TS flag is inferred as 1.
[0409] b. In one example, for a non-BDPCM / QR-BDPCM encoding / decoding block, the TS flag is inferred as 0.
[0410] 2. It is proposed that a scaling matrix can be applied to a BDPCM / QR-BDPCM encoding / decoding block.
[0411] a. Alternatively, it is proposed that a scaling matrix is not allowed for a BDPCM / QR-BDPCM encoding / decoding block.
[0412] b. In one example, how to select a scaling matrix for a BDPCM / QR-BDPCM encoding / decoding block can be performed in the same way as for a transform skip encoding / decoding block.
[0413] c. In one example, how to select a scaling matrix for a BDPCM / QR-BDPCM encoding / decoding block can depend on whether one or more transforms are applied to the block.
[0414] i. In one example, if one or more transforms are applied to a BDPCM / QR-BDPCM encoding / decoding block, then a scaling matrix can be allowed.
[0415] ii. In one example, if one or more transforms are applied to a BDPCM / QR-BDPCM encoding / decoding block, then how to select a scaling matrix for a BDPCM / QR-BDPCM encoding / decoding block can be performed in the same way as for an intra encoding / decoding block.
[0416] 3. Whether and / or how to filter samples / edges between two adjacent blocks by a loop filter (deblocking filter) and / or other post-reconstruction filters can depend on whether any one or both of the two adjacent blocks are encoded in a transform skip (TS) mode.
[0417] a. Whether and / or how to filter samples / edges between two adjacent blocks by a loop filter (deblocking filter) and / or other post-reconstruction filters can depend on whether the two adjacent blocks are encoded in a TS, BDPCM, QR-BDPCM, or palette mode.
[0418] b. In one example, the derivation of the boundary filtering strength can depend on the (multiple) TS mode flags of one or both of the two adjacent blocks.
[0419] c. In one example, for samples located at a TS encoded block, the deblocking filter / sample adaptive offset / adaptive loop filter / other types of loop filters / other post-reconstruction filters can be disabled.
[0420] i. In one example, if two adjacent blocks are coded using the transform skip mode, then no edge filtering is required between these two blocks.
[0421] ii. In one example, if one of two adjacent blocks is coded using the transform skip mode and the other is not, then no sample filtering is required at the samples at the TS-coded block.
[0422] d. Alternatively, for samples at the TS-coded block, different filters (e.g., smoother filters) may be allowed.
[0423] e. During the processing of the in-loop filter (e.g., deblocking filter) and / or other post-reconstruction filters (e.g., bilateral filter, diffusion filter), blocks coded using PCM / BDPCM / QR-BDPCM and / or other types of modes that do not apply transforms may be processed in the same way as those coded using the TS mode (e.g., as mentioned above).
[0424] 4. Blocks coded using the IBC mode and the inter prediction mode may share the same scaling matrix.
[0425] 5. The scaling matrix selection may depend on the transform matrix type.
[0426] a. In one example, the selection of the scaling matrix may depend on whether the block uses the default transform, e.g., DCT2.
[0427] b. In one example, the scaling matrix may be signaled separately for multiple transform matrix types.
[0428] 6. The scaling matrix selection may depend on the motion information of the block.
[0429] a. In one example, the scaling matrix selection may depend on whether the block is coded using the sub-block coding mode (e.g., affine model).
[0430] b. In one example, the scaling matrix may be signaled separately for the affine mode and the non-affine mode.
[0431] c. In one example, the scaling matrix selection may depend on whether the block is coded using the affine intra prediction mode.
[0432] 7. It is proposed to enable the scaling matrix for the quadratic transform coded blocks, instead of disabling the scaling matrix for the quadratic transform coded blocks. Assume that the transform block size is represented by K×L, and a quadratic transform is applied to the upper left M×N block.
[0433] a. A scaling matrix can be signaled in the slice header / strip header / PPS / VPS / SPS for a second transformation or / and a reduced second transformation or / and a rotation transformation.
[0434] b. In one example, the scaling matrix can be selected according to whether a second transformation is applied.
[0435] i. In one example, elements of the scaling matrix for the top-left M×N block can be signaled separately for whether a second transformation is applied.
[0436] c. Alternatively, in addition, the scaling matrix can be applied only to regions where a second transformation is not applied.
[0437] d. In one example, the scaling matrix can still be applied to the remaining part except for the top-left M×N region.
[0438] e. Alternatively, the scaling matrix can be applied only to regions where a second transformation is applied.
[0439] 8. Signaling a scaling matrix for non-square blocks is proposed, replacing deriving the scaling matrix for non-square blocks from square blocks.
[0440] a. In one example, a scaling matrix can be enabled for non-square blocks that may be predicted and decoded using a prediction from square blocks.
[0441] 9. It is proposed to disable the use of the scaling matrix for some positions and enable the use of the scaling matrix for the remaining positions within a block.
[0442] a. For example, for a block that includes more than M*N positions, only the top-left M*N region can use the scaling matrix.
[0443] b. For example, for a block that includes more than M*N positions, only the top M*N region can use the scaling matrix.
[0444] c. For example, for a block that includes more than M*N positions, only the left M*N region can use the scaling matrix.
[0445] 10. How many elements in the signaling scaling matrix can depend on whether zeroing is applied.
[0446] a. In one example, for a 64×64 transformation, assuming only the top-left M×N transformation coefficients are retained and all the remaining coefficients are zeroed. Then the number of elements to be signaled can be derived as M / 8*N / 8.
[0447] 11. For a transformation with zeroing, it is proposed to prohibit signaling elements in the scaling matrix that are in the zeroing region. For a K×L transformation, assuming only the top-left M×N transformation coefficients are retained and all the remaining coefficients are zeroed.
[0448] a. In one example, K = L = 64 and M = N = 32.
[0449] b. In one example, signaling of elements corresponding to positions outside the upper left M×N region in the scaling matrix is skipped.
[0450] Figure 10 An example is shown where only the selected elements within the dashed region (e.g., the M×N region) are signaled.
[0451] c. In one example, the downsampling ratio for selecting elements in the scaling matrix can be determined by K and / or L.
[0452] i. For example, the transform block is divided into multiple sub-regions, and each sub-region has a size of Uw*Uh. One element within each sub-region located in the upper left M×N region can be signaled.
[0453] ii. Alternatively, in addition, the number of elements to be decoded can depend on M and / or N.
[0454] 1) In one example, for such a K×L transform with zeroing, the number of elements to be decoded is different from that of an M×N transform block without zeroing.
[0455] d. In one example, the downsampling ratio for selecting elements in the scaling matrix can be
[0456] determined by M and / or N, rather than by K and L.
[0457] i. For example, the M×N region is divided into multiple sub-regions. One element within each region (M / Uw, N / Uh) can be signaled.
[0458] ii. Alternatively, in addition, for such a K×L transform with zeroing, the number of elements to be decoded is the same as that of an M×N transform block without zeroing.
[0459] e. In one example, K = L = 64, M = N = 32, Uw = Uh = 8.
[0460] 12. It is proposed to use only one quantization matrix for certain block sizes (e.g., for small-sized blocks).
[0461] a. In one example, it may not be allowed for all blocks smaller than W×H (regardless of block type) to use two or more quantization matrices.
[0462] b. In one example, it may not be allowed for all blocks with a width smaller than a threshold to use two or more quantization matrices.
[0463] c. In one example, blocks with all heights less than a threshold may not be allowed to use two or more quantization matrices.
[0464] d. In one example, quantization matrices may not be applied to small-sized blocks.
[0465] 13. The above bullet points may apply to other encoding and decoding methods that do not apply a transform (or do not apply the identity transform).
[0466] a. In one example, by replacing "TS / BDPCM / QR-BDPCM" with a "palette", the above bullet points may apply to palette-mode encoded and decoded blocks.
[0467] 5. Embodiments
[0468] 5.1 Embodiment #1 Regarding the Deblocking Filter
[0469] Modifications imposed on the VVC working draft version 5 are highlighted by bold italic text. One or more of the highlighted conditions may be added.
[0470] 8.8.2 Deblocking Filtering Process
[0471] 8.8.2.1 Overview
[0472] The input to this process is the reconstructed picture before deblocking, i.e., the array recPicture L , and the array recPicture when ChromaArrayType is not equal to 0 Cb and recPicture Cr .
[0473] The output of this process is the modified reconstructed picture after deblocking, i.e., the array recPicture L and the array recPicture when ChromaArrayType is not equal to 0 Cb and recPicture Cr .
[0474] First, the vertical edges in the picture are filtered. Then, the samples modified by the vertical edge filtering process are used as input to filter the horizontal edges in the picture. Based on the coding unit, the vertical and horizontal edges in the CTB of each CTU are processed independently. Starting from the edge on the left side of the coding block, the vertical edges of the coding block in the coding unit are filtered, in its geometric order, towards the right side of the coding block, filtering through each edge. Starting from the edge above the coding block, the horizontal edges of the coding block in the coding unit are filtered, in its geometric order, towards the bottom of the coding block, filtering through each edge.
[0475] Note - Although the filtering process is specified based on pictures in this specification, an equivalent effect can be obtained by implementing the filtering process based on coding / decoding units, provided that the decoder appropriately considers the processing dependency order to produce the same output values.
[0476] Apply the deblocking filtering process to all coding / decoding sub-block edges and transform block edges of the picture, except for the following types of edges:
[0477] - Edges at the picture boundary,
[0478] - Edges that coincide with the virtual boundary of the picture when pps_loop_filter_across_virtual_boundaries_disabled_flag is equal to 1,
[0479] - Edges that coincide with the tile boundary when loop_filter_across_bricks_enabled_flag is equal to 0,
[0480] - Edges that coincide with the top or left boundary of the slice when slice_loop_filter_across_slices_enabled_flag is equal to 0 or slice_deblocking_filter_disabled_flag is equal to 1,
[0481] - Edges within the slice when slice_deblocking_filter_disabled_flag is equal to 1,
[0482] - Edges that do not correspond to the 8x8 sample grid boundary of the considered component,
[0483] – Edges within the chrominance component for which inter prediction is used on both sides of the edge,
[0484] – Edges of the chrominance transform block that are not edges of the associated transform unit.
[0485] – Edges of the luma transform block across coding / decoding units with an IntraSubPartitionsSplit value not equal to ISP_NO_SPLIT.
[0486]
[0487] 5.2 Example #2 Regarding the Scaling Matrix
[0488] This section provides an example of item 11.d in section 4.
[0489] The modifications applied to JVET-N0847 are highlighted in bold italic text, and the text removed is marked. One or more of the highlighting conditions can be added.
[0490] The elements ScalingFactor[6][matrixId][][] of the quantization matrix with size 64×64 are derived as follows:
[0491] ScalingFactor[6][6][matrixId][x*4 +k][y*4 +j] = ScalingList[6][matrixId][i](7 - XX)
[0492] where i = 0..63, j = 0..3 , k = 0..3 matrixId = 0,3, x = DiagScanOrder[3][3][i][0] and y = DiagScanOrder[3][3][i][1]
[0493]
[0494]
[0495]
[0496] ScalingFactor[6][6][matrixId][0][0] = scaling_list_dc_coef_minus8[2][matrixId]+8 (7 - XX)
[0497] where matrixId = 0,3
[0498]
[0499] For samples that meet the following conditions, the quantization matrix with rectangular dimensions must be zeroed out
[0500]
[0501] - The decoded tu is not encoded / decoded by the default transform mode, (1<<sizeIdW) == 32 and x > 16
[0502] - The decoded tu is not encoded / decoded by the default transform mode, (1<<sizeIdH) == 32 and y > 16
[0503] 5.3 Examples of Scaling Matrices #3
[0504] This section provides examples for bullets 9 and 11.c in Section 4.
[0505] The modifications applied to JVET-N0847 are highlighted in bold italic text, and the text with markings removed is used. One or more of the highlighted conditions can be added.
[0506] 7.3.2.11 Scaling List Data Syntax
[0507]
[0508]
[0509] The following derives the elements ScalingFactor[6][matrixId][][] of the quantization matrix with dimensions 64x64:
[0510] ScalingFactor[6][6][matrixId][x*8+k][y*8+j] = ScalingList[6][matrixId][i](7 - XX)
[0511] where i = 0.. , j = 0..7, k = 0..7, matrixId = 0, 3, x = DiagScanOrder[3][3][i][0] and y = DiagScanOrder[3][3][i][1]
[0512]
[0513]
[0514]
[0515] ScalingFactor[6][6][matrixId][0][0] = scaling_list_dc_coef_minus8[2][matrixId] + 8 (7 - XX)
[0516] where matrixId = 0, 3
[0517]
[0518] For samples that meet the following conditions, the quantization matrix with rectangular dimensions must be zeroed out
[0519]
[0520] - The decoded TU is not encoded / decoded by the default transform mode, (1 << sizeIdW) == 32 and x > 16
[0521] - The decoded TU is not encoded / decoded by the default transform mode, (1 << sizeIdH) == 32 and y > 16
[0522] 5.4 Example #4 Regarding the Scaling Matrix
[0523] This section provides an example where a scaling matrix is not allowed for the QR - BDPCM encoding / decoding block.
[0524] The modifications applied to JVET - N0847 are highlighted in bold italic text, and the text with the markup removed. One or more of the highlighted conditions can be added.
[0525] 8.7.3 Scaling Process of Transform Coefficients
[0526] The inputs to this process are:
[0527] – The luminance position (xTbY, yTbY), which specifies the top - left sample of the current luminance transform block relative to the top - left luminance sample of the current picture,
[0528] – The variable nTbW that specifies the width of the transform block,
[0529] – The variable nTbH that specifies the height of the transform block,
[0530] – The variable cIdx that specifies the color component of the current block,
[0531] – The variable bitDepth that specifies the bit - depth of the current color component.
[0532] The output of this process is an (nTbW) x (nTbH) array d of scaled transform coefficients with elements d[x][y].
[0533] The quantization parameter qP is derived as follows:
[0534] – If cIdx is equal to 0, then the following applies:
[0535] qP = Qp′ Y (8 - 1019)
[0536] – Otherwise, if cIdx is equal to 1, then the following applies:
[0537] qP = Qp′Cb (8-1020)
[0538] – Otherwise (cIdx equals 2), then the following applies:
[0539] qP = Qp′ Cr (8-1021)
[0540] Derive the variable rectNonTsFlag as follows:
[0541] rectNonTsFlag = (((Log2(nTbW)+Log2(nTbH)) & 1) == 1 &&
[0542] (8-1022)
[0543] transform_skip_flag[xTbY][yTbY] == 0)
[0544] Derive the variables bdShift, rectNorm, and bdOffset as follows:
[0545] bdShift = bitDepth + ((rectNonTsFlag? 8:0) +
[0546] (8-1023)
[0547] (Log2(nTbW)+Log2(nTbH)) / 2) - 5 + dep_quant_enabled_flag
[0548] rectNorm = rectNonTsFlag? 181:1 (8-1024)
[0549] bdOffset = (1 << bdShift) >> 1
[0550] (8-1025)
[0551] Specify the list levelScale[] as levelScale[k] = {40, 45, 51, 57, 64, 72}, where k = 0..5.
[0552] For the derivation of the scaled transform coefficient d[x][y] (where x = 0..nTbW-1, y = 0..nTbH-1), the following applies:
[0553] – Derive the intermediate scaling factor m[x][y] as follows:
[0554] – If one or more of the following conditions are true, then set m[x][y] equal to 16:
[0555] – The scaling_list_enabled_flag is equal to 0.
[0556] – The transform_skip_flag[xTbY][yTbY] is equal to 1.
[0557]
[0558] – Otherwise, the following applies:
[0559] – m[x][y] = ScalingFactor[sizeIdW][sizeIdH][matrixId][x][y]
[0560] (8 - XXX)
[0561] where sizeIdW is set to be equal to Log2(nTbW), sizeIdH is set to be equal to Log2(nTbH), and matrixId is specified in Table 7 - 14.
[0562] – Derive the scaling factor ls[x][y] as follows:
[0563] - If the dep_quant_enabled_flag is equal to 1, then the following applies:
[0564] ls[x][y] = (m[x][y] * levelScale[(qP + 1) % 6]) << ((qP + 1) / 6) (8 - 1026)
[0565] – Otherwise (dep_quant_enabled_flag is equal to 0), the following applies:
[0566] ls[x][y] = (m[x][y] * levelScale[qP % 6]) << (qP / 6) (8 - 1027) – Derive the value of dnc[x][y] as follows:
[0567] dnc[x][y] = (8 - 1028)
[0568] (TransCoeffLevel[xTbY][yTbY][cIdx][x][y] * ls[x][y] * rectNorm + bdOffset) >> bdShift
[0569] – Derive the scaled transform coefficient d[x][y] as follows:
[0570] d[x][y] = Clip3(CoeffMin, CoeffMax, dnc[x][y]) (8 - 1029)
[0571] 5.5 Example #5 Regarding the Semantics of the Transform Skip Flag
[0572] transform_skip_flag[x0][y0] specifies whether to apply a transform to a luma transform block. The array indices x0, y0 specify the position (x0, y0) of the top - left luma sample of the transform block under consideration relative to the top - left luma sample of the picture. transform_skip_flag[x0][y0] being equal to 1 specifies not to apply a transform to the luma transform block. transform_skip_flag[x0][y0] being equal to 0 specifies that the decision of whether to apply a transform to the luma transform block will depend on other syntax elements. When transform_skip_flag[x0][y0] does not exist it is inferred that transform_skip_flag[x0][y0] is equal to 0. When transform_skip_flag[x0][y0] does not exist it is inferred that transform_skip_flag[x0][y0] is equal to 0.
[0573] Figure 11 is a block diagram of a video processing apparatus 1100. The apparatus 1100 can be used to implement one or more of the methods described herein. The apparatus 1100 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1100 can include one or more processors 1102, one or more memories 1104, and video processing hardware 1106. The (one or more) processors 1102 can be configured to implement one or more of the methods described in this document. The (one or more) memories 1104 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1106 can be used to implement some of the techniques described in this document in hardware circuits.
[0574] The following solutions can be implemented as preferred solutions in some embodiments.
[0575] The following solutions can be implemented together with the additional techniques described in the items (e.g., item 1) listed in the previous section.
[0576] 1. A video processing method (e.g., Figure 12The method shown in (1200) includes: performing a conversion between a codec representation of a video block of a video and the video block, determining (1202) whether to enable a transform skip mode for the conversion based on codec mode information; and performing (1204) the conversion based on the determination, where, in the transform skip mode, application of a transform to at least some coefficients representing the video block is skipped during the conversion.
[0577] 2. The method according to Solution 1, wherein the transform skip mode is determined to be enabled because the codec mode information indicates block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR - BDPCM).
[0578] 3. The method according to Solution 1, wherein a flag indicating the transform skip mode in the codec representation is not parsed.
[0579] 4. The method according to Solution 1, wherein parsing of a flag indicating the transform skip mode in the codec representation is skipped.
[0580] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 2).
[0581] 5. A video processing method includes: determining to use a scaling matrix for the conversion because block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR - BDPCM) mode is used for the conversion between a codec representation of a video block and the video block; and performing the conversion using the scaling matrix, where the scaling matrix is used to scale at least some coefficients representing the video block during the conversion.
[0582] 6. The method according to Solution 5, wherein the conversion includes applying the scaling matrix according to a mode depending on the number of transforms to be applied to the coefficients during the conversion.
[0583] 7. A video processing method includes: determining to prohibit using a scaling matrix for the conversion because block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR - BDPCM) mode is used for the conversion between a codec representation of a video block and the video block; and performing the conversion using the scaling matrix, where the scaling matrix is used to scale at least some coefficients representing the video block during the conversion.
[0584] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 3).
[0585] 8. A video processing method, comprising: performing a conversion between a codec representation of a video block of a video and the video block, determining the applicability of a loop filter based on whether transform skip mode is enabled for the conversion; and performing the conversion based on the applicability of the loop filter, wherein, in the transform skip mode, application of a transform to at least some coefficients representing the video block is skipped during the conversion.
[0586] 9. The method according to solution 8, wherein the loop filter comprises a deblocking filter.
[0587] 10. The method according to any one of solutions 8 - 9, further comprising determining the strength of the loop filter based on the transform skip mode of the video block and another transform skip mode of a neighboring block.
[0588] 11. The method according to any one of solutions 8 - 9, wherein the determination comprises determining that the loop filter is not applicable due to a reason for disabling the transform skip mode for the video block.
[0589] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., items 4 and 5).
[0590] 12. A video processing method, comprising: selecting a scaling matrix for a conversion between a video block of a video and a codec representation of the video block, such that the same scaling matrix is used for conversions based on inter - frame codec and intra - block copy codec; and performing the conversion using the selected scaling matrix, wherein the scaling matrix is used to scale at least some coefficients of the video block.
[0591] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 6).
[0592] 13. A video processing method, comprising: selecting a scaling matrix for the conversion based on a transform matrix selected for a conversion between a video block of a video and a codec representation of the video block; and performing the conversion using the selected scaling matrix, wherein the scaling matrix is used to scale at least some coefficients of the video block, and wherein the transform matrix is used to transform at least some coefficients of the video block during the conversion.
[0593] 14. The method according to solution 13, wherein the selection of the scaling matrix is based on whether the conversion of the video block uses a sub - block codec mode.
[0594] 15. The method according to solution 14, wherein the sub - block codec mode is an affine codec mode.
[0595] 16. The method according to solution 15, wherein the scaling matrix for the affine coding / decoding mode is different from another scaling matrix of another video block, wherein the transformation of the other video block does not use the affine coding / decoding mode.
[0596] The following solution can be implemented together with the additional techniques described in the items (e.g., item 7) listed in the previous section.
[0597] 17. A video processing method, comprising: selecting a scaling matrix for a transformation between a video block of a video and a coded / decoded representation of the video block based on a quadratic transformation matrix selected for the transformation; and performing the transformation using the selected scaling matrix, wherein the scaling matrix is used to scale at least some coefficients of the video block, and wherein the quadratic transformation matrix is used to transform at least some residual coefficients of the video block during the transformation.
[0598] 18. The method according to solution 17, wherein the quadratic transformation matrix is applied to the upper left MxN part of the video block, and wherein the scaling matrix is applied to a range exceeding the upper left MxN part of the video block.
[0599] 19. The method according to solution 17, wherein the quadratic transformation matrix is applied to the upper left MxN part of the video block, and wherein the scaling matrix is applied only to the upper left MxN part of the video block.
[0600] 20. The method according to any of solutions 17 - 19, wherein a syntax element in the coded / decoded representation indicates the scaling matrix.
[0601] The following solution can be implemented together with the additional techniques described in the items (e.g., item 8) listed in the previous section.
[0602] 21. A video processing method, comprising: determining a scaling matrix used in a transformation between a video block having a non-square shape and a coded / decoded representation of the video block, wherein a syntax element in the coded / decoded representation signals the scaling matrix; and performing the transformation based on the scaling matrix, wherein the scaling matrix is used to scale at least some coefficients of the video block during the transformation.
[0603] 22. The method according to solution 21, wherein the syntax element is predictive coded with the scaling matrix by a scaling matrix of a previous square block.
[0604] The following solution can be implemented together with the additional techniques described in the items (e.g., item 9) listed in the previous section.
[0605] 23. A video processing method, comprising: determining a scaling matrix that will be partially applied during the conversion between the coded representation of a video block and the video block; and performing the conversion by partially applying the scaling matrix such that the scaling matrix is applied at a first set of positions of the video block and disabled at the remaining positions of the video block.
[0606] 24. The method according to solution 23, wherein the first set of positions comprises the upper left M*N positions of the video block.
[0607] 25. The method according to solution 23, wherein the first set of positions corresponds to the top M*N positions of the video block.
[0608] 26. The method according to solution 23, wherein the first set of positions comprises the left M*N positions of the video block.
[0609] The following solution can be implemented together with the additional techniques described in the items listed in the previous section (e.g., items 10 and 11).
[0610] 27. A video processing method, comprising: determining a scaling matrix that will be applied during the conversion between the coded representation of a video block and the video block; and performing the conversion based on the scaling matrix, wherein the coded representation signals the number of elements of the scaling matrix, and wherein the number depends on the application of coefficient zeroing in the conversion.
[0611] 28. The method according to solution 27, wherein the conversion includes zeroing all positions other than the upper left MxN positions of the video block, and wherein the number is M / 8*N / 8.
[0612] The following solution can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 11).
[0613] 29. The method according to any one of solutions 27-28, wherein the number depends on the transformation matrix used during the conversion.
[0614] 30. The method according to solution 29, wherein the transformation matrix has dimensions KxL, and wherein only the top MxN coefficients are not zeroed.
[0615] 31. The method according to any one of solutions 27-30, wherein the scaling matrix is applied by subsampling according to a factor determined by K or L.
[0616] The following solution can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 12).
[0617] 32. A video processing method, comprising: determining a single quantization matrix to be used based on the size of a video block of a specific type during conversion between the video block and its codec representation; and performing the conversion using the quantization matrix.
[0618] 33. The method according to solution 32, wherein the size of the video block is less than WxH, where W and H are integers.
[0619] 34. The method according to any one of solutions 32 - 33, wherein the width of the video block is less than a threshold.
[0620] 35. The method according to any one of solutions 32 - 33, wherein the height of the video block is less than a threshold.
[0621] 36. The method according to solution 32, wherein the quantization matrix is an identity quantization matrix that does not affect the quantization value.
[0622] 37. The method according to any one of solutions 1 to 36, wherein the conversion includes encoding the video into a codec representation.
[0623] 38. The method according to any one of solutions 1 to 36, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0624] 39. A video decoding device, comprising a processor configured to implement the method according to one or more of solutions 1 to 38.
[0625] 40. A video encoding device, comprising a processor configured to implement the method according to one or more of solutions 1 to 38.
[0626] 41. A computer program product having computer code stored thereon, the code causing the processor to implement the method according to any one of solutions 1 to 38 when executed by the processor.
[0627] 42. The method, device, or system described in this document.
[0628] Some embodiments of the disclosed technology include making a determination or decision to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a determination or decision, the conversion from video blocks to the bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the determination or decision.
[0629] Some embodiments of the disclosed technology include making a determination or decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in converting video blocks to the bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode disabled based on the determination or decision.
[0630] Figure 15 is a block diagram showing an exemplary video codec system 100 in which the techniques of the present disclosure can be utilized. As Figure 15 shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110, which may be referred to as a video encoding device, generates encoded video data. The destination device 120, which may be referred to as a video decoding device, may decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video codec 114, and an input / output (I / O) interface 116.
[0631] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. The video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly transmitted to target device 120 via I / O interface 116 over network 130a. The encoded video data may also be stored on storage medium / server 130b for access by target device 120.
[0632] Target device 120 may include I / O interface 126, video decoder 124, and display device 122.
[0633] Target device 120 may include I / O interface 126, video decoder 124, and display device 122.
[0634] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain the encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120, and target device 120 is configured to interface with an external display device.
[0635] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0636] Figure 16 is a block diagram showing an example of video encoder 200, and video encoder 200 may be Figure 15 the video encoder 114 in the system 100 shown.
[0637] Video codec 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 16 an example, video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0638] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0639] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0640] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in the examples for the purpose of explanation. Figure 18 for the purpose of explanation.
[0641] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0642] Mode selection unit 203 may select, for example, one of an intra or inter coding mode based on an error result, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block to be used as a reference picture. In some examples, mode selection unit 203 may select an intra and inter prediction combination (CIIP) mode, in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy).
[0643] To perform inter prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine the predicted video block for the current video block based on the motion information and decoded samples of a picture from buffer 213 other than the picture associated with the current video block.
[0644] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0645] In some examples, the motion estimation unit 204 may perform uni - directional prediction on a current video block, and the motion estimation unit 204 may search for a reference picture in list 0 or list 1 of reference video blocks for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0646] In other examples, the motion estimation unit 204 may perform bi - directional prediction on a current video block. The motion estimation unit 204 may search for a reference picture in list 0 of reference video blocks for the current video block and may also search for another reference picture in list 1 of reference video blocks for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0647] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.
[0648] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0649] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0650] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0651] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector predication (AMVP) and Merge mode signaling.
[0652] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0653] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0654] In other examples, such as in the skip mode, the current video block may not have residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0655] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0656] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0657] The dequantization unit 210 and the inverse transform unit 211 may apply dequantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0658] After the reconstruction unit 212 reconstructs the video block, a loop filter operation may be performed to reduce video block artifacts in the video block.
[0659] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0660] Figure 13 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 114 in the system 100 shown Figure 15 as such.
[0661] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 13 an example, the video decoder 300 includes multiple functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0662] In Figure 13 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process that is generally opposite to the encoding process described for the video encoder 200 ( Figure 16 ).
[0663] The entropy decoding unit 301 may obtain an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and based on the entropy-decoded video data, the motion compensation unit 302 may determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and Merge modes.
[0664] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0665] The motion compensation unit 302 may use the interpolation filter used by the video encoder 20 during video block encoding to calculate interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.
[0666] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0667] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 303 applies an inverse transform.
[0668] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces the decoded video for presentation on a display device.
[0669] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation or the coded / decoded representation of the current video block may, for example, correspond to bits juxtaposed or distributed at different positions in the bitstream, as defined by the syntax. For example, a video block may be encoded according to transform and coded / decoded error residual values and may also be coded / decoded using bits in the headers and other fields in the bitstream. Furthermore, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solutions. Similarly, the encoder may determine whether to include certain syntax fields and may accordingly generate the coded / decoded representation by including or excluding the syntax fields from the coded / decoded representation.
[0670] Figure 14FIG. 0 is a block diagram illustrating an example video processing system 2000 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 2000. System 2000 may include an input 2002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 2002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0671] System 2000 may include a codec component 2004, which may implement various codec or encoding methods described in this document. The codec component 2004 may reduce the average bit rate of the video from the output of input 2002 to the output of the codec component 2004 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 2004 may be stored or transmitted via a connected communication (represented by component 2006). Component 2008 may use the bitstream (or coded) representation of the video stored or communicated at input 2002 to generate pixel values or a displayable video to be sent to a display interface 2010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, although a particular video processing operation is referred to as a "codec" operation or tool, it should be understood that the codec tool or operation is used at the encoder and the corresponding decoding tool or operation that will reverse the codec result will be performed by the decoder.
[0672] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0673] Figure 17is a flowchart of an exemplary method 1700 for video processing. Method 1700 includes: performing (1702) a conversion between a video block of a video and a coded representation of the video, where the coded representation conforms to format rules, where the format rules specify the applicability of a transform skip mode to the video block based on coding conditions of the video block, where the format rules specify omitting a syntax element indicating the applicability of the transform skip mode from the coded representation, and where the transform skip mode includes skipping applying a forward transform to at least some coefficients before encoding into the coded representation or skipping applying an inverse transform to at least some coefficients during decoding before decoding from the coded representation.
[0674] Figure 18 is a flowchart of an exemplary method 1800 for video processing. Method 1800 includes: determining (1802) whether to use a loop filter or a post - reconstruction filter for a conversion between two adjacent video blocks of a video and a coded representation of the video, depending on whether a forward transform or an inverse transform is used for the conversion, where the forward transform includes skipping applying the forward transform to at least some coefficients before encoding into the coded representation or skipping applying the inverse transform to at least some coefficients during decoding before decoding from the coded representation; and performing (1804) the conversion based on the use of the loop filter or the post - reconstruction filter.
[0675] Figure 19 is a flowchart of an exemplary method 1900 for video processing. Method 1900 includes: determining (1902) to use a scaling tool for a conversion between a video block of a video and a coded representation of the video because a block differential pulse codec modulation (BDPCM) codec tool or a quantization residual BDPCM (QR - BDPCM) codec tool is used for the conversion; and performing (1904) the conversion when the scaling tool is used, where a syntax element in the coded representation indicates the use of the scaling tool, and where the use of the scaling includes scaling at least some coefficients representing the video block during encoding or de - scaling at least some coefficients from the coded representation during decoding.
[0676] Figure 20is a flowchart of an exemplary method 2000 for video processing. Method 2000 includes: for the conversion between a video block of a video and the coded representation of the video, determining (2002) to prohibit the use of a scaling tool due to a block differential pulse coding modulation (BDPCM) coding / decoding tool or a quantization residual BDPCM (QR-BDPCM) coding / decoding tool for the conversion of the video block; and performing (2004) the conversion without using the scaling tool, wherein the use of the scaling includes: scaling at least some coefficients representing the video block during encoding or de-scaling at least some coefficients from the coded representation during decoding.
[0677] Figure 21 is a flowchart of an exemplary method 2100 for video processing. Method 2100 includes: for the conversion between a video block of a video and the coded representation of the video, selecting (2102) a scaling matrix based on a transform matrix selected for the conversion, wherein the scaling matrices are used to scale at least some coefficients of the video blocks, and wherein the transform matrices are used to transform the at least some coefficients of the video blocks during the conversion; and performing (2104) the conversion using the scaling matrices.
[0678] Figure 22 is a flowchart of an exemplary method 2200 for video processing. Method 2200 includes: determining (2202) whether to apply a scaling matrix according to a rule based on whether to apply a quadratic transform matrix to a portion of a video block of a video, wherein the scaling matrix is used to scale at least some coefficients of the video block, and wherein the quadratic transform matrix is used to transform at least some residual coefficients of the portion of the video block during the conversion; and performing (2204) the conversion between the video block of the video and the bitstream representation of the video using the selected scaling matrix.
[0679] Figure 23 is a flowchart of an exemplary method 2300 for video processing. Method 2300 includes: determining (2302) a scaling matrix used in the conversion between a video block of a video having a non-square shape and the coded representation of the video, wherein a syntax element in the coded representation signals the scaling matrix, and wherein the scaling matrix is used to scale at least some coefficients of the video block during the conversion; and performing (2304) the conversion based on the scaling matrix.
[0680] Figure 24is a flowchart of an exemplary method 2400 for video processing. Method 2400 includes performing (2402) a conversion between a video block of a video and an encoded / decoded representation of the video, wherein, based on a rule, the video block includes a first number of positions at which a scaling matrix is applied during the conversion, and the video block further includes a second number of positions at which the scaling matrix is not applied during the conversion.
[0681] Figure 25 is a flowchart of an exemplary method 2500 for video processing. Method 2500 includes: determining (2502) a scaling matrix to be applied during a conversion between a video block of a video and an encoded / decoded representation of the video; and performing the conversion based on the scaling matrix, wherein the encoded / decoded representation indicates a number of elements in the scaling matrix, and wherein the number depends on whether coefficient zeroing is applied to the video block.
[0682] Figure 26 is a flowchart of an exemplary method 2600 for video processing. Method 2600 includes: performing (2602) a conversion between a video block of a video and an encoded / decoded representation of the video according to a rule, wherein, after applying a K×L transformation matrix to transform coefficients of the video block and after zeroing all transform coefficients except for upper-left M×N transform coefficients, the video block is represented into the encoded / decoded representation, wherein the encoded / decoded representation is configured to exclude signaling of elements of a scaling matrix at positions corresponding to the zeroed ones, wherein the scaling matrix is used to scale the transform coefficients.
[0683] Figure 27 is a flowchart of an exemplary method 2700 for video processing. Method 2700 includes determining (2702) based on a rule during a conversion between a video block of a video and an encoded / decoded representation of the video whether a single quantization matrix will be used based on the size of the video block, wherein all video blocks having the size use the single quantization matrix; and performing the conversion using the quantization matrix.
[0684] The following three sections describe exemplary video processing techniques with the following numbers:
[0685] Section A
[0686] 1. A video processing method, comprising: performing a conversion between a video block of a video and a codec representation of the video, wherein the codec representation conforms to format rules, wherein the format rules specify the applicability of a transform skip mode to the video block based on codec conditions of the video block, wherein the format rules specify omitting a syntax element indicating the applicability of the transform skip mode from the codec representation, and wherein the transform skip mode includes skipping application of a forward transform to at least some coefficients before encoding into the codec representation, or skipping application of an inverse transform to at least some coefficients before decoding from the codec representation during decoding.
[0687] 2. The method according to Example 1, wherein the transform skip mode is determined to be enabled because the codec conditions of the video block indicate use of block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR-BDPCM) for the video block.
[0688] 3. The method according to Example 1, wherein the transform skip mode is determined to be disabled because the codec conditions of the video block indicate use of non-block differential pulse codec modulation (non-BDPCM) or non-quantized residual BDPCM (non-QR-BDPCM) for the video block.
[0689] 4. A video processing method, comprising: for conversion between two adjacent video blocks of a video and a codec representation of the video, determining whether to use a loop filter or a post-reconstruction filter for the conversion based on whether to use a forward transform or an inverse transform for the conversion, wherein the forward transform includes skipping application of the forward transform to at least some coefficients before encoding into the codec representation, or skipping application of the inverse transform to at least some coefficients before decoding from the codec representation during decoding; and performing the conversion based on use of the loop filter or the post-reconstruction filter.
[0690] 5. The method according to Example 4, wherein the forward transform or the inverse transform includes a transform skip mode or block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR-BDPCM) or a palette mode, and wherein use of the loop filter or the post-reconstruction filter for the two adjacent video blocks is based on whether the two adjacent video blocks use a transform skip mode or block differential pulse codec modulation (BDPCM) or quantized residual BDPCM (QR-BDPCM) or a palette mode.
[0691] 6. The method according to Example 4, wherein the forward transform or the inverse transform includes a transform skip mode, and wherein derivation of a boundary filtering strength depends on one or more syntax elements indicating whether the transform skip mode is enabled for one or both of the two adjacent video blocks.
[0692] 7. The method according to Example 4, wherein the forward transform or inverse transform includes a transform skip mode, and wherein the deblocking filter, sample adaptive offset, adaptive loop filter, or post-reconstruction filter is disabled in response to samples located at the two adjacent video blocks being encoded or decoded using the transform skip mode.
[0693] 8. The method according to Example 7, wherein the forward transform or inverse transform includes a transform skip mode, and wherein the loop filter and post-reconstruction filter are not applied to an edge between the two adjacent video blocks in response to the transform skip mode being enabled for the two adjacent video blocks.
[0694] 9. The method according to Example 7, wherein the forward transform or inverse transform includes a transform skip mode, and wherein the loop filter and post-reconstruction filter are not applied to samples between the two adjacent video blocks in response to the transform skip mode being enabled for one of the two adjacent video blocks.
[0695] 10. The method according to Example 4, wherein the forward transform or inverse transform includes a transform skip mode, and wherein the samples are filtered using a filter other than the loop filter or post-reconstruction filter in response to the transform skip mode being enabled for the two adjacent video blocks.
[0696] 11. The method according to Example 10, wherein the filter includes a smoother filter.
[0697] 12. The method according to Example 4, wherein the video includes video blocks encoded using pulse code modulation (PCM) or block differential pulse code modulation (BDPCM) or quantized residual BDPCM (QR-BDPCM) or another type of mode that does not apply a forward transform or inverse transform to the video block, and wherein whether to use the loop filter or post-reconstruction filter for the conversion of the video block is determined in the same manner as for the two adjacent video blocks when the transform skip mode is enabled for the two adjacent video blocks.
[0698] 13. The method according to any of Examples 4-12, wherein the loop filter includes a deblocking filter.
[0699] 14. The method according to any of Examples 4-12, wherein the post-reconstruction filter includes a bilateral filter or a diffusion filter.
[0700] 15. The method according to any of Examples 1 to 14, wherein the conversion includes encoding the video block into the codec representation.
[0701] 16. The method according to any one of Examples 1 to 14, wherein the conversion includes decoding the codec representation to generate pixel values of the video block.
[0702] 17. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of Examples 1 to 16.
[0703] 18. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method described in any one of Examples 1 to 17.
[0704] Section B
[0705] 1. A video processing method, comprising: determining a factor of a scaling tool based on a codec mode of a video block for conversion between the video block of a video and a codec representation of the video; and performing the conversion using the scaling tool, wherein use of the scaling tool includes: scaling at least some coefficients representing the video block during encoding or unscaling at least some coefficients from the codec representation during decoding.
[0706] 2. The method according to Example 1, further comprising: determining a factor of the scaling tool based on a predefined value in response to using a block differential pulse codec modulation (BDPCM) codec tool or a quantization residual BDPCM (QR - BDPCM) codec tool for the conversion of the video block.
[0707] 3. The method according to Example 2, wherein the factor of the scaling tool for a video block to which the BDPCM codec tool or the QR - BDPCM codec tool is applied is the same as the factor of the scaling tool for a video block to which a transform skip mode is applied, and wherein the transform skip mode includes skipping application of a forward transform to at least some coefficients before encoding into the codec representation, or skipping application of an inverse transform to at least some coefficients before decoding from the codec representation during decoding.
[0708] 4. The method according to Example 2, wherein the conversion includes determining a factor of the scaling tool based on one or more transforms applied to at least some coefficients of the video block during the conversion.
[0709] 5. The method according to Example 4, wherein the scaling tool is allowed for the conversion in response to application of the one or more transforms to at least some coefficients of the video block.
[0710] 6. The method according to Example 4, wherein the technique for determining the factor of the scaling tool is the same as that for an intra - codec block in response to application of the one or more transforms to at least some coefficients of the video block.
[0711] 7. The method according to Example 1, wherein, for a video block of the video encoded and decoded using an intra block copy mode and an inter mode, the factors of the scaling matrix are determined in the same manner.
[0712] 8. A video processing method, comprising: determining to prohibit the use of a scaling tool due to a block differential pulse codec modulation (BDPCM) codec tool or a quantization residual BDPCM (QR-BDPCM) codec tool for the conversion of a video block of a video to and from an encoded and decoded representation of the video; and performing the conversion without using the scaling tool, wherein the use of the scaling includes: scaling at least some coefficients representing the video block during encoding or unscaling at least some coefficients from the encoded and decoded representation during decoding.
[0713] 9. A video processing method, comprising: selecting a scaling matrix based on a transform matrix selected for the conversion of a video block of a video to and from an encoded and decoded representation of the video, wherein the scaling matrices are used to scale at least some coefficients of the video blocks, and wherein the transform matrices are used to transform the at least some coefficients of the video blocks during the conversion; and performing the conversion using the scaling matrices.
[0714] 10. The method according to Example 9, wherein the selection of the scaling matrix is based on whether the conversion of the video block uses a default transform mode.
[0715] 11. The method according to Example 10, wherein the default transform mode includes a discrete cosine transform 2 (DCT2).
[0716] 12. The method according to Example 9, wherein each scaling matrix is signaled separately for a plurality of transform matrices.
[0717] 13. The method according to Example 9, wherein the selection of the scaling matrix is based on the motion information of the video block.
[0718] 15. The method according to Example 13, wherein the selection of the scaling matrix is based on whether the conversion of the video block uses a sub-block codec mode.
[0719] 15. The method according to Example 14, wherein the sub-block codec mode includes an affine codec mode.
[0720] 16. The method according to Example 15, wherein the scaling matrix for the affine codec mode is signaled differently compared to the scaling matrix of another video block whose conversion uses a non-affine mode.
[0721] 17. The method according to Example 13, wherein the selection of the scaling matrix is based on whether the video block is encoded or decoded using an affine intra prediction mode.
[0722] 18. The method according to any one of Examples 1 to 17, wherein the transformation includes encoding the (one or more) video blocks into the codec representation.
[0723] 19. The method according to any one of Examples 1 to 17, wherein the transformation includes decoding the codec representation to generate pixel values of the (one or more) video blocks.
[0724] 20. A video decoding apparatus, comprising a processor configured to perform the method described in one or more of Examples 1 to 19.
[0725] 21. A video encoding apparatus, comprising a processor configured to perform the method described in one or more of Examples 1 to 19.
[0726] 22. A computer program product having computer code stored thereon, the code causing the processor to perform the method described in any one of Examples 1 to 19 when executed by the processor.
[0727] Section C
[0728] 1. A video processing method, comprising: determining whether to apply a scaling matrix based on a rule according to whether a quadratic transformation matrix is applied to a portion of a video block of a video, wherein the scaling matrix is used to scale at least some coefficients of the video block, and wherein the quadratic transformation matrix is used to transform at least some residual coefficients of the portion of the video block during the transformation; and performing a transformation between the video block of the video and a bitstream representation of the video using the selected scaling matrix.
[0729] 2. The method according to Example 1, wherein the rule specifies applying the scaling matrix to the MxN upper left portion of the video block in response to applying the quadratic transformation matrix to the MxN upper left portion of the video block having a KxL transform block size.
[0730] 3. The method according to Example 1, wherein the scaling matrix is signaled in the bitstream representation.
[0731] 4. The method according to Example 3, wherein the scaling matrix is signaled in a slice group header, a strip header, a picture parameter set (PPS), a video parameter set (VPS), or a sequence parameter set (SPS) for the quadratic transformation matrix or for a reduced quadratic transformation or a rotation transformation.
[0732] 5. The method according to Example 1, wherein the bitstream representation includes a first syntax element indicating whether to apply the scaling matrix, and wherein the bitstream representation includes a second syntax element indicating whether to apply the secondary transformation matrix.
[0733] 6. The method according to Example 1, wherein the rule specifies to apply the scaling matrix only to a portion of the video block to which the secondary transformation matrix is not applied.
[0734] 7. The method according to Example 1, wherein the rule specifies to apply the scaling matrix to a portion of the video block other than the top-left MxN portion to which the secondary transformation matrix is applied.
[0735] 8. The method according to Example 1, wherein the rule specifies to apply the scaling matrix only to a portion of the video block to which the secondary transformation matrix is applied.
[0736] 9. A video processing method, comprising: determining, for a video block having a non-square shape, a scaling matrix used in a conversion between the video block of the video and a codec representation of the video, wherein a syntax element in the codec representation signals the scaling matrix, and wherein the scaling matrix is used to scale at least some coefficients of the video block during the conversion; and performing the conversion based on the scaling matrix.
[0737] 10. The method according to Example 9, wherein the syntax element is predictive coded for the scaling matrix by another scaling matrix of a previous square block of the video.
[0738] 11. A video processing method, comprising: performing a conversion between a video block of a video and a codec representation of the video, wherein, based on a rule, the video block includes a first number of positions at which a scaling matrix is applied during the conversion, and the video block further includes a second number of positions at which the scaling matrix is not applied during the conversion.
[0739] 12. The method according to Example 11, wherein the first number of positions includes the top-left M*N positions of the video block, and wherein the video block includes more than M*N positions.
[0740] 13. The method according to Example 11, wherein the first number of positions includes the M*N positions at the top of the video block, and wherein the video block includes more than M*N positions.
[0741] 14. The method according to Example 11, wherein the first number of positions includes the M*N positions on the left side of the video block, and wherein the video block includes more than M*N positions.
[0742] 15. A video processing method includes: determining a scaling matrix to be applied during the conversion between a video block of a video and the codec representation of the video; and performing the conversion based on the scaling matrix, wherein the codec representation indicates the number of elements in the scaling matrix, and wherein the number depends on whether coefficient zeroing is applied to the video block.
[0743] 16. The method according to example 15, wherein for a 64x64 transform, the conversion includes zeroing all positions of the video block except for the top-left MxN positions, and wherein the number of elements of the scaling matrix is M / 8 * N / 8.
[0744] 17. A video processing method includes: performing a conversion between a video block of a video and the codec representation of the video according to a rule, wherein after applying a KxL transform matrix to the transform coefficients of the video block and after zeroing all transform coefficients except for the top-left MxN transform coefficients, the video block is represented into the codec representation, wherein the codec representation is configured to exclude signaling of elements of the scaling matrix at positions corresponding to the zeroed positions, wherein the scaling matrix is used to scale the transform coefficients.
[0745] 18. The method according to example 17, wherein signaling of elements of the scaling matrix in a region outside the top-left MxN coefficients is skipped.
[0746] 19. The method according to example 17, wherein the scaling matrix is applied by subsampling according to a ratio determined by K and / or L.
[0747] 20. The method according to example 19, wherein the video block is divided into a plurality of sub-regions, wherein the size of each sub-region is Uw * Uh, and wherein one element of the scaling matrix in each sub-region of the region including the top-left MxN coefficients of the video block is signaled in the codec representation.
[0748] 21. The method according to example 19, wherein the number of elements of the scaling matrix indicated in the codec representation is based on the aforementioned M and / or N.
[0749] 22. The method according to example 19, wherein a first number of elements of the scaling matrix indicated in the codec representation for the KxL transform matrix is different from a second number of elements of the scaling matrix indicated in the codec representation for the top-left MxN coefficients without applying zeroing.
[0750] 23. The method according to example 17, wherein the scaling matrix is applied by subsampling according to a ratio determined by M and / or N.
[0751] 24. The method according to Example 23, wherein a region including the upper-left MxN coefficients is divided into a plurality of sub-regions, wherein the size of each sub-region is Uw*Uh, and wherein one element within each sub-region is signaled in the codec representation.
[0752] 25. The method according to Example 23, wherein the number of elements of the scaling matrix indicated in the codec representation for the K×L transform matrix is the same as that indicated in the codec representation for the upper-left MxN coefficients without zeroing.
[0753] 26. The method according to any of Examples 17 to 25, wherein K = L = 64, and wherein M = N = 32.
[0754] 27. The method according to any of Examples 20 or 24, wherein K = L = 64, wherein M = N = 32, and wherein Uw = Uh = 8.
[0755] 28. A video processing method, comprising: determining, based on a rule, whether to use a single quantization matrix based on the size of a video block during conversion between the video block of a video and the codec representation of the video, wherein all video blocks having the size use the single quantization matrix; and performing the conversion using the quantization matrix.
[0756] 29. The method according to Example 28, wherein the rule stipulates that only a single quantization matrix is allowed in response to the size of the video block being less than WxH, where W and H are integers.
[0757] 30. The method according to Example 28, wherein the rule stipulates that only a single quantization matrix is allowed in response to the width of the video block being less than a threshold.
[0758] 31. The method according to Example 28, wherein the rule stipulates that only a single quantization matrix is allowed in response to the height of the video block being less than a threshold.
[0759] 32. The method according to Example 28, wherein the rule stipulates that a single quantization matrix is not applied to video blocks having a size associated with small-sized video blocks.
[0760] 33. The method according to any of Examples 1 to 32, wherein a palette mode is applied to the video block.
[0761] 34. The method according to any of Examples 1 to 33, wherein the conversion includes encoding the video into the codec representation.
[0762] 35. The method according to any one of Examples 1 to 33, wherein the conversion includes decoding the codec representation to generate pixel values of the video block.
[0763] 36. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of Examples 1 to 35.
[0764] 37. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of Examples 1 to 35.
[0765] 38. A computer program product having computer code stored thereon, the code causing the processor to implement the method described in any one of Examples 1 to 35 when executed by the processor.
[0766] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a substance composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including (e.g.) programmable processors, computers, or multiple processors or computers. In addition to the hardware, the apparatus may also include code for creating an execution environment for the computer program under consideration, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver device.
[0767] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). The computer program can be deployed to execute on one or more computers, which can be located at one site or distributed across multiple sites and interconnected by a communication network.
[0768] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs, thereby performing functions by operating on input data and generating output. These processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0769] For example, processors suitable for executing a computer program include general and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally speaking, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including (e.g.) semiconductor storage devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0770] Although this patent document contains many details, it should not be construed as limiting any subject matter or the scope of any claims, but rather as a description of specific features of particular embodiments of a particular technology. Certain features described in the context of individual embodiments of this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although certain features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be excluded from that combination, and the claimed combination may cover a sub-combination or a variation of a sub-combination.
[0771] Similarly, although the operations are described in a particular order in the figures, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order to obtain the desired result, or that all illustrated operations be performed. Additionally, the partitioning of various system components in the embodiments described in this patent document should not be construed as required in all embodiments.
[0772] Only a few embodiments and examples are described, and other embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: For the conversion between the first video block of the video and the bitstream of the video, due to the differential codec tool being applied to the first video block, determining to prohibit the use of an explicit scaling matrix in the scaling tool, where the explicit scaling matrix represents a scaling matrix whose information is explicitly signaled; and Performing the conversion based on the determination, wherein the scaling tool includes: During encoding, scaling at least some coefficients representing the first video block; or During decoding, descaling at least some coefficients from the bitstream, wherein the second video block encoded and decoded using the intra-block copy mode in the video region and the third video block encoded and decoded using the inter-frame mode in the video region share the same set of explicit scaling matrices, wherein the video region is a stripe or a picture.
2. The method according to claim 1, wherein, For the conversion of the first video block, a flat scaling matrix is used in the scaling tool.
3. The method according to claim 2, wherein, The value of the flat scaling matrix is equal to 16.
4. The method according to claim 1, wherein, For the second video block encoded and decoded using the intra-block copy mode and the third video block encoded and decoded using the inter-frame mode, the value of the explicit scaling matrix is determined in the same way.
5. The method according to claim 1, wherein, The set of explicit scaling matrices is derived from the scaling list data syntax structure in the bitstream.
6. The method according to claim 5, wherein, The scaling list data syntax structure includes at least one of a first syntax element indicating whether the value of the scaling list is predicted from a reference scaling list, a second syntax element indicating the reference scaling list, a third syntax element indicating the value of the DC position of the explicit scaling matrix, or a fourth syntax element indicating the difference between the current matrix coefficient and the previous matrix coefficient.
7. The method according to claim 1, wherein, The differential codec tool includes the block differential pulse codec modulation mode.
8. The method according to claim 1, wherein, The conversion includes encoding the first video block into the bitstream.
9. The method according to claim 1, wherein, The conversion includes decoding the first video block from the bitstream.
10. The method according to claim 1, further comprising: For the conversion between the video block of the video and the codec representation of the video, determining the factor of the scaling tool based on the codec mode of the video block; and Performing the conversion using the scaling tool, wherein the use of the scaling tool includes: During encoding, scaling at least some coefficients representing the video block, or During decoding, descaling at least some coefficients from the codec representation.
11. The method according to claim 10, further comprising: Determining the factor of the scaling tool based on a predetermined value in response to using the block differential pulse coding modulation (BDPCM) codec tool or the quantization residual BDPCM (QR-BDPCM) codec tool for the conversion of the video block.
12. The method according to claim 11, wherein, a factor of a scaling tool for a video block to which a BDPCM codec tool or a QR - BDPCM codec tool is applied is the same as a factor of a scaling tool for a video block to which a transform skip mode is applied, and wherein, the transform skip mode includes skipping applying a forward transform to at least some coefficients during encoding into the codec representation, or skipping applying an inverse transform to at least some coefficients during decoding before decoding from the codec representation.
13. The method according to claim 11, wherein, the transformation includes determining a factor of the scaling tool based on one or more transforms applied to at least some coefficients of the video block during the transformation.
14. The method according to claim 13, wherein, in response to applying the one or more transforms to at least some coefficients of the video block, the scaling tool is allowed for the transformation.
15. The method according to claim 13, wherein, in response to applying the one or more transforms to at least some coefficients of the video block, a technique for determining a factor of the scaling tool is the same as a technique for an intra - coded block.
16. The method according to claim 10, wherein, for a video block of the video encoded and decoded using an intra - block copy mode and an inter - frame mode, a factor of the scaling matrix is determined in the same manner.
17. The method according to claim 1, further comprising: determining a prohibition on using a scaling tool for a conversion between a video block of a video and a codec representation of the video due to a block differential pulse code modulation (BDPCM) codec tool or a quantization residual BDPCM (QR - BDPCM) codec tool for the conversion of the video block; and performing the conversion without using the scaling tool, wherein, the use of the scaling tool includes: scaling at least some coefficients representing the video block during encoding, or descaling at least some coefficients from the codec representation during decoding.
18. The method according to claim 1, further comprising: selecting a scaling matrix based on a transform matrix selected for the conversion for a conversion between a video block of a video and a codec representation of the video, wherein, the scaling matrix is used to scale at least some coefficients of the video block, and wherein, the transform matrix is used to transform at least some coefficients of the video block during the conversion; and performing the conversion using the scaling matrix.
19. The method according to claim 18, wherein, selecting the scaling matrix is based on whether the conversion of the video block uses a default transform mode.
20. The method according to claim 19, wherein, the default transform mode includes discrete cosine transform 2 (DCT2).
21. The method according to claim 18, wherein, signaling the scaling matrix separately for multiple transform matrices.
22. The method according to claim 18, wherein, selecting the scaling matrix is based on motion information of the video block.
23. The method according to claim 22, wherein, Select the scaling matrix based on whether the transformation of the video block uses a sub-block codec mode.
24. The method according to claim 23, wherein, the sub-block codec mode includes an affine codec mode.
25. The method according to claim 24, wherein, signaling the scaling matrix of the affine codec mode differently compared to the scaling matrix of another video block whose transformation uses a non-affine codec mode.
26. The method according to claim 22, wherein, select the scaling matrix based on whether the video block is coded using an affine intra prediction mode.
27. The method according to any one of claims 10 to 26, wherein, the transformation includes encoding the video block into the codec representation.
28. The method according to any one of claims 10 to 26, wherein, the transformation includes decoding the codec representation to generate pixel values of the video block.
29. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: For the transformation between a first video block of a video and the bitstream of the video, determine to prohibit the use of an explicit scaling matrix in a scaling tool due to a differential codec tool being applied to the first video block, where the explicit scaling matrix represents a scaling matrix whose information is explicitly signaled; and Perform the transformation based on the determination, where the scaling tool includes: During encoding, scale at least some coefficients representing the first video block; or During decoding, descale at least some coefficients from the bitstream, where a second video block coded using an intra block copy mode in a video region and a third video block coded using an inter mode in the video region share the same set of explicit scaling matrices, where the video region is a strip or a picture.
30. The apparatus according to claim 29, wherein, For the transformation of the first video block, use a flat scaling matrix in the scaling tool, and wherein the value of the flat scaling matrix is equal to 16.
31. The apparatus according to claim 29, wherein, For a second video block coded using an intra block copy mode and a third video block coded using an inter mode, determine the value of the explicit scaling matrix in the same manner; where the set of explicit scaling matrices is derived from a scaling list data syntax structure in the bitstream; where the scaling list data syntax structure includes at least one of a first syntax element indicating whether the value of the scaling list is predicted from a reference scaling list, a second syntax element indicating the reference scaling list, a third syntax element indicating the value of the DC position of the explicit scaling matrix, or a fourth syntax element indicating the difference between the current matrix coefficient and the previous matrix coefficient; and where the differential codec tool includes a block differential pulse codec modulation mode.
32. A non-transitory computer-readable storage medium storing instructions that cause a processor to: For the conversion between the first video block of a video and the bitstream of the video, due to a differential coding / decoding tool being applied to the first video block, it is determined that the use of an explicit scaling matrix in a scaling tool is prohibited, wherein, the explicit scaling matrix represents a scaling matrix whose information is explicitly signaled; and perform the conversion based on the determination, wherein the scaling tool includes: during encoding, scale at least some coefficients representing the first video block; or during decoding, descale at least some coefficients from the bitstream, wherein a second video block encoded / decoded using an intra-block copy mode in a video region and a third video block encoded / decoded using an inter mode in the video region share the same set of explicit scaling matrices, wherein the video region is a slice or a picture.
33. The non-transitory computer-readable storage medium according to claim 32, wherein, for the conversion of the first video block, use a flat scaling matrix in the scaling tool, and wherein the value of the flat scaling matrix is equal to 16.
34. The non-transitory computer-readable storage medium according to claim 32, wherein, for a second video block encoded / decoded using an intra-block copy mode and a third video block encoded / decoded using an inter mode, determine the value of the explicit scaling matrix in the same manner; wherein the set of explicit scaling matrices is derived from a scaling list data syntax structure in the bitstream; wherein the scaling list data syntax structure includes at least one of a first syntax element indicating whether the value of the scaling list is predicted from a reference scaling list, a second syntax element indicating the reference scaling list, a third syntax element indicating the value of the DC position of the explicit scaling matrix, or a fourth syntax element indicating the difference between the current matrix coefficient and the previous matrix coefficient; and wherein the differential coding / decoding tool includes a block differential pulse coding / decoding modulation mode.
35. A method for storing a bitstream of a video, comprising: for a first video block of a video, due to a differential coding / decoding tool being applied to the first video block, determine that the use of an explicit scaling matrix in a scaling tool is prohibited, wherein the explicit scaling matrix represents a scaling matrix whose information is explicitly signaled; perform the bitstream based on the determination; and store the bitstream in a non-transitory computer-readable storage medium, wherein the scaling tool includes: during encoding, scale at least some coefficients representing the first video block; or during decoding, descale at least some coefficients from the bitstream, wherein a second video block encoded / decoded using an intra-block copy mode in a video region and a third video block encoded / decoded using an inter mode in the video region share the same set of explicit scaling matrices, wherein the video region is a slice or a picture.
36. A video decoding device, comprising a processor configured to implement the method according to any one of claims 10 to 28.
37. A video encoding device, comprising a processor configured to implement the method according to any one of claims 10 to 28.
38. A computer-readable medium having computer code stored thereon, the code, when executed by a processor, causing the processor to perform the method recited in any one of claims 10 to 28.