Separable secondary transform processing of coded video
By introducing secondary transformation technology into video encoding and decoding, the problem of increasing bandwidth demand for digital video has been solved, enabling more efficient video compression and encoding/decoding, and improving video quality and encoding/decoding efficiency.
Patent Information
- Application Number
- CN202080083999.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-02
- Filing Date
- 2020-12-02
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-12-02
AI Technical Summary
Digital video consumes a significant amount of bandwidth on the Internet and other digital communication networks. As the number of connected user devices increases, bandwidth demand will continue to grow, and existing video encoding and decoding technologies struggle to effectively optimize video compression efficiency.
The system employs a quadratic transformation technique, including a separable quadratic transformation tool, a scan region-based transformation tool, transformation matrix selection, dimension reduction transformation, zeroing rules, and conditional transformations, to convert between video blocks and bitstream representations, thereby optimizing the video encoding and decoding process.
By optimizing the video encoding and decoding process, video compression efficiency has been improved, bitrate requirements have been reduced, and video quality and encoding/decoding efficiency have been enhanced.
Smart Images

Figure CN115066899B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application is related to and claims priority to International Patent Application PCT / CN2019 / 122366, filed December 2, 2019, under the Paris Convention under applicable patent laws and / or rules. The entire disclosure of the above application is incorporated by reference as part of the disclosure of this application for all purposes. TECHNICAL FIELD
[0003] This patent document relates to video coding and decoding techniques, devices, and systems. BACKGROUND
[0004] Despite advances in video compression, digital video accounts for the largest bandwidth use on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is expected that bandwidth demand for digital video usage will continue to grow. SUMMARY
[0005] This document describes various embodiments and techniques that use a secondary transform (also referred to as a low-frequency non-separable transform) during decoding or encoding of a video or image.
[0006] In one example aspect, a video processing method is disclosed. The method includes determining, for a conversion between a video unit of a video and a bitstream representation of the video, whether a separable secondary transform (SST) tool is enabled for the video unit. The method also includes performing the conversion based on the determination.
[0007] In another example aspect, a video processing method is disclosed. The method includes determining, for a conversion between a block of a video and a bitstream representation of the video, a manner of indicating a transform tool or a transform matrix used by the transform tool based on a right-bottom position (SRx, SRy) of a scan region. The method also includes performing the conversion based on the determination.
[0008] In another example aspect, a video processing method is disclosed. The method includes determining, for a conversion between a block of a video and a bitstream representation of the video, a transform matrix used in a separable secondary transform (SST) tool based on a characteristic of the block. The SST tool provides a set of transform matrices available for use. The method also includes performing the conversion based on the determination.
[0009] In another example aspect, a video processing method is disclosed. The method includes determining a constraint rule for selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the constraint rule. The secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block. The secondary transform with reduced dimension is applied in a particular order with a primary transform during the conversion.
[0010] In another example aspect, another video processing method is disclosed. The method includes determining a constraint rule for selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the constraint rule. The secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block. The secondary transform with reduced dimension is applied in a particular order with a primary transform during the conversion.
[0011] In yet another example aspect, another video processing method is disclosed. The method includes determining a zero-out rule for selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the zero-out rule. The secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block. The zero-out rule specifies a maximum number of coefficients used by the secondary transform with reduced dimension.
[0012] In yet another example aspect, another video processing method is disclosed. The method includes determining a zero-out rule for selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the zero-out rule. The secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block. The zero-out rule specifies a maximum number of coefficients used by the secondary transform with reduced dimension.
[0013] In yet another example aspect, another video processing method is disclosed. The method includes determining a condition for selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the condition. The secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block. The condition is signaled in the bitstream representation.
[0014] In yet another example aspect, another video processing method is disclosed. The method includes selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The conversion includes selectively applying a position-dependent intra prediction combination (PDPC) based on a coexistence rule.
[0015] In yet another example aspect, another video processing method is disclosed. The method includes applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The application controls usage of neighboring samples for intra prediction during the conversion.
[0016] In yet another example aspect, another video processing method is disclosed. The method includes selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The selective application controls usage of a quantization matrix during the conversion.
[0017] In yet another example aspect, another video processing method is disclosed. The method includes, for a conversion between a current video block of a video and a bitstream representation of the video, determining whether to use a separable secondary transform (SST) for the conversion based on a coding condition; and performing the conversion according to the determination.
[0018] In yet another example aspect, a video encoder is disclosed. The video encoder includes a processor configured to implement one or more of the above-described methods.
[0019] In yet another example aspect, a video decoder is disclosed. The video decoder includes a processor configured to implement one or more of the above-described methods.
[0020] In yet another example aspect, a computer readable medium is disclosed. The medium includes code stored thereon for implementing one or more of the above-described methods.
[0021] These and other aspects are described in the detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 An example of an encoder block diagram is shown.
[0023] Figure 2 An example of 67 intra prediction modes is shown.
[0024] Figures 3A-3B An example of reference samples for wide angle intra prediction is shown.
[0025] Figure 4 An example illustration of discontinuity issue in case of direction more than 45 degrees.
[0026] Figures 5A-5D An example illustration of samples used for PDPC application to diagonal and adjacent angular intra modes is shown.
[0027] Figure 6 An example of 4x8 and 8x4 block split.
[0028] Figure 7 An example of split for all blocks except 4x8, 8x4 and 4x4.
[0029] Figure 8 Splitting a block of 4x8 samples into two independent decodable regions.
[0030] Figure 9 An example order of processing pixel rows to maximize throughput for 4xN blocks with vertical predictors is shown.
[0031] Figure 10 An example of secondary transform is shown.
[0032] Figure 11 An example of proposed reduced secondary transform (RST) is shown.
[0033] Figure 12 An example of forward and inverse (or inverse) reduced transform is shown.
[0034] Figure 13 An example of forward RST 8x8 process with 16x48 matrix is shown.
[0035] Figure 14 An example of scanning positioning non-zero elements 17 to 64 is shown.
[0036] Figure 15 An illustration of sub-block transform modes SBT-V and SBT-H.
[0037] Figure 16 A block diagram of an example hardware platform for implementing the techniques described in this document is shown.
[0038] Figure 17 A flow diagram of an example method of video processing is shown.
[0039] Figure 18AAn example of coefficient coding based on scan region is illustrated.
[0040] Figure 18B Another example of coefficient coding based on scan region is illustrated.
[0041] Figure 19 is a block diagram of an example video processing system in which the disclosed technology can be implemented.
[0042] Figure 20 is a flowchart of an example method of video processing according to the present technology.
[0043] Figure 21 is a flowchart of another example method of video processing according to the present technology.
[0044] Figure 22 is a flowchart of another example method of video processing according to the present technology.
[0045] Figure 23 is a block diagram illustrating an example video coding system.
[0046] Figure 24 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0047] Figure 25 is a block diagram illustrating a decoder according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0048] Section headings are used in this document for ease of understanding and do not limit the embodiments disclosed in a section to only that section. Furthermore, while certain embodiments are described with reference to general video coding or other specific video codecs, the disclosed technology is also applicable to other video coding technologies. Moreover, while some embodiments describe video encoding steps in detail, it will be appreciated that corresponding decoding steps to undo the encoding will be implemented by a decoder. Furthermore, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at a different compression bit rate.
[0049] I. SUMMARY
[0050] This patent document relates to video coding technology. In particular, it is a related transform in video coding. It can be applied to existing video coding standards (such as HEVC), or the yet to be finalized standard (Versatile Video Coding). It can also be applicable to future video coding standards or video codecs.
[0051] 2. Preliminary Discussion
[0052] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, both organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding are utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Test Model (JEM). In April 2018, the Joint Video Expert Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created, which is dedicated to the VVC standard, aiming to reduce bitrates by 50% compared to HEVC.
[0053] 2.1 Color Space and Chroma Subsampling
[0054] A color space, also called a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a tuple of numbers, typically 3 or 4 values or color components (e.g. RGB). Basically, a color space is an elaborate exposition of a coordinate system and a subspace.
[0055] For video compression, the most commonly used color spaces are YCbCr and RGB.
[0056] YCbCr, Y'CbCr or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, CB and CR are the blue-difference and red-difference chroma components. Y' (with the apostrophe) is not the same as Y, which is luminance, meaning that the light intensity is based on gamma-corrected RGB primary non-linear encoding.
[0057] Chroma subsampling is the practice of encoding an image with lower resolution for chroma information than for luma by exploiting the lower acuity of the human visual system to color differences than to luminance.
[0058] 2.1.1 Format 4:4:4
[0059] Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used for high-end film scanners and film post-production.
[0060] 2.1.2 Format 4:2:2
[0061] The two chroma components are sampled at half the luma sampling rate: the horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by a third, with little visual difference.
[0062] 2.1.3 Format 4:2:0
[0063] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are only sampled on every other line in this scheme. The data rate is therefore the same. Cb and Cr are subsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, with different horizontal and vertical subsampling.
[0064] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are co-located between pixels in the vertical direction (interstitially co-located).
[0065] In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are interstitially co-located, halfway between alternate luma samples.
[0066] In 4:2:0 DV, Cb and Cr are co-located horizontally. In the vertical direction, they are co-located on alternate lines.
[0067] 2.2 Encoding and decoding process of typical video codecs
[0068] Figure 1 An example of the encoder block diagram of VVC is shown, which contains three in-loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a pre-defined filter, SAO and ALF exploit the original samples of the current picture by adding an offset and by applying a finite impulse response (FIR) filter, respectively, to reduce the mean square error between the original and the reconstructed samples, where the side information of the encoding signals the offset and the filter coefficients. ALF is located at the last processing stage of each picture and can be seen as a tool that tries to capture and fix artifacts created by previous stages.
[0069] 2.3 Intra mode coding with 67 intra prediction modes
[0070] To capture arbitrary edge directions present in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The additional directional modes are depicted as dashed arrows in Figure 2 while the planar and DC modes remain the same. These denser directional intra prediction modes are applied to all block sizes and for both luma and chroma intra prediction.
[0071] The conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in a clockwise direction, as shown in Figure 2 In VTM2, several of the conventional angular intra prediction modes are used for wide angle intra prediction modes for non-square blocks adaptively. The replaced modes are signaled using the original method and remapped to the indices of the wide angle modes after parsing. The total number of intra prediction modes (e.g. 67) is unchanged and the intra mode coding is unchanged.
[0072] In HEVC, each intra coded block has a square shape with each side length being a power of 2. Therefore, generating the intra predictor using the DC mode does not require division operations. In VTM2, the blocks can have rectangular shapes, which requires division operations in general case for each block. To avoid the division operations for DC prediction, only the longer side is used to calculate the average for non-square blocks.
[0073] 2.4 Wide angle intra prediction for non-square blocks
[0074] The conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in a clockwise direction. In VTM2, several of the conventional angular intra prediction modes are used for wide angle intra prediction modes for non-square blocks adaptively. The replaced modes are signaled using the original method and remapped to the indices of the wide angle modes after parsing. The total number of intra prediction modes (e.g. 67) for a certain block is unchanged and the intra mode coding is unchanged.
[0075] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined as shown in Figures 3A-3B
[0076] The number of modes for the replaced modes in the wide angle directional modes depends on the aspect ratio of the block. The replaced intra prediction modes for the wide angle modes are shown in Table 1.
[0077] Table 1 Intra prediction modes for wide angle mode replacement
[0078] Conditions Replacing Intra Prediction Modes W / H == 2 Modes 2, 3, 4, 5, 6, 7 W / H > 2 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H == 1 None H / W == 1 / 2 Modes 61, 62, 63, 64, 65, 66 H / W <1 / 2 Modes 57, 58, 59, 60, 61, 62, 63, 64, 65, 66
[0079] As Figure 4 As shown, in the case of wide-angle intra prediction, two vertically neighboring prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and a side smoothing are applied to wide-angle prediction to reduce the negative impact of the added gap Δp α .
[0080] 2.5 Position dependent intra prediction combination
[0081] In VTM2, the intra prediction result of the planar mode is further modified by a position dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that invokes a combination of HEVC-style intra prediction with unfiltered boundary reference samples and with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, left-bottom angular mode and its eight neighboring angular modes, and top-right angular mode and its eight neighboring angular modes.
[0082] The prediction sample pred(x, y) is predicted using a linear combination of reference samples and the intra prediction mode (DC, planar, angular) according to the following equation:
[0083] pred(x, y) = (wL x R -1,y + wT x R x,-1 - wTL x R -1,-1 + (64 - wL - wT + wTL) x pred(x, y) + 32) » 6
[0084] where R x,-1 , R -1,y denote the reference samples located at the top and left of the current sample (x, y), respectively, and R -1,-1 denotes the reference sample located at the top-left corner of the current block.
[0085] If PDPC is applied to the DC, planar, horizontal and vertical intra modes, no additional boundary filter is needed, which is necessary in the case of the HEVC DC mode boundary filter or the horizontal / vertical mode edge filter.
[0086] Figures 5A-5D The definition of the reference samples (R x,-1 , R -1,y and R -1,-1 ) applied to PDPC on various prediction modes is illustrated. The prediction sample pred(x', y') is located at (x', y') within the prediction block. The coordinates x of the reference sample R x,-1 is given by x = x' + y' + 1, and the coordinates y of the reference sample R -1,yThe coordinates y are similarly given by: y = x' + y' + 1.
[0087] Figures 5A to 5D A definition of the samples used by the PDPC applied to the diagonal and adjacent angular intra modes is provided.
[0088] The PDPC weights depend on the prediction mode as shown in Table 2.
[0089] Table 2. Example of PDPC weights according to the prediction mode
[0090]
[0091]
[0092] 2.6 Intra subblock partitioning (ISP)
[0093] In some embodiments, ISP is proposed to split a luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions depending on the block size dimension as shown in Table 3. Figure 6 and 7 An example is shown for both possibilities. All sub-partitions satisfy the condition of having at least 16 samples.
[0094] Table 3. Number of sub-partitions depending on the block size
[0095]
[0096] Figure 6 An example is shown for the splitting of 4x8 and 8x4 blocks.
[0097] Figure 7 An example is shown for the splitting of all blocks except 4x8, 8x4 and 4x4.
[0098] For each of these sub-partitions, the residual signal is generated by entropy decoding the coefficients sent by the encoder, then inverse vector quantization and inverse transform of them. Then, the sub-partition is intra predicted and finally the corresponding reconstructed samples are obtained by adding the residual signal to the predicted signal. Therefore, the reconstructed values of each sub-partition will be used to generate the prediction of the next sub-partition, which will repeat the process, and so on. All sub-partitions share the same intra mode.
[0099] Based on the intra mode and partitioning used, two different categories of processing order are used, called normal order and reverse order. In normal order, the first sub- partition to be processed is the one containing the top-left sample of the CU, then it continues either down (for horizontal partitioning) or to the right (for vertical partitioning). As a result, the reference samples used to generate the sub- partition prediction signal are only located on the top-left side of the line. On the other hand, the reverse processing order either starts from the sub-partition containing the bottom-left sample of the CU and continues up, or from the sub-partition containing the top-right sample of the CU and continues left.
[0100] 2.7 Block differential pulse-code modulation coding (BDPCM)
[0101] Due to the shape of the horizontal (resp. vertical) predictor using the left (resp. top) pixel of the current pixel to predict, the most throughput efficient way to process a block is to process all the pixels of a column (resp. line) in parallel and sequentially process these columns (resp. lines). To increase the throughput, the following procedure is introduced: when the predictor selected on the block is vertical, split the block of width 4 into two halves with a horizontal edge; when the predictor selected on the block is horizontal, split the block of height 4 into two halves with a vertical edge.
[0102] When a block is split, a sample from one region is not allowed to use a pixel from the other region to compute the prediction: if this happens, the predicted pixel is replaced by the reference pixel in the prediction direction. This is illustrated in Figure 8 for different positions of the current pixel X in a 4x8 block with vertical prediction.
[0103] Figure 8 An example of splitting a block of 4x8 samples into two independently decodable regions is shown.
[0104] Thanks to this property, it is now possible to process a 4x4 block in 2 loops, a 4x8 or 8x4 block in 4 loops, and so on, as illustrated in Figure 9
[0105] An example of the order in which to process the pixel rows to maximize the throughput of a 4xN block with vertical predictors is shown. Figure 9 Table 4 summarizes the number of loops required to process a block, depending on the block size. It is easy to show that any block with both dimensions greater than or equal to 8 can be processed with 8 pixels per loop or more loops.
[0106] Table 4 Worst case throughput for blocks of size 4xN, Nx4
[0107] Table 4 Worst case throughput for blocks of size 4xN, Nx4
[0108]
[0109] 2.8 Quantized residual domain BDPCM
[0110] In some embodiments, a quantized residual domain BDPCM (denoted as RBDPCM in the following) is proposed. Similar to intra prediction, an entire block is intra predicted by copying samples on the prediction direction (horizontal or vertical prediction). The residual is quantized and the delta between the quantized residual and its predictor (horizontal or vertical) is coded.
[0111] For a block of size M (rows) x N (columns), let r i,j (0≤i≤M-1, 0≤j≤N-1) be the prediction residual after intra prediction is performed horizontally (copying left neighbor pixel values of the prediction block line by line) or vertically (copying the top neighbor line to each line in the prediction block) using unfiltered samples from above or left block boundary samples. Let Q(r i, )(0≤i≤M-1, 0≤j≤N-1) denote the quantized version of the residual r i, , where the residual is the difference between the original block and the predicted block values. A block DPCM is then applied to the quantized residual samples, resulting in a modified MxN array with elements When vertical BDPCM is signaled:
[0112]
[0113] For horizontal prediction, similar rules apply, the residual quantized samples are obtained by:
[0114]
[0115] The residual quantized samples are sent to the decoder.
[0116] At the decoder side, the above calculations are reversed to produce Q(r i, )(0≤i≤M-1, 0≤j≤N-1). For the vertical prediction case,
[0117]
[0118] For the horizontal case,
[0119]
[0120] The inverse quantized residual Q -1 (Q(r i, ) is added to the intra block predicted values to produce the reconstructed sample values.
[0121] The main benefit of this scheme is that the inverse DPCM can be done on the fly during the parsing of the coefficients, by adding a predictor only at the parsing of the coefficients (it can also be performed after the parsing).
[0122] Transform skip is always used for quantized residual domain BDPCM.
[0123] 2.9 Multiple Transform Set (MTS) in VVC
[0124] In VTM4, large block size transforms up to 64x64 size are enabled, which is mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with size (width or height, or both width and height) equal to 64, high frequency transform coefficients are zeroed out, leaving only lower frequency coefficients. For example, for a transform block of MxN, M is the block width and N is the block height, when M is equal to 64, only the left 32 columns of transform coefficients are kept. Similarly, when N is equal to 64, only the top 32 rows of transform coefficients are kept. When transform skip mode is used for large blocks, the entire block is used without zeroing out any values.
[0125] In addition to DCT-II already adopted in HEVC, Multiple Transform Selection (MTS) scheme is used for residual coding of inter and intra coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII.
[0126] The following table shows the basis functions of the selected DST / DCT.
[0127]
[0128] In order to maintain the orthogonality of the transform matrices, the quantization of the transform matrices is more accurate than in HEVC. In order to keep the mid value of the coefficients of the transform in the 16-bit range, after the horizontal and vertical transform, all coefficients will have 10 bits.
[0129] In order to control the MTS scheme, separate enabling flags are signaled at the SPS level for intra and inter respectively. When MTS is enabled at the SPS, CU level flags are signaled to indicate whether MTS is applied or not. Here, MTS is only applied to luma. The MTS CU level flag is signaled when the following conditions are met.
[0130] - both width and height are less than or equal to 32
[0131] - CBF flag is equal to 1
[0132] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 3-10. In terms of transform matrix precision, an 8-bit primary transform kernel is used. Thus, all transform kernels used in HEVC are kept the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8 use an 8-bit primary transform kernel.
[0133]
[0134] To reduce the complexity of large size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, high frequency transform coefficients are zeroed. Only coefficients within the 16x16 lower frequency region are kept.
[0135] As in HEVC, the transform skip mode can be used to code the residual of a block. To avoid redundancy in syntax coding, the transform skip flag is not signaled when the CU level MTS CU flag is not equal to 0. The block size limit for transform skip is the same as MTS in JEM4, which indicates that transform skip is applicable to CUs with both block width and height equal to or smaller than 32.
[0136] 2.10 Example Reduced Secondary Transform (RST)
[0137] 2.10.1 Example Non-Separable Secondary Transform (NSST)
[0138] In some embodiments, a secondary transform, also referred to as a non-separable transform, is applied between the forward primary transform and quantization (at the encoder) and between dequantization and the inverse primary transform (at the decoder side). As shown in Figure 10 The execution of the 4x4 (or 8x8) secondary transform depends on the block size, as shown. For example, a 4x4 secondary transform is applied for small blocks (e.g., min(width, height) < 8) and an 8x8 secondary transform is applied for large blocks (e.g., min(width, height) > 4) per 8x8 block.
[0139] Figure 10 An example of the secondary transform in JEM is shown.
[0140] The application of the non-separable transform is described below with an example of an input. To apply the non-separable transform, a 4x4 input block X
[0141] The application of the non-separable transform is described below with an example of an input. To apply the non-separable transform, a 4x4 input block X
[0142] is first represented as a vector
[0143]
[0144] The non-separable transform is computed as where denotes the vector of transform coefficients, while T is a 16x16 transform matrix. The 16x1 vector of coefficients is then reorganized into 4x4 blocks using the scan order (horizontal, vertical or diagonal) of the block. Coefficients with smaller indices are placed together with smaller scan indices in the 4x4 coefficient block. There are 35 transform sets in total, each using 3 non-separable transform matrices (kernels). The mapping from intra prediction mode to transform set is pre-defined. For each transform set, the selected non-separable quadratic transform candidate is further specified by an explicit signaled quadratic transform index. After the transform coefficients, an index is signaled in the bitstream for each intra CU once.
[0145] 2.10.2 Example Reduced Secondary Transform (RST) / Low Frequency Non- Separable Transform (LFNST) Figure 11 Figure 11
[0146] Reduced quadratic transform (RST), also known as low-frequency non-separable transform (LFNST), is introduced as 4 transform sets (instead of 35 transform sets) mapping. In some embodiments, 16x64 (can be further reduced to 16x48) and 16x16 matrices are used for 8x8 and 4x4 blocks, respectively. For convenience, the 16x64 (can be further reduced to 16x48) transform is denoted as RST8x8, and the 16x16 transform is denoted as RST4x4. Figure 12 An example of RST is shown.
[0147] Figure 12 An example of the proposed reduced quadratic transform (RST) is shown.
[0148] RST computation
[0149] The main idea of a reduced transform (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N (R
[0150] The RT matrix is an R x N matrix, as follows:
[0151]
[0152] where the R rows of the transform are R bases of the N-dimensional space. The inverse transform matrix of an RT is the transpose of its forward transform. Examples of forward and inverse RTs are shown inFigure 13 are depicted.
[0153] Figure 13 Examples of forward and inverse reduced transforms are shown.
[0154] In some embodiments, RST8x8 with a reduction factor of 4 (1 / 4 size) is applied. Thus, instead of a traditional 8x8 non-separable transform matrix size of 64x64, a 16x64 direct matrix is used. In other words, a 64x16 inverse RST matrix is used at the decoder side to generate the core (primary) transform coefficients in the 8x8 left-top region. The forward RST8x8 uses a 16x64 (or 8x64 for 8x8 blocks) matrix such that it produces non-zero coefficients only in the left-top 4x4 region within the given 8x8 region. In other words, if RST is applied, the 8x8 region except the left-top 4x4 region will have only zero coefficients. For RST4x4, a 16x16 (or 8x16 for 4x4 blocks) direct matrix multiplication is applied.
[0155] Inverse RST is conditionally applied when the following two conditions are met:
[0156] a. Block size is greater than or equal to a given threshold (W>=4 && H>=4)
[0157] b. Transform skip mode flag is zero
[0158] If the width (W) and height (H) of the transform coefficient block are both greater than 4, RST8x8 is applied to the left-top 8x8 region of the transform coefficient block. Otherwise, RST4x4 is applied to the left-top min(8,W)xmin(8,H) region of the transform coefficient block.
[0159] If the RST index is equal to 0, no RST is applied. Otherwise, RST is applied with the kernel selected using the RST index. The RST selection method and coding of the RST index will be explained later.
[0160] In addition, RST is applied to intra CUs in intra and inter slices, as well as to both luma and chroma. If dual tree is enabled, the RST index for luma and chroma is signaled separately. For inter slices (dual tree disabled), a single RST index is signaled and used for both luma and chroma.
[0161] In some embodiments, Intra Sub-Partition (ISP) is adopted as a new intra prediction mode. When ISP mode is selected, RST is disabled and RST index is not signaled, because the performance improvement is marginal even if RST is applied to each feasible partitioned block. In addition, disabling RST for the residual of ISP prediction can reduce the encoding complexity.
[0162] RST selection
[0163] The RST matrix is selected from four transform sets, each consisting of two transforms. Which transform set is applied is determined by the intra prediction mode, as follows:
[0164] (1) If one of the three CCLM modes is indicated, transform set 0 is selected.
[0165] (2) Otherwise, transform set selection is performed according to the following table:
[0166] Transform set selection table
[0167]
[0168]
[0169] The index to the table (denoted as IntraPredMode) ranges from [-14, 83], which is the transform mode index for wide-angle intra prediction.
[0170] Reduced dimension RST matrix
[0171] As a further simplification, a 16x48 matrix is applied instead of a 16x64 with the same transform set configuration, each matrix taking 48 input data from three 4x4 blocks in the left-top 8x8 block (excluding the right-bottom 4x4 block) Figure 14 ).
[0172] Figure 14 An example of the forward RST 8x8 process with 16x48 matrix is shown.
[0173] RST signaling
[0174] The forward RST 8x8 (where R=16) uses a 16x64 matrix such that it produces non-zero coefficients only in the left-top 4x4 region within the given 8x8 region. In other words, if RST is applied, the 8x8 region excluding the left-top 4x4 region generates only zero coefficients. As a result, when any non-zero element is detected within the 8x8 block region excluding the left-top 4x4 Figure 15 , the RST index is not coded because this implies that RST is not applied. In this case, the RST index is inferred to be zero.
[0175] Figure 15 An example of scanning for non-zero elements from 17 to 64 is shown.
[0176] Zero-out range
[0177] In general, any coefficient in a 4x4 subblock can be non-zero before inverse RST is applied to the 4x4 subblock. However, in some cases, some coefficients in a 4x4 subblock must be zero before inverse RST is applied to the subblock.
[0178] Let nonZeroSize be a variable. Any coefficient with index not less than nonZeroSize must be zero when rearranged into a one-dimensional array before inverse RST.
[0179] When nonZeroSize is equal to 16, there is no zeroing constraint for coefficients in the top-left 4x4 subblock.
[0180] In some embodiments, when the current block size is 4x4 or 8x8, nonZeroSize is set equal to 8. For other block dimensions, nonZeroSize is set equal to 16.
[0181] Example description of RST
[0182] In the following tables and descriptions, bold italicized text is used to indicate changes that can be made to the current syntax to accommodate certain embodiments described in this document.
[0183] Sequence parameter set RBSP syntax
[0184]
[0185]
[0186] Residual coding syntax
[0187]
[0188] Coded unit syntax
[0189]
[0190]
[0191] Sequence parameter set RBSP semantics
[0192] …
[0193]
[0194] …
[0195] Coded unit semantics
[0196] …
[0197]
[0198] Transform process of scaling transform coefficients
[0199] Generally
[0200] The input of the process is:
[0201] - a luma position (xTbY, yTbY) specifying the left-top sample of the current luma transform block relative to the left-top luma sample of the current picture,
[0202] - a variable nTbW specifying the width of the current transform block,
[0203] - a variable nTbH specifying the height of the current transform block,
[0204] - a variable cldx specifying the color component of the current block,
[0205] - an (nTbW) x (nTbH) array d[x][y] of scaling transform coefficients, with x = 0..nTbW-1, y = 0..nTbH-1.
[0206] The output of the process is an (nTbW) x (nTbH) array r[x][y] of residual samples, with x = 0..nTbW-1, y = 0..nTbH-1.
[0207]
[0208]
[0209] Secondary transform process
[0210]
[0211]
[0212] Secondary transform matrix derivation process
[0213]
[0214]
[0215]
[0216]
[0217] 2.11 Clipping of dequantized in HEVC
[0218] In HEVC, the scaled transform coefficient d' is computed as d' = Clip3(coeffMin, coeffMax, d), where d is the scaled transform coefficient before clipping.
[0219] For luma components, coeffMin = CoeffMinY; coeffMax = CoeffMaxY. For chroma components, coeffMin = CoeffMinC; coeffMax = CoeffMaxC; where
[0220] CoeffMinY = -(1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15))
[0221] CoeffMinC = -(1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))
[0222] CoeffMaxY = (1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) - 1
[0223] CoeffMaxC = (1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) - 1
[0224] extended_precision_processing_flag is a syntax element signaled in the SPS.
[0225] 2.12 Affine linear weighted intra prediction (ALWIP, also known as matrix-based intra prediction (MIP))
[0226] In some embodiments, two tests are conducted. In test 1, ALWIP is designed with memory limit of 8K bytes and at most 4 multiplications per sample. Test 2 is similar to test 1, but further simplifies the design in terms of memory requirement and model architecture.
[0227] * A single set of matrices and offset vectors for all block shapes.
[0228] * Reducing the number of modes for all block shapes to 19.
[0229] * Reducing the memory requirement to 5760 10-bit values, i.e., 7.20 kilobytes.
[0230] The linear interpolation of the prediction samples is implemented in a single step in each direction (instead of iterative interpolation in the first test).
[0231] 2.13 Sub-block transform
[0232] For inter-predicted CUs with cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is decoded. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, one part of the residual block is coded with an inferred adaptive transform while the other part of the residual block is zeroed. SBT is not applicable to combined inter-intra modes.
[0233] In sub-block transform, the luma transform blocks in SBT-V and SBT-H are applied with a transform depending on the position (the chroma TBs always use DCT-2). Two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transforms of each SBT position are specified in Figures 18A-18B , for example. The horizontal and vertical transforms of SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of a residual TU is larger than 32, the corresponding transform is set to DCT-2. Thus, the sub-block transform jointly specifies the TU slice of a residual block, cbf, and the horizontal and vertical transforms, which can be considered as a kind of syntax shortcut for the case where the main residual of a block is located at one side of the block.
[0234] Improvements to Separable Secondary Transform (SST) is the specification of sub-block transform modes SBT-V and SBT-H.
[0235] 2.14 Separable quadratic transform in AVS
[0236] In some embodiments, a 4x4 separable quadratic transform (SST) is applied to all luma blocks coded in intra mode after the main transform if the main transform is DCT2.
[0237] When applying SST on a block at the encoder, the left-top 4x4 sub-block of the transformed block after the main transform (denoted as L) is further transformed as L' = T' x L x T,
[0238] where T is the quadratic transform matrix.
[0239] L' is then quantized together with the other parts of the transformed block.
[0240] When applying SST on a block at the decoder, the left-top 4x4 sub-block of the transformed block after dequantization (denoted as M) is further inverse-transformed as
[0241] M' = S' x M x S,
[0242] where S is the inverse of the secondary transform matrix. In particular, S' = T.
[0243] M' is then input to the primary inverse transform together with the other parts of the transform block.
[0244] 2.15 Scan Region-based Coefficient Coding (SRCC)
[0245] SRCC has been adopted by AVS-3. In the SRCC case, the scan and signaling is as shown in Figure 18A the signaled right-bottom position (SRx, SRy), and only the coefficients within the rectangle (e.g., scan region) with four corners (0, 0), (SRx, 0), (0, SRy), (SRx, SRy). All coefficients outside the rectangle are zero.
[0246] 3. Examples of problems solved by the embodiments
[0247] The current design has the following problems:
[0248] (1) The clipping and shifting / rounding operations in MTS / RST can not be optimal.
[0249] (2) The RST applied to two adjacent 4x4 blocks can be costly.
[0250] (3) The RST can be applied differently for different color components.
[0251] (4) The RST can not be suitable for screen content coding.
[0252] (5) The interaction of RST with other coding tools is unclear.
[0253] (6) The transform matrix of RST can be stored more efficiently.
[0254] (7) How to apply the quantization matrix on RST is unclear.
[0255] 4. Example embodiments and techniques
[0256] The embodiments listed below should be considered as examples to explain the general concepts. The embodiments should not be interpreted in a narrow way. Furthermore, the embodiments can be combined in any way.
[0257] In the following description, coded information can include prediction modes (e.g., intra / inter / IBC modes), motion vectors, reference pictures, inter prediction directions, intra prediction modes, combined intra-inter prediction (CIIP) modes, ISP modes, affine intra modes, adopted transform cores, transform skip flags, etc., e.g., information needed when encoding a block.
[0258] In the following discussion, SatShift(x, n) is defined as
[0259]
[0260] Shift(x, n) is defined as Shift(x, n) = (x + offset0) » n.
[0261] In one example, offset0 and / or offset1 is set to (1 « n) » 1 or (1 « (n - 1)). In another example, offset0 and / or offset1 is set to 0.
[0262] In another example, offset0 = offset1 = ((1 « n) » 1) - 1 or ((1 « (n - 1)) - 1).
[0263] Clip3(min, max, x) is defined as
[0264]
[0265] 1. After inverse RST, the output value should be clipped to the range [MinCoef, MaxCoef] inclusive, where MinCoef and / or MaxCoef are two integer values that can vary.
[0266] a. In one example, assuming the coefficients after dequantization are clipped to the range [QMinCoef, QMaxCoef] inclusive, MinCoef can be set equal to QMinCoef and / or MaxCoef can be set equal to QMaxCoef.
[0267] b. In one example, MinCoef and / or MaxCoef can depend on the color component.
[0268] i. In one example, MinCoef and / or MaxCoef can depend on the bit depth of the corresponding color component.
[0269] c.In one example, MinCoef and / or MaxCoef can depend on block shape (e.g., square or non-square) and / or block dimension.
[0270] d.In one example, the selection of the value or candidate values of MinCoef and / or MaxCoef can be signaled such as in SPS, PPS, slice header / tile group header / CTU / CU.
[0271] e.In one example, for luma component, MinCoef and / or MaxCoef can be derived as:
[0272] MinCoef = -(1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15))
[0273] MaxCoef = (1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) - 1
[0274] where BitDepthY is the bit depth of luma component, and extended_precision_processing_flag can be signaled such as in SPS.
[0275] f.In one example, for one component, MinCoef and / or MaxCoef can be derived as:
[0276] MinCoef = -(1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))
[0277] MaxCoef = (1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) - 1,
[0278] where BitDepthC is the bit depth of chroma component, and extended_precision_processing_flag can be signaled such as in SPS.
[0279] g.In some embodiments, MinCoef is -(1 « 15) and MaxCoef is (1 « 15) - 1.
[0280] h. In one example, the consistent bitstream should satisfy that the transform coefficients after forward RST should be within a given range.
[0281] 2. It is proposed that the way of applying forward RST and / or inverse RST on the MxN sub-blocks of coefficients can depend on the number of sub-blocks on which forward RST and / or inverse RST is applied, e.g. M=N=4.
[0282] a. In one example, the zero-out range can depend on the sub-block index on which RST is applied.
[0283] i. Alternatively, the zero-out range can depend on the number of sub-blocks on which RST is applied.
[0284] b. In one example, when there are S sub-blocks on which forward RST and / or inverse RST is applied, the way of applying forward RST and / or inverse RST on the first and second sub-blocks of coefficients in the entire coefficient block can be different, where S>1, e.g. S=2.
[0285] For example, the first MxN sub-block can be the left-top MxN sub-block.
[0286] i. In one example, the nonZeroSize described in section 2.10 can be different for the first MxN sub-block of coefficients (denoted as nonZeroSizeO) and the second MxN sub-block of coefficients (denoted as nonZeroSizei).
[0287] 1) In one example, nonZeroSizeO can be greater than nonZeroSizei. For example, nonZeroSizeO=16 and nonZeroSizei=8.
[0288] ii. In one example, when there is only one MxN sub-block to be applied forward RST and / or inverse RST, or there are more than one MxN sub-block to be applied forward RST and / or inverse RST, the nonZeroSize as described in section 2.10 can be different.
[0289] 1) In one example, if there are more than one MxN sub-block to be applied forward RST and / or inverse RST, the nonZeroSize can be equal to 8.
[0290] 3. It is proposed that if the current block size is 4xH or Wx4 (where H>8 and W>8), only one MxN sub-block of coefficients (e.g. the left-top MxN sub-block) is applied forward RST and / or inverse RST. For example, M=N=4.
[0291] a. In one example, if H > T1 and / or W > T2, forward RST and / or inverse RST is applied to only one MxN sub-block of coefficients. For example, T1 = T2 = 16.
[0292] b. In one example, if H < T1 and / or W < T2, forward RST and / or inverse RST is applied to only one MxN sub-block of coefficients. For example, T1 = T2 = 32.
[0293] c. In one example, for all H > 8 and / or W > 8, forward RST and / or inverse RST is applied to only one MxN sub-block of coefficients.
[0294] d. In one example, if the current block size is MxH or WxN (where H >= N and W >= M), forward RST and / or inverse RST is applied to only one MxN sub-block (e.g. the left-top MxN sub-block). For example, M = N = 4.
[0295] 4. RST can be applied to non-square regions. Assume the region size is represented by KxL, where K is not equal to L.
[0296] a. Alternatively, in addition, zeroing can be applied to the transformed coefficients after forward RST, so as to meet the maximum number of non-zero coefficients.
[0297] i. In one example, if a transformed coefficient is outside the left-top MxM region, where M is not larger than K and M is not larger than L, the transformed coefficient can be set to 0.
[0298] 5. It is proposed that the coefficients in two adjacent MxN sub-blocks can be included in a single forward RST and / or inverse RST. For example, M = N = 4.
[0299] a. In one example, one or several operations can be performed at the encoder. The operations can be performed in order.
[0300] i. The coefficients in two adjacent MxN sub-blocks are rearranged into a one-dimensional vector with 2xMxN elements.
[0301] ii. A forward RST with a transform matrix having 2xMxN columns and MxN rows (or MxN columns and 2xMxN rows) is applied to the one-dimensional vector.
[0302] iii. The transformed one-dimensional vector with MxN elements is rearranged into a first MxN sub-block (such as the left-top sub-block).
[0303] iv. All coefficients in the second MxN sub-block can be set to zero.
[0304] b. In one example, one or several of the following operations can be performed at the decoder.
[0305] The operations can be performed in order.
[0306] i. The coefficients in the first MxN sub-block, such as the top-left sub-block, are rearranged into a one-dimensional vector with MxN elements.
[0307] ii. The inverse RST with a transform matrix of MxN columns and 2xMxN rows (or 2xMxN columns and MxN rows) is applied to the one-dimensional vector.
[0308] iii. The transformed one-dimensional vector with 2xMxN elements is rearranged into two adjacent MxN sub-blocks.
[0309] c. In one example, a block can be divided into K (K>1) sub-blocks, and the primary and secondary transforms can be performed at the sub-block level.
[0310] 6. The zero-out range (e.g., nonZeroSize described in Section 2.10) can depend on the color component.
[0311] a. In one example, the range for luma and chroma components can be different for the same block dimension.
[0312] 7. The zero-out range (e.g., nonZeroSize described in Section 2.10) can depend on the coding information.
[0313] a. In one example, it can depend on the coding mode, such as intra or non-intra mode.
[0314] b. In one example, it can depend on the coding mode, such as intra or inter or IBC mode.
[0315] c. In one example, it can depend on the reference picture / motion information.
[0316] 8. The zero-out range (e.g., nonZeroSize described in Section 2.10) for a specific block dimension can depend on the quantization parameter (QP).
[0317] a. In one example, it is assumed that nonZeroSize equals nonZeroSizeA when QP equals QPA, and nonZeroSize equals nonZeroSizeB when QP equals QPB. If QPA is not smaller than QPB, then nonZeroSizeA is not larger than nonZeroSizeB.
[0318] b. Different transform / inverse transform matrices can be used for different nonZeroSize.
[0319] 9. It is proposed that the zero-out range (e.g., nonZeroSize described in section 2.10) can be signaled such as in SPS, PPS, picture header, slice header, tile group header, CTU row, CTU, CU, or any video data unit.
[0320] a. Alternatively, multiple ranges can be defined. And an indication of which candidate nonZeroSize to select can be signaled such as in SPS, PPS, picture header, slice header, tile group header, CTU row, CTU, and CU.
[0321] 10. Whether and / or how to apply RST can depend on the color format and / or the use of separate plane coding and / or color component.
[0322] a. In one example, RST can not be applied to chroma components (such as Cb and / or Cr).
[0323] b. In one example, if the color format is 4:0:0, RST can not be applied to chroma components.
[0324] c. In one example, if separate plane coding is used, RST can not be applied to chroma components.
[0325] d. In one example, the nonZeroSize for a specific block dimension can depend on the color component.
[0326] i. In one example, for the same block dimension, the nonZeroSize on chroma components can be smaller than the nonZeroSize on luma components.
[0327] 11. It is proposed that when a single coding structure tree is used to code a component, the RST control information (such as whether to apply RST, and / or which set of transform matrices to select) can be signaled separately for luma and chroma components.
[0328] 12. Whether and how to apply RST can depend on the coding information (such as coding mode) of the current block and / or the neighboring blocks.
[0329] a. In one example, RST can not be used for one or more specific intra prediction modes.
[0330] i. For example, RST can not be used for LM mode.
[0331] ii. For example, RST can not be used for LM-T mode.
[0332] iii. For example, RST cannot be used for LM-A mode.
[0333] iv. For example, RST cannot be used for wide angle intra prediction mode.
[0334] v. For example, RST cannot be used for BDPCM mode or / and DPCM mode or / and RBDPCM mode.
[0335] vi. For example, RST cannot be used for ALWIP mode.
[0336] vii. For example, RST cannot be used for certain specific angular intra prediction modes (such as DC, Planar, Vertical, Horizontal, etc.).
[0337] viii. For example, RST can be used for luma component in LM mode or / and LM-T mode or / and LM-A mode but cannot be used for chroma component.
[0338] ix. For example, RST can not be used for chroma component when joint chroma residual coding is applied.
[0339] b. If RST cannot be applied, the syntax element indicating the information related to RST in the current block can not be signaled.
[0340] 13. It is proposed that RST can be applied to non-intra coded blocks.
[0341] a. In one example, RST can be applied to inter coded blocks.
[0342] b. In one example, RST can be applied to intra block copy (IBC) coded blocks.
[0343] c. In one example, RST can be applied to combined inter-intra prediction (CIIP) coded blocks.
[0344] 14. It is proposed that RST can be controlled at different levels.
[0345] a. For example, the information indicating whether RST (such as a control flag) is applicable can be signaled in PPS, slice header, picture header, tile group header, tile, CTU row, CTU.
[0346] b. Whether RST is applicable can depend on the standard profile / level / tier.
[0347] 15. It is proposed that whether to apply position dependent intra prediction combination (PDPC) can depend on whether RST is applied.
[0348] a. In one example, if RST is applied for the current block, PDPC can not be applied.
[0349] b. In one example, if RST is applied for the current block, PDPC can be applied.
[0350] c. Alternatively, whether RST is applied can depend on whether PDPC is applied.
[0351] i. In one example, RST is not applied when PDPC is applied.
[0352] ii. If RST cannot be applied, a syntax element indicating information related to RST in the current block can not be signaled.
[0353] 16. It is proposed that whether to filter the neighboring samples for intra prediction can depend on whether RST is applied.
[0354] a. In one example, if RST is applied for the current block, the neighboring samples can not be filtered.
[0355] b. In one example, if RST is applied for the current block, the neighboring samples can be filtered.
[0356] c. Alternatively, whether RST is applied can depend on whether the neighboring samples for intra prediction are filtered.
[0357] i. In one example, RST is not applied when the neighboring samples for intra prediction are filtered.
[0358] ii. In one example, RST is not applied when the neighboring samples for intra prediction are not filtered.
[0359] iii. If RST cannot be applied, a syntax element indicating information related to RST in the current block can not be signaled.
[0360] 17. It is proposed that RST can be applied when the current block is coded with transform skip.
[0361] a. For example, the primary transform is skipped, but a second transform can still be applied.
[0362] b. The secondary transform matrix used in transform skip mode can be different from that used in non-transform skip mode.
[0363] 18. It is proposed that the transform matrix for RST can be stored with a bit width smaller than 8. For example, the transform matrix for RST can be stored with a bit width of 6 or 4.
[0364] 19. It is proposed that the transform matrix for RST can be stored in a predictive manner.
[0365] a. In one example, a first element in a first transform matrix for RST can be predicted by a second element in the first transform matrix for RST.
[0366] i. For example, a difference between the two elements can be stored.
[0367] ii. For example, the difference can be stored in a bit width less than 8, such as 6 or 4.
[0368] b. In one example, a first element in a first transform matrix for RST can be predicted by a second element in a second transform matrix for RST.
[0369] i. For example, a difference between the two elements can be stored.
[0370] ii. For example, the difference can be stored in a bit width less than 8, such as 6 or 4.
[0371] 20. It is proposed that the first transform matrix for RST can be derived from the second transform matrix for RST.
[0372] a. In one example, partial elements of the second transform matrix for RST can be picked up to construct the first transform matrix for RST.
[0373] b. In one example, the first transform matrix for RST is derived by rotating or flipping all or a portion of the second transform matrix for RST.
[0374] c. In one example, the first transform matrix for RST is derived by down- sampling or up-sampling the second transform matrix for RST.
[0375] 21. It is proposed that syntax elements indicating RST-related information in a current block can be signaled before signaling residuals (which can be transformed).
[0376] a. In one example, signaling of RST-related information can not depend on counting of non-zero or zero coefficients when parsing residuals.
[0377] b. In one example, non-zero or zero coefficients can not be counted when parsing residuals.
[0378] c. In one example, a coding block flag (cbf) flag for a sub-block set to all zeros by RST can not be signaled and inferred to be 0.
[0379] d. In one example, a significance flag for a coefficient set to zero by RST can not be signaled and inferred to be 0.
[0380] e. The scan order for parsing the residual block can depend on whether and how RST is applied.
[0381] i. In one example, coefficients set to zero by RST can not be scanned.
[0382] f. The arithmetic coding context for parsing the residual block can depend on whether and how RST is applied.
[0383] 22. Whether and how to apply a quantization matrix can depend on whether and how RST is applied.
[0384] a. In one example, different quantization matrices can be applied whether or not RST is applied.
[0385] b. Alternatively, whether and how RST is applied can depend on whether and how a quantization matrix is applied.
[0386] i. In one example, RST can not be applied when a quantization matrix is applied to the block.
[0387] 23. It is proposed that RST can be applied to quantized coefficients / residuals.
[0388] a. In one example, RST can be applied to residuals when transform skip is used.
[0389] b. In one example, RST can be applied to quantized transform coefficients of a block.
[0390] 24. It is proposed that RST can be applied to sub-block transform blocks.
[0391] a. In one example, RST can be applied to the top-left coefficient generated by sub-block transform.
[0392] 25. It is proposed that how and / or whether RST is applied can depend on the number of TUs in a CU.
[0393] a. For example, how and / or whether RST is applied can depend on whether the number of TUs in the CU is greater than 1.
[0394] i. In one example, RST is not applied if the number of TUs in the CU is greater than 1.
[0395] ii. In one example, RST is applied only to one of the multiple TUs in the CU if the number of TUs in the CU is greater than 1.
[0396] 1) In one example, RST is applied only to the first TU in the CU if the number of TUs in the CU is greater than 1.
[0397] 2) In one example, if the number of TUs in a CU is greater than 1, RST is only applied to the last TU in the CU.
[0398] iii. In one example, if the number of TUs in a CU is greater than 1, RST is independently applied to each TU of the CU.
[0399] 1) Alternatively, when the number of TUs in a CU is greater than 1, whether to apply RST to the first TU of the CU can be determined independently of whether to apply RST to the second TU of the CU.
[0400] 2) In one example, when the number of TUs in a CU is greater than 1, whether to apply RST to a TU of the CU can depend on the number of non-zero coefficients (denoted as NZ) of the TU, but not on the number of non-zero coefficients of other TUs of the CU.
[0401] a) In one example, if NZ is less than a threshold T (e.g., T = 2), RST is not applied to the TU.
[0402] b) If it is determined that RST is not applied to the TU, (the) syntax element(s) used to indicate whether RST is applied can not be signaled for the TUs of the CU.
[0403] b. For example, how and / or whether to apply RST can depend on whether the TU size is equal to the CU size.
[0404] i. In one example, when the CU size is greater than the TU size, RST is disabled.
[0405] c. It is proposed to use the decoding information of the first TU or the last TU in the decoding order of the CU to decide the usage of RST and / or the signaling of syntax elements related to RST.
[0406] i. In one example, if the number of non-zero coefficients of the first or last TU is less than a threshold T (e.g., T = 2), RST is not applied to the CU.
[0407] ii. In one example, if the number of non-zero coefficients of a sub-region (e.g., the left-top 4x4) within the first or last TU is less than a threshold T (e.g., T = 2), RST is not applied to the CU.
[0408] 26. It is proposed to have a flag for a TU to control whether RST is applied.
[0409] a. Whether RST is applied to a TU can depend on the flag of the TU.
[0410] i. When the flag is not present or not derived, it can be derived as false.
[0411] ii. Alternatively, when the flag is not present or not derived, it can be derived as true.
[0412] b. When the CU contains only one TU, the flag of the TU can be equal to the CU RST flag, which can be derived on the fly (e.g., based on coefficient information).
[0413] c. When the number of TUs in the CU is greater than 1, the flag of the last TU in the CU can be derived from the CU RST flag, which can be derived on the fly (e.g., based on coefficient information), and the flags of all other TUs can be set to false.
[0414] i. Alternatively, when the number of TUs in the CU is greater than 1, the flag of the last TU in the CU can be derived from the CU RST flag, and the flags of all other TUs can be set to true.
[0415] 27. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether and / or how to apply RST to a first component of a block can be different from whether and / or how to apply RST to a second component of the block. That is, separate control of applying RST to different color components.
[0416] a. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a first component of a block can be determined independently of whether to apply RST to a second component of the block.
[0417] i. In one example, when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a component of a block can depend on the decoded information (e.g., the number of non-zero coefficients (denoted as NZ)) of that component of the block, but not on the decoded information of any other component of the block.
[0418] 1) In one example, if NZ is less than a threshold T (e.g., T = 2), RST is not applied to the component of the block.
[0419] 2) If it is determined that RST is not applied to the component of the block, the syntax element(s) indicating whether to apply RST can not be signaled for the component of the block.
[0420] b. In one example, for the case of a single tree, whether to enable RST and / or how to apply RST can be determined independently for the luma and chroma components.
[0421] 28. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a first component of a block can be determined by a second component of the block.
[0422] a. In one example, when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to the first component of the block can be determined by the number of non-zero coefficients of the second component of the block.
[0423] i. In one example, if NZ (e.g., the number of non-zero coefficients of the second component of the block or a sub-region (e.g., the left-top 4x4) of the block) is less than a threshold T (e.g., T = 2), RST is not applied to the first component of the block.
[0424] ii. If it is determined that RST is not applied to the first component of the block, the syntax element(s) indicating whether to apply RST can not be signaled for the components of the block.
[0425] iii. In one example, the first component is Cb or Cr, and the second component is Y.
[0426] iv. In one example, the first component is R or B, and the second component is G.
[0427] 29. In one example, whether to apply bullet 25 and / or bullet 26 and / or bullet 27 can depend on the width and height (denoted as W and H) of the CU and / or TU and / or block and / or the maximum transform block size.
[0428] a. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T or H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.
[0429] b. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T and H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.
[0430] c. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T and H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.
[0431] Figure 17
[0432] 30. In one example, it can be determined to enable or disable SST for a video unit.
[0433] a. For example, the determination can be made based on signaling in a video syntax structure associated with the video unit.
[0434] i. In one example, the signaling such as a flag can be coded with at least one context in arithmetic coding.
[0435] ii. In one example, the signaling can be conditionally skipped based on coding / decoding information such as block dimension, coding block flag (cbf), and coding mode of the current block.
[0436] 1) In one example, the signaling can be skipped when cbf is equal to 0.
[0437] b. For example, the determination can be made based on inference without signaling associated with the video unit.
[0438] i. The inference can depend on information of the video unit, such as coding mode, intra prediction mode, type of primary transform, and dimension or size of the video unit.
[0439] c. For example, the video unit can be a block, such as a coding block or a transform block. The video syntax structure can be a coding unit (CU) or a transform unit (TU).
[0440] d. For example, the video unit can be a picture. The video syntax structure can be a picture header or a PPS.
[0441] e. For example, the video unit can be a slice. The video syntax structure can be a slice header.
[0442] f. For example, the video unit can be a slice. The video syntax structure can be a sequence header or a SPS.
[0443] g. The video syntax structure can be VPS / DPS / APS / Tile Group / Tile / CTU row / CTU.
[0444] 31. In one example, whether to disable or enable SST can be based on block dimension.
[0445] h. For example, SST can be disabled if at least one of the block width or height is less than (or not greater than) Tmin.
[0446] i. For example, SST can be disabled if both the block width and height are less than Tmin.
[0447] j. For example, SST can be disabled if at least one of the block width or height is greater than (or not less than) Tmax.
[0448] k. For example, SST can be disabled if both the block width and height are greater than (or not less than) Tmax.
[0449] l. For example, Tmin can be 2 or 4.
[0450] m. For example, Tmax can be 32, 64 or 128.
[0451] n. In one example, SST can be disabled based on the block width or / and height of the first color component.
[0452] i. For example, the first color component can be the luma color component.
[0453] ii. For example, the first color component can be the R color component.
[0454] o. In one example, SST can be disabled based on the block width or / and height of all color components.
[0455] p. Alternatively, in addition, the indication of the omission of the use of SST and / or the related signaling of other side information is omitted when SST is disabled.
[0456] q. In one example, based on the block dimension, SST can be enabled for the first color component and disabled for the second color component.
[0457] 32. In one example, a set of SSTs can be employed, and the selection of the SST matrix for a block can depend on decoded information, such as the block dimension.
[0458] r. Alternatively, in addition, the same decoded / signaled SST index or the same on / off control flag can be interpreted differently, such as different matrices corresponding to different block dimensions.
[0459] s. For example, different SSTs in the set can have different dimensions, such as a 4x4 SST, an 8x8 SST or a 16x16 SST,
[0460] t. For example, a 4x4 SST can be applied to the block in case of condition C4, an 8x8 SST can be applied to the block in case of condition C8.
[0461] i. Alternatively, in addition, a 4x4 SST can be applied to the block in case of condition C4, an 8x8 SST can be applied to the block in case of condition C8,... an NxN SST can be applied to the block in case of condition CN, where N is an integer.
[0462] u. In one example, condition C4 is that at least one of the block width and height is equal to 4.
[0463] v. In one example, condition C4 is that both the block width and height are equal to 4.
[0464] w. In one example, condition C4 is that the smaller one of the block width and height is equal to 4.
[0465] x. In one example, condition C8 is that the smaller of the block width and height is not less than 8.
[0466] y. In one example, condition C8 is that at least one of the block width and height is equal to 8.
[0467] z. In one example, condition C8 is that both the block width and height are equal to 8.
[0468] aa. In one example, condition C8 is that at least one of the block width and height is greater than or equal to 8.
[0469] bb. In one example, condition C8 is that both the block width and height are greater than or equal to 8.
[0470] cc. In one example, condition CN is that at least one of the block width and height is equal to N.
[0471] dd. In one example, condition CN is that both the block width and height are equal to N.
[0472] ee. In one example, condition CN is that at least one of the block width and height is greater than or equal to N.
[0473] ff. In one example, condition CN is that both the block width and height are greater than or equal to N.
[0474] gg. In one example, an NxN SST can be applied to the left-top NxN sub-block of the transform block.
[0475] hh. In one example, SST can be applied horizontally or vertically or both, depending on the block dimensions.
[0476] ii. In one example, different SST matrices can be selected for different color components.
[0477] i. For example, the above rules can be applied independently for different color components.
[0478] jj. In one example, one same SST matrix can be selected for all color components.
[0479] i. For example, the above rules can be applied for a first color component, and the selected SST matrix can be applied for all color components.
[0480] 1) In one example, the first color component can be the luma component.
[0481] 2) In one example, the first color component can be the Cb or Cr component.
[0482] 3) Alternatively, in addition, if the selected SST matrix is not applicable to the second color component, then SST is disabled for the second color component.
[0483] kk. In one example, SST can be allowed only if the selected SST matrices for all color components (by independently applying the above rules to different color components) are the same.
[0484] i. Alternatively, in addition, if SST is not allowed, then the indication of the use of SST and / or the related signaling of other side information is omitted.
[0485] 33. In one example, NxN SST can be applied to at least one NxN subblock that is not the same as the left-top NxN subblock.
[0486] ll. For example, NxN SST can be applied to the NxN subblock that is right-adjacent to the left-top NxN subblock.
[0487] mm. For example, NxN SST can be applied to the NxN subblock that is bottom-adjacent to the left-top NxN subblock.
[0488] 34. In one example, a first SST can be applied as a horizontal transform to a transform block, and a second SST can be applied as a vertical transform to the transform block, where the first SST and the second SST can be different.
[0489] nn. For example, the first SST and the second SST can have different dimensions.
[0490] oo. Assuming the first SST is an NxN SST, the second SST is an MxM SST, and the dimension of the transformed block is WxH, the following rules can be applied:
[0491] i. If W is equal to Wl, then N is set equal to Wl, where Wl is an integer such as 4 or 8.
[0492] ii. If W is greater than or not less than W2, then N is set equal to W2, where W2 is an integer such as 4 or 8.
[0493] iii. If H is equal to HI, then M is set equal to HI, where HI is an integer such as 4 or 8.
[0494] iv. If H is greater than or not less than H2, then M is set equal to H2, where H2 is an integer such as 4 or 8.
[0495] 35. In one example, one SST in a set of SSTs can be used for a block, where there are multiple SSTs in the set that have the same dimension.
[0496] pp. In one example, the signaling message is to indicate which SST to select.
[0497] qq. In one example, the selection of which SST is inferred without signaling. The inference can depend on
[0498] i. Block dimension.
[0499] ii. Intra prediction mode.
[0500] iii. Transformed quantized / unquantized coefficients.
[0501] iv. Color component.
[0502] v. Type of primary transform.
[0503] 36. In one example, a different SST can be applied if the primary transform is different.
[0504] rr. For example, the SST used in association with DCT2 can be different from the SST used in association with DST7.
[0505] 37. In one example, SST can be applied to chroma components.
[0506] ss. In one example, different SST matrices can be applied to different color components, such as Y, Cb, and Cr.
[0507] tt. In one example, different rules of whether and how to apply SST can follow different color components.
[0508] uu. In one example, separate control of two color components can be applied.
[0509] i. In one example, an indication of the use of SST and / or an indication of the matrix can be signaled for each of the two color components.
[0510] 38. The indication of the use of SST and / or the indication of the SST matrix can be signaled according to a condition check of the right-bottom positioning of the scan region. The right-bottom positioning is denoted by (SRx, SRy), such as depicted in Figure 19 -B.
[0511] vv. In one example, the indication of the use of SST and / or the indication of the SST matrix can be omitted when SRx is greater than or not less than Kx and / or when SRy is greater than or not less than Ky.
[0512] ww. In one example, the indication of the use of SST and / or the indication of the SST matrix can be omitted when SRx is less than or not greater than K'x and / or when SRy is less than or not greater than K'y.
[0513] xx. Alternatively, in addition, when the indication is not signaled, it can be inferred that SST is disabled.
[0514] yy. Alternatively, in addition, when the indication is not signaled, it can be inferred that the default SST.
[0515] i. In one example, the default SST can be set to a K*L transform.
[0516] ii. In one example, the default SST can be determined from decoded information, such as block dimensions.
[0517] zz. Alternatively, in addition, the above methods can also be applied to other non-separable primary transforms.
[0518] FIG. 1600 is a block diagram of a video processing device 1600. The device 1600 can be used to implement one or more methods described herein. The device 1600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 1600 can include one or more processors 1602, one or more memories 1604, and video processing hardware 1606. The processor(s) 1602 can be configured to implement one or more methods described in the present document. The memory(ies) 1604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 1606 can be used to implement, in hardware circuitry, some of the techniques described in the present document.
[0519] Figure 20 is a flowchart of an example method 1700 of video processing. The method 1700 includes determining (1702) a constraint rule for selectively applying a secondary transform having a reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block. The method 1700 includes performing (1704) the conversion by applying the secondary transform having the reduced dimension according to the constraint rule. The secondary transform having the reduced dimension has a dimension that is reduced from a dimension of the current video block. The secondary transform having the reduced dimension is applied in a particular order with a primary transform during the conversion.
[0520] Additional embodiments and techniques are described in the following examples.
[0521] 1. A method of video processing, comprising determining a constraint rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality according to the constraint rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block, and wherein the secondary transform with reduced dimensionality is applied in a specific order with a primary transform during the conversion.
[0522] 2. The method of example 1, wherein the conversion comprises encoding the current video block into the bitstream representation, and wherein the specific order comprises first applying the primary transform in a forward direction, then selectively applying the secondary transform with reduced dimensionality in the forward direction, then quantizing an output of the secondary transform with reduced dimensionality in the forward direction.
[0523] 3. The method of example 1, wherein the conversion comprises decoding the current video block from the bitstream representation, and wherein the specific order comprises first applying dequantization to the bitstream representation, then selectively applying the secondary transform with reduced dimensionality in an inverse direction, then applying the primary transform in the inverse direction to an output of the secondary transform with reduced dimensionality in the inverse direction.
[0524] 4. The method of any of examples 1-3, wherein the constraint rule specifies clipping a range of the output of the secondary transform with reduced dimensionality in the inverse direction to a range [MinCoef, MaxCoef] inclusive, wherein MinCoef and / or MaxCoef are two integer values that are a function of a condition of the current video block.
[0525] 5. The method of example 4, wherein the condition of the current video block is a type of color or luma component represented by the current video block.
[0526] 6. The method of example 1, wherein the constraint rule specifies applying the secondary transform with reduced dimensionality to one or more MxN sub-blocks of the current video block and zeroing remaining sub-blocks of the current video block.
[0527] 7. The method of example 1, wherein the constraint rule specifies applying the secondary transform with reduced dimensionality differently to different sub-blocks of the current video block.
[0528] 8. The method of any of examples 1-5, wherein the constraint rule specifies applying the secondary transform with reduced dimensionality to exactly one MxN sub-block of the current video block due to a size of the current video block being 4xH or Wx4, wherein H is a height in integer pixels and W is a width in integer pixels.
[0529] 9. The method of example 8, wherein H > 8 or W > 8.
[0530] 10. The method of any of examples 1 to 9, wherein the current video block is a non-square region of a video.
[0531] 11. The method of example 2 or 3, wherein the constraint rule specifies zeroing transform coefficients of the primary transform in the forward direction, or padding zero coefficients to the output of the secondary transform in the reverse direction.
[0532] Further embodiments of examples 1-5 are described in item 1 of section 4. Further embodiments of examples 6-7 are described in item 2 of section 4. Further embodiments of examples 8-9 are described in item 3 of section 4. Further embodiments of examples 10-11 are described in item 4 of section 4.
[0533] 12. A video processing method comprising: determining a constraint rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and a neighboring video region and pixels of the current video block and pixels of the neighboring region, and performing the conversion by applying the secondary transform with reduced dimensionality according to the constraint rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block and the neighboring video region, and wherein the secondary transform with reduced dimensionality is applied in a specific order with a primary transform during the conversion.
[0534] 13. The method of example 12, wherein the neighboring video region comprises a left-top block of the current video block.
[0535] 14. The method of example 12, wherein the current video block and the neighboring video region correspond to sub-blocks of a parent video block.
[0536] Further embodiments of examples 12-14 are described in item 5 of section 4.
[0537] 15. A video processing method comprising: determining a zero-out rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality according to the zero-out rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block; wherein the zero-out rule specifies a maximum number of coefficients used by the secondary transform with reduced dimensionality.
[0538] 16. The method of example 15, wherein the maximum number of coefficients is a function of a component identification of the current video block.
[0539] 17. The method of example 16, wherein the maximum number of coefficients is different for luma video blocks and chroma video blocks.
[0540] 18. The method of any of examples 15 to 17, wherein the zero-out rule specifies a zero-out range as a function of coding information of the current video block.
[0541] 19. The method of any of examples 15 to 17, wherein the zero-out rule specifies a zero-out range as a function of a quantization parameter of the current video block.
[0542] 20. The method of any of examples 15 to 19, wherein the zero-out range is indicated in the bitstream representation by a field included at a sequence parameter set level, or a picture parameter set level, or a picture header or slice header, or a tile group header, or a coding tree unit row, or a coding tree unit, or a coding unit, or at a video data unit level.
[0543] Further embodiments of examples 15-17 are described in clause 4 item 6. Further embodiments of example 18 are described in clause 4 item 7. Further embodiments of example 19 are described in clause 4 item 8. Further embodiments of example 20 are described in clause 4 item 9.
[0544] 21. A video processing method, comprising: determining a condition for selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to the condition; wherein the reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block; and wherein the condition is signaled in the bitstream representation.
[0545] 22. The method of example 21, wherein the condition is a color format or a use of separate plane coding or based on a color identification of the current video block.
[0546] Further embodiments of examples 21-22 are described in clause 4 item 10.
[0547] 23. The method of any of examples 21 to 22, wherein the condition is signaled in the bitstream representation separately for chroma and luma components.
[0548] Further embodiments of example 23 are described in clause 4 item 11.
[0549] 24. The method of any of examples 21 to 23, wherein the condition depends on coding information of the current video block and a neighboring video region.
[0550] 25. The method of example 24, wherein the condition excludes application of the current video block coded using a particular intra prediction mode.
[0551] Other embodiments of examples 24-25 are described in item 12 of section 4.
[0552] 26. The method of example 24, wherein the condition specifies application of the current video block inter coded.
[0553] 27. The method of example 24, wherein the condition specifies application of the current video block coded using an intra block copy mode.
[0554] Other embodiments of examples 25-26 are described in item 13 of section 4.
[0555] 28. The method of example 21, wherein the condition is signaled in a bitstream representation at a level such that all blocks within the level conform to the condition, wherein the level is a sequence parameter set level, or a picture parameter set level, or a picture header, or a slice header, or a tile group header, or a coding tree unit row, or a coding tree unit, or a coding unit, or at a video data unit level.
[0556] Other embodiments of example 28 are described in item 14 of section 4.
[0557] 29. The method of example 21, wherein the condition is that the current video block is coded using a transform skip mode.
[0558] Other embodiments of example 29 are described in item 17 of section 4.
[0559] 30. A video processing method, comprising: selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition; wherein the reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block; and wherein the conversion includes selectively applying a position-dependent intra prediction combination (PDPC) based on a coexistence rule.
[0560] 31. The method of example 30, wherein the coexistence rule excludes application of the PDPC to the current video block due to the application of the secondary transform.
[0561] 32. The method of example 30, wherein the coexistence rule specifies application of the PDPC to the current video block due to the application of the secondary transform.
[0562] 33. The method of example 30, wherein the selectively applying the secondary transform is performed on the current video block using the PDPC.
[0563] Further embodiments of examples 30-33 are described in clause 4, item 15.
[0564] 34. A method of video processing, comprising: applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition; wherein the reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block; and wherein applying controls use of neighboring samples for intra prediction during the conversion.
[0565] Further embodiments of example 34 are described in clause 4, item 16.
[0566] 35. A method of video processing, comprising: selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to a condition; wherein the reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block; and wherein selectively applying controls use of a quantization matrix during the conversion.
[0567] 36. The method of example 35, wherein the use of the quantization matrix occurs only as a result of applying the secondary transform.
[0568] Further embodiments of examples 35-36 are described in clause 4, item 22.
[0569] 37. The method of any of examples 1-36, wherein the primary transform and the secondary transform are stored as transform matrices having a bit-width less than 8.
[0570] 38. The method of any of examples 1-36, wherein the primary transform and the secondary transform are stored as prediction transform matrices.
[0571] 39. The method of any of examples 1-36, wherein the primary transform is derivable from the secondary transform using a first rule, or wherein the secondary transform is derivable from the primary transform using a second rule.
[0572] 40. The method of any of examples 1-36, wherein the bitstream representation includes information about the secondary transform or the primary transform prior to residual information for the current video block.
[0573] Further embodiments of examples 37-40 are described in clause 4, items 18, 19, 20, and 21.
[0574] 41. The method of example 1, wherein the constraint rule for selectively applying the secondary transform depends on a number of transform units in a coding unit of the current video block.
[0575] 42. The method of example 41, wherein the constraint rule specifies that the secondary transform is applied due to the number of transform units in the coding unit being greater than one.
[0576] 43. The method of example 1, wherein a flag in the bitstream representation indicates whether a secondary transform with reduced dimensionality is applied to the conversion.
[0577] 44. The method of example 1, wherein the current video block comprises more than one component video block, and wherein the constraint rule specifies that the secondary transform with reduced dimensionality is differentially applicable for different component video blocks.
[0578] 45. The method of example 44, wherein the constraint rule specifies applicability of the secondary transform with reduced dimensionality for a first component video block based on how the constraint rule applies to a second component video block.
[0579] 46. The method of any of examples 44-45, wherein the constraint rule further depends on a dimensionality of the current video block.
[0580] Other embodiments of examples 47-53 are described in, for example, items 30 to 38 of section 4.
[0581] 47. A method of video processing, comprising: for a conversion between a current video block of a video and a bitstream representation of the video, determining whether to use a separable secondary transform (SST) for the conversion based on a coding condition; and performing the conversion in accordance with the determination.
[0582] 48. The method of example 47, wherein the coding condition corresponds to a syntax element in the bitstream representation.
[0583] 49. The method of example 48, wherein the coding condition comprises a size of the current video block.
[0584] 50. The method of any of examples 47-49, wherein, when it is determined to use the SST, the conversion uses a selected SST that is selected from a set of SSTs based on another coding condition.
[0585] 51. The method of example 50, wherein the another coding condition comprises a dimensionality of the current video block.
[0586] 52. The method of any of examples 1-51, wherein the conversion comprises decoding and parsing the bitstream representation to generate the video.
[0587] 53. The method of any of examples 1-51, wherein the converting comprises encoding the video into a bitstream representation.
[0588] 54. A video processing apparatus comprising a processor configured to implement any one or more of examples 1 to 53.
[0589] 55. A computer readable medium having code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any one or more of examples 1 to 53.
[0590] It should be appreciated that the disclosed techniques can be embodied in a video encoder or decoder to improve compression efficiency using techniques including using a secondary transform with reduced dimensionality.
[0591] As Figure 21 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0592] The system 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored or transmitted via connected communication (as represented by component 1906). The bitstream (or coded) representation of the video received at the input 1902 can be used by a component 1908 for generating pixel values or displayable video to a display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder, while corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0593] Examples of a peripheral bus interface or a display interface can include a universal serial bus (USB) or a high-definition multimedia interface (HDMI) or Displayport, etc. Examples of a storage interface include a SATA (serial advanced technology attachment), a PCI, an IDE interface, etc. The technology described in this document can be embodied in various electronic devices such as a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0594] Figure 22 is a flowchart of another example method of video processing in accordance with the present technology. The method 2000 includes, at operation 2010, for a conversion between a video unit of a video and a bitstream representation of the video, determining whether a separable secondary transform (SST) tool is enabled for the video unit. The method 2000 includes, at operation 2020, performing the conversion based on the determination.
[0595] In some embodiments, the determination is based on a syntax structure associated with the video unit. In some embodiments, the syntax structure is included in the bitstream representation, and the syntax structure is coded with at least one context used in an arithmetic coding.
[0596] In some embodiments, the syntax structure is omitted from the bitstream representation based on a characteristic of the video unit. In some embodiments, the characteristic of the video unit includes at least a dimension of the video unit, a coding mode of the video unit, or a syntax flag associated with the video unit. In some embodiments, the syntax structure is omitted from the bitstream representation in a case where a coding block flag associated with the video unit is equal to zero.
[0597] In some embodiments, the video unit includes a coding block or a transform block, and the syntax structure includes a coding unit or a transform unit. In some embodiments, the video unit includes a picture, and the syntax structure includes a picture header or a picture parameter set. In some embodiments, the video unit includes a slice, and the syntax structure includes a slice header, a sequence header, or a sequence parameter set. In some embodiments, the syntax structure includes a video parameter set, a decoder parameter set, an adaptation parameter set, a tile group, a tile, a coding tree unit row, or a coding tree unit.
[0598] In some embodiments, the determining is based on a characteristic of the video unit. In some embodiments, the characteristic of the video unit comprises at least a coding mode of the video unit, an intra prediction mode of the video unit, a type of primary transform of the video unit, or a dimension of the video unit. In some embodiments, the video unit comprises a block of the video, and the characteristic of the video unit comprises a dimension of the block. In some embodiments, the SST tool is disabled in a case that at least one of a width or a height of the block is less than or equal to a threshold Tmin. In some embodiments, Tmin is equal to 2 or 4. In some embodiments, the SST tool is disabled in a case that at least one of a width or a height of the block is greater than or equal to a threshold Tmax. In some embodiments, Tmax is equal to 32, 64, or 128. In some embodiments, the determining is based on the width of the block and a height of a color component of the block. In some embodiments, the color component comprises a luma component or an R color component. In some embodiments, the determining is based on the width of the block and a height of all color components of the block.
[0599] In some embodiments, in a case that the SST tool is determined to be disabled, signaling of usage of the SST tool or information related to the SST tool is omitted in the bitstream representation. In some embodiments, the SST is enabled for a first color component of the video unit and disabled for a second color component of the video unit.
[0600] Figure 23 is a flowchart of another example method of video processing in accordance with the present technology. The method 2100 includes, at operation 2110, for a conversion between a video unit of a video and a bitstream representation of the video, determining, based on a right-bottom positioning (SRx, SRy) of a scan region, a manner of indicating usage of a transform tool or a transform matrix used by the transform tool. The method 2100 also includes, at operation 2120, performing the conversion based on the determining.
[0601] In some embodiments, the transform tool comprises at least a separable secondary transform (SST), a non-separable secondary transform, or a primary transform. In some embodiments, in a case that SRx is greater than or equal to a first threshold and / or SRy is greater than or equal to a second threshold, an indication of usage of the transform tool or an indication of the transform matrix is omitted in the bitstream representation. In some embodiments, in a case that SRx is less than or equal to a third threshold and / or SRy is less than or equal to a fourth threshold, the indication of usage of the transform tool or the indication of the transform matrix is omitted in the bitstream representation. In some embodiments, in a case that the indication of usage of the transform tool or the indication of the transform matrix is omitted in the bitstream representation, the transform tool is considered to be disabled. In some embodiments, in a case that the indication of usage of the transform tool or the indication of the transform matrix is omitted in the bitstream representation, a default transform matrix is used. In some embodiments, coefficients outside the scan region are zero.
[0602] Figure 23 is a flowchart of another example method of video processing in accordance with the present technology. The method 2200 includes, at operation 2210, for a conversion between a block of a video and a bitstream representation of the video, determining a transform matrix used in a separable secondary transform (SST) tool based on a characteristic of the block. The SST tool provides a set of available transform matrices. The method 2200 also includes, at operation 2220, performing the conversion based on the determining.
[0603] In some embodiments, the characteristic of the block comprises at least a dimension of the block, an intra prediction mode of the block, quantized / unquantized coefficients after application of a transform, a color component of the block, or a type of primary transform of the block. In some embodiments, a same syntax element associated with the use of the SST tool indicates different transform matrices for blocks of different dimensions. In some embodiments, the set of available transform matrices comprises at least a 4x4 matrix, an 8x8 matrix, an NxN matrix, or an N’xN’ matrix, where N and N’ are integers. The transform matrix is determined according to a condition with respect to the characteristic of the block. In some embodiments, the transform matrix is determined to be a 4x4 matrix in a case that the condition specifies that at least one of a width or a height of the block is equal to 4. In some embodiments, the transform matrix is determined to be an 8x8 matrix in a case that the condition specifies that at least one of the width or the height of the block is equal to 8.
[0604] In some embodiments, the transform matrix is determined to be an 8x8 matrix in a case that the condition specifies that at least one of the width or the height of the block is greater than or equal to 8. In some embodiments, the transform matrix is determined to be an NxN matrix in a case that the condition specifies that at least one of the width or the height of the block is equal to N. In some embodiments, the transform matrix is determined to be an NxN matrix in a case that the condition specifies that at least one of the width or the height of the block is greater than or equal to N. In some embodiments, the transform matrix is determined to be an NxN matrix to be applied to a left-top NxN sub-block of the block. In some embodiments, the transform matrix is determined to be an NxN matrix to be applied to an NxN sub-block of the block that is different from the left-top NxN sub-block. In some embodiments, the NxN sub-block is adjacent to the top NxN sub-block rightward or bottomward.
[0605] In some embodiments, the manner in which the SST tool is applied is based on a characteristic of the block, the manner including applying the SST tool horizontally and / or vertically. In some embodiments, a first transform matrix is determined to be usable as a horizontal transform for the block, a second transform matrix is determined to be usable as a vertical transform for the block, and the first transform matrix and the second transform matrix are different. In some embodiments, the first transform matrix and the second transform matrix have different dimensions. In some embodiments, the first transform matrix has a size of NxN, the second transform matrix has a size of MxM, and the block has a dimension of WxH, and determining M or N is based on at least one of W or H. In some embodiments, in a case that W is equal to Wl, N is determined to be Wl, where Wl is equal to 4 or 8. In some embodiments, in a case that W is greater than or equal to W2, N is determined to be W2, where W2 is 4 or 8. In some embodiments, in a case that H is equal to HI, M is determined to be HI, where HI is equal to 4 or 8. In some embodiments, in a case that H is greater than or equal to H2, M is determined to be H2, where H2 is equal to 4 or 8. In some embodiments, the set of transform matrices includes at least two matrices having the same dimension.
[0606] In some embodiments, the transform matrix is indicated using a syntax element in the bitstream representation. In some embodiments, the transform matrix is derived without any indication in the bitstream representation. In some embodiments, different transform matrices are applied to different color components of the block. In some embodiments, for each color component of the block, the transform matrix is determined independently based on a characteristic of the block. In some embodiments, the use of the transform matrix for each component of the block is signaled separately in the bitstream. In some embodiments, the same transform matrix is applied to all color components of the block. In some embodiments, the transform matrix is first determined to be applied to a first color component of the block based on a characteristic of the block, and then it is determined that the same transform matrix is applied to all remaining color components of the block. In some embodiments, the first color component includes a luma component, a Cb component, or a Cr component. In some embodiments, in a case that the transform matrix is not applied to a second color component of the remaining color components, the SST tool is disabled for the second color component.
[0607] In some embodiments, the SST tool is enabled in a case that the transform matrices applicable to different color components of the block are the same. In some embodiments, in a case that it is determined to disable the SST tool, signaling of the use of the SST tool or information related to the SST tool in the bitstream representation is omitted. In some embodiments, different transform matrices are applied for different primary transform types. In some embodiments, the different primary transform types include at least DCT2 or DST7.
[0608] In some embodiments, performing the conversion includes generating the bitstream representation based on the video block. In some embodiments, performing the conversion includes generating the video block from the bitstream representation.
[0609] Figure 24 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0610] As shown in Figure 23 , video coding system 100 can include a source device 110 and a destination device 120. Source device 110 (which can be referred to as a video encoding device) generates encoded video data. Destination device 120 (which can be referred to as a video decoding device) can decode the encoded video data generated by source device 110.
[0611] Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0612] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via I / O interface 116, by network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by destination device 120.
[0613] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0614] I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.
[0615] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.
[0616] Figure 24 is a block diagram illustrating an example of a video encoder 200, which can be Figure 19 video encoder 114 in system 100 illustrated in FIG. 1.
[0617] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 25 In an example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared amongst the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0618] The functional components of video encoder 200 can include partition unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.
[0619] In other instances, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0620] Furthermore, some components, such as motion estimation unit 204 and motion compensation unit 205, can be highly integrated, but are represented separately in Figure 23 examples for the purpose of explanation.
[0621] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0622] The mode selection unit 203 can select one of the coding modes (intra or inter), for example, based on the error results, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for the block.
[0623] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples for pictures other than the picture associated with the current video block from the buffer 213.
[0624] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0625] In some examples, the motion estimation unit 204 can perform single directional prediction for a current video block, and the motion estimation unit 204 can search the reference pictures of List 0 or List 1 for a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index indicating the reference picture containing the reference video block in List 0 or List 1, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0626] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search the reference pictures in list 0 for a reference video block for the current video block, and also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 can then generate a reference index that indicates the reference picture in list 0 and list 1 that contains the reference video block, and a motion vector that indicates a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as the motion information for the current video block. The motion compensation unit 205 can generate the predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0627] In some examples, the motion estimation unit 204 can output the full set of motion information for use in decoding processing by the decoder.
[0628] In some examples, the motion estimation unit 204 can not output the full set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0629] In one instance, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[0630] In another instance, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0631] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0632] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0633] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0634] In other examples, the residual data for the current video block can not exist for the current video block, such as in a skip mode, and residual generation unit 207 can not perform the subtraction operation.
[0635] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0636] After transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video.
[0637] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0638] After reconstruction unit 212 reconstructs the video block, in-loop filtering operations can be performed to reduce video block artifacts in the video block.
[0639] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0640] Figure 25 is a block diagram illustrating an example of a video decoder 300 that can be Figure 25 video decoder 114 in the system 100 described in
[0641] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 24In examples of video decoder 300, video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0642] In In examples of video decoder 300, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to video encoder 200
[0643] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy encoded video data and, from the entropy decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, motion compensation unit 302 can determine such information by performing AMVP and Merge modes.
[0644] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for an interpolation filter to use at sub-pixel precision can be included in the syntax elements.
[0645] Motion compensation unit 302 can use the interpolation filter used by video encoder 20 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from received syntax information and use the interpolation filter to generate the prediction block.
[0646] Motion compensation unit 302 can use some of the syntax information to determine the following: block sizes used to encode frames and / or slices of the encoded video sequence, information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.
[0647] The intra prediction unit 303 can form a prediction block from spatial neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0648] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0649] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode when processing video blocks, but can not necessarily have to modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from video blocks to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, a decoder will process a bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.
[0650] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode when converting video blocks to a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder will process a bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.
[0651] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation (or vice versa). The bitstream representation of a current video block can correspond, for example, to bits co-located within the bitstream or bits distributed at different locations within the bitstream as defined by the syntax. For example, a macroblock can be encoded according to transformed and encoded error residual values, and can also be encoded using bits in the header and bits in other fields in the bitstream.
[0652] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include, in addition to a hardware component, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0653] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0654] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA or an ASIC.
[0655] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0656] Although the present patent document contains many details, these should not be construed as limiting the scope of any subject matter or of any embodiment, but as merely describing features that can be typical of some embodiments. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0657] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to implement such a methodology. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0658] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of video processing, comprising: determining, for a first conversion between a video unit of a video and a bitstream of the video, whether a separable secondary transform (SST) tool is enabled for the video unit; and performing the first conversion based on the determination; wherein the determination is based on a characteristic of the video unit, wherein the SST is enabled for a first color component of the video unit and disabled for a second color component of the video unit. the determination is based on a syntax structure associated with the video unit.
2. The method of claim 1, wherein, the syntax structure is included in the bitstream, and wherein the syntax structure is coded with at least one context used in an arithmetic coding.
3. The method of claim 2, wherein, the syntax structure is omitted in the bitstream based on a characteristic of the video unit.
4. The method of claim 2, wherein, the characteristic of the video unit comprises at least a dimension of the video unit, a coding mode of the video unit, or a syntax flag associated with the video unit.
5. The method of claim 4, wherein, the syntax structure is omitted in the bitstream in a case where a coded block flag (CBF) associated with the video unit is equal to 0.
6. The method of claim 4, wherein, the video unit comprises a coding block or a transform block, and wherein the syntax structure comprises a coding unit or a transform unit.
7. The method of any one of claims 2 to 6, wherein, the video unit comprises a picture, and wherein the syntax structure comprises a picture header or a picture parameter set.
8. The method of any one of claims 2 to 6, wherein, the video unit comprises a slice, and wherein the syntax structure comprises a slice header, a sequence header, or a sequence parameter set.
9. The method of any one of claims 2 to 6, wherein, the syntax structure comprises a video parameter set, a decoder parameter set, an adaptation parameter set, a tile group, a tile, a coding tree unit row, or a coding tree unit.
10. The method of any one of claims 2 to 6, wherein, the characteristic of the video unit comprises at least a coding mode of the video unit, an intra prediction mode of the video unit, a type of primary transform of the video unit, or a dimension of the video unit.
11. The method of claim 1, wherein, the video unit comprises a block of the video, and wherein the characteristic of the video unit comprises a dimension of the block.
12. The method of claim 1, wherein, the SST tool is disabled in a case where at least one of a width or a height of the block is less than or equal to a threshold Tmin.
13. The method of claim 12, wherein, 14. The method of claim 13, wherein Tmin is equal to 2 or 4. the SST tool is disabled in a case where at least one of a width or a height of the block is greater than or equal to a threshold Tmax.
15. The method of claim 14, wherein, 16. The method of claim 15, wherein Tmax is equal to 32, 64, or 128. the determination is based on a width of the block and a height of a color component of the block.
17. The method of any one of claims 12 to 16, wherein, the color component comprises a luma component or an R color component.
18. The method of claim 17, wherein, the determination is based on a width of the block and a height of all color components of the block.
19. The method of any one of claims 12 to 16, wherein, in a case where it is determined to disable the SST tool, signaling usage of the SST tool or information related to the SST tool is omitted in the bitstream.
20. The method of any one of claims 1 to 4, 11-16, and 18, wherein, 21. The method of video processing of claim 1, further comprising: determining, for a second conversion between a video unit of a video and a bitstream of the video, a manner of indicating usage of a transform tool or a transform matrix used by the transform tool based on a right bottom position of a scan region (SRx, SRy); and performing said second conversion based on said determining.
22. The method of claim 21, wherein, The transform tool comprises at least a separable secondary transform SST, a non-separable secondary transform or a primary transform.
23. The method of claim 21 or 22, wherein, In case SRx is greater than or equal to a first threshold and / or SRy is greater than or equal to a second threshold, an indication of a use of the transform tool or an indication of the transform matrix is omitted in the bitstream.
24. The method of claim 21 or 22, wherein, In case SRx is less than or equal to a third threshold and / or SRy is less than or equal to a fourth threshold, an indication of a use of the transform tool or an indication of the transform matrix is omitted in the bitstream.
25. The method of claim 21 or 22, wherein, In case the indication of the use of the transform tool or the transform matrix is omitted in the bitstream, the transform tool is considered disabled.
26. The method of claim 21 or 22, wherein, In case the indication of the use of the transform tool or the transform matrix is omitted in the bitstream, a default transform matrix is used.
27. The method of claim 21 or 22, wherein, Coefficients located outside the scanning region are zero.
28. The video processing method of claim 1, further comprising: determining, for a third conversion between a block of a video and a bitstream of the video, a transform matrix used in a separable secondary transform SST tool based on a characteristic of the block, wherein the SST tool provides a set of available transform matrices; and performing the third conversion based on the determining.
29. The method of claim 28, wherein, The characteristic of the block comprises at least a dimension of the block, an intra prediction mode of the block, quantized / unquantized coefficients after application of a transform, a color component of the block or a type of primary transform of the block.
30. The method of claim 28 or 29, wherein, A same syntax element associated with the use of the SST tool indicates different transform matrices for blocks of different dimensions.
31. The method of claim 28 or 29, wherein, The set of available transform matrices comprises at least a 4x4 matrix, an 8x8 matrix,... or an NxN matrix, wherein N is an integer, and wherein the transform matrix is determined according to a condition with respect to the characteristic of the block.
32. The method of claim 31, wherein, In case the condition specifies that at least one of the width or the height of the block is equal to 4, the transform matrix is determined to be the 4x4 matrix.
33. The method of claim 31, wherein, In case the condition specifies that at least one of the width or the height of the block is equal to 8, the transform matrix is determined to be the 8x8 matrix.
34. The method of claim 31, wherein, In case the condition specifies that at least one of the width or the height of the block is greater than or equal to 8, the transform matrix is determined to be the 8x8 matrix.
35. The method of claim 31, wherein, In case the condition specifies that at least one of the width or the height of the block is equal to N, the transform matrix is determined to be the NxN matrix.
36. The method of claim 31, wherein, In case the condition specifies that at least one of the width or the height of the block is greater than or equal to N, the transform matrix is determined to be the NxN matrix.
37. The method of claim 28 or 29, wherein, The transform matrix is determined to be the NxN matrix to be applied to a left-top NxN sub-block of the block.
38. The method of claim 28 or 29, wherein, The transform matrix is determined to be an NxN matrix to be applied to an NxN sub-block of the block different from a left-top NxN sub-block.
39. The method of claim 38, wherein, The NxN sub-block is adjacent right or bottom to a top NxN sub-block.
40. The method of claim 28 or 29, wherein, A manner in which the SST tool is applied is based on the characteristic of the block, the manner comprising a horizontal and / or vertical application of the SST tool.
41. The method of claim 40, wherein, A first transform matrix is determined to be usable as a horizontal transform for the block, and a second transform matrix is determined to be usable as a vertical transform for the block, and wherein the first transform matrix and the second transform matrix are different.
42. The method of claim 41, wherein, The first transform matrix and the second transform matrix have different dimensions.
43. The method of claim 41, wherein, The first transform matrix has a size of N x N, the second transform matrix has a size of M x M, and the block has a dimension of W x H, and wherein determining M or N is based on at least one of W or H.
44. The method of claim 43, wherein, In case W is equal to W1, N is determined to be W1, wherein W1 is equal to 4 or 8.
45. The method of claim 43, wherein, In case W is greater than or equal to W2, N is determined to be W2, wherein W2 is 4 or 8.
46. The method of claim 43, wherein, In case H is equal to H1, M is determined to be H1, wherein H1 is equal to 4 or 8.
47. The method of claim 43, wherein, In case H is greater than or equal to H2, M is determined to be H2, wherein H2 is equal to 4 or 8.
48. The method of claim 28 or 29, wherein, The set of available transform matrices comprises at least two matrices having the same dimension.
49. The method of claim 48, wherein, The transform matrix is indicated using a syntax element in the bitstream.
50. The method of claim 48, wherein, The transform matrix is derived without any indication in the bitstream.
51. The method of claim 28 or 29, wherein, Different transform matrices are applied to different color components of the block.
52. The method of claim 51, wherein, For each color component of the block, a transform matrix is independently determined based on a characteristic of the block.
53. The method of claim 52, wherein, The use of the transform matrix for each component of the block is separately signaled in the bitstream.
54. The method of claim 28 or 29, wherein, The same transform matrix is applied to all color components of the block.
55. The method of claim 54, wherein, A first color component of the block comprises a luma component, a Cb component, or a Cr component.
56. The method of claim 54, wherein, In case the transform matrix is not applied to a second color component of the remaining color components, the SST tool is disabled for the second color component.
57. The method of claim 55, wherein, In case the transform matrices applicable to different color components of the block are the same, the SST tool is enabled.
58. The method of claim 28 or 29, wherein, In case the SST tool is determined to be disabled, signaling of the use of the SST tool or information related to the SST tool in the bitstream is omitted.
59. The method of claim 28 or 29, wherein, Different transform matrices are applied for different primary transform types.
60. The method of claim 28 or 29, wherein, The different primary transform types comprise at least DCT2 or DST7.
61. The method of claim 60, wherein, Performing the first conversion comprises generating the bitstream based on the video.
62. The method of any one of claims 1 to 6 and 11 to 16, wherein, Performing the first conversion comprises generating the video from the bitstream.
63. The method of any one of claims 1 to 6 and 11 to 16, wherein, 64. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement a method according to any of claims 1 to 63.
65. A computer program product comprising a computer program which, when executed by a processor, implements a method according to any of claims 1 to 63.
66. A computer readable medium having stored thereon a bitstream of a video, the bitstream being generated according to a method of any of claims 1 to 63.
Citation Information
Patent Citations
Primary transform and secondary transform in video coding
US20180103252A1
Method and apparatus for encoding / decoding video signal using secondary transform
US20190356915A1