Codec video processing using enhanced secondary transform

By introducing a secondary transformation technique during video encoding and optimizing the conversion rules between video blocks and bitstreams, the problem of high bandwidth requirements in video encoding is solved, achieving more efficient video data compression and decoding, and making it applicable to existing and future video encoding and decoding standards.

CN115606182BActive Publication Date: 2026-02-06DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180024692.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-25
Filing Date
2021-03-25
Publication Date
2026-02-06
Estimated Expiration
2041-03-25

AI Technical Summary

Technical Problem

Existing video coding technologies have high bandwidth requirements when processing video data, making it difficult to effectively compress and decode, especially in high-resolution and high-frame-rate video data, resulting in excessive network bandwidth consumption.

Method used

The method employs a quadratic transform technique, which involves applying a quadratic transform during the conversion between video blocks and bitstreams. The use of the quadratic transform is determined by rules. Combined with motion-compensated interpolation filters and interleaving prediction modes, a quadratic transform with reduced dimensionality is selectively applied. The method controls the use of adjacent samples in quantization matrix and intra-frame prediction to optimize the video encoding and decoding process.

Benefits of technology

By optimizing the video encoding and decoding process, the bandwidth requirements for video data are reduced, encoding efficiency is improved, network load is reduced, and it is applicable to existing and future video encoding and decoding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115606182B_ABST
    Figure CN115606182B_ABST
Patent Text Reader

Abstract

A method of video processing includes performing a conversion between a video block of a video and a bitstream of the video according to a rule. The rule specifies whether or how to indicate, in the bitstream, a use of a secondary transform within a video unit. The secondary transform is applied before quantization or after inverse quantization.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This is the national phase of International Patent Application No. PCT / CN2021 / 082962, filed on March 25, 2021, claiming priority and benefit to International Patent Application No. PCT / CN2020 / 081174, filed on March 25, 2020. For all purposes required by law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology

[0004] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document describes various embodiments and techniques for using quadratic transform (also known as low-frequency non-separable transform) during the decoding or encoding of video or images.

[0006] In one example aspect, a method for video processing is disclosed. This method includes performing conversions between video blocks and a video bitstream according to rules. The rules specify whether or how a quadratic transform is used within a video unit in the bitstream. The quadratic transform is applied between forward master transform and quantization, or between inverse quantization and inverse master transform.

[0007] In another example, a method for video processing is disclosed. This method includes performing conversions between video blocks and the video bitstream according to rules. These rules specify the use of separable quadratic transforms within the video blocks based on syntax elements associated with those blocks. The separable quadratic transforms are applied between the forward master transform and quantization, or between inverse quantization and the inverse master transform.

[0008] In another example, a method for video processing is disclosed. This method includes performing a conversion between video blocks and a bitstream of video according to rules. The rules specify the selection of a quadratic transform from a plurality of separable quadratic transforms to be applied to the video blocks. The quadratic transform is applied to rows or columns of the video blocks.

[0009] In another example, a method for video processing is disclosed. This method includes determining one or more interpolation filters for motion compensation of video blocks in a video based on conditions, and performing a conversion between the video blocks and the video bitstream according to the determination.

[0010] In another example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video block of a video and a bitstream representation of the video according to a rule. The video block is coded using an interweaved prediction mode in which the video block is partitioned into sub-blocks using a first pattern and a second pattern, and a final prediction is determined as a weighted sum of two auxiliary predictions having the first pattern and the second pattern. The rule specifies that the two auxiliary predictions having the first pattern and the second pattern include a single prediction and a bi-prediction, where the first pattern and the second pattern are different.

[0011] In another example aspect, a method of video processing is disclosed. The method includes determining a constraint rule for selectively applying a secondary transform having a reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the secondary transform having the reduced dimension according to the constraint rule. The secondary transform having the reduced dimension has a dimension reduced from a dimension of the current video block. The secondary transform having the reduced dimension is applied together with a primary transform in a particular order during the conversion.

[0012] In another example aspect, another method of video processing is disclosed. The method includes determining a constraint rule for selectively applying a secondary transform having a reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block and a bitstream representation of a neighboring video region and pixels of the neighboring region and performing the conversion by applying the secondary transform having the reduced dimension according to the constraint rule. The secondary transform having the reduced dimension has a dimension reduced from a dimension of the current video block and the neighboring video region. The secondary transform having the reduced dimension is applied together with a primary transform in a particular order during the conversion.

[0013] In yet another example aspect, another method of video processing is disclosed. The method includes determining a zero-out rule for selectively applying a secondary transform having a reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the secondary transform having the reduced dimension according to the zero-out rule. The secondary transform having the reduced dimension has a dimension reduced from a dimension of the current video block. The zero-out rule specifies a maximum number of coefficients used by the secondary transform having the reduced dimension.

[0014] In yet another example aspect, another method of video processing is disclosed. The method includes determining a condition for selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the reduced-dimension secondary transform according to the condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The condition is signaled in the bitstream representation.

[0015] In yet another example aspect, another method of video processing is disclosed. The method includes determining a condition for selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the reduced-dimension secondary transform according to the condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The condition is signaled in the bitstream representation.

[0016] In yet another example aspect, another method of video processing is disclosed. The method includes selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the reduced-dimension secondary transform according to the condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The conversion includes selectively applying position dependent intra prediction combination (PDPC) based on a coexistence rule.

[0017] In yet another example aspect, another method of video processing is disclosed. The method includes applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the reduced-dimension secondary transform according to the condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The application controls use of neighboring samples for intra prediction during the conversion.

[0018] In yet another example aspect, another method of video processing is disclosed. The method includes selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block and performing the conversion by applying the reduced-dimension secondary transform according to the condition. The reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block. The selective application controls use of a quantization matrix during the conversion.

[0019] In yet another example aspect, another method of video processing is disclosed. The method includes determining, for a conversion between a current video block of a video and a bitstream representation of the video, whether to use a separable secondary transform (SST) for the conversion based on a coding condition; and performing the conversion according to the determination.

[0020] In yet another example aspect, a video encoder is disclosed. The video encoder includes a processor configured to implement one or more of the above-described methods.

[0021] In yet another example aspect, a video decoder is disclosed. The video decoder includes a processor configured to implement one or more of the above-described methods.

[0022] In yet another example aspect, a computer readable medium is disclosed. The medium includes code for implementing one or more of the above-described methods stored on the medium.

[0023] These and other aspects are described in the present document. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 An example of an encoder block diagram is shown.

[0025] Figure 2 An example of 67 intra prediction modes is shown.

[0026] Figure 3A An example of reference samples for wide-angle intra prediction is shown.

[0027] Figure 3B Another example of reference samples for wide-angle intra prediction is shown.

[0028] Figure 4 An example illustration of discontinuity issues in case of directions exceeding 45 degrees.

[0029] Figure 5A An example illustration of samples used for PDPC applied to diagonal and neighboring angular intra modes is shown.

[0030] Figure 5B Another example illustration of samples used for PDPC applied to diagonal and neighboring angular intra modes is shown.

[0031] Figure 5C Another example illustration of samples used for PDPC applied to diagonal and neighboring angular intra modes is shown.

[0032] Figure 5D Yet another example illustration of samples used for PDPC applied to diagonal and neighboring angular intra modes is shown.

[0033] Figure 6 An example of 4x8 and 8x4 block partitioning is shown.

[0034] Figure 7 An example of partitioning for all blocks except 4x8, 8x4 and 4x4 is shown.

[0035] Figure 8 A 4x8 block of samples is divided into two independent decodable regions.

[0036] Figure 9 An example order of pixel row processing is shown to maximize the throughput of 4xN blocks with vertical predictors.

[0037] Figure 10 An example of a secondary transform is shown.

[0038] Figure 11 An example of the proposed reduced secondary transform (RST) is shown.

[0039] Figure 12 An example of the forward and inverse (or reverse) reduced transform is shown.

[0040] Figure 13 An example of a forward RST 8x8 process with a 16x48 matrix is shown.

[0041] Figure 14 An example of scanning positions 17 to 64 for non-zero elements is shown.

[0042] Figure 15 is an illustration of sub-block transform modes SBT-V and SBT-H.

[0043] Figure 16 is a block diagram of an example hardware platform for implementing the techniques described in this document.

[0044] Figure 17 is a flowchart of an example method of video processing.

[0045] Figure 18A An example of scan region based coefficient coding is shown.

[0046] Figure 18B Another example of scan region based coefficient coding is shown.

[0047] Figure 19 is a block diagram of an example video processing system in which the disclosed techniques can be implemented.

[0048] Figure 20 An example of a simplified affine motion model is shown.

[0049] Figure 21 An example of an affine MV per sub-block is shown.

[0050] Figure 22 An example of an MVP for AF INTER is shown.

[0051] Figure 23AAn example of a candidate for AF_MERGE is shown.

[0052] Figure 23B Another example of a candidate for AF_MERGE is shown.

[0053] Figure 24A An example of a block partitioned by 4x4 sub-blocks is shown.

[0054] Figure 24B An example of a block partitioned by 4x4 sub-blocks with offset is shown.

[0055] Figure 25 An example of weighting values in sub-blocks is shown.

[0056] Figure 26A An example of a partition pattern is shown.

[0057] Figure 26B Another example of a partition pattern is shown.

[0058] Figure 27A An example of weighting values in sub-blocks is shown.

[0059] Figure 27B Another example of weighting values in sub-blocks is shown.

[0060] Figure 27C Another example of weighting values in sub-blocks is shown.

[0061] Figure 27D Yet another example of weighting values in sub-blocks is shown.

[0062] Figure 28 is a block diagram illustrating an example video coding system.

[0063] Figure 29 is a block diagram illustrating an encoder in accordance with some embodiments of the disclosure.

[0064] Figure 30 is a block diagram illustrating a decoder in accordance with some embodiments of the disclosure.

[0065] Figure 31 is a flowchart representation of a method for video processing in accordance with the present technology.

[0066] Figure 32 is a flowchart representation of another method for video processing in accordance with the present technology.

[0067] Figure 33 is a flowchart representation of another method for video processing in accordance with the present technology.

[0068] Figure 34 is a flowchart representation of another method for video processing in accordance with the present technology.

[0069] Figure 35 is a flowchart representation of yet another method for video processing according to the present technology. DETAILED DESCRIPTION

[0070] The use of section headings in this document is for convenience only and does not limit the disclosure to the sections in which the embodiments are described. Moreover, although certain embodiments are described with reference to the versatile video coding or other specific video codecs, the disclosed technology is applicable to other video coding technologies as well. Furthermore, although some embodiments describe video coding steps in detail, it will be understood that the corresponding steps of decoding a codec would be implemented by a decoder. Moreover, the term video processing includes video coding or compression, video decoding or decompression, and transcoding of video pixels from one compressed format representation to another compressed format representation or at different compression bit rates.

[0071] 1. SUMMARY

[0072] This patent document relates to video coding technology. In particular, it is about related transforms in video coding. It can be applied to existing video coding standards, such as HEVC, or standards under finalization (versatile video coding). It can also be applicable to future video coding standards or video codecs.

[0073] 2. PRELIMINARY DISCUSSION

[0074] Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure where temporal prediction is used in combination with transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and applied to the reference software named Joint Exploration Model (JEM). In April 2018, the Joint Video Team (JVT) was formed between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to develop the VVC standard, targeting at 50% bitrate reduction compared to HEVC.

[0075] 2.1 Color Space and Chroma Subsampling

[0076] A color space, also called a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a tuple of numbers, usually 3 or 4 values or color components (e.g. RGB). Fundamentally, a color space is a refinement of a coordinate system and subspace.

[0077] For video compression, the most commonly used color spaces are YCbCr and RGB.

[0078] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces, used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue-difference and red-difference chrominance components. Y' (prime) is not the same as Y, which is luminance, meaning that the light intensity is non-linearly encoded based on the RGB primaries with gamma correction.

[0079] Chroma subsampling is a way of encoding an image by implementing a lower resolution for the chrominance information than for the luminance information, exploiting the fact that the human visual system is less acute in its discrimination of color differences than in its discrimination of luminance.

[0080] 2.1.1 Format 4:4:4

[0081] Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used for high-end film scanners and in film post-production.

[0082] 2.1.2 Format 4:2:2

[0083] The two chroma components are sampled at half the luminance sampling rate: horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by a factor of three with little visual difference.

[0084] 2.1.3 Format 4:2:0

[0085] In 4:2:0, the horizontal sampling is twice that of 4:1:1, but in this scheme the vertical resolution is halved because the Cb and Cr channels are only sampled on every other line. Thus, the data rate is the same. Cb and Cr are subsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme with different horizontal and vertical addressing.

[0086] In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels in the vertical direction (in the gaps).

[0087] In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in the gap positions between alternate luma samples.

[0088] In 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternate lines.

[0089] 2.2 Encoding and decoding process of a typical video codec

[0090] Figure 1 An example of the encoder block diagram of VVC is shown, which contains three in-loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF which uses a pre-defined filter, SAO and ALF utilize the original samples of the current picture, with the coding side information of the signaled offsets and filter coefficients, to reduce the mean square error between the original and reconstructed samples by adding offsets and applying a Finite Impulse Response (FIR) filter, respectively. ALF is located at the last processing stage of each picture and can be regarded as a tool that tries to capture and fix the artifacts produced by the previous stages.

[0091] 2.3 Intra mode coding with 67 intra prediction modes

[0092] To capture arbitrary edge directions presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The additional directional modes are depicted as dashed arrows in Figure 2 , while the planar mode and DC mode remain unchanged. These denser directional intra prediction modes are applied to all block sizes and for both luma and chroma intra prediction.

[0093] As shown in Figure 2 , the traditional angular intra prediction directions are defined in the clockwise direction as 45 degrees to -135 degrees. In VTM2, for non-square blocks, several of the traditional angular intra prediction modes are replaced adaptively by wide-angle intra prediction modes. The replaced modes are signaled using the original method and remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes is unchanged, e.g., 67, and the intra mode coding is unchanged.

[0094] In HEVC, each intra coded block has a square shape and its length of each side is a power of 2. Therefore, generating an intra predictor using the DC mode does not require a division operation. In VVV2, a block can have a rectangular shape, which in general case requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average for non-square blocks.

[0095] 2.4 Wide-angle intra prediction for non-square blocks

[0096] Conventional angular intra prediction directions are defined in a clockwise direction as 45 degrees to -135 degrees. In VTM2, for non-square blocks, several conventional angular intra prediction modes are replaced adaptively by wide-angle intra prediction modes. The replaced modes are signaled using the original method and remapped to the index of the wide-angle modes after parsing. The total number of intra prediction modes for a particular block is unchanged, e.g., 67, and the intra mode coding is unchanged.

[0097] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined as shown in Figures 3A to 3B

[0098] In wide-angle directional modes, the number of modes that are replaced depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 1.

[0099] Table 1 Intra prediction modes replaced by wide-angle modes

[0100]

[0101] As shown in Figure 4 , in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and a lateral smoothing are applied to the wide-angle prediction to reduce the negative impact of the increased gap α .

[0102] 2.5 Position Dependent Intra Prediction Combination

[0103] In VTM2, the intra prediction result of the planar mode is further modified by a position dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that invokes a combination of unfiltered boundary reference samples and filtered boundary reference samples of HEVC-style intra prediction. PDPC is applied to the following non-signaled intra modes: planar, DC, horizontal, vertical, bottom-left mode and its eight neighboring angular modes, and top-right mode and its eight neighboring angular modes.

[0104] Using a linear combination of the intra prediction mode (DC, planar, angular) and the reference samples, the prediction sample pred(x, y) is predicted according to the following equation:

[0105] pred(x, y) = (wL x R -1,y + wT x R x,-1 - wTL x R -1,-1 + (64 - wL - wT + wTL) x pred(x, y) + 32) » 6

[0106] where R x,-1 , R -1,y ​denote the reference samples located at the top and left side of the current sample (x, y), respectively, and R -1,-1 denotes the reference sample located at the top-left corner of the current block.

[0107] If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filter is needed, which is necessary in the case of the HEVC DC mode boundary filter or the horizontal / vertical mode edge filter.

[0108] Figures 5A to 5D The definition of the reference samples (R x,-1 , R -1,y , and R -1,-1 ) applied by PDPC on various prediction modes is shown. The prediction sample pred(x’, y’) is located at (x’, y’) within the prediction block. The coordinate x of the reference sample R x,-1 is given by x = x’ + y’ + 1, and the coordinate y of the reference sample R -1,y is similarly given by y = x’ + y’ + 1.

[0109] Figures 5A to 5D The definition of the samples used by PDPC applied to diagonal and adjacent angle intra modes is provided.

[0110] The PDPC weights depend on the prediction mode, as shown in Table 2.

[0111] Table 2. Example of PDPC weights according to prediction mode

[0112] Predictive patterns wT wL wTL diagonal top right 16 >> ((y' << 1) » shift) 16 >> ((x' << 1) » shift) 0 diagonal bottom left 16 >> ((y' << 1) » shift) 16 >> ((x' << 1) » shift) 0 adjacent diagonal top right 32 >> ((y' << 1) » shift) 0 0 Adjacent diagonal bottom left 0 32 >> ((x' << 1) » shift) 0

[0113] 2.6 Intra Sub-Block Partitioning (ISP)

[0114] In some embodiments, ISP is proposed to partition a luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions according to the block size dimensions, as shown in Table 3. Figure 6 and Figure 7 An example of these two possibilities is shown. All sub-partitions satisfy the condition of having at least 16 samples.

[0115] Table 3. Number of sub-partitions depending on block size

[0116]

[0117]

[0118] Figure 6 An example of the partitioning of 4x8 and 8x4 blocks is shown.

[0119] Figure 7 An example of the partitioning of all blocks except 4x8, 8x4 and 4x4 is shown.

[0120] For each of these sub-partitions, a residual signal is generated by entropy decoding the coefficients sent by the encoder, then inverse quantizing and inverse transforming them. The sub-partition is then intra predicted and the corresponding reconstructed samples are finally obtained by adding the residual signal to the predicted signal. The reconstructed values of each sub-partition will thus be available to generate the prediction for the next sub-partition, which will repeat the process and so on. All sub-partitions share the same intra mode.

[0121] Depending on the intra mode and partitioning used, two different processing orders are used, called normal and reverse order. In normal order, the first sub-partition to be processed is the one containing the top-left sample of the CU, then it continues either down (horizontal partitioning) or to the right (vertical partitioning). As a result, the reference samples used to generate the sub-partition prediction signal are only located to the left and above the line. On the other hand, the reverse processing order either starts from the sub-partition containing the bottom-left sample of the CU and continues upwards, or from the sub-partition containing the top-right sample of the CU and continues to the left.

[0122] 2.7 Block Differential Pulse Code Modulation Coding (BDPCM)

[0123] Due to the shape of the horizontal (vertical) predictor, which uses the left (top) pixel to predict the current pixel, the most efficient way to process a block in terms of throughput is to process all the pixels of a column (row) in parallel and the columns (rows) sequentially. To improve the throughput, we introduce the following procedure: when the predictor selected on the block is vertical, a block of width 4 is split into two halves with horizontal boundaries; when the predictor selected on the block is horizontal, a block of height 4 is split into two halves with vertical boundaries.

[0124] When a block is split, a sample from one region is not allowed to use a pixel from the other region to compute the prediction: if this happens, the predicted pixel is replaced by the reference pixel in the prediction direction. This is illustrated in Figure 8 for different positions of the current pixel X in a 4x8 block with vertical prediction.

[0125] Figure 8 An example of splitting a block of 4x8 samples into two independent decodable regions is shown.

[0126] Thanks to this property, it is now possible to process a 4x4 block in 2 cycles, a 4x8 or 8x4 block in 4 cycles, and so on, as illustrated in Figure 9 .

[0127] Figure 9 An example of processing order for the pixels rows to maximize the throughput of a 4xN block with vertical predictor is shown.

[0128] Table 4 summarizes the number of cycles required to process a block, depending on the block size. It is shown that any block with both dimensions greater than or equal to 8 can be processed with 8 pixels or more per cycle is not significant.

[0129] Table 4 Worst case throughput for blocks of size 4 x N, N x 4

[0130]

[0131] 2.8 Quantized residual domain BDPCM

[0132] In some embodiments, a quantized residual domain BDPCM (hereinafter RBDPCM) is proposed. Intra prediction is performed on the entire block by copying samples in a similar prediction direction (horizontal or vertical prediction) as intra prediction. The residual is quantized and the difference between the quantized residual and its predictor (horizontal or vertical) quantized value is coded.

[0133] For a block of size M (rows) x N (columns), let r i,j , 0≤i≤M-1, 0≤j≤N-1 be the prediction residual after intra prediction using unfiltered samples from above or left block boundary samples (horizontal (copying left neighboring pixel value on the predicted block row-wise) or vertical (copying top neighboring row to each row in the predicted block) respectively). Let Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1 denote the quantized version of the residual r i,j , where the residual is the difference between the original block value and the predicted block value. Block DPCM is then applied to the quantized residual samples resulting in a modified M x N array with elements When vertical BDPCM is signaled:

[0134]

[0135] For horizontal prediction, similar rules apply and the residual quantized samples are obtained from

[0136]

[0137] The residual quantized samples are sent to the decoder.

[0138] At the decoder side, the above calculations are reversed to produce Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1. For the vertical prediction case,

[0139]

[0140] ​For the horizontal case,

[0141]

[0142] Inverse quantized residual Q -1 (Q(r i,j )) is added to the intra-predicted value to produce the reconstructed sample value.

[0143] The main advantage of this approach is that the inverse DPCM can be done on-the-fly during the coefficient parsing, requiring only the addition of the predictor at the coefficient parsing time, but it can also be performed after the parsing.

[0144] Transform skipping is usually used for quantized residual domain BDPCM.

[0145] 2.9 Multiple Transform Set (MTS) in VVC

[0146] In VTM4, large block size transforms with maximum size 64x64 are enabled, which is mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with size (width or height, or both width and height) equal to 64, the high frequency transform coefficients are zeroed out, thus only low frequency coefficients are kept. For example, for an MxN transform block, with M the block width and N the block height, when M is equal to 64, only the left 32 columns of transform coefficients are kept. Similarly, when N is equal to 64, only the first 32 rows of transform coefficients are kept. When transform skipping mode is used for large blocks, the whole block is used without zeroing out any values.

[0147] In addition to the DCT-II already employed in HEVC, a Multiple Transform Selection (MTS) scheme is used for residual coding of inter and intra coded blocks. It uses multiple selected transforms among DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. The following table shows the basis functions of the selected DST / DCT.

[0148]

[0149] In order to maintain the orthogonality of the transform matrices, the quantization of the transform matrices is more accurate than in HEVC. In order to keep the mid value of the transformed coefficients in the 16-bit range, all coefficients should be 10-bit after the horizontal and vertical transform.

[0150] In order to control the MTS scheme, separate enabling flags are signaled at SPS level for intra and inter respectively. When MTS is enabled at SPS, a CU flag is signaled to indicate whether MTS is applied or not. Here, MTS is only applied to luma. The MTS CU level flag is signaled when the following conditions are met.

[0151] - both width and height are less than or equal to 32

[0152] - CBF flag is equal to 1

[0153] If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two flags are additionally signaled to indicate the transform type in the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Tables 3 to 10. In terms of transform matrix precision, an 8-bit primary transform kernel is used. Therefore, all transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8 use an 8-bit primary transform kernel.

[0154]

[0155] To reduce the complexity of large size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, the high frequency transform coefficients are set to zero. Only the coefficients within the 16x16 low frequency region are kept.

[0156] As in HEVC, the residual of a block can be coded with transform skip mode. To avoid the redundancy of syntax coding, the transform skip flag is not signaled when the CU level MTS CU flag is not equal to zero. The block size limit for transform skip is the same as that of MTS in JEM4, which indicates that transform skip is applicable to CUs with both block width and block height equal to or smaller than 32.

[0157] 2.10 Example Reduced Secondary Transform (RST)

[0158] 2.10.1 Example of Inseparable Quadratic Transformation (NSST)

[0159] In some embodiments, a secondary transform, also known as a non-separable transform, is applied between the forward primary transform and quantization (at the encoder) and between the inverse quantization and inverse primary transform (at the decoder side). As shown in Figure 10 A 4x4 (or 8x8) secondary transform is performed according to the block size, for example. A 4x4 secondary transform is applied to small blocks (e.g., min(width, height) < 8), while an 8x8 secondary transform is applied to larger blocks of each 8x8 block (e.g., min(width, height) > 4).

[0160] Figure 10 An example of the secondary transform in JEM is shown.

[0161] The application of the non-separable transform is described below with an example of an input. To apply the non-separable transform, a 4x4 input block X

[0162]

[0163] First, the vector is denoted as

[0164]

[0165] The non-separable transform is computed as where denotes the transform coefficient vector, T is a 16x16 transform matrix. The 16x1 coefficient vector is then reorganized into 4x4 blocks using the scan order (horizontal, vertical or diagonal) of the block. The coefficients with smaller indices will be placed in the 4x4 coefficient block with smaller scan indices. There are 35 transform sets, each using 3 non-separable transform matrices (kernels). The mapping from intra prediction modes to transform sets is predefined. For each transform set, the selected non-separable secondary transform candidate is further specified by an explicitly signaled secondary transform index. After the transform coefficients, the index is signaled once in the bitstream within each CU.

[0166] 2.10.2 Example Reduced Quadratic Transform (RST) / Low-Frequency Inseparable Transform (LFNST)

[0167] Reduced secondary transform (RST), also known as low-frequency non-separable transform (LFNST), is introduced as 4 transform sets (instead of 35 transform sets) mapping. In some embodiments, 16x64 (can be further reduced to 16x48) and 16x16 matrices are used for 8x8 and 4x4 blocks, respectively. For convenience of notation, the 16x64 (can be further reduced to 16x48) transform is denoted as RST8x8, while the 16x16 transform is denoted as RST4x4. Figure 11 An example of RST is shown.

[0168] Figure 11 An example of the proposed reduced secondary transform (RST) is shown.

[0169] RST computation

[0170] The main idea of reduced transform (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N (R

[0171] The RT matrix is an R x N matrix as shown below:

[0172]

[0173] where the R rows of the transform are R bases of the N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. Figure 12 Examples of forward and inverse RT are depicted in

[0174] Figure 12 Examples of forward and inverse reduced transform are shown.

[0175] In some embodiments, RST 8x8 with a reduction factor of 4 (1 / 4 size) is applied. Thus, instead of a traditional 8x8 non-separable transform matrix size of 64x64, a 16x64 direct matrix is used. In other words, a 64x16 inverse RST matrix is used at the decoder side to produce the core (primary) transform coefficients in the 8x8 top-left corner region. The forward RST 8x8 uses a 16x64 (or 8x64 for 8x8 blocks) matrix, so it only produces non-zero coefficients in the top-left 4x4 region within the given 8x8 region. In other words, if RST is applied, the 8x8 region except the top-left 4x4 region will have only zero coefficients. For RST 4x4, a 16x16 (or 8x16 for 4x4 blocks) direct matrix multiplication is applied.

[0176] The inverse RST is conditionally applied when the following two conditions are met:

[0177] a. The block size is greater than or equal to a given threshold (W>=4 && H>=4)

[0178] b. The transform skip mode flag is equal to zero

[0179] If the width (W) and height (H) of the transform coefficient block are both greater than 4, RST 8x8 is applied to the top-left 8x8 region of the transform coefficient block. Otherwise, RST 4x4 is applied on the top-left min(8,W)xmin(8,H) region of the transform coefficient block.

[0180] If the RST index is equal to 0, no RST is applied. Otherwise, RST is applied, where the kernel is selected using the RST index. The RST selection method and coding of the RST index will be explained later.

[0181] In addition, RST is applied within and across CUs in slices, as well as for luma and chroma. If dual tree is enabled, the RST index for luma and chroma is signaled separately. For across-slice (dual tree disabled), a single RST index is signaled and used for luma and chroma.

[0182] In some embodiments, Intra Sub-Partition (ISP) is adopted as a new intra prediction mode. When ISP mode is selected, RST is disabled and RST index is not signaled, because even if RST is applied to each feasible partitioned block, the performance improvement is marginal. In addition, disabling RST on ISP prediction residual can reduce the encoding complexity.

[0183] RST selection

[0184] One RST matrix is selected from four transform sets, each of which consists of two transforms. Which transform set is applied is determined by the intra prediction mode, as follows:

[0185] (1) If one of the three CCLM modes is indicated, transform set 0 is selected.

[0186] (2) Otherwise, transform set selection is performed as per the following table:

[0187] Transform set selection table

[0188]

[0189]

[0190] The index into the table is denoted as IntraPredMode, which ranges from [-14, 83], which is the transform mode index for wide-angle intra prediction.

[0191] Reduced dimension RST matrices

[0192] As a further simplification, a 16x48 matrix is applied instead of a 16x64 with the same transform set configuration, each of which takes 48 input data from three 4x4 blocks in the top-left 8x8 block (excluding the right-bottom 4x4 block) Figure 13 ).

[0193] Figure 13 An example of the forward RST 8x8 process with a 16x48 matrix is shown.

[0194] RST signaling

[0195] Forward RST 8x8 with R=16 uses a 16x64 matrix, so it only produces non-zero coefficients in the top-left 4x4 region within the given 8x8 region. In other words, if RST is applied, the 8x8 region except the top-left 4x4 region only produces zero coefficients. As a result, when any non-zero element is detected within the 8x8 block region except the top-left 4x4 (as shown in Figure 14 ), the RST index is not coded because it means that RST is not applied. In this case, the RST index is inferred to be zero.

[0196] Figure 14 An example for non-zero element scan positions 17 to 64 is shown.

[0197] Zero-out range

[0198] In general, any coefficient in a 4x4 sub-block can be non-zero before applying inverse RST on the 4x4 sub-block. However, there are restrictions in some cases that some coefficients in a 4x4 sub-block must be zero in order for inverse RST to be applied on the sub-block.

[0199] Let nonZeroSize be a variable. It is required that any coefficient with index not less than nonZeroSize must be zero when rearranging the coefficients into a one-dimensional array before inverse RST.

[0200] When nonZeroSize is equal to 16, there is no zeroing constraint on the coefficients in the top-left 4x4 sub-block.

[0201] In some embodiments, when the current block size is 4x4 or 8x8, nonZeroSize is set to be equal to 8. For other block sizes, nonZeroSize is set to be equal to 16.

[0202] Example description of RST

[0203] In the following tables and descriptions, bold italicized text is used to indicate changes that can be made to the current syntax to accommodate certain embodiments described in this document.

[0204] Sequence parameter set RBSP syntax

[0205]

[0206] Residual coding syntax

[0207]

[0208]

[0209] Coded unit syntax

[0210]

[0211]

[0212] Sequence parameter set RBSP semantics

[0213]

[0214]

[0215]

[0216] Coded unit semantics

[0217]

[0218]

[0219] Transform process of scaling transform coefficients

[0220] Generally

[0221] The input of the process is:

[0222] - a luma position (XTbY, yTbY) specifying the top-left sample of the current luma transform block relative to the top-left luma sample of the current picture,

[0223] - a variable nTbW specifying the width of the current transform block,

[0224] - a variable nTbH specifying the height of the current transform block,

[0225] - a color component variable cldx specifying the current block,

[0226] - an (nTbW) x (nTbH) array d[x][y] of scaling transform coefficients, with x = 0..nTbW-1, y = 0..nTbH-1.

[0227] The output of the process is an (nTbW) x (nTbH) array r[x][y] of residual samples, with x = 0..nTbW-1, y = 0..nTbH-1.

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235] 2.11 Clipping of dequantization in HEVC

[0236] In HEVC, the scaled transform coefficients d' are computed as d' = Clip3(coeffMin, coeffMax, d), where d is the scaled transform coefficient before clipping.

[0237] For the luma component, coeffMin = CoeffMinY; coeffMax = CoeffMaxY. For the chroma components, coeffMin = CoeffMinC; coeffMax = CoeffMaxC; where

[0238] CoeffMinY = -(1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15))

[0239] CoeffMinC = -(1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))

[0240] CoeffMaxY = (1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) - 1

[0241] CoeffMaxC = (1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) - 1

[0242] extended_precision_processing_flag is a syntax element signaled in the SPS.

[0243] 2.12 Affine Linear Weighted Intra Prediction (ALWIP, also known as Matrix-based Intra Prediction, MIP)

[0244] In some embodiments, two tests are conducted. In Test 1, ALWIP is designed with a memory limit of 8K bytes, and up to 4 multiplications per sample point. Test 2 is similar to Test 1, but further simplifies the design in terms of memory requirement and model architecture.

[0245] * A single set of matrices and offset vectors for all block shapes.

[0246] * The number of modes is reduced to 19 for all block shapes.

[0247] * The memory requirement is reduced to 5760 10-bit values, i.e., 7.20 kilobytes.

[0248] * The linear interpolation of the prediction samples is performed in a single step in each direction, instead of the iterative interpolation in the first test.

[0249] 2.13 Sub-block Transform

[0250] For inter predicted CUs with cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is decoded. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is coded with an inferred adaptive transform and another portion of the residual block is zeroed. SBT is not applied to combined inter-intra modes.

[0251] In sub-block transform, position dependent transforms are applied to luma transform blocks in SBT-V and SBT-H (DCT-2 is always used for chroma TBs). Two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transforms for each SBT position are specified in Figure 15 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of a residual TU is larger than 32, the corresponding transform is set to DCT-2. Thus, sub-block transform jointly specifies the TU tiling, cbf, and horizontal and vertical transforms for a residual block, which can be considered as a syntax shortcut for the case where the main residual of a block is located at one side of the block.

[0252] Figure 15 is an illustration of sub-block transform modes SBT-V and SBT-H.

[0253] 2.14 Separable secondary transform in AVS

[0254] In some embodiments, if the primary transform is DCT2, a 4x4 separable secondary transform (SST) is applied to all luma blocks coded with intra mode after the primary transform.

[0255] When the SST is applied to a block at the encoder, the top-left 4x4 sub-block of the transformed block after the primary transform (denoted as L) is further transformed as L' = T' x L x T,

[0256] where T is the secondary transform matrix.

[0257] L' is then quantized together with the other parts of the transformed block.

[0258] When the SST is applied to a block at the decoder, the top-left 4x4 sub-block of the dequantized transformed block (denoted as M) is further inverse transformed as

[0259] M' = S' x M x S,

[0260] where S is the inverse secondary transform matrix. Specifically, S' = T.

[0261] M' is then input to the primary inverse transform together with the other parts of the transformed block.

[0262] 2.15 Scan Region-based Coefficient Coding (SRCC)

[0263] SRCC has been adopted by AVS-3. With SRCC, as shown in Figures 18A to 18B the bottom-right position (SRx, SRy) is signaled, and only the coefficients within the rectangle (e.g., scan region) with four corners (0, 0), (SRx, 0), (0, SRy), (SRx, SRy) are scanned and signaled. All the coefficients outside the rectangle are zero.

[0264] 2.16 Affine Prediction

[0265] In HEVC, only translational motion model is applied to the motion compensated prediction (MCP). However, in the real world, there are many kinds of motion, such as zooming-in / zooming-out, rotation, perspective motion, and other irregular motions. In JEM, a simplified affine transform motion compensated prediction is applied. As shown in Figure 20 the affine motion field of a block is described by two control point motion vectors.

[0266] The motion vector field (MVF) of a block is described by the following equation:

[0267]

[0268] where (v0x, v0y) is the motion vector of the top-left control point and (v1x, v1y) is the motion vector of the top-right control point.

[0269] To further simplify the motion compensated prediction, a subblock-based affine transform prediction is applied. The subblock size M x N is derived as in Equation (2), where MvPreis the motion vector fractional precision (1 / 16 in JEM), (v2x, v2y) is the motion vector of the bottom-left control point, which is calculated according to Equation 1.

[0270]

[0271] After being derived by Equation (2), M and N should be down-adjusted to be the divisors of w and h, respectively, if necessary.

[0272] To derive the motion vector of each M x N subblock, the motion vector of the center sample of each subblock is calculated according to Equation (1), as shown in Figure 21 and rounded to 1 / 16 fractional precision. Then a motion compensated interpolation filter is applied to generate the prediction of each subblock with the derived motion vector.

[0273] After MCP, the high-precision motion vector of each subblock is rounded and saved at the same precision as the normal motion vector.

[0274] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. AF_INTER mode can be applied to CUs with both width and height greater than 8. Signaling in the bitstream informs the affine flag at the CU level to indicate whether AF_INTER mode is used. In this mode, adjacent blocks are used to construct motion vector pairs {(v0,v1)|v0={v...}. A ,v B ,v c},v1={vD,v E The candidate list. For example... Figure 22 As shown, v0 is selected from the motion vectors of blocks A, B, or C. The motion vectors from neighboring blocks are scaled based on the relationship between the reference list and the POC of the references of neighboring blocks, the POC of the reference of the current CU, and the POC of the current CU. The method for selecting v1 from neighboring blocks D and E is similar. If the number of candidates in the candidate list is less than 2, the list is populated by copying the motion vector pairs consisting of each AMVP candidate. When the candidate list is greater than 2, the candidates are first sorted based on the consistency of adjacent motion vectors (the similarity between the two motion vectors in a candidate pair), and only the top two candidates are retained. An RD cost check is used to determine which motion vector pair is selected as the control point motion vector prediction (CPMVP) for the current CU. The index indicating the position of the CPMVP in the candidate list is signaled in the bitstream. After determining the CPMVP of the current affine CU, affine motion estimation is applied to find the control point motion vector (CPMV). The difference between the CPMV and CPMVP is then signaled in the bitstream.

[0275] When the CU is applied in AF_MERGE mode, it obtains the first block encoded / decoded in affine mode from the valid adjacent reconstructed blocks. The selection order of candidate blocks is from left, top, top right, bottom left to top left, as follows: Figure 23A As shown. If the adjacent lower left block A is encoded and decoded in affine mode, as... Figure 23B As shown, the motion vectors v2, v3, and v4 of the top-left, top-right, and bottom-left corners of the CU containing block A are derived. Then, the motion vector v0 of the top-left corner of the current CU is calculated based on v2, v3, and v4. Next, the motion vector v1 of the top-right corner of the current CU is calculated.

[0276] After deriving the CPMVv0 and v1 of the current CU, the MVF of the current CU is generated according to the simplified affine motion model equation (1). In order to identify whether the current CU is encoded or decoded in AF_MERGE mode, when at least one adjacent block is encoded or decoded in affine mode, the affine flag is signaled in the bit stream.

[0277] 2.17 Intertwined Prediction

[0278] To solve the dilemma of affine motion compensation (AMC), interleaved prediction is proposed to achieve a finer granularity of MVs without adding much complexity.

[0279] First, the coding block is divided into sub-blocks with two different division patterns. The first division pattern is the same as in BMS-1.1, as shown in Figure 24A , while the second division pattern also divides the coding block into 4x4 sub-blocks but with a 2x2 offset, as shown in Figure 24B .

[0280] Second, AMC generates two auxiliary predictions with these two division patterns. The MV of each sub-block in the division pattern is derived from mv0 and mv1 by equation (1).

[0281] The final prediction is calculated as a weighted sum of the two auxiliary predictions, formulated as:

[0282]

[0283] As shown in Figure 25 , the auxiliary prediction samples located at the center of the sub-block are associated with a weight value of 3, while the auxiliary prediction samples located at the boundary of the sub-block are associated with a weight value of 1.

[0284] In this contribution, the proposed interleaved prediction is only applied to the luma component of the affine coded block with single prediction, thus the bandwidth is unchanged in the worst case.

[0285] 3. Examples of problems solved by the embodiments

[0286] The current design has the following problems:

[0287] (1) The clipping and shifting / rounding operations in MTS / RST can not be optimal.

[0288] (2) The RST applied to the two adjacent 4x4 blocks can be costly.

[0289] (3) Different approaches can be taken for RST for different color components.

[0290] (4) RST can not be suitable for screen content coding.

[0291] (5) The interaction of RST with other coding tools is not clear.

[0292] (6) The RST transform matrix can be stored more efficiently.

[0293] (7) It is not clear how to apply the quantization matrix on RST.

[0294] 4. Example embodiments and techniques

[0295] The embodiments listed below should be considered as examples to explain general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0296] In the following description, the encoding / decoding information may include prediction modes (e.g., intra / inter / IBC modes), motion vectors, reference pictures, inter-frame prediction directions, intra-frame prediction modes, CIIP (Combined Intra-Inter-Frame Prediction) modes, ISP modes, affine intra-frame modes, the transform kernel used, transform skip flags, etc., for example, information required when coding blocks.

[0297] In the following discussion, SatShift(x, n) is defined as

[0298]

[0299] Shift(x, n) is defined as Shift(x, n) = (x + offset 0) >> n.

[0300] In one example, offset0 and / or offset1 are set to (1 < 0.05).<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.

[0301] In another example, offset0 = offset1 = ((1<<n)> >1)-1 or ((1<<(n-1)))-1.

[0302] Clip3(min, max, x) is defined as

[0303]

[0304] 1. After the reverse RST, the output value should be clipped to the range [MinCoef, MaxCoef] (inclusive), where MinCoef and / or MaxCoef are two integer values ​​that can vary.

[0305] a. In one example, assuming the dequantized coefficients are clipped to the range [QMinCoef, QMaxCoef] (inclusive of QMinCoef and QMaxCoef), then MinCoef can be set to equal QMinCoef and / or MaxCoef can be set to equal QMaxCoef.

[0306] b. In one example, MinCoef and / or MaxCoef may depend on the color components.

[0307] i. In one example, MinCoef and / or MaxCoef can depend on the bit depth of the corresponding color component.

[0308] c. In one example, MinCoef and / or MaxCoef can depend on the block shape (e.g., square or non-square) and / or block size.

[0309] d. In one example, the selection of the value or candidate values of MinCoef and / or MaxCoef can be signaled (such as in SPS, PPS, slice header / tile group header / CTU / CU).

[0310] e. In one example, for luma component, MinCoef and / or MaxCoef can be derived as:

[0311] MinCoef = -(1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15))

[0312] MaxCoef = (1 « (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) - 1

[0313] where BitDepthY is the bit depth of luma component, and extended_precision_processing_flag can be signaled such as in SPS.

[0314] f. In one example, for chroma component, MinCoef and / or MaxCoef can be derived as:

[0315] MinCoef = -(1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))

[0316] MaxCoef = (1 « (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) - 1

[0317] where BitDepthC is the bit depth of chroma component, and extended_precision_processing_flag can be signaled such as in SPS.

[0318] g. In some embodiments, MinCoef is -(1 « 15) and MaxCoef is (1 « 15) - 1.

[0319] h. In one example, the consistent bitstream should satisfy that the transform coefficients after forward RST should be within a given range.

[0320] 2. It is proposed that the way forward RST and / or inverse RST is applied on MxN coefficient subblocks can depend on the number of subblocks on which forward RST and / or inverse RST is applied, e.g. M=N=4.

[0321] a. In one example, the zero-out range can depend on the subblock index on which RST is applied.

[0322] i. Optionally, the zero-out range can depend on the number of subblocks on which RST is applied.

[0323] b. In one example, when there are S subblocks on which forward RST and / or inverse RST is applied in the entire coefficient block, the way forward RST and / or inverse RST is applied on the first and second subblock of coefficients can be different, where e.g. S>1, e.g. S=2. For example, the first MxN subblock can be the top-left MxN subblock.

[0324] i. In one example, nonZeroSize as described in section 2.10 can be different for the first MxN coefficient subblock (denoted as nonZeroSizeO) and for the second MxN coefficient subblock (denoted as nonZeroSizei).

[0325] 1) In one example, nonZeroSizeO can be larger than nonZeroSizei. For example, nonZeroSizeO=16 and nonZeroSizei=8.

[0326] ii. In one example, when only one MxN subblock is to apply forward RST and / or inverse RST, or when more than one MxN subblock is to apply forward RST and / or inverse RST, nonZeroSize as described in section 2.10 can be different.

[0327] 1) In one example, nonZeroSize can be equal to 8 if there are more than one MxN subblock to apply forward RST and / or inverse RST.

[0328] 3. If the current block size is 4xH or Wx4, where H>8 and W>8, it is proposed that forward RST and / or inverse RST is applied only on one MxN coefficient subblock, such as the top-left MxN subblock, e.g. M=N=4.

[0329] a. In one example, if H > T1 and / or W > T2, e.g., T1 = T2 = 16, then only one MxN coefficient sub-block is applied with forward RST and / or inverse RST.

[0330] b. In one example, if H < T1 and / or W < T2, e.g., T1 = T2 = 32, then only one MxN coefficient sub-block is applied with forward RST and / or inverse RST.

[0331] c. In one example, for all H > 8 and / or W > 8, only one MxN coefficient sub-block is applied with forward RST and / or inverse RST.

[0332] d. In one example, if the current block size is MxH or WxN, where H >= N and W >= M, e.g., M = N = 4, then only one MxN sub-block, such as the top-left MxN sub-block, is applied with forward RST and / or inverse RST,

[0333] 4. RST can be applied to non-square regions. Assume the region size is denoted by KxL, where K is not equal to L.

[0334] a. Optionally, in addition, zeroing can be applied to the transform coefficients after the forward RST, so as to meet the maximum number of non-zero coefficients.

[0335] i. In one example, if a transform coefficient is outside the top-left MxM region, where M is not larger than K and M is not larger than L, then the transform coefficient can be set to 0.

[0336] 5. It is proposed that the coefficients in two adjacent MxN sub-blocks can involve a single forward RST and / or inverse RST, e.g., M = N = 4.

[0337] a. In one example, one or several operations can be performed at the encoder. The operations can be performed in order.

[0338] i. Re-arrange the coefficients in two adjacent MxN sub-blocks into a one-dimensional vector with 2xMxN elements

[0339] ii. Apply a forward RST with a transform matrix having 2xMxN columns and MxN rows (or MxN columns and 2xMxN rows) on the one-dimensional vector.

[0340] iii. Re-arrange the transformed one-dimensional vector with MxN elements into a first MxN sub-block, such as the top-left sub-block.

[0341] iv. All coefficients in the second MxN sub-block can be set to zero.

[0342] b. In one example, one or several of the following operations can be performed at the decoder. The operations can be performed in order.

[0343] i. The coefficients in the first MxN subblock (such as the top-left subblock) are rearranged into a one-dimensional vector with MxN elements

[0344] ii. The inverse RST with a transform matrix of MxN columns and 2xMxN rows (or 2xMxN columns and MxN rows) is applied on the one-dimensional vector.

[0345] iii. The transformed one-dimensional vector with 2xMxN elements is rearranged into two adjacent MxN subblocks.

[0346] c. In one example, a block can be split into K (K>1) subblocks, and the primary transform and secondary transform can be performed at the subblock level.

[0347] 6. The zero-out range (e.g., nonZeroSize as described in section 2.10) can depend on the color component.

[0348] a. In one example, for the same block size, the range can be different for luma and chroma components.

[0349] 7. The zero-out range (e.g., nonZeroSize as described in section 2.10) can depend on the coding information.

[0350] a. In one example, the range can depend on the coding mode, such as intra or non-intra mode.

[0351] b. In one example, the range can depend on the coding mode, such as intra or inter or IBC mode.

[0352] c. In one example, the range can depend on the reference picture / motion information.

[0353] 8. The proposed zero-out range (e.g., nonZeroSize as described in section 2.10) for a particular block size can depend on the quantization parameter (QP).

[0354] a. In one example, it is assumed that when QP is equal to QPA, nonZeroSize is equal to nonZeroSizeA, and when QP is equal to QPB, nonZeroSize is equal to nonZeroSizeB. If QPA is not smaller than QPB, then nonZeroSizeA is not larger than nonZeroSizeB.

[0355] b. Different transform / inverse transform matrices can be used for different nonZeroSize.

[0356] 9. It is proposed that the zero-out range (e.g., nonZeroSize as described in Section 2.10) can be signaled such as in SPS, PPS, picture header, slice header, tile group header, CTU row, CTU, CU, or any video data unit.

[0357] a. Optionally, multiple ranges can be defined. And an indication of which candidate nonZeroSize to select can be signaled such as in SPS, PPS, picture header, slice header, tile group header, CTU row, CTU, and CU.

[0358] 10. Whether and / or how to apply RST can depend on the color format, and / or the use of separate plane coding, and / or the color component.

[0359] a. In one example, RST can not be applied to chroma components (such as Cb and / or Cr).

[0360] b. In one example, if the color format is 4:0:0, RST can not be applied to chroma components.

[0361] c. In one example, if separate plane coding is used, RST can not be applied to chroma components.

[0362] d. In one example, the nonZeroSize for a particular block size can depend on the color component.

[0363] i. In one example, for the same block size, the nonZeroSize on chroma components can be smaller than the nonZeroSize on luma components.

[0364] 11. It is proposed that when coding luma and chroma components with a single coding structure tree, the RST control information (such as whether to apply RST, and / or which set of transform matrices to select) can be signaled separately for luma and chroma components.

[0365] 12. Whether and how to apply RST can depend on the coding information (such as coding mode) of the current block and / or neighboring blocks.

[0366] a. In one example, RST cannot be used for one or more particular intra prediction modes.

[0367] i. For example, RST cannot be used for LM mode.

[0368] ii. For example, RST cannot be used for LM-T mode.

[0369] iii. For example, RST cannot be used for LM-A mode.

[0370] iv. For example, RST cannot be used for wide-angle intra prediction modes.

[0371] v. For example, RST cannot be used for BDPCM mode or / and DPCM mode or / and RBDPCM mode.

[0372] vi. For example, RST cannot be used for ALWIP mode.

[0373] vii. For example, RST cannot be used for certain specific angular intra prediction modes (such as DC, Planar, Vertical, Horizontal, etc.).

[0374] viii. For example, RST can be used for luma component, but cannot be used for chroma component in LM mode or / and LM-T mode or / and LM-A mode.

[0375] ix. For example, RST can not be used for chroma component when joint chroma residual coding is applied.

[0376] b. If RST cannot be applied, the syntax element indicating the information related to RST in the current block can not be signaled.

[0377] 13. It is proposed that RST can be applied to non-intra coded blocks.

[0378] a. In one example, RST can be applied to inter coded blocks.

[0379] b. In one example, RST can be applied to Intra Block Copy (IBC) coded blocks.

[0380] c. In one example, RST can be applied to blocks coded with Combined Intra Inter Prediction (CIIP).

[0381] 14. It is proposed that RST can be controlled at different levels.

[0382] a. For example, the information indicating whether RST such as a control flag is applicable can be signaled in PPS, slice header, picture header, tile group header, tile, CTU row, CTU.

[0383] b. Whether RST is applicable can depend on the standard profile / level / layer.

[0384] 15. It is proposed that whether to apply Position Dependent Intra Prediction Combination (PDPC) can depend on whether RST is applied.

[0385] a. In one example, if RST is applied for the current block, PDPC can not be applied.

[0386] b. In one example, if RST is applied for the current block, PDPC can be applied.

[0387] c. Optionally, whether RST is applied can depend on whether PDPC is applied.

[0388] i. In one example, RST is not applied when PDPC is applied.

[0389] ii. If RST cannot be applied, a syntax element indicating information related to RST in the current block can not be signaled.

[0390] 16. It is proposed that whether to filter neighboring samples for intra prediction can depend on whether RST is applied.

[0391] a. In one example, if RST is applied for the current block, the neighboring samples can not be filtered.

[0392] b. In one example, if RST is applied for the current block, the neighboring samples can be filtered.

[0393] c. Optionally, whether RST is applied can depend on whether the neighboring samples for intra prediction are filtered.

[0394] i. In one example, RST is not applied when the neighboring samples for intra prediction are filtered.

[0395] ii. In one example, RST is not applied when the neighboring samples for intra prediction are not filtered.

[0396] iii. If RST cannot be applied, a syntax element indicating information related to RST in the current block can not be signaled.

[0397] 17. It is proposed that RST can be applied when the current block is coded with transform skip.

[0398] a. For example, the primary transform is skipped, but a second transform can still be applied.

[0399] b. The secondary transform matrix used in the transform skip mode can be different from the secondary transform matrix used in the no transform skip mode.

[0400] 18. It is proposed that the transform matrix for RST can be stored with a bit width smaller than 8. For example, the transform matrix for RST can be stored with a bit width of 6 or 4.

[0401] 19. It is proposed that the transform matrix for RST can be stored in a predictive manner.

[0402] a. In one example, a first element in a first transform matrix for RST can be predicted by a second element in the first transform matrix for RST.

[0403] i. For example, a difference between the two elements can be stored.

[0404] ii. For example, the difference can be stored in less than 8 bits wide, such as 6 or 4.

[0405] b. In one example, a first element in a first transform matrix for RST can be predicted by a second element in a second transform matrix for RST.

[0406] i. For example, a difference between the two elements can be stored.

[0407] ii. For example, the difference can be stored in less than 8 bits wide, such as 6 or 4.

[0408] 20. It is proposed that the first transform matrix for RST can be derived from the second transform matrix for RST.

[0409] a. In one example, partial elements of the second transform matrix for RST can be picked up to construct the first transform matrix for RST.

[0410] b. In one example, the first transform matrix for RST is derived by rotating or flipping all or part of the second transform matrix for RST.

[0411] c. In one example, the first transform matrix for RST is derived by down- or up- sampling the second transform matrix for RST.

[0412] 21. It is proposed that syntax elements indicating information related to RST in the current block can be signaled before signaling the residual (which can be transformed).

[0413] a. In one example, the signaling of the information related to RST can not depend on counting non-zero or zero coefficients when parsing the residual.

[0414] b. In one example, non-zero or zero coefficients can not be counted when parsing the residual.

[0415] c. In one example, a coding block flag (cbf) flag for a sub-block set to all zeros by RST can not be signaled and inferred to be 0.

[0416] d. In one example, a significance flag for a coefficient set to zero by RST can not be signaled and inferred to be 0.

[0417] e. The scan order of parsing the residual block can depend on whether and how RST is applied.

[0418] i. In one example, no coefficients with RST set to zero can be scanned.

[0419] f. The arithmetic coding context for parsing the residual block can depend on whether and how RST is applied.

[0420] 22. Whether and how to apply a quantization matrix can depend on whether and how RST is applied.

[0421] a. In one example, different quantization matrices can be applied whether or not RST is applied.

[0422] b. Optionally, whether and how RST is applied can depend on whether and how a quantization matrix is applied.

[0423] i. In one example, RST can not be applied when a quantization matrix is applied on a block.

[0424] 23. It is proposed that RST can be applied to quantized coefficients / residuals.

[0425] a. In one example, RST can be applied to residuals when transform skip is used.

[0426] b. In one example, RST can be applied to quantized transform coefficients of a block.

[0427] 24. It is proposed that RST can be applied to sub-block transform blocks.

[0428] a. In one example, RST can be applied to the top-left coefficient generated by a sub-block transform.

[0429] 25. It is proposed that how and / or whether RST is applied can depend on the number of TUs in a CU.

[0430] a. For example, how and / or whether RST is applied can depend on whether the number of TUs in a CU is greater than 1.

[0431] i. In one example, if the number of TUs in a CU is greater than 1, RST is not applied.

[0432] ii. In one example, if the number of TUs in a CU is greater than 1, RST is only applied to one of the multiple TUs in the CU.

[0433] 1) In one example, if the number of TUs in a CU is greater than 1, RST is only applied to the first TU in the CU.

[0434] 2) In one example, if the number of TUs in a CU is greater than 1, RST is only applied to the last TU in the CU.

[0435] iii. In one example, if the number of TUs in a CU is independently greater than 1, RST is applied to each TU of the CU.

[0436] 1) Optionally, when the number of TUs in a CU is greater than 1, whether to apply RST to a first TU of the CU can be determined independently of whether to apply RST to a second TU of the CU.

[0437] 2) In one example, whether to apply RST to a TU of a CU can depend on the number of non-zero coefficients (denoted as NZ) of the TU, but not on the number of non-zero coefficients of other TUs of the CU when the number of TUs in the CU is greater than 1.

[0438] a) In one example, if NZ is less than a threshold T (e.g., T = 2), RST is not applied to the TU.

[0439] b) If it is determined that RST is not applied to a TU of a CU, the syntax element(s) indicating whether to apply RST can not be signaled for the TU of the CU.

[0440] b. For example, how and / or whether to apply RST can depend on whether the TU size is equal to the CU size.

[0441] i. In one example, RST is disabled when the CU size is greater than the TU size.

[0442] c. It is proposed to use the decoding information of the first TU or the last TU in the decoding order of the CU to decide the usage of RST and / or the signaling of RST related syntax elements.

[0443] i. In one example, RST is not applied to a CU if the number of non-zero coefficients of the first or last TU is less than a threshold T (e.g., T = 2).

[0444] ii. In one example, RST is not applied to a CU if the number of non-zero coefficients of a sub-region (e.g., top-left 4x4) within the first or last TU is less than a threshold T (e.g., T = 2).

[0445] 26. It is proposed to set a flag for a TU to control whether to apply RST.

[0446] a. Whether to apply RST to a TU can depend on the flag of the TU.

[0447] i. When the flag is not present or has not been derived, it can be derived as false.

[0448] ii. Optionally, when the flag is not present or has not been derived, it can be derived as true.

[0449] b. When the CU contains only one TU, the flag for the TU can be equal to the CU RST flag that can be derived dynamically (e.g., based on coefficient information).

[0450] c. When the number of TUs in the CU is greater than 1, the flag for the last TU in the CU can be derived from the CU RST flag that can be derived dynamically (e.g., based on coefficient information), and the flags for all other TUs can be set to false.

[0451] i. Optionally, when the number of TUs in the CU is greater than 1, the flag for the last TU in the CU can be derived from the CU RST flag, and the flags for all other TUs can be set to true.

[0452] 27. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether and / or how to apply RST to a first component of a block can be different from whether and / or how to apply RST to a second component of the block. That is, separate control of applying RST to different color components.

[0453] a. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a first component of a block can be determined independently of whether to apply RST to a second component of the block.

[0454] i. In one example, whether to apply RST to a component of a block can depend on the decoded information (e.g., the number of non-zero coefficients (denoted as NZ)) of the component of the block, but not on the decoded information of any other component of the block when the number of components is greater than 1 and a single coding tree is used.

[0455] 1) In one example, if NZ is less than a threshold T (e.g., T = 2), RST is not applied to the component of the block.

[0456] 2) If it is determined that RST is not applied to the component of the block, the syntax element(s) indicating whether to apply RST can not be signaled for the component of the block.

[0457] b. In one example, for the single tree case, whether to enable RST and / or how to apply RST can be determined independently for the luma and chroma components.

[0458] 28. It is proposed that when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a first component of a block can be determined by a second component of the block.

[0459] a. In one example, when the number of components is greater than 1 and a single coding tree is used, whether to apply RST to a first component of a block can be determined by the number of non-zero coefficients of a second component of the block.

[0460] i. In one example, if NZ (e.g., the number of non-zero coefficients of a second component of a block or a sub-region (e.g., top-left 4x4) of a block) is less than a threshold T (e.g., T = 2), then RST is not applied on the first component of the block.

[0461] ii. If it is determined that RST is not applied on the first component of the block, then syntax element(s) indicating whether RST is applied can not be signaled for the component of the block.

[0462] iii. In one example, the first component is Cb or Cr and the second component is Y.

[0463] iv. In one example, the first component is R or B and the second component is G.

[0464] 29. In one example, whether to apply bullet 25 and / or bullet 26 and / or bullet 27 can depend on the width and height (denoted as W and H) of the CU and / or TU and / or block and / or the maximum transform block size.

[0465] a. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T or H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.

[0466] b. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T and H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.

[0467] c. In one example, bullet 25 and / or bullet 26 and / or bullet 27 is applied only when W > T and H > T. In one example, T can be equal to 64. In an alternative example, T can be equal to the maximum transform size.

[0468] Improvements to Separable Quadratic Transform (SST)

[0469] 30. In one example, for a video unit, it can be determined to enable or disable SST.

[0470] a. For example, the determination can be made based on signaling in a video syntax structure associated with the video unit.

[0471] i. In one example, the signaling (such as a flag) can be coded with at least one context in an arithmetic coding.

[0472] ii.In one example, the signaling can be conditionally skipped based on coding / decoding information such as block size, coding block flag (cbf), and coding mode of the current block.

[0473] 1) In one example, the signaling can be skipped when cbf is equal to zero.

[0474] b. For example, the determination can be made based on inference without signaling associated with the video unit.

[0475] i. The inference can depend on information of the video unit, for example, coding mode, intra prediction mode, type of primary transform, and size or dimension of the video unit,

[0476] c. For example, the video unit can be a block, such as a coding block or a transform block. The video syntax structure can be a coding unit (CU) or a transform unit (TU).

[0477] d. For example, the video unit can be a picture. The video syntax structure can be a picture header or a PPS.

[0478] e. For example, the video unit can be a slice. The video syntax structure can be a slice header.

[0479] f. For example, the video unit can be a slice. The video syntax structure can be a sequence header or a SPS.

[0480] g. The video syntax structure can be VPS / DPS / APS / tile group / tile / CTU row / CTU.

[0481] 31. In one example, whether to disable or enable SST can be based on block size.

[0482] h. For example, SST can be disabled if at least one of the block width or height is less than (or not greater than) Tmin.

[0483] i. For example, SST can be disabled if both the block width and height are less than Tmin.

[0484] j. For example, SST can be disabled if at least one of the block width or height is greater than (or not less than) Tmax.

[0485] k. For example, SST can be disabled if both the block width and height are greater than (or not less than) Tmax.

[0486] l. For example, Tmin can be 2 or 4.

[0487] m. For example, Tmax can be 32, 64, or 128.

[0488] n. In one example, SST can be disabled based on the block width or / and height of the first color component.

[0489] i. For example, the first color component can be the luma color component.

[0490] ii. For example, the first color component can be the R color component.

[0491] o. In one example, SST can be disabled based on the block width or / and height of all color components.

[0492] p. Optionally, in addition, when SST is disabled, the related signaling of the usage indication of SST and / or other side information is omitted.

[0493] q. In one example, based on the block size, SST can be enabled on the first color component and disabled on the second color component.

[0494] 32. In one example, a set of SSTs can be utilized and the selection of the SST matrix for a block can depend on the decoding information such as the block size.

[0495] r. Optionally, in addition, the same decoded / signaled SST index or the same on / off control flag can be interpreted in different ways (such as different matrices corresponding to different block sizes).

[0496] s. For example, different SSTs in the set can have different sizes, such as 4x4 SST, 8x8 SST, or 16x16 SST,

[0497] t. For example, 4x4 SST can be applied to blocks with condition C4, 8x8 SST can be applied to blocks with condition C8.

[0498] i. Optionally, in addition, 4x4 SST can be applied to blocks with condition C4, 8x8 SST can be applied to blocks with condition C8,..., NxN SST can be applied to blocks with condition CN, where N is an integer.

[0499] u. In one example, condition C4 is that at least one of the block width and height is equal to 4.

[0500] v. In one example, condition C4 is that both the block width and height are equal to 4.

[0501] w. In one example, condition C4 is that the smaller value of the block width and height is equal to 4.

[0502] x. In one example, condition C8 is that the smaller value of the block width and height is not less than 8.

[0503] y. In one example, condition C8 is that at least one of the block width and height is equal to 8.

[0504] z. In one example, condition C8 is that both the block width and height are equal to 8.

[0505] aa. In one example, condition C8 is that at least one of the block width and height is greater than or equal to 8.

[0506] bb. In one example, condition C8 is that both the block width and height are greater than or equal to 8.

[0507] cc. In one example, condition CN is that at least one of the block width and height is equal to N.

[0508] dd. In one example, condition CN is that both the block width and height are equal to N.

[0509] ee. In one example, condition CN is that at least one of the block width and height is greater than or equal to N.

[0510] ff. In one example, condition CN is that both the block width and height are greater than or equal to N.

[0511] gg. In one example, NxN SST can be applied to the top-left NxN sub-block of the transform block.

[0512] hh. In one example, SST can be applied horizontally or vertically, or both, depending on the block size.

[0513] ii. In one example, different SST matrices can be selected for different color components.

[0514] i. For example, the above rules can be applied independently for different color components.

[0515] jj. In one example, one same SST matrix can be selected for all color components.

[0516] i. For example, the above rules can be applied for a first color component, and the selected SST matrix can be applied for all color components.

[0517] 1) In one example, the first color component can be the luma component.

[0518] 2) In one example, the first color component can be the Cb or Cr component.

[0519] 3) Optionally, in addition, if the selected SST matrix is not applicable for a second color component, then SST is disabled for the second color component.

[0520] kk. In one example, SST can be allowed only if the selected SST matrices for all color components are the same (by independently applying the above rules to different color components).

[0521] i. Optionally, in addition, if SST is not allowed, the signaling of the indication of the use of SST and / or other side information is omitted.

[0522] 33. In one example, the NxN SST can be applied to at least one NxN subblock that is not the same as the top-left NxN subblock.

[0523] ll. For example, the NxN SST can be applied to an NxN subblock that is right-adjacent to the top-left NxN subblock.

[0524] mm. For example, the NxN SST can be applied to an NxN subblock that is bottom-adjacent to the top-left NxN subblock.

[0525] 34. In one example, a first SST can be applied as a horizontal transform to the transformed block, and a second SST can be applied as a vertical transform to the transformed block, where the first SST and the second SST can be different.

[0526] nn. For example, the first SST and the second SST can have different sizes.

[0527] oo. Assuming the first SST is an NxN SST and the second SST is an MxM SST, and the size of the transformed block is WxH, the following rules can be applied:

[0528] i. If W is equal to Wl, then N is set equal to Wl, where Wl is an integer such as 4 or 8.

[0529] ii. If W is greater than or not less than W2, then N is set equal to W2, where W2 is an integer such as 4 or 8.

[0530] iii. If H is equal to HI, then M is set equal to HI, where HI is an integer such as 4 or 8.

[0531] iv. If H is greater than or not less than H2, then M is set equal to H2, where H2 is an integer such as 4 or 8.

[0532] 35. In one example, one of a set of SSTs can be used for a block, where there is more than one SST in the set that has the same size.

[0533] pp. In one example, a message is signaled to indicate which one to use.

[0534] qq. In one example, the choice of which to infer without signaling. The inference can depend on

[0535] i. Block size.

[0536] ii. Intra prediction mode.

[0537] iii. Transformed quantized / non-quantized coefficients.

[0538] iv. Color component.

[0539] v. Type of primary transform.

[0540] 36. In one example, different SSTs can be applied if the primary transform is different.

[0541] rr. For example, the SST used in association with DCT2 can be different from the SST used in association with DST7.

[0542] 37. In one example, SST can be applied to chroma components.

[0543] ss. In one example, different SST matrices can be applied to different color components, e.g., Y, Cb, and Cr.

[0544] tt. In one example, different color components can follow different rules of whether and how to apply SST.

[0545] uu. In one example, separate control of two color components can be applied.

[0546] i. In one example, an indication of use of SST and / or matrix can be signaled for each of the two color components.

[0547] 38. An indication of use of SST and / or an indication of SST matrix can be signaled depending on a condition check of the right lower position of the scanned region. Denote the right lower position by (SRx, SRy), such as Figure 18A - depicted in -B.

[0548] vv. In one example, an indication of use of SST and / or an indication of SST matrix can be omitted when SRx is greater than or not less than Kx and / or when SRy is greater than or not less than Ky.

[0549] ww. In one example, an indication of use of SST and / or an indication of SST matrix can be omitted when SRx is less than or not greater than K'x and / or when SRy is less than or not greater than K'y.

[0550] xx. Optionally, in addition, when no indication is signaled, SST can be inferred to be disabled.

[0551] yy. Optionally, in addition, the default SST can be inferred when no signaling indication is signaled.

[0552] i. In one example, the default SST can be set to K*L transform.

[0553] ii. In one example, the default SST can be determined according to the decoding information such as block size.

[0554] zz. Optionally, in addition, the above methods can also be applied to other non-separable secondary / primary transforms.

[0555] Related to Enhanced Quadratic Transform (EST)

[0556] 39. Implicit and explicit signaling of whether to apply secondary transform can be used within a video data unit (e.g., picture).

[0557] a. In one example, whether to use implicit or explicit method depends on the coding mode information, e.g., whether DT (derived tree) is applied.

[0558] i. In one example, for intra coded blocks without DT (derived tree) and / or intra coded blocks without PCM mode, explicit signaling method can be applied.

[0559] 1) In one example, a flag can be signaled to indicate whether to apply secondary transform.

[0560] ii. In one example, for intra coded blocks with DT (derived tree), implicit signaling method can be applied, where secondary transform is always applied.

[0561] iii. In one example, for intra coded blocks with PCM mode and / or non-intra coded blocks, implicit signaling method can be applied, where secondary transform is never applied.

[0562] 40. Multiple ways of implicit signaling of whether to apply or not apply secondary transform can be used within a video data unit (e.g., picture).

[0563] a. In one example, which implicit method to use can depend on the coding mode information, e.g., whether DT (derived tree) is applied.

[0564] iv. In one example, for intra coded blocks without DT (derived tree) and / or intra coded blocks without PCM mode, transform coefficient dependent implicit method can be applied.

[0565] 1) In one example, the parity of the number of odd transform coefficients in a block and / or the parity of the number of even transform coefficients in a block can be used to determine whether to apply a secondary transform.

[0566] v. In one example, for intra coded blocks with DT (derived tree), an implicit signaling approach can be applied, where a secondary transform is always applied.

[0567] vi. In one example, for intra coded blocks with PCM mode and / or non-intra coded blocks, an implicit signaling approach can be applied, where a secondary transform is never applied.

[0568] 41. A separable secondary transform can or can not be applied to a block depending on a syntax element (SE) such as a flag. For example, if the SE associated with the block is 1, a separable secondary transform can be applied to the block; if the SE associated with the block is 1, a separable secondary transform can not be applied to the block.

[0569] a. The signaling of the SE for a block can be conditional on information of the block.

[0570] vii. For example, the SE is signaled only if the block is a luma block.

[0571] viii. For example, the SE is not signaled if DT is used.

[0572] ix. For example, the SE is not signaled if the primary transform is not DCT2.

[0573] b. When the SE is not signaled, it can be set equal to a default value such as 0.

[0574] 42. One of a plurality of separable secondary transforms can be selected for the block in the row and / or column secondary transform.

[0575] a. In one example, the selection can be signaled.

[0576] b. In one example, the selection can not be signaled, but rather derived from information of the block.

[0577] x. For example, if the width of the block is not greater than 4 (or equal to 4, or less than 8), a 4x4 separable secondary transform can be applied to the rows (e.g., the top four rows or eight rows) of the block.

[0578] xi. For example, if the height of the block is not greater than 4 (or equal to 4, or less than 8), a 4x4 separable secondary transform can be applied to the columns (e.g., the leftmost four columns or eight columns) of the block.

[0579] xii. For example, if the width of the block is greater than 4 (or not less than 8), an 8x8 separable quadratic transform can be applied to the rows (e.g., the top four rows or the top eight rows) of the block.

[0580] xiii. For example, if the height of the block is greater than 4 (or not less than 8), an 8x8 separable quadratic transform can be applied to the columns (e.g., the left four columns or the left eight columns) of the block.

[0581] xiv. For example, if the height of the block is not greater than 4 (or equal to 4, or less than 8), a 4x4 or 8x8 separable quadratic transform can be applied to the top four rows of the block.

[0582] xv. For example, if the height of the block is greater than 4 (or not less than 8), a 4x4 or 8x8 separable quadratic transform can be applied to the top eight rows of the block.

[0583] xvi. For example, if the width of the block is not greater than 4 (or equal to 4, or less than 8), a 4x4 or 8x8 separable quadratic transform can be applied to the left four columns of the block.

[0584] xvii. For example, if the width of the block is greater than 4 (or not less than 8), a 4x4 or 8x8 separable quadratic transform can be applied to the left eight columns of the block.

[0585] 43. The transform kernel matrix of the 8x8 separable quadratic transform can be defined as

[0586]

[0587] It is related to affine prediction and interleaved prediction.

[0588] 44. The interpolation filter used for motion compensation can be different depending on whether condition A is satisfied.

[0589] a. For example, condition A indicates that affine prediction is applied.

[0590] b. For example, condition A indicates that interweaving prediction is applied.

[0591] c. For example, condition A indicates that affine prediction is applied and bi-prediction is used.

[0592] d. For example, condition A indicates that interweaving prediction is applied and bi-prediction is used.

[0593] e. Two interpolation filters being different can mean that they have different numbers of filter coefficients.

[0594] f. Two interpolation filters being different can mean that they have at least one different filter coefficient.

[0595] g.In one example, when condition A is satisfied, the interpolation filter is defined as

[0596] { 0, 0, 0, 64, 0, 0, 0, 0},

[0597] { 0, 1, -3, 63, 4, -2, 1, 0},

[0598] { 0, 2, -5, 62, 8, -3, 0, 0},

[0599] { 0, 3, -8, 60, 13, -4, 0, 0},

[0600] { 0, 4, -10, 58, 17, -5, 0, 0},

[0601] { 0, 3, -11, 52, 26, -8, 2, 0},

[0602] { 0, 2, -9, 47, 31, -10, 3, 0},

[0603] { 0, 3, -11, 45, 34, -10, 3, 0},

[0604] { 0, 3, -11, 40, 40, -11, 3, 0},

[0605] { 0, 3, -10, 34, 45, -11, 3, 0},

[0606] { 0, 3, -10, 31, 47, -9, 2, 0},

[0607] { 0, 2, -8, 26, 52, -11, 3, 0},

[0608] { 0, 0, -5, 17, 58, -10, 4, 0},

[0609] { 0, 0, -4, 13, 60, -8, 3, 0},

[0610] { 0, 0, -3, 8, 62, -5, 2, 0},

[0611] { 0, 1, -2, 4, 63, -3, 1, 0}

[0612] h.In one example, when condition A is satisfied, the interpolation filter is defined as

[0613] { 0, 0, 0, 64, 0, 0, 0, 0},

[0614] { 0, 1, -3, 63, 4, -1, 0, 0},

[0615] { 0, 2, -6, 62, 8, -3, 1, 0},

[0616] {0, 2, -8, 60, 13, -5, 2, 0},

[0617] {0, 2, -9, 57, 18, -6, 2, 0},

[0618] {0, 3, -11, 53, 24, -8, 3, 0},

[0619] {0, 3, -10, 49, 29, -9, 2, 0},

[0620] {0, 3, -11, 45, 34, -10, 3, 0},

[0621] {0, 3, -11, 40, 40, -11, 3, 0},

[0622] {0, 3, -10, 34, 45, -11, 3, 0},

[0623] {0, 2, -9, 29, 49, -10, 3, 0},

[0624] {0, 3, -8, 24, 53, -11, 3, 0},

[0625] {0, 2, -6, 18, 57, -9, 2, 0},

[0626] {0, 2, -5, 13, 60, -8, 2, 0},

[0627] {0, 1, -3, 8, 62, -6, 2, 0},

[0628] {0, 0, -1, 4, 63, -3, 1, 0}

[0629] 45. When inter prediction is applied, different patterns can be used for single prediction and bi-prediction.

[0630] a. For single prediction, the second partition pattern partitions the block into 4x4 sub-blocks with an offset of 2x2. One example is pattern 4 shown in Figure 26A

[0631] b. For single prediction, the second partition pattern partitions the block into 8x8 sub-blocks with an offset of 4x4. One example is pattern 6 shown in Figure 26B

[0632] 46. In one embodiment, there are two possible weighting values Waand Wb, satisfying Wa+Wb=2N. Exemplary weighting values {Wa, Wb} are {3, 1}, {7, 1}, {5, 3}, {13, 3}, etc.

[0633] ​​a. If the weighting value wi associated with the prediction sample Pi generated by the first partition pattern and the weighting value w2 associated with the prediction sample P2 generated by the second partition pattern are the same (both equal to Waor Wb), the final prediction P of the sample is calculated as P = (Pi + P2) » 1 or P = (Pi + P2 + 1) » 1.

[0634] b. If the weighting value wi associated with the prediction sample Pi generated by the first partition pattern and the weighting value w2 associated with the prediction sample P2 generated by the second partition pattern are different ({wi, w2} = {Wa, Wb} or {wi, w2} = {Wb, Wa}), the final prediction P of the sample is calculated as P = (wi x Pi + w2 x P2 + offset) » N, where the offset can be 1 « (N - 1) or 0.

[0635] c. Figures 27A to 27D Exemplary weighting values for 8x8, 8x4, 4x8, 4x4 subblocks when a block is partitioned into 8x8 subblocks are shown in Table 1.

[0636] FIG. 1600 is a block diagram of a video processing device 1600. The device 1600 can be used to implement one or more methods described herein. The device 1600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 1600 can include one or more processors 1602, one or more memories 1604, and video processing hardware 1606. The processor(s) 1602 can be configured to implement one or more methods described in this document. The memory (memories) 1604 can be used for storing data and code used during operation of the device 1600. The video processing hardware 1606 can be used to implement, in hardware circuitry, some of the techniques described in this document.

[0637] Figure 17 is a flowchart of an example method 1700 of video processing. The method 1700 includes determining (1702) a constraint rule for selectively applying a secondary transform having a reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block. The method 1700 includes performing (1704) the conversion by applying the secondary transform having the reduced dimension according to the constraint rule. The secondary transform having the reduced dimension has a dimension that is reduced from a dimension of the current video block. The secondary transform having the reduced dimension is applied together with a primary transform in a particular order during the conversion.

[0638] Additional embodiments and techniques are described in the following examples.

[0639] 1. A method of video processing, comprising determining a constraint rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality according to the constraint rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block, and wherein the secondary transform with reduced dimensionality is applied together with a primary transform in a particular order during the conversion.

[0640] 2. The method of example 1, wherein the conversion comprises encoding the current video block into the bitstream representation, and wherein the particular order comprises first applying the primary transform in a forward direction, then selectively applying the secondary transform with reduced dimensionality in the forward direction, then quantizing an output of the secondary transform with reduced dimensionality in the forward direction.

[0641] 3. The method of example 1, wherein the conversion comprises decoding the current video block from the bitstream representation, and wherein the particular order comprises first applying inverse quantization to the bitstream representation, then selectively applying the secondary transform with reduced dimensionality in an inverse direction, then applying the primary transform to an output of the secondary transform with reduced dimensionality in the inverse direction.

[0642] 4. The method of any of examples 1-3, wherein the constraint rule specifies that a range of the output of the secondary transform with reduced dimensionality in the inverse direction is clipped to a range of [MinCoef, MaxCoef] (inclusive of MinCoef, MaxCoef), wherein MinCoef and / or MaxCoef are two integer values that are a function of a condition of the current video block.

[0643] 5. The method of example 4, wherein the condition of the current video block is a type of color or luma component represented by the current video block.

[0644] 6. The method of example 1, wherein the constraint rule specifies that the secondary transform with reduced dimensionality is applied to one or more MxN sub-blocks of the current video block, and remaining sub-blocks of the current video block are zeroed.

[0645] 7. The method of example 1, wherein the constraint rule specifies that the secondary transform with reduced dimensionality is applied differently to different sub-blocks of the current video block.

[0646] 8. The method of any of examples 1-5, wherein, due to a size of the current video block being 4xH or Wx4, where H is a height in integer pixels and W is a width in integer pixels, the constraint rule specifies that the secondary transform with reduced dimensionality is applied to exactly one MxN sub-block of the current video block.

[0647] 9. The method of example 8, wherein H > 8 or W > 8.

[0648] 10. The method of any of examples 1 to 9, wherein the current video block is a non-square region of a video.

[0649] 11. The method of example 2 or 3, wherein the constraint rule specifies zeroing transform coefficients of the primary transform in a forward direction, or padding zero coefficients to an output of the secondary transform in a reverse direction.

[0650] Further embodiments of examples 1-5 are described in section 4, item 1. Further embodiments of examples 6-7 are described in section 4, item 2. Further embodiments of examples 8-9 are described in section 4, item 3. Further embodiments of examples 10-11 are described in section 4, item 4.

[0651] 12. A video processing method, comprising: determining a constraint rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and a neighboring video region and pixels of the current video block and pixels of the neighboring region, and performing the conversion by applying the secondary transform with reduced dimensionality according to the constraint rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block and the neighboring video region, and wherein the secondary transform with reduced dimensionality is applied together with a primary transform in a particular order during the conversion.

[0652] 13. The method of example 12, wherein the neighboring video region comprises an upper-left block of the current video block.

[0653] 14. The method of example 12, wherein the current video block and the neighboring video region correspond to sub-blocks of a parent video block.

[0654] Further embodiments of examples 12-14 are described in section 4, item 5.

[0655] 15. A video processing method, comprising: determining a zero-out rule for selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality according to the zero-out rule; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block; wherein the zero-out rule specifies a maximum number of coefficients used by the secondary transform with reduced dimensionality.

[0656] 16. The method of example 15, wherein the maximum number of coefficients is a function of a component identification of the current video block.

[0657] 17. The method of example 16, wherein the maximum number of coefficients is different for luma video blocks and chroma video blocks.

[0658] 18. The method of any of examples 15 to 17, wherein the zero-out rule specifies that the zero-out range is a function of coding information of the current video block.

[0659] 19. The method of any of examples 15 to 17, wherein the zero-out rule specifies that the zero-out range is a function of a quantization parameter of the current video block.

[0660] 20. The method of any of examples 15 to 19, wherein the zero-out range is indicated in the bitstream representation by a field included at a sequence parameter set level, or a picture parameter set level, or a picture header, or a slice header, or a tile group header, or a coding tree unit row, or a coding tree unit, or a coding unit, or a video data unit level.

[0661] Further embodiments of examples 15-17 are described in section 4 item 6. Further embodiments of example 18 are described in section 4 item 7. Further embodiments of example 19 are described in section 4 item 8. Further embodiments of example 20 are described in section 4 item 9.

[0662] 21. A video processing method, comprising: determining a condition for selectively applying a reduced-dimension secondary transform during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the reduced-dimension secondary transform according to the condition; wherein the reduced-dimension secondary transform has a dimension reduced from a dimension of the current video block; and wherein the condition is signaled in the bitstream representation.

[0663] 22. The method of example 21, wherein the condition is a use of a color format or separate plane coding or based on a color identification of the current video block.

[0664] Further embodiments of examples 21-22 are described in section 4 item 10.

[0665] 23. The method of any of examples 21 to 22, wherein the condition is signaled in the bitstream representation separately for chroma and luma components.

[0666] Further embodiments of example 23 are described in section 4 item 11.

[0667] 24. The method of any of examples 21 to 23, wherein the condition depends on coding information of the current video block and a neighboring video region.

[0668] 25. The method of example 24, wherein the condition excludes application to the current video block coded using a particular intra prediction mode.

[0669] Further embodiments of examples 24-25 are described in section 4 item 12.

[0670] 26. The method of example 24, wherein the condition specifies that the current video block is coded using intra block copy mode.

[0671] 27. The method of example 24, wherein the condition specifies that the current video block is coded using intra block copy mode.

[0672] Other embodiments of examples 25-26 are described in section 4, item 13.

[0673] 28. The method of example 21, wherein the condition is signaled in a bitstream representation at a level such that all blocks within the level satisfy the condition, wherein the level is a sequence parameter set level, or a picture parameter set level, or a picture header, or a slice header, or a tile group header, or a coding tree unit row, or a coding tree unit, or a coding unit, or a video data unit level.

[0674] Other embodiments of example 28 are described in section 4, item 14.

[0675] 29. The method of example 21, wherein the condition is that the current video block is coded using a transform skip mode.

[0676] Other embodiments of example 29 are described in section 4, item 17.

[0677] 30. A video processing method, comprising: selectively applying a secondary transform with reduced dimension during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimension according to the condition; wherein the secondary transform with reduced dimension has a dimension reduced from a dimension of the current video block; and wherein the conversion includes selectively applying position dependent intra prediction combination (PDPC) based on a coexistence rule.

[0678] 31. The method of example 30, wherein the coexistence rule excludes applying PDPC to the current video block as a result of applying the secondary transform.

[0679] 32. The method of example 30, wherein the coexistence rule specifies applying PDPC to the current video block as a result of applying the secondary transform.

[0680] 33. The method of example 30, wherein the selectively applying the secondary transform is performed for the current video block using PDPC.

[0681] Other embodiments of examples 30-33 are described in section 4, item 15.

[0682] 34. A method of video processing, comprising selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality in accordance with the condition; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block; and wherein the applying controls usage of neighboring samples for intra prediction during the conversion.

[0683] Other embodiments of example 34 are described in section 4, item 16.

[0684] 35. A method of video processing, comprising selectively applying a secondary transform with reduced dimensionality during a conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the secondary transform with reduced dimensionality in accordance with the condition; wherein the secondary transform with reduced dimensionality has a dimension reduced from a dimension of the current video block; and wherein the selective applying controls usage of a quantization matrix during the conversion.

[0685] 36. The method of example 35, wherein the usage of the quantization matrix occurs only due to the application of the secondary transform.

[0686] Other embodiments of examples 35-36 are described in section 4, item 22.

[0687] 37. The method of any of examples 1-36, wherein the primary transform and the secondary transform are stored as transform matrices having a bit-width less than 8.

[0688] 38. The method of any of examples 1-36, wherein the primary transform and the secondary transform are stored as prediction transform matrices.

[0689] 39. The method of any of examples 1-36, wherein the primary transform is derivable from the secondary transform using a first rule, or wherein the secondary transform is derivable from the primary transform using a second rule.

[0690] 40. The method of any of examples 1-36, wherein the bitstream representation includes information about the secondary transform or the primary transform prior to residual information for the current video block.

[0691] Other embodiments of examples 37-40 are described in section 4, items 18, 19, 20, 21.

[0692] 41. The method of example 1, wherein a constraint rule for selectively applying the secondary transform depends on a number of transform units in a coding unit of the current video block.

[0693] 42. The method of example 41, wherein, due to the number of transform units in the coding unit being greater than one, the constraint rule specifies that a secondary transform with reduced dimensionality is applied.

[0694] 43. The method of example 1, wherein a flag in the bitstream representation indicates whether a secondary transform with reduced dimensionality is applied to the conversion.

[0695] 44. The method of example 1, wherein the current video block comprises more than one component video block, and wherein the constraint rule specifies applicability of a secondary transform with reduced dimensionality differently for different component video blocks.

[0696] 45. The method of example 44, wherein the constraint rule specifies applicability of a secondary transform with reduced dimensionality for a first component video block based on how the constraint rule applies to a second component video block.

[0697] 46. The method of any of examples 44-45, wherein the constraint rule further depends on a dimension of the current video block.

[0698] Other embodiments of examples 47-53 are described in, for example, items 30-38 of section 4.

[0699] 47. A method of video processing, comprising: for a conversion between a current video block of a video and a bitstream representation of the video, determining whether to use a separable secondary transform (SST) for the conversion based on a coding condition; and performing the conversion in accordance with the determination.

[0700] 48. The method of example 47, wherein the coding condition corresponds to a syntax element in the bitstream representation.

[0701] 49. The method of example 48, wherein the coding condition comprises a size of the current video block.

[0702] 50. The method of any of examples 47-49, wherein, when it is determined to use the SST, the conversion uses a selected SST from a set of SSTs based on another coding condition.

[0703] 51. The method of example 50, wherein the another coding condition comprises a size of the current video block.

[0704] 52. The method of any of examples 1-51, wherein the conversion comprises decoding and parsing the bitstream representation to generate the video.

[0705] 53. The method of any of examples 1-51, wherein the conversion comprises encoding the video into the bitstream representation.

[0706] 54. A video processing apparatus comprising a processor configured to implement one or more of the examples of examples 1 to 53.

[0707] 55. A computer readable medium having code stored thereon, when executed by a processor, causes the processor to implement a method recited in any one or more of examples 1 to 53.

[0708] It will be appreciated that the disclosed technology can be embodied in a video encoder or decoder to improve compression efficiency using techniques including using reduced dimension secondary transforms.

[0709] Figure 19 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0710] System 1900 can include a codec component 1904, which can implement various coding or encoding methods described in this document. Codec component 1904 can reduce the average bitrate of video from input 1902 to the output of codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via a communication connected by component 1906. Component 1908 can use a stored or communicated bitstream (or coded) representation of the video received at input 1902 to generate pixel values or displayable video sent to a display interface 1910. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder, and a decoder will perform corresponding decoding tools or operations that reverse the results of the coding.

[0711] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices, such as mobile telephones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0712] Figure 28 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.

[0713] As Figure 28 shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 generates encoded video data, and can be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110, and can be referred to as a video decoding device.

[0714] Source device 110 can include a video source 112, a video encoder 114, and an input / output (EO) interface 116.

[0715] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. EO interface 116 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 by EO interface 116 through network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by destination device 120.

[0716] Destination device 120 can include an EO interface 126, a video decoder 124, and a display device 122.

[0717] EO interface 126 can include a receiver and / or a modem. EO interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.

[0718] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.

[0719] Figure 29 is shown to be Figure 28 A block diagram of an example of a video encoder 200 of the video encoder 114 in the system 100 shown in FIG. 1.

[0720] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 29 In an example, the video encoder 200 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor(s) can be configured to perform any or all of the techniques described in this disclosure.

[0721] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202, a prediction unit 1602 can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0722] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0723] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately in Figure 19 examples for the purposes of explanation.

[0724] The partition unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0725] The mode select unit 203 can select one of the coding modes (intra or inter), for example, based on error results, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode select unit 203 can select a combined intra and inter prediction (CIIP) mode in which the prediction is based on both an inter prediction signal and an intra prediction signal. The mode select unit 203 can also select a resolution of motion vectors (e.g., sub-pixel or integer-pixel precision) for the block in the case of inter prediction.

[0726] To perform inter prediction on a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 (except for the picture associated with the current video block).

[0727] For example, motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0728] In some examples, motion estimation unit 204 can perform uni-prediction on the current video block, and motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 or list 1. Motion estimation unit 204 can then generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0729] In other examples, motion estimation unit 204 can perform bi-prediction on the current video block, motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0, and can also search for another reference video block for the current video block in a reference picture in list 1. Motion estimation unit 204 can then generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates a spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and the motion vector for the current video block as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.

[0730] In some examples, motion estimation unit 204 can output a complete set of motion information for a decoding process of a decoder.

[0731] In some examples, motion estimation unit 204 can not output a complete set of motion information for the current video. Instead, motion estimation unit 204 can signal the motion information for the current video block with reference to motion information of another video block. For example, motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.

[0732] In one example, the motion estimation unit 204 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0733] In another example, the motion estimation unit 204 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0734] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0735] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0736] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the samples in the current video block.

[0737] In other examples, such as in skip mode, there can be no residual data for the current video block, and the residual generation unit 207 can not perform the subtraction operation.

[0738] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0739] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0740] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated from prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.

[0741] After reconstruction unit 212 reconstructs a video block, loop filtering operations can be performed to reduce video block artifacts in the video block.

[0742] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0743] Figure 30 is shown to illustrate that video decoder 300 of video decoder 114 in system 100 can be Figure 28 A block diagram illustrating an example of video decoder 300 of video decoder 114 in system 100 is shown.

[0744] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 30 In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0745] In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure. Figure 30 In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure. Figure 29 ) described with respect to video encoder 200.

[0746] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data, and motion compensation unit 302 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information from the entropy decoded video data. For example, motion compensation unit 302 can determine such information by performing AMVP and merge modes.

[0747] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel precision can be included in syntax elements.

[0748] Motion compensation unit 302 can calculate interpolated values for sub-integer pixels of the reference block using the interpolation filter used by video encoder 20 during encoding of the video block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 20 from the received syntax information and use the interpolation filter to generate the prediction block.

[0749] Motion compensation unit 302 can use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information used to decode the encoded video sequence.

[0750] Intra prediction unit 303 can use intra prediction modes, e.g., received in the bitstream, to form the prediction block from spatial neighboring blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0751] Reconstruction unit 306 can add the residual block to the corresponding prediction block generated by motion compensation unit 202 or intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0752] Figure 31 is a flowchart representation of a method 3100 for video processing in accordance with the present technology. The method 3100 includes, at operation 3110, performing a conversion between a video block of a video and a bitstream of the video according to a rule. The rule specifies whether or how to indicate a use of a secondary transform within a video unit in the bitstream. The secondary transform is applied before quantization or after inverse quantization.

[0753] In some embodiments, the secondary transform is applied between a forward primary transform and quantization or between inverse quantization and an inverse primary transform. In some embodiments, the video unit comprises a video picture of the video. In some embodiments, the video unit comprises a video sequence of the video. In some embodiments, the secondary transform comprises an enhanced secondary transform.

[0754] In some embodiments, the implicit or explicit indication of the use of the secondary transform is based on a coding mode of the video block. In some embodiments, the video block is coded in an intra coding mode. In the case that the video block is coded without using a derived tree block partitioning or a pulse code modulation (PCM) coding tool, the use of the secondary transform is explicitly indicated in the bitstream. In some embodiments, a syntax element is used to explicitly indicate the use of the secondary transform. In some embodiments, the video block is coded in an intra coding mode. In the case that the video block is coded using a derived tree block partitioning, the use of the secondary transform is implicitly indicated in the bitstream. In some embodiments, the secondary transform is always applied within a video unit. In some embodiments, the video block is coded in an intra coding mode. In the case that the video block is coded using a pulse code modulation (PCM) coding tool, the use of the secondary transform is implicitly indicated in the bitstream. In some embodiments, the secondary transform is always excluded within a video unit.

[0755] In some embodiments, the use of the secondary transform is implicitly indicated using one or more implicit methods. In some embodiments, the determination of the one or more implicit methods is based on coding information of the video block. In some embodiments, the video block is coded in an intra coding mode. In the case that the video block is coded without using a derived tree block partitioning or a pulse code modulation (PCM) coding tool, the one or more implicit methods are determined based on applicable transform coefficients. In some embodiments, the use of the secondary transform is indicated based on the parity of odd transform coefficients and / or the parity of even coefficients in the video block.

[0756] Figure 32 is a flowchart representation of a method 3200 for video processing according to the present technology. The method 3200 includes, at operation 3210, performing a conversion between a video block of a video and a bitstream of the video according to a rule. The rule specifies that a use of a separable secondary transform in the video block is determined based on a syntax element associated with the video block. The separable secondary transform is applied between a forward primary transform and quantization or between inverse quantization and an inverse primary transform.

[0757] In some embodiments, the separable secondary transform is applied to the video block if the value of the syntax element is 1. In some embodiments, the separable secondary transform is disabled in the video block if the value of the syntax element is 1. In some embodiments, the indication of the syntax element is conditioned on coding information of the video block. In some embodiments, the syntax element is indicated if the video block is a luma block. In some embodiments, the syntax element is omitted in the bitstream if a derived tree block partition is used in the video block. In some embodiments, the syntax element is omitted in the bitstream if a primary transform is not a discrete cosine transform type-II (DCT-2). In some embodiments, a default value of the syntax element is inferred to be 0 if the syntax element is omitted in the bitstream.

[0758] Figure 33 is a flowchart representation of a method 3300 for video processing in accordance with the present technology. The method 3300 includes, at operation 3310, performing a conversion between a video block of a video and a bitstream of the video according to a rule. The rule specifies a selection of one secondary transform from a plurality of separable secondary transforms to be applied to the video block. The secondary transform is applied to a row of the video block or a column of the video block.

[0759] In some embodiments, the selection of the secondary transform is indicated in the bitstream. In some embodiments, the selection of the secondary transform is derived based on coding information of the video block. In some embodiments, a 4x4 separable secondary transform is applied to a row of the video block if a width of the video block is less than or equal to N, N being a positive integer. In some embodiments, a 4x4 separable secondary transform is applied to a column of the video block if a height of the video block is less than or equal to N, N being a positive integer. In some embodiments, an 8x8 separable secondary transform is applied to a row of the video block if the width of the video block is greater than N, N being a positive integer. In some embodiments, an 8x8 separable secondary transform is applied to a column of the video block if the height of the video block is greater than N, N being a positive integer. In some embodiments, a 4x4 or 8x8 separable secondary transform is applied to a top four rows of the video block if the height of the video block is less than or equal to N, N being a positive integer. In some embodiments, a 4x4 or 8x8 separable secondary transform is applied to a top eight rows of the video block if the height of the video block is greater than N, N being a positive integer. In some embodiments, a 4x4 or 8x8 separable secondary transform is applied to a leftmost four columns of the video block if the width of the video block is less than or equal to N, N being a positive integer. In some embodiments, a 4x4 or 8x8 separable secondary transform is applied to a leftmost eight columns of the video block if the width of the video block is greater than N, N being a positive integer. In some embodiments, N is 4 or 8.

[0760] In some embodiments, the core matrix of the 8x8 separable quadratic transform is defined as:

[0761] Figure 34 FIG. 34 is a flowchart representation of a method 3400 for video processing in accordance with the techniques herein. The method 3400 includes, at operation 3410, determining one or more interpolation filters for motion compensation of a video block of a video based on a condition. The method 3400 includes, at operation 3420, performing a conversion between a video block of the video and a bitstream of the video in accordance with the determination.

[0762] In some embodiments, the condition comprises whether affine prediction is applied to the video block. In some embodiments, the condition comprises whether interweave prediction is applied to the video block. In some embodiments, the condition comprises whether affine prediction and bi-prediction are applied to the video block. In some embodiments, the condition comprises whether interweave prediction and bi-prediction are applied to the video block.

[0763] In some embodiments, the two interpolation filters having different numbers of filter coefficients are different. In some embodiments, the two interpolation filters having at least one different filter coefficient are different. In some embodiments, the one or more interpolation filters comprise at least one of: {0, 0, 0, 64, 0, 0, 0, 0}, {0, 1, -3, 63, 4, -2, 1, 0}, {0, 2, -5, 62, 8, -3, 0, 0}, {0, 3, -8, 60, 13, -4, 0, 0}, {0, 4, -10, 58, 17, -5, 0, 0}, {0, 3, -11, 52, 26, -8, 2, 0}, {0, 2, -9, 47, 31, -10, 3, 0}, {0, 3, -11, 45, 34, -10, 3, 0}, {0, 3, -11, 40, 40, -11, 3, 0}, {0, 3, -10, 34, 45, -11, 3, 0}, {0, 3, -10, 31, 47, -9, 2, 0}, {0, 2, -8, 26, 52, -11, 3, 0}, {0, 0, -5, 17, 58, -10, 4, 0}, {0, 0, -4, 13, 60, -8, 3, 0}, {0, 0, -3, 8, 62, -5, 2, 0}, or {0, 1, -2, 4, 63, -3, 1, 0} in the case where the condition is satisfied.

[0764] {0, 0, 0, 64, 0, 0, 0, 0}, {0, 1, -3, 63, 4, -1, 0, 0}, {0, 2, -6, 62, 8, -3, 1, 0}, {0, 2, -8, 60, 13, -5, 2, 0}, {0, 2, -9, 57, 18, -6, 2, 0}, {0, 3, -11, 53, 24, -8, 3, 0}, {0, 3, -10, 49, 29, -9, 2, 0}, {0, 3, -11, 45, 34, -10, 3, 0}, {0, 3, -11, 40, 40, -11, 3, 0}, {0, 3, -10, 34, 45, -11, 3, 0}, {0, 2, -9, 29, 49, -10, 3, 0}, {0, 3, -8, 24, 53, -11, 3, 0}, {0, 2, -6, 18, 57, -9, 2, 0}, {0, 2, -5, 13, 60, -8, 2, 0}, {0, 1, -3, 8, 62, -6, 2, 0}, or {0, 0, -1, 4, 63, -3, 1, 0}.

[0765] Figure 35 is a flowchart representation of a method 3500 for video processing in accordance with the present technology. The method 3500 includes, at operation 3510, performing a conversion between a video block of a video and a bitstream of the video according to a rule. The video block is coded using an interweaved prediction mode in which the video block is partitioned into sub-blocks using a first pattern and a second pattern, and a final prediction is determined as a weighted sum of two auxiliary predictions having the first pattern and the second pattern. The rule specifies that the two auxiliary predictions having the first pattern and the second pattern include a single prediction and a bi-prediction, where the first pattern and the second pattern are different.

[0766] In some embodiments, the first pattern for the single prediction mode includes a 4x4 sub-block having a 2x2 offset at a lower-left corner of the video block. In some embodiments, the first pattern for the single prediction mode includes an 8x8 sub-block having a 4x4 offset at a lower-left corner of the video block.

[0767] In some embodiments, the two applicable weighting values Waand Wbfor the weighting and satisfy Wa+ Wb= 2N, where N is a positive integer greater than 1. In some embodiments, a first weight wl is associated with a first prediction sample Pl generated by a first pattern, and a second weight w2 is associated with a second prediction sample P2 generated by a second pattern. In the case where wl and w2 are the same, wl and w2 are either Waor Wb, the final prediction is calculated as P = (Pl + P2) » 1 or (Pl + P2 + 1) » 1. In some embodiments, a first weight wl is associated with a first prediction sample Pl generated by a first pattern, and a second weight w2 is associated with a second prediction sample P2 generated by a second pattern. In the case where wl and w2 are different, where the offset is equal to 1 « (N - 1) or 0, the final prediction is calculated as P = (wl x Pl + wl x P2 + offset) » N.

[0768] In some embodiments, the weighting values for an 8x8 subblock are illustrated by the following matrix:

[0769] Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb Wb

[0770] In some embodiments, the weighting values for an 8x4 subblock are illustrated by the following matrix:

[0771] Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb Wb Wb Wa Wa Wa Wa Wb Wb

[0772] In some embodiments, the weighting values for a 4x8 subblock are illustrated by the following matrix:

[0773] Wb Wb Wb Wb Wb Wb Wb Wb Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wb Wb Wb Wb Wb Wb Wb Wb

[0774] In some embodiments, the weighting values for a 4x4 subblock are illustrated by the following matrix:

[0775] Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa Wa

[0776] In some embodiments, the conversion includes encoding a video into a bitstream. In some embodiments, the conversion includes decoding a bitstream to generate a video.

[0777] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in the processing of a video block, but can not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of a video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, a decoder will process a bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of a video to a video block will be performed using the video processing tool or mode enabled based on a decision or determination.

[0778] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode in the conversion of a video block to a bitstream representation of a video. In another example, when a video processing tool or mode is disabled, a decoder will process a bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on a decision or determination.

[0779] In this document, the term “video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, a bitstream representation of a current video block can correspond to bits as defined by syntax, located at the same position or scattered at different positions within the bitstream. For example, a macroblock can be encoded according to a transform and coded error residual, and can also use bits in a header and other fields in the bitstream.

[0780] The disclosures and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structural equivalents of such as disclosed in this document, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for encoding information to be transmitted to a suitable receiver apparatus.

[0781] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0782] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0783] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0784] Although the present patent document contains many details, these should not be construed as limiting the scope of any subject matter or potentially patentable content in any way. Rather, these contain descriptions of features that can be particular to a particular embodiment of technology. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Also, while the features described above can be described as acting in certain combinations and initially claimed as such, in some cases one or more features from a claimed combination can be removed and the claimed combination can be directed to a subcombination or variation of a subcombination.

[0785] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, or that all illustrated operations be performed, to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0786] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: Perform the conversion between video blocks and the video bitstream according to the rules. The rule specifies whether an implicit or explicit indication is applied to indicate the use of secondary transforms within a video block in the bitstream, wherein the use of secondary transforms includes applying the secondary transform and not applying the secondary transform, wherein the secondary transform is applied before quantization or after dequantization, and wherein the secondary transform is applied between the forward master transform and quantization or between dequantization and the inverse master transform. Specifically, the implicit or explicit indication method is determined based on the encoding / decoding mode of the video block. The implicit indication method is used to indicate the use of the secondary transformation, and the explicit indication method is used to indicate the use of the secondary transformation. Specifically, the method of determining whether to use the implicit indication method or the explicit indication method is based on whether the video block is encoded and decoded using a derived tree partitioning.

2. The method according to claim 1, wherein the video unit includes a video image of the video.

3. The method according to claim 1, wherein the video unit comprises a video sequence of the video.

4. The method of claim 1, wherein the quadratic transformation comprises an enhanced quadratic transformation.

5. The method of claim 1, wherein the video blocks within the video unit are implicitly indicated using a variety of methods.

6. The method of claim 1, wherein it is determined whether the implicit indication or the explicit indication is applied to indicate the use of the secondary transform based on whether a pulse codec modulation (PCM) codec tool is applied to the video block.

7. The method of claim 6, wherein in response to applying the PCM codec tool to the video block, the implicit indication method without using the secondary transformation is applied to the video block.

8. The method of claim 1, wherein in response to applying a non-intra-predictive coding / decoding mode to the video block, the implicit indication mode without using the secondary transform is applied to the video block.

9. The method of claim 1, wherein the syntax elements are used for the explicit indication of the use of the secondary transformation.

10. The method according to any one of claims 1 to 9, wherein the conversion comprises encoding the video into the bitstream.

11. The method according to any one of claims 1 to 9, wherein the conversion comprises decoding the video from the bitstream.

12. The method of claim 1, wherein the video block is encoded and decoded in an intra-frame encoding / decoding mode, and wherein the use of the secondary transformation is explicitly indicated in the bitstream when the video block is encoded and decoded without using a derived tree partitioning or pulse codec modulation (PCM) encoding / decoding tools.

13. The method of claim 1, wherein the video block is encoded and decoded in an intra-frame encoding / decoding mode, and wherein, when the video block is encoded and decoded using a derivative tree partition, the use of the secondary transformation is implicitly indicated in the bitstream.

14. The method of claim 6, wherein the quadratic transformation is always applied within the video unit.

15. The method of claim 1, wherein the video block is encoded and decoded in an intra-frame encoding / decoding mode, and wherein, when the video block is encoded and decoded using a Pulse Code-Code-Modulation (PCM) encoding / decoding tool, the use of the secondary transformation is implicitly indicated in the bitstream.

16. The method of claim 15, wherein the secondary transformation is always excluded within the video unit.

17. The method of claim 16, wherein one or more implicit methods are used to implicitly indicate the use of the quadratic transformation.

18. The method of claim 17, wherein the determination of the one or more implicit methods is based on the encoding / decoding information of the video block.

19. The method of claim 18, wherein the video block is encoded and decoded in an intra-frame encoding / decoding mode, and wherein the one or more implicit methods are determined based on applicable transform coefficients when the video block is encoded and decoded without using derived tree partitioning or pulse codec modulation (PCM) encoding / decoding tools.

20. The method of claim 19, wherein the use of the secondary transformation is indicated based on the parity of the odd transform coefficients and / or the parity of the even coefficients in the video block.

21. An apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform the conversion between video blocks and the video bitstream according to the rules. in, The rule specifies whether an implicit or explicit indication is applied to indicate the use of the secondary transform within the video block in the bitstream, wherein the use of the secondary transform includes applying the secondary transform and not applying the secondary transform, wherein the secondary transform is applied before quantization or after dequantization, and wherein the secondary transform is applied between the forward master transform and quantization or between dequantization and the inverse master transform. Specifically, the implicit or explicit indication method is determined based on the encoding / decoding mode of the video block. The implicit indication method is used to indicate the use of the secondary transformation, and the explicit indication method is used to indicate the use of the secondary transformation. Specifically, the method of determining whether to use the implicit indication method or the explicit indication method is based on whether the video block is encoded and decoded using a derived tree partitioning.

22. The apparatus of claim 21, wherein the video unit comprises a video sequence of the video.

23. A non-transitory computer-readable storage medium storing instructions that cause a processor to: Perform the conversion between video blocks and the video bitstream according to the rules. in, The rule specifies whether an implicit or explicit indication is applied to indicate the use of the secondary transform within the video block in the bitstream, wherein the use of the secondary transform includes applying the secondary transform and not applying the secondary transform, wherein the secondary transform is applied before quantization or after dequantization, and wherein the secondary transform is applied between the forward master transform and quantization or between dequantization and the inverse master transform. Specifically, the implicit or explicit indication method is determined based on the encoding / decoding mode of the video block. The implicit indication method is used to indicate the use of the secondary transformation, and the explicit indication method is used to indicate the use of the secondary transformation. Specifically, the method of determining whether to use the implicit indication method or the explicit indication method is based on whether the video block is encoded and decoded using a derived tree partitioning.

24. The non-transitory computer-readable storage medium according to claim 23, The determination of whether the implicit or explicit indication is applied to indicate the use of the secondary transform is based on whether a Pulse Codec Modulation (PCM) codec tool is applied to the video block, wherein in response to applying the PCM codec tool to the video block, the implicit indication method without using the secondary transform is applied to the video block; or In response to applying a non-intra-frame predictive coding mode to the video block, the implicit indication mode that does not use the secondary transform is applied to the video block.

25. A method for storing a bitstream of video, comprising: The bitstream of the video is generated from the video blocks of the video according to the rules. The bitstream is stored in a non-transitory computer-readable recording medium. The rule specifies whether an implicit or explicit indication is applied to indicate the use of secondary transforms within a video block in the bitstream, wherein the use of secondary transforms includes applying the secondary transform and not applying the secondary transform, wherein the secondary transform is applied before quantization or after dequantization, and wherein the secondary transform is applied between the forward master transform and quantization or between dequantization and the inverse master transform. Specifically, the implicit or explicit indication method is determined based on the encoding / decoding mode of the video block. The implicit indication method is used to indicate the use of the secondary transformation, and the explicit indication method is used to indicate the use of the secondary transformation. Specifically, the method of determining whether to use the implicit indication method or the explicit indication method is based on whether the video block is encoded and decoded using a derived tree partitioning.

26. The method of claim 25, wherein it is determined whether the implicit indication or the explicit indication is applied to indicate the use of the secondary transform based on whether a Pulse Codec Modulation (PCM) codec tool is applied to the video block, wherein in response to applying the PCM codec tool to the video block, the implicit indication without using the secondary transform is applied to the video block; or In response to applying a non-intra-frame predictive coding mode to the video block, the implicit indication mode that does not use the secondary transform is applied to the video block.

27. The method according to claim 1, further comprising: The rule further specifies that the use of separable quadratic transforms in the video block is determined based on the syntax elements associated with the video block, wherein the separable quadratic transform is applied between the forward master transform and quantization or between the inverse quantization and the inverse master transform.

28. The method of claim 27, wherein the separable quadratic transformation is applied to the video block when the value of the syntax element is 1.

29. The method of claim 27, wherein the separable quadratic transformation is disabled in the video block when the value of the syntax element is 1.

30. The method of claim 27, wherein the indication of the syntax element is adjusted based on the encoding / decoding information of the video block.

31. The method of claim 30, wherein the syntax element is indicated when the video block is a luminance block.

32. The method of claim 30, wherein, in the case of using derived tree block partitioning in the video block, the syntax element is omitted in the bitstream.

33. The method of claim 30, wherein if the main transform is not a Discrete Cosine Transform Type-II (DCT-2), the syntax element is omitted in the bitstream.

34. The method of claim 27, wherein if the syntax element is omitted in the bitstream, the default value of the syntax element is inferred to be 0.

35. The method according to claim 1, further comprising: The rule further specifies the selection of a secondary transformation from a plurality of separable secondary transformations to be applied to the video block, wherein the secondary transformation is applied to a row or a column of the video block.

36. The method of claim 35, wherein the selection of the secondary transformation is indicated in the bit stream.

37. The method of claim 35, wherein the selection of the secondary transform is derived based on the encoding and decoding information of the video block.

38. The method of claim 37, wherein when the width of the video block is less than or equal to N, where N is a positive integer, a 4×4 separable quadratic transformation is applied to the rows of the video block.

39. The method of claim 37, wherein when the height of the video block is less than or equal to N, where N is a positive integer, a 4×4 separable quadratic transformation is applied to the columns of the video block.

40. The method of claim 37, wherein when the width of the video block is greater than N, and N is a positive integer, an 8×8 separable quadratic transformation is applied to the rows of the video block.

41. The method of claim 37, wherein when the height of the video block is greater than N, and N is a positive integer, an 8×8 separable quadratic transformation is applied to the columns of the video block.

42. The method of claim 37, wherein when the height of the video block is less than or equal to N, where N is a positive integer, a 4×4 or 8×8 separable quadratic transformation is applied to the top four rows of the video block.

43. The method of claim 37, wherein when the height of the video block is greater than N, and N is a positive integer, a 4×4 or 8×8 separable quadratic transformation is applied to the top eight rows of the video block.

44. The method of claim 37, wherein when the width of the video block is less than or equal to N, where N is a positive integer, a 4×4 or 8×8 separable quadratic transformation is applied to the leftmost four columns of the video block.

45. The method of claim 37, wherein when the width of the video block is greater than N, and N is a positive integer, a 4×4 or 8×8 separable quadratic transformation is applied to the leftmost eight columns of the video block.

46. ​​The method of claim 35, wherein N is 4 or 8.

47. The method according to claim 35, wherein the core matrix of the 8×8 separable quadratic transformation is defined as: 。 48. The method according to any one of claims 12 to 20 and 27 to 47, wherein the conversion comprises encoding the video into the bitstream.

49. The method according to any one of claims 12 to 20 and 27 to 47, wherein the conversion comprises decoding the bitstream to generate the video.

50. An apparatus for processing video data, comprising a processor configured to implement the method of any one of claims 2, 4-20, and 27-49.

51. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any one of claims 2-20 and 27-49.

Citation Information

Patent Citations

  • Non-separable secondary transform for video coding

    CN108141596A

  • Image decoding device and image encoding device

    CN109076223A