Use of a secondary transform related to block size in coded and decoded video
By applying dimensionality reduction quadratic transformation and constraint rules in the video encoding process, the problems of unoptimized transformation matrix cropping and low storage efficiency in the prior art are solved, and video encoding efficiency and adaptability are improved.
Patent Information
- Application Number
- CN202080042498.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-15
- Filing Date
- 2020-06-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-06-15
AI Technical Summary
When processing video data, existing video encoding technologies have problems such as unoptimized transformation matrix cropping and shift/rounding operations, high dimensionality reduction quadratic transformation costs, unsuitable RST applications for screen content encoding, unclear interaction between RST and other encoding and decoding tools, low storage efficiency of transformation matrix, and unclear application of quantization matrix.
The transformation process is optimized by applying the dimensionality reduction quadratic transformation between the forward main transformation and the quantization step or between the inverse and inverse main transformation, combining constraint rules and zeroing rules, and applying the main transformation and dimensionality reduction quadratic transformation in a specific order.
It improves video encoding efficiency, reduces the storage requirements of transformation matrix, optimizes the application cost of RST, and clarifies the use of quantization matrix to adapt to the needs of different encoding and decoding modes and color components.
Smart Images

Figure CN113994696B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] According to applicable patent laws and / or the rules applicable to the Paris Convention, this application timely claims the priority and benefits of International Patent Application No. PCT / CN2019 / 091434, filed on June 15, 2019. For all purposes of the law, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background art
[0004] Despite the progress in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth required for digital video usage is expected to continue to grow. Summary of the invention
[0005] This document describes various embodiments and techniques in which a secondary transform (also known as a low - frequency non - separable transform) is used during the decoding or encoding of video or images.
[0006] In one exemplary aspect, a video processing method is disclosed. The method includes: for the conversion between a video coding / decoding unit and the bit - stream representation of the video, determining, based on rules associated with one or more transform units in the coding / decoding unit, to use a reduced - dimensional secondary transform (RST) tool for the coding / decoding unit. The RST tool includes, during encoding, applying a forward secondary transform between a forward primary transform and a quantization step, or during decoding, applying an inverse secondary transform between an inverse quantization step and an inverse primary transform. The sizes of the forward secondary transform and the inverse secondary transform are smaller than the size of the coding / decoding unit. The method further includes performing the conversion based on the determination.
[0007] In another exemplary aspect, a video processing method is disclosed. The method includes: for the conversion between a video block including one or more color - component blocks and the bit - stream representation of the video, determining, according to rules, a manner of applying a reduced - dimensional secondary transform (RST) tool to one or more color components of the block. The RST tool includes, during encoding, applying a forward secondary transform between a forward primary transform and a quantization step, or during decoding, applying an inverse secondary transform between an inverse quantization step and an inverse primary transform. The sizes of the forward secondary transform and the inverse secondary transform are smaller than the size of the coding / decoding unit. The method further includes performing the conversion based on the determination.
[0008] In another example aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and a bitstream representation of the video, determining that output values from a reduced-dimensional inverse quadratic transform are inclusively constrained within a range [min, max]. The inverse quadratic transform can be applied to a block between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block, and min and max are integer values. The method further includes performing the conversion based on the determination.
[0009] In another example aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and a bitstream representation of the video, determining a manner of applying a reduced-dimensional quadratic transform to sub-blocks of the block based on the number of sub-blocks to which a quadratic transform can be applied. The quadratic transform can be applied to a block between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0010] In another example aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and a bitstream representation of the video, determining that a reduced-dimensional quadratic transform is applied to a single sub-block of the block when the dimension of the block satisfies a condition. The quadratic transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0011] In another example aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and a bitstream representation of the video, determining that a reduced-dimensional quadratic transform can be applied to a region in a block of dimension K×L. K and L are positive integers and K is not equal to L. The quadratic transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0012] In another example aspect, a method for video processing is disclosed. The method includes: for the conversion between a block of a video and a bitstream representation of the video, determining a non-zero range based on characteristics of the block. The non-zero range corresponds to a range outside of which coefficients associated with the reduced-dimensional quadratic transform are set to zero. The quadratic transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0013] In another example aspect, a method for video encoding is disclosed. The method includes: determining that a dimensionality-reduced quadratic transform is applicable to two adjacent sub-blocks of a video block. Each of the two adjacent sub-blocks has dimensions of M×N, where M and N are positive integers. The quadratic transform is performed between the forward primary transform and the quantization step. The reduced dimension is the dimension reduced from the dimension of the block. The method further includes generating an encoded / decoded representation of the video based on the determination.
[0014] In another example aspect, a method for video decoding is disclosed. The method includes: determining that a dimensionality-reduced quadratic transform is applicable to two adjacent sub-blocks of a video block. Each of the two adjacent sub-blocks has dimensions of M×N, where M and N are positive integers. The quadratic transform is performed between the inverse quantization step and the inverse primary transform. The reduced dimension is the dimension reduced from the dimension of the block. The method further includes generating blocks of the video by parsing the encoded / decoded representation of the video according to the determination.
[0015] In another example aspect, a method for video processing is disclosed. The method includes: determining, for the conversion between a block of a video and a bitstream representation of the video, whether to apply a dimensionality-reduced quadratic transform to the block according to rules based on characteristics associated with the block. The quadratic transform is performed between the forward primary transform and the quantization step, or between the inverse quantization step and the inverse primary transform. The reduced dimension is the dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0016] In another example aspect, a method for video processing is disclosed. The method includes: determining bit-precision constraints for coefficients of one or more transform matrices of a dimensionality-reduced quadratic transform applied to a block for the conversion between the block of the video and the bitstream representation of the video. The quadratic transform is performed between the forward primary transform and the quantization step, or between the inverse quantization step and the inverse primary transform. The reduced dimension is the dimension reduced from the dimension of the block. The method further includes performing the conversion based on the determination.
[0017] In another example aspect, a method for video processing is disclosed. The method includes: determining a constraint rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the constraint rule. The dimensionality-reduced quadratic transform has a dimension reduced from the dimension of the current video block. The dimensionality-reduced quadratic transform and the primary transform are applied together in a specific order during the conversion.
[0018] In another example aspect, another method for video processing is disclosed. The method includes: determining a constraint rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and an adjacent video region, and the pixels of the current video block and the adjacent region, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the constraint rule. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block and the adjacent video region. The dimensionality-reduced quadratic transform and the primary transform are applied together in a specific order during the conversion.
[0019] In yet another example aspect, another method for video processing is disclosed. The method includes: determining a nulling rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the nulling rule. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The nulling rule specifies the maximum number of coefficients used by the dimensionality-reduced quadratic transform.
[0020] In yet another example aspect, another method for video processing is disclosed. The method includes: determining a nulling rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the nulling rule. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The nulling rule specifies the maximum number of coefficients used by the dimensionality-reduced quadratic transform.
[0021] In yet another example aspect, another method for video processing is disclosed. The method includes: determining a condition for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the condition. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The condition is signaled in the bitstream representation.
[0022] In yet another example aspect, another method for video processing is disclosed. The method includes: selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to a condition. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The conversion includes selectively applying a position-dependent intra prediction combination (PDPC) based on a coexistence rule.
[0023] In yet another example aspect, another method of video processing is disclosed. The method includes: applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of the current video block and the pixels of the current video block, and performing the conversion by conditionally applying the dimensionality-reduced quadratic transform. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The application controls the use of neighboring samples for intra prediction during the conversion.
[0024] In yet another example aspect, another method of video processing is disclosed. The method includes: selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of the current video block and the pixels of the current video block, and performing the conversion by conditionally applying the dimensionality-reduced quadratic transform. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The selective application controls the use of the quantization matrix during the conversion.
[0025] In yet another example aspect, a video encoder is disclosed. The video encoder includes a processor configured to implement one or more of the above methods.
[0026] In yet another example aspect, a video decoder is disclosed. The video decoder includes a processor configured to implement one or more of the above methods.
[0027] In yet another example aspect, a computer-readable medium is disclosed. The medium includes code for implementing one or more of the above methods stored on the medium.
[0028] These and other aspects are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 An example of an encoder block diagram is shown.
[0030] Figure 2 An example of 67 intra prediction modes is shown.
[0031] Figures 3A - 3B An example of reference samples for wide-angle intra prediction is shown.
[0032] Figure 4 It is an example of the discontinuity problem in the case of a direction exceeding 45 degrees.
[0033] Figures 5A - 5D An example of samples used by PDPC applied to diagonal and adjacent corner intra modes is shown.
[0034] Figure 6 It is an example of the partitioning of 4×8 and 8×4 blocks.
[0035] Figure 7An example of the partitioning of all blocks other than 4×8, 8×4, and 4×4.
[0036] Figure 8 Partition a 4×8 sample block into two independently decodable regions.
[0037] Figure 9 An example sequence showing the processing of pixel rows using a vertical predictor to maximize the throughput of 4×N blocks.
[0038] Figure 10 An example of a quadratic transform is shown.
[0039] Figure 11 An example of the proposed reduced-dimensional quadratic transform (RST) is shown.
[0040] Figure 12 Examples of forward and inverse (or reverse) reduced-dimensional transforms are shown.
[0041] Figure 13 An example of forward RST 8×8 processing with a 16×48 matrix is shown.
[0042] Figure 14 An example of non-zero elements in scan positions 17 to 64 is shown.
[0043] Figure 15 Diagrams of sub-block transform modes SBT-V and SBT-H.
[0044] Figure 16 A block diagram of an example hardware platform for implementing the techniques described in this document.
[0045] Figure 17 A flowchart of an example method for video processing.
[0046] Figure 18 A block diagram of an example video processing system in which the disclosed techniques may be implemented.
[0047] Figure 19 A flowchart of an example method for video processing according to the present technology.
[0048] Figure 20 A flowchart of another example method for video processing according to the present technology.
[0049] Figure 21 A flowchart of another example method for video processing according to the present technology.
[0050] Figure 22 A flowchart of another example method for video processing according to the present technology.
[0051] Figure 23It is a flowchart of another example method for video processing according to the present technology.
[0052] Figure 24A It is a flowchart of an example method for video encoding according to the present technology.
[0053] Figure 24B It is a flowchart of an example method for video decoding according to the present technology.
[0054] Figure 25 It is a flowchart of another example method for video processing according to the present technology.
[0055] Figure 26 It is a flowchart of another example method for video processing according to the present technology.
[0056] Figure 27 It is a flowchart of another example method for video processing according to the present technology.
[0057] Figure 28 It is a flowchart of yet another example method for video processing according to the present technology. Detailed implementation
[0058] In this document, chapter titles are used for easy understanding, and the embodiments disclosed in a chapter are not limited to that chapter. In addition, although some embodiments are described with reference to general video coding or other specific video codecs, the disclosed technology is also applicable to other video coding technologies. In addition, although some embodiments describe video encoding steps in detail, it will be understood that the corresponding decoding steps inverse to the encoding will be implemented by a decoder. In addition, the term video processing covers video encoding or compression, video decoding or decompression, and video transcoding, i.e., the representation of video pixels from one compressed format to another compressed format or at a different compression bit rate.
[0059] 1. Overview
[0060] This patent document relates to video coding and decoding technology. Specifically, it relates to the transformation in video coding and decoding. It can be applied to existing video coding and decoding standards, such as HEVC, or the finalized standard (General Video Coding). It may also be applicable to future video coding and decoding standards or video codecs.
[0061] 2. Preliminary discussion
[0062] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Starting from H.262, video coding standards are based on a hybrid video coding structure that utilizes temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the "Joint Exploration Model" (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the Versatile Video Coding (VVC) standard, which aims to reduce the bitrate by 50% compared to HEVC.
[0063] 2.1 Color Space and Chroma Subsampling
[0064] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a digital tuple, typically 3 or 4 values or color components (e.g., RGB). Basically, a color space is a refinement of a coordinate system and subspace.
[0065] For video compression, the most commonly used color spaces are YCbCr and RGB.
[0066] YcbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also known as YCBCR or Y'CBCR, is a series of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue-difference and red-difference chrominance components. Y′ (with the prime symbol) is distinguished from Y (luminance), which means that the light intensity is non-linearly encoded based on gamma-corrected RGB primaries.
[0067] Chroma subsampling is the practice of encoding an image by taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance, and achieving a lower resolution for chrominance information than for luminance information.
[0068] 2.1.1 Format 4:4:4
[0069] Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production.
[0070] 2.1.2 Format 4:2:2
[0071] The two chroma components are sampled at half the sampling rate of the luminance (halving the horizontal chroma resolution). This reduces the bandwidth of the uncompressed video signaling by one-third with little visual difference.
[0072] 2.1.3 Format 4:2:0
[0073] In 4:2:0, the horizontal sampling is twice that of 4:1:1, but since the Cb and Cr channels are only sampled on every other line in this scheme, the vertical resolution is halved. Thus, the data rate is the same. Cb and Cr are subsampled by a factor of 2 in the horizontal and vertical directions respectively. There are three variants of the 4:2:0 scheme with different horizontal and vertical positions.
[0074] In MPEG-2, Cb and Cr are cosited in the horizontal direction. Cb and Cr are located between pixels (interstitial positions) in the vertical direction.
[0075] In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are in interstitial positions, midway between alternating luminance samples.
[0076] In 4:2:0 DV, Cb and Cr are cosited in the horizontal direction. In the vertical direction, they are both located on alternating lines.
[0077] 2.2 Encoding and decoding streams of typical video codecs
[0078] Figure 1 An example of the encoder block diagram of VVC is shown, which includes three loop filter blocks: the deblocking filter (DF), sample adaptive offset (SAO), and ALF (adaptive loop filter). Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by increasing the offset and applying a finite impulse response (FIR) filter respectively, where the auxiliary information for encoding and decoding signals the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and repair the artifacts created by the previous stage.
[0079] 2.3 Intra-mode encoding and decoding with 67 intra-prediction modes
[0080] To capture any edge direction presented in natural videos, the number of intra prediction modes in the direction frame is extended from 33 used in HEVC to 65. The additional direction modes are indicated by dashed arrows in Figure 2 and the planar and DC modes remain unchanged. These denser intra prediction modes are applicable to all block sizes as well as both luma and chroma intra prediction.
[0081] As Figure 2 shown, the conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VTM2, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original method and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0082] In HEVC, each intra-coded block has a square shape and the length of each side is a power of 2. Therefore, no division operation is required to generate the intra predictor using the DC mode. In VTM2, blocks can have a rectangular shape and in general, a division operation must be used for each block. To avoid using division for DC prediction, only the longer side is used to calculate the average of non-square blocks.
[0083] 2.4 Wide-Angle Intra Prediction for Non-Square Blocks
[0084] The conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VTM2, for non-square blocks, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled using the original method and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes for a particular block remains unchanged (e.g., 67), and the intra mode coding and decoding remain unchanged.
[0085] To support these prediction directions, as Figures 3A - 3B shown, an upper reference of length 2W+1 and a left reference of length 2H+1 are defined.
[0086] The number of modes of the replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 1.
[0087] Table 1: Replacing Intra Prediction Modes with Wide-Angle Modes
[0088] Condition Replaced Intra - prediction Mode W / H == 2 Modes 2, 3, 4, 5, 6, 7 W / H > 2 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H == 1 None H / W == 1 / 2 Modes 61, 62, 63, 64, 65, 66 H / W < 1 / 2 Modes 57, 58, 59, 60, 61, 62, 63, 64, 65, 66
[0089] As Figure 4As shown, in the case of wide-angle intra prediction, two vertically adjacent predicted sample points can use two non-adjacent reference sample points. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the increased gap Δp α of the negative impact.
[0090] 2.5 Intra Prediction Combinations Related to Position
[0091] In VTM2, the intra prediction result of the planar mode is further corrected using the position-dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that calls a combination of unfiltered boundary reference sample points and HEVC-style intra prediction with filtered boundary reference sample points. PDPC is applied to the following non-signaled intra modes: planar, DC, horizontal, vertical, lower left angular mode and its eight adjacent angular modes, upper right angular mode and its eight adjacent angular modes.
[0092] The predicted sample point pred(x, y) is predicted using a linear combination of the intra prediction mode (DC, planar, angular) and reference sample points according to the following equation:
[0093] pred(x,y) = (wL × R -1,y + wT × R x,-1 – wTL × R -1,-1 +(64 – wL – wT + wTL) × pred(x,y) + 32) >> 6
[0094] where, R x,-1 , R -1,y represent the reference sample points located above and to the left of the current sample point (x, y) respectively, and R -1,-1 represents the reference sample point located at the upper left corner of the current block.
[0095] If PDPC is applied to the DC, planar, horizontal, and vertical intra modes, no additional boundary filters are required, as required by the DC mode boundary filter or the horizontal / vertical mode edge filter in the case of HEVC.
[0096] Figures 5A - 5D Illustrates the definition of the reference sample points (R x,-1 , R- 1,y and R -1,-1 ) applied to various prediction modes. The predicted sample point pred(x’, y’) is located at (x', y') within the prediction block. The coordinate x of the reference sample point R x,-1 is given by x = x’ + y’ + 1, and the coordinate y of the reference sample point R -1,y is similarly given by y = x’ + y’ + 1.
[0097] Figures 5A to 5DProvides the definition of the samples used by PDPC applied to diagonal and adjacent angle intra modes
[0098] The PDPC weights depend on the prediction mode and are shown in Table 2.
[0099] Table 2: Example of PDPC weights according to the prediction mode
[0100] Prediction Mode wT wL wTL Upper - right Diagonal 16 >> ((y' << 1) >> shift) 16 >> ((x' << 1) >> shift) 0 Lower - left Diagonal 16 >> ((y' << 1) >> shift) 16>>((x’<<1)>>shift) 0 Upper - right Adjacent Diagonal 32 >> ((y' << 1) >> shift) 0 0 Lower - left Adjacent Diagonal 0 32 >> ((x' << 1) >> shift) 0
[0101] 2.6 Intra Sub - block Partitioning (ISP)
[0102] In some embodiments, as shown in Table 3, ISP is proposed to divide the block for intra - luminance prediction vertically or horizontally into 2 or 4 sub - partitions according to the block size dimension. Figure 6 and Figure 7 Shows examples of two possibilities. All sub - partitions meet the condition of having at least 16 samples.
[0103] Table 3 Number of sub - partitions depending on the block size
[0104] Block Size Number of Sub - divisions 4×4 Not Divided 4×8 and 8×4 2 All Other Cases 4
[0105] Figure 6 Shows examples of the partitioning of 4×8 and 8×4 blocks.
[0106] Figure 7 Shows examples of the partitioning of all blocks other than 4×8, 8×4, and 4×4.
[0107] For each of these sub - partitions, the residual signal is generated by entropy - decoding the coefficients sent by the encoder, then inverse - quantizing and inverse - transforming them. Then, intra - prediction is performed on the sub - partition, and finally, the corresponding reconstructed samples are obtained by adding the residual signal and the prediction signal. Thus, the reconstructed values of each sub - partition will be available for generating the prediction of the next one, repeating this process, and so on. All sub - partitions share the same intra mode.
[0108] Based on the partitioning and intra mode utilized, two different types of processing orders are used, which are called the normal order and the reverse order. In the normal order, the first sub - partition to be processed is the one containing the top - left sample of the CU, and then it continues down (for horizontal partitioning) or to the right (for vertical partitioning). As a result, the reference samples used to generate the sub - partition prediction signals are only on the left and above the line. On the other hand, the reverse processing order starts from the sub - partition containing the bottom - left sample of the CU and then continues up; or, starts from the sub - partition containing the top - right sample of the CU and then continues to the left.
[0109] 2.7 Block Differential Pulse Code - Modulation Coding (BDPCM)
[0110] Due to the shape of the horizontal (as opposed to vertical) predictor, which uses the left (A) (as opposed to the upper (B)) pixel for the prediction of the current pixel, the most efficient way to process the throughput of this block is to process all pixels in a column (as opposed to a row) in parallel and process these columns (as opposed to rows) sequentially. To increase throughput, we introduce the following processing: when the predictor selected on this block is vertical, a block with a width of 4 is divided into two halves with horizontal boundaries, and when the predictor selected on this block is horizontal, a block with a height of 4 is divided into two halves with vertical boundaries.
[0111] When dividing the block, it is not allowed for samples from one region to use pixels from another region to calculate the prediction: if this occurs, the predicted pixel is replaced with a reference pixel in the prediction direction. Figure 8 The different positions of the current pixel X in a 4×8 block for vertical prediction are shown.
[0112] Figure 8 An example of dividing a block of 4×8 samples into two independently decodable regions is shown.
[0113] Due to this feature, it is now possible to process a 4×4 block in 2 cycles and a 4×8 or 8×4 block in 4 cycles, and so on, as Figure 9 shown.
[0114] Figure 9 An example of the order of processing pixel rows using a vertical predictor to maximize the throughput of a 4×N block is shown.
[0115] Table 4 summarizes the number of cycles required to process a block depending on the block size. Simply put, any block with a size greater than or equal to 8 can be processed at 8 pixels or more per cycle.
[0116] Table 4 Worst-case throughput for blocks of size 4×N, N×4
[0117]
[0118] 2.8 Quantized Residual Domain BDPCM
[0119] In some embodiments, quantized residual domain BDPCM (hereinafter referred to as RBDPCM) is proposed. Intra-frame prediction of the entire block is performed by sample replication in a prediction direction (horizontal or vertical prediction) similar to intra-frame prediction. The residuals are quantized, and the delta between the quantized residuals and the quantized values of their predictors (horizontal or vertical) is encoded and decoded.
[0120] For a block of size M (rows) × N (columns), let r i,j, where \(0\leq i\leq M - 1\) and \(0\leq j\leq N - 1\), is the prediction residual after performing intra prediction using unfiltered samples from the upper or left block boundaries, horizontally (copying the left - adjacent pixel values across rows for the prediction block) or vertically (copying the upper - adjacent rows to each row in the prediction block). Let \(Q(r i,j )\), where \(0\leq i\leq M - 1\) and \(0\leq j\leq N - 1\), denote the quantized version of the residual \(r i,j \), where the residual is the difference between the original block and the prediction block values. Then, block DPCM is applied to the quantized residual samples, resulting in a modified \(M\times N\) array with elements When signaling vertical BDPCM:
[0121]
[0122] For horizontal prediction, similar rules apply, and the residual quantized samples are obtained as follows:
[0123]
[0124] The residual quantized samples are sent to the decoder.
[0125] On the decoder side, the above calculations are reversed to generate \(Q(r i,j )\), where \(0\leq i\leq M - 1\) and \(0\leq j\leq N - 1\). For the case of vertical prediction,
[0126]
[0127] For the horizontal case,
[0128]
[0129] The inverse - quantized residual \(Q -1 (Q(r(Q(r i,j ))) is added to the intra - block prediction values to generate the reconstructed sample values.
[0130] The main benefit of this scheme is that inverse DPCM can be done on - the - fly during coefficient parsing by simply adding the predictor when parsing the coefficients, or it can be performed after parsing.
[0131] Transform skip is always used in the quantized - residual domain BDPCM.
[0132] 2.9 Multiple Transform Sets (MTS) in VVC
[0133] In VTM 4, the transformation with a maximum block size of 64×64 is enabled, which is mainly used for videos with higher resolutions, such as 1080p and 4K sequences. For a transform block with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are set to zero so that only the low-frequency coefficients are retained. For example, for a transform block of M×N, where M is the block width and N is the block height, when M equals 64, only the transform coefficients in the left 32 columns are retained. Similarly, when N equals 64, only the transform coefficients in the top 32 rows are retained. When the transform skip mode is used for large blocks, the entire block is used without setting any values to zero.
[0134] In addition to the DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding and decoding of inter-frame and intra-frame coding blocks. It uses multiple selected transforms of DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. The following table shows the basic functions of the selected DST / DCT.
[0135]
[0136] To maintain the orthogonality of the transform matrix, the transform matrix is quantized more precisely compared to the transform matrix in HEVC. To keep the intermediate values of the transformed coefficients within the 16-bit range, all coefficients should have 10 bits after horizontal and vertical transforms.
[0137] To control the MTS scheme, separate enable flags are specified for intra-frame and inter-frame at the SPS level. When MTS is enabled on the SPS, the CU-level flag is signaled to indicate whether MTS is applied. Here, MTS only applies to luminance. The MTS CU-level flag is signaled when the following conditions are met.
[0138] - Both the width and height are less than or equal to 32
[0139] - The CBF flag equals 1
[0140] If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 3-10. When it comes to the transform matrix precision, an 8-bit primary transform core is used. Therefore, all the transform cores used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform cores (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) all use the 8-bit primary transform core.
[0141]
[0142] To reduce the complexity of large-size DST-7 and DCT-8, the high-frequency transform coefficients of DST-7 and DCT-8 with a size (width or height, or both width and height) equal to 32 are set to zero. Only the coefficients within the 16×16 low-frequency region are retained.
[0143] As in HEVC, the residual of a block can be encoded and decoded in transform skip mode. To avoid redundancy in syntax encoding and decoding, when the CU-level MTS_CU_flag is not equal to 0, the transform skip flag is not signaled. The block size limit for transform skip is the same as that of MTS in JEM4, which means that transform skip is applicable to a CU when both the width and height of the block are equal to or less than 32.
[0144] 2.10 Examples of Reduced-Dimension Second-Order Transform (RST)
[0145] 2.10.1 Examples of Non-Separable Second-Order Transform (NSST)
[0146] In some embodiments, a second-order transform (also known as non-separable second-order transform) is applied between the forward primary transform and quantization (on the encoder side), and between the inverse quantization and inverse primary transform (on the decoder side). As Figure 10 shown, a 4×4 (or 8×8) second-order transform is performed according to the block size. For example, for each 8×8 block, a 4×4 second-order transform is applied to the smaller blocks (e.g., min(width, height) < 8), and an 8×8 second-order transform is applied to the larger blocks (e.g., min(width, height) > 4).
[0147] Figure 10 Examples of the second-order transform in JEM are shown.
[0148] The application of the non-separable transform is described below by taking the input as an example. To apply the non-separable transform, first, the following 4×4 input block X
[0149]
[0150] Represented as a vector
[0151]
[0152] The calculation of the inseparable transformation is in indicates the transform coefficient vector, and T is the 16×16 transform matrix. The 16×1 coefficient vector is then converted to Reorganized into 4×4 blocks. Coefficients with smaller indices will be placed in the 4×4 coefficient block with smaller scan indices. There are a total of 35 transform sets, each using 3 inseparable transform matrices (kernels). The mapping from intra prediction modes to transform sets is predefined. For each transform set, the selected inseparable secondary transform candidate is further specified by a secondary transform index that is explicitly signaled. In the bitstream, the index is signaled once per intra CU after the transform coefficients.
[0153] 2.10.2 Example of Reducing Dimensionality with Quadratic Transform (RST) / Low-Frequency Non-separable Transform (LFNST)
[0154] A reduced-dimensional quadratic transform (RST), also known as a low-frequency non-separable transform (LFNST), is introduced to map to 4 transform sets (instead of 35 transform sets). In some embodiments, 16×64 (can be further reduced to 16×48) and 16×16 matrices are used for 8×8 and 4×4 blocks, respectively. For convenience, the 16×64 (can be further reduced to 16×48) transform is denoted as RST8×8 and the 16×16 transform is denoted as RST4×4. Figure 11 An example of RST is shown.
[0155] Figure 11 An example of the proposed reduced-dimensionality quadratic transform (RST) is shown.
[0156] RST calculation
[0157] The main idea of dimensionality reduction transformation (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N(R <N)是降维因子。
[0158] The RT matrix is an R×N matrix as follows:
[0159]
[0160] The R rows of the transformation are the R bases of the N-dimensional space. The inverse transformation matrix of RT is the transposed matrix of its forward transformation. Figure 12 Examples of forward and reverse RT are shown in .
[0161] Figure 12 Examples of forward and inverse dimensionality reduction transforms are shown.
[0162] In some embodiments, an RST 8×8 with a dimensionality reduction factor of 4 (1 / 4 size) is applied. Thus, instead of using 64×64 (which is the size of a conventional 8×8 non-separable transform matrix), a 16×64 direct matrix is used. In other words, a 64×16 inverse RST matrix is used on the decoder side to generate core (primary) transform coefficients in the upper left 8×8 region. The forward RST 8×8 uses a 16×64 (or 8×64 for an 8×8 block) matrix such that non-zero coefficients are generated only in the upper left 4×4 region within a given 8×8 region. In other words, if RST is applied, the 8×8 region except for the upper left 4×4 region will have only zero coefficients. For RST 4×4, a 16×16 (or 8×16 for a 4×4 block) direct matrix multiplication is applied.
[0163] The inverse RST is conditionally applied when the following two conditions are met:
[0164] a. The block size is greater than or equal to a given threshold (W >= 4 and H >= 4)
[0165] b. The transform skip mode flag is equal to zero
[0166] If the width (W) and height (H) of the transform coefficient block are both greater than 4, the RST 8×8 is applied to the upper left 8×8 region of the transform coefficient block. Otherwise, the RST 4×4 is applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block.
[0167] If the RST index is equal to 0, the RST is not applied. Otherwise, the RST is applied and its kernel is selected using the RST index. The RST selection method and the encoding and decoding of the RST index will be described later.
[0168] In addition, the RST is applied to intra CUs in intra and inter slices, and is applied to both luminance and chrominance. If dual-tree is enabled, the RST indices for luminance and chrominance are signaled separately. For inter slices (dual-tree disabled), a single RST index is signaled and used for both luminance and chrominance.
[0169] In some embodiments, intra-subdivision (ISP) is adopted as a new intra prediction mode. When the ISP mode is selected, since the performance improvement is weak even if the RST is applied to each feasible segmentation block, the RST is disabled and the RST index is not signaled. In addition, disabling the RST for the residuals of ISP prediction can reduce the encoding complexity.
[0170] RST Selection
[0171] Select the RST matrix from four transform sets, each transform set consisting of two transforms. Determine which transform sets are applied according to the intra prediction mode as follows:
[0172] (1) If one of the three CCLM modes is indicated, select transform set 0.
[0173] (2) Otherwise, perform transform set selection according to the following table:
[0174] Transform Set Selection Table
[0175]
[0176] The index (denoted as IntraPredMode) for accessing this table ranges from [-14, 83], and this range is the transform mode index for wide-angle intra prediction.
[0177] Reduced-dimensional RST matrix
[0178] As a further simplification, apply a 16×48 matrix with the same transform set configuration instead of a 16×64 matrix. Each matrix obtains 48 input data from three 4×4 blocks in the upper-left 8×8 block (except for the lower-right 4×4 block) ( Figure 13 ).
[0179] Figure 13 An example of forward RST 8×8 processing with a 16×48 matrix is shown.
[0180] RST Signaling
[0181] Forward RST 8×8 with R = 16 uses a 16×64 matrix, so it only produces non-zero coefficients in the upper-left 4×4 area within a given 8×8 area. In other words, if RST is applied, the 8×8 area except for the upper-left 4×4 area only generates zero coefficients. As a result, when any non-zero elements are detected in the 8×8 block area except for the upper-left 4×4 ( Figure 14 shown in), the RST index is not encoded or decoded because this means that RST is not applied. In this case, the RST index is inferred to be zero.
[0182] Figure 14 An example of non-zero elements in scan positions 17 to 64 is shown.
[0183] Zeroing Range
[0184] Normally, any coefficient in a 4×4 sub-block may be non-zero before applying the inverse RST to the 4×4 sub-block. However, in some cases, some coefficients in the 4×4 sub-block must be zero before applying the inverse RST to the sub-block.
[0185] Let nonZeroSize be a variable. When rearranging the indices into a one-dimensional array before inverse RST, any coefficient whose index is not less than nonZeroSize must be zero.
[0186] When nonZeroSize is equal to 16, there are no zeroing constraints for the coefficients in the top-left 4×4 sub-block.
[0187] In some embodiments, when the current block size is 4×4 or 8×8, nonZeroSize is set to be equal to 8. For other block dimensions, nonZeroSize is set to be equal to 16.
[0188] Example description of RST
[0189] In the following tables and descriptions, bold italic text is used to indicate changes that can be made to the current syntax to accommodate certain embodiments described in this document.
[0190] Sequence parameter set RBSP syntax
[0191]
[0192]
[0193] Residual encoding and decoding syntax
[0194]
[0195] Encoding and decoding unit syntax
[0196]
[0197]
[0198] Sequence parameter set RBSP semantics
[0199]
[0200] Encoding and decoding unit semantics
[0201]
[0202] Transformation processing of transform coefficients for scaling
[0203] Overview
[0204] The input to this processing is:
[0205] – Luminance position (xTbY, yTbY), specifying the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture;
[0206] – The variable nTbW, which specifies the width of the current transform block,
[0207] – The variable nTbH, which specifies the height of the current transform block,
[0208] – The variable cIdx, which specifies the color component of the current block,
[0209] – The (nTbW)×(nTbH) array d[x][y] of scaling transform coefficients, where x = 0..nTbW-1 and y = 0..nTbH-1.
[0210] The output of this process is the (nTbW)×(nTbH) array r[x][y] of residual samples for x = 0..nTbW-1 and y = 0..nTbH-1.
[0211]
[0212]
[0213] Second transform process
[0214]
[0215] Second transform matrix derivation process
[0216]
[0217]
[0218] 2.11 Clipping of inverse quantization in HEVC
[0219] In HEVC, the scaled transform coefficient d' is calculated as d' = Clip3(coeffMin, coeffMax, d), where d is the scaled transform coefficient before clipping.
[0220] For the luma component, coeffMin = CoeffMinY; coeffMax = CoeffMaxY. For the chroma component, coeffMin = CoeffMinC; coeffMax = CoeffMaxC; where,
[0221] CoeffMinY = -(1<<(extended_precision_processing_flag? Max(15, BitDepthY+6) : 15))
[0222] CoeffMinC = -(1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))
[0223] CoeffMaxY = (1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) - 1
[0224] CoeffMaxC = (1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) – 1
[0225] extended_precision_processing_flag is a syntax element signaled in the SPS.
[0226] 2.12 Affine Linear Weighted Intra Prediction (ALWIP, also known as Matrix-based Intra Prediction, MIP)
[0227] In some embodiments, two tests are conducted. In Test 1, the memory limit for ALWIP is 8K bytes, with a maximum of 4 multiplications per sample. Test 2 is similar to Test 1, but further simplifies the design in terms of memory requirements and model architecture.
[0228] * A single set of matrices and offset vectors for all block shapes.
[0229] * For all block shapes, the number of modes is reduced to 19.
[0230] * The memory requirement is reduced to 5760 10-bit values, i.e., 7.20KB.
[0231] * The linear interpolation of the predicted samples is done in one step in each direction, instead of iterative interpolation as in the first test.
[0232] 2.13 Sub-Block Transform
[0233] For an inter-predicted CU where cu_cbf equals 1, cu_sbt_flag can be signaled to indicate whether to decode the entire residual block or a sub-part of the residual block. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded / decoded by an inferred adaptive transform while the other part of the residual block is zeroed. SBT is not applicable to combined inter-intra modes.
[0234] In the sub-block transform, a position-dependent transform is applied to the luminance transform blocks in SBT-V and SBT-H (always using the chrominance TB of DCT-2). Two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transforms for each SBT position are specified in Figure 15 . For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the corresponding transform is set to DCT-2. Thus, the sub-block transform jointly specifies the TU slices, cbf, and horizontal and vertical transforms of the residual block, which can be regarded as a syntax shortcut for the case where the main residual of the block is on one side of the block.
[0235] Figure 15 FIGS. to
[0236] are diagrams of the sub-block transform modes SBT-V and SBT-H.
[0236] 3. Examples of Problems Solved by the Embodiments
[0237] The current design has the following problems:
[0238] (1) The cropping and shift / rounding operations in MTS / RST may not be optimal.
[0239] (2) The RST applied to two adjacent 4×4 blocks may be costly.
[0240] (3) The RST can be performed in different ways for different color components.
[0241] (4) The RST may not perform well when used for screen content coding and decoding.
[0242] (5) The interaction between the RST and other coding and decoding tools is not clear.
[0243] (6) The transform matrix of the RST can be stored more efficiently.
[0244] (7) It is not clear how to apply the quantization matrix to the RST.
[0245] 4. Example Embodiments and Techniques
[0246] The following list of embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way.
[0247] In the following description, the coding and decoding information may include a prediction mode (e.g., intra / inter / IBC mode), a motion vector, a reference picture, an inter - prediction direction, an intra - prediction mode, a CIIP (Combined Intra - Inter Prediction) mode, an ISP mode, an affine intra - mode, a transform core used, a transform skip flag, etc., for example, the information required when encoding a block.
[0248] In the following discussion, SatShift(x, n) is defined as
[0249]
[0250] SatShift(x, n) is defined as Shift(x,n)=(x + offset0)>>n.
[0251] In one example, offset0 and / or offset1 are set to (1<<n)>>1 or r(1<<(n - 1)). In another example, offset0 and / or offset1 are set to 0.
[0252] In another example, offset0 = offset1 = ((1<<n)>>1)-1 or ((1<<(n - 1)))-1.
[0253] Clip3(min, max, x) is defined as
[0254]
[0255] 1. After inverse RST, the output value should be clipped to the range [MinCoef, MaxCoef] (including MinCoef and MaxCoef), where MinCoef and / or MaxCoef are two variable integer values.
[0256] a. In one example, assuming that the coefficient after inverse quantization is clipped to the range [QMinCoef, QMaxCoef] (including QMinCoef and QMaxCoef), then MinCoef can be set equal to QMinCoef, and / or MaxCoef can be set equal to QMaxCoef.
[0257] b. In one example, MinCoef and / or MaxCoef may depend on the color component.
[0258] i. In one example, MinCoef and / or MaxCoef may depend on the bit - depth of the corresponding color component.
[0259] c. In one example, MinCoef and / or MaxCoef may depend on the shape of the block (e.g., square or non-square) and / or the size of the block.
[0260] d. In one example, the values of MinCoef and / or MaxCoef or the selection of candidate values may be signaled in, such as, SPS, PPS, slice header / picture group header / CTU / CU.
[0261] e. In one example, for the luminance component, MinCoef and / or MaxCoef may be derived as:
[0262] MinCoef = -(1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15))
[0263] MaxCoef = (1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)) – 1
[0264] where BitDepthY is the bit depth of the luminance component, and extended_precision_processing_flag may be signaled, such as, in SPS.
[0265] f. In one example, for a component, MinCoef and / or MaxCoef may be derived as:
[0266] MinCoef = -(1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15))
[0267] MaxCoef = (1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)) – 1,
[0268] where BitDepthC is the bit depth of the chrominance component, and extended_precision_processing_flag may be signaled, such as, in SPS.
[0269] g. In some embodiments, MinCoef is -(1 << 15), and MaxCoef is (1 << 15) - 1.
[0270] h. In one example, the consistency bitstream should satisfy that the transform coefficients after the forward RST should be within a given range.
[0271] 2. It is proposed that the way of applying the forward RST and / or the inverse RST to the M×N sub-blocks of the coefficients can depend on the number of sub-blocks to which the forward RST and / or the inverse RST are applied. For example, M = N = 4.
[0272] a. In one example, the zeroing range can depend on the sub-block index to which the RST is applied.
[0273] i. Alternatively, the zeroing range can depend on the number of sub-blocks to which the RST is applied.
[0274] b. In one example, when there are S sub-blocks (where S > 1, for example, S = 2) in the entire coefficient block to which the forward RST and / or the inverse RST are applied, the way of applying the forward RST and / or the inverse RST to the first and second sub-blocks of the coefficients can be different. For example, the first M×N sub-block can be the upper-left M×N sub-block.
[0275] i. In one example, for the first M×N sub-block of the coefficients (denoted as nonZeroSize0) and for the second M×N sub-block of the coefficients (denoted as nonZeroSize1), the nonZeroSize as described in Section 2.10 can be different.
[0276] 1) In one example, nonZeroSize0 can be greater than nonZeroSize1. For example, nonZeroSize0 = 16, and nonZeroSize1 = 8.
[0277] ii. In one example, when there is only one M×N sub-block to which the forward RST and / or the inverse RST are applied, or when there are more than one M×N sub-blocks to which the forward RST and / or the inverse RST are applied, the nonZeroSize as described in Section 2.10 can be different.
[0278] 1) In one example, if there are more than one M×N sub-blocks to which the forward RST and / or the inverse RST are applied, nonZeroSize can be equal to 8.
[0279] 3. It is recommended that if the current block size is 4×H or W×4, where H > 8 and W > 8, for example, M = N = 4, only apply the forward RST and / or the inverse RST to one M×N sub-block of the coefficients (such as the upper-left M×N sub-block).
[0280] a. In one example, if H > T1 and / or W > T2, for example, T1 = T2 = 16, then only the forward RST and / or the inverse RST is applied to an M×N sub-block of the coefficients.
[0281] b. In one example, if H < T1 and / or W < T2, for example, T1 = T2 = 32, then only the forward RST and / or the inverse RST is applied to an M×N sub-block of the coefficients.
[0282] c. In one example, for all H > 8 and / or W > 8, the forward RST and / or the inverse RST is only applied to an M×N sub-block of the coefficients.
[0283] d. In one example, if the current block size is M×H or W×N, where H >= N and W >= M, for example, M = N = 4, then the forward RST and / or the inverse RST is only applied to an M×N sub-block (e.g., the upper-left M×N sub-block).
[0284] 4. RST can be applied to a non-square region. Assume the region size is represented by K×L, where K is not equal to L.
[0285] a. Alternatively, in addition, zeroing can be applied to the transform coefficients after the forward RST to satisfy the maximum number of non-zero coefficients.
[0286] i. In one example, if the transform coefficients are outside the upper-left M×M region, where M is not greater than K and M is not greater than L, then they can be set to 0.
[0287] 5. It is recommended that the coefficients in two adjacent M×N sub-blocks can be included in a single forward RST and / or inverse RST. For example, M = N = 4.
[0288] a. In one example, one or more of the following operations can be performed on the encoder side. The following operations can be performed in sequence.
[0289] i. Rearrange the coefficients in two adjacent M×N sub-blocks into a one-dimensional vector with 2×M×N elements.
[0290] ii. Apply the forward RST with the transform matrix to the one-dimensional vector, where the transform matrix has 2×M×N columns and M×N rows (or M×N columns and 2×M×N rows).
[0291] iii. Rearrange the transformed one-dimensional vector with M×N elements into the first M×N sub-block (such as the upper-left sub-block).
[0292] iv. All coefficients in the second M×N sub-block can be set to zero.
[0293] b. In one example, one or more of the following operations may be performed on the decoder side. The following operations may be performed sequentially.
[0294] i. Rearrange the coefficients in the first M×N sub-block (such as the upper left sub-block) into a one-dimensional vector with M×N elements.
[0295] ii. Apply the inverse RST with the transformation matrix to the one-dimensional vector, where the transformation matrix has M×N columns and 2×M×N rows (or 2×M×N columns and M×N rows).
[0296] iii. Rearrange the transformed one-dimensional vector with 2×M×N elements into two adjacent M×N sub-blocks.
[0297] c. In one example, a block may be divided into K (K>1) sub-blocks, and both the primary transformation and the secondary transformation may be performed at the sub-block level.
[0298] 6. The zeroing range (e.g., nonZeroSize as described in Section 2.10) may depend on the color component.
[0299] a. In one example, for the same block size, the ranges of the luminance and chrominance components may be different.
[0300] 7. The zeroing range (e.g., nonZeroSize as described in Section 2.10) may depend on the codec information.
[0301] a. In one example, it may depend on the codec mode, such as the intra or non-intra mode.
[0302] b. In one example, it may depend on the codec mode, such as the intra or inter or IBC mode.
[0303] c. In one example, it may depend on the reference picture / motion information.
[0304] 8. It is recommended that the zeroing range (e.g., nonZeroSize as described in Section 2.10) for a specific block size may depend on the quantization parameter (QP).
[0305] a. In one example, assume that nonZeroSize is equal to nonZeroSizeA when QP is equal to QPA, and nonZeroSize is equal to nonZeroSizeB when QP is equal to QPB. If QPA is not less than QPB, then nonZeroSizeA is not greater than nonZeroSizeB.
[0306] b. Different transformation / inverse transformation matrices may be used for different nonZeroSizes.
[0307] 9. It is recommended that the zeroing range (e.g., nonZeroSize as described in Section 2.10) can be signaled, such as in SPS, PPS, picture header, slice header, slice group header, CTU row, CTU, CU, or any data unit in the video.
[0308] a. Alternatively, multiple ranges can be defined. And an indication of which candidate nonZeroSize is selected can be signaled, such as in SPS, PPS, picture header, slice header, slice group header, CTU row, CTU, and CU.
[0309] 10. Whether and / or how to apply RST can depend on the color format and / or separate plane coding and decoding and / or the use of color components.
[0310] a. In one example, RST may not be applied to chrominance components (such as Cb and / or Cr).
[0311] b. In one example, if the color format is 4:0:0, RST may not be applied to chrominance components.
[0312] c. In one example, if separate plane coding and decoding is used, RST may not be applied to chrominance components.
[0313] d. In one example, the nonZeroSize for a specific block dimension can depend on the color component.
[0314] i. In one example, for the same block dimension, the nonZeroSize on the chrominance component can be smaller than the nonZeroSize on the luminance component.
[0315] 11. It is recommended that when coding and decoding luminance and chrominance components with a single coding structure tree, RST control information (such as whether to apply RST, and / or which set of transform matrices to select) can be signaled to the luminance and chrominance components respectively.
[0316] 12. Whether and how to apply RST can depend on the coding information (such as coding mode) of the current block and / or adjacent blocks.
[0317] a. In one example, RST cannot be used for one or more specific intra prediction modes.
[0318] i. For example, RST cannot be used for the LM mode.
[0319] ii. For example, RST cannot be used for the LM-T mode.
[0320] iii. For example, RST cannot be used for the LM-A mode.
[0321] iv. For example, RST cannot be used for wide-angle intra prediction mode.
[0322] v. For example, RST cannot be used for BDPCM mode or / and DPCM mode or / and RBDPCM mode.
[0323] vi. For example, RST cannot be used for ALWIP mode.
[0324] vii. For example, RST cannot be used for certain specific angular intra prediction modes (such as, DC, planar, vertical, horizontal, etc.).
[0325] vx. For example, in LM mode or / and LM-T mode or / and LM-A mode, RST can be used for the luminance component, but not for the chrominance component.
[0326] ix. For example, when applying joint chrominance residual coding and decoding, RST may not be used for the chrominance component.
[0327] b. If RST cannot be applied, the syntax element used to indicate the information related to RST in the current block may not be signaled.
[0328] 13. It is recommended to apply RST to non-intra coded blocks.
[0329] a. In one example, RST can be applied to inter coded blocks.
[0330] b. In one example, RST can be applied to blocks coded with intra block copy (IBC).
[0331] c. In one example, RST can be applied to blocks coded with combined inter-intra prediction (CIIP).
[0332] 14. It is recommended to perform different levels of control on RST.
[0333] a. For example, information indicating whether RST (such as, a control flag) is applicable can be signaled in the PPS, slice header, picture header, slice group header, slice, CTU row, CTU.
[0334] b. Whether RST is applicable may depend on the standard profile / level / layer.
[0335] 15. It is proposed that whether to apply position-dependent intra prediction combination (PDPC) can depend on whether RST is applied.
[0336] a. In one example, if RST is applied to the current block, PDPC may not be applied.
[0337] b. In one example, if RST is applied to the current block, PDPC can be applied.
[0338] c. Alternatively, whether to apply RST can depend on whether PDPC is applied.
[0339] i. In one example, when PDPC is applied, RST is not applied.
[0340] ii. If RST cannot be applied, the syntax element used to indicate the information related to RST in the current block may not be signaled.
[0341] 16. It is recommended that whether to filter the neighboring samples used for intra prediction can depend on whether RST is applied.
[0342] a. In one example, if RST is applied to the current block, the neighboring samples may not be filtered.
[0343] b. In one example, if RST is applied to the current block, the neighboring samples may be filtered.
[0344] c. Alternatively, whether to apply RST can depend on whether the neighboring samples used for intra prediction are filtered.
[0345] i. In one example, when the neighboring samples used for intra prediction are filtered, RST is not applied.
[0346] ii. In one example, when the neighboring samples used for intra prediction are not filtered, RST is not applied.
[0347] iii. If RST cannot be applied, the syntax element used to indicate the information related to RST in the current block may not be signaled.
[0348] 17. It is recommended that when the current block is encoded and decoded using transform skip, RST can be applied.
[0349] a. For example, the main transform is skipped, but the secondary transform can still be applied.
[0350] b. The secondary transform matrix used in the transform skip mode may be different from the secondary transform matrix used in the non - transform skip mode.
[0351] 18. It is recommended that the transform matrix for RST can be stored with a bit width less than 8. For example, the transform matrix for RST can be stored with a bit width of 6 or 4.
[0352] 19. It is recommended that the transform matrix for RST can be stored in a predictive manner.
[0353] a. In one example, the first element in the first transformation matrix for RST can be predicted by the second element in the first transformation matrix for RST.
[0354] i. For example, the difference between the two elements can be stored.
[0355] ii. For example, the difference can be stored with a bit width less than 8 (such as 6 or 4).
[0356] b. In one example, the first element in the first transformation matrix for RST can be predicted by the second element in the second transformation matrix for RST.
[0357] i. For example, the difference between the two elements can be stored.
[0358] ii. For example, the difference can be stored with a bit width less than 8 (such as 6 or 4).
[0359] 20. It is proposed that the first transformation matrix for RST can be derived from the second transformation matrix for RST.
[0360] a. In one example, some elements of the second transformation matrix for RST can be picked up to construct the first transformation matrix for RST.
[0361] b. In one example, the first transformation matrix for RST is obtained by rotating or flipping all or part of the second transformation matrix for RST.
[0362] c. In one example, the first transformation matrix for RST is derived by downsampling or upsampling on the second transformation matrix for RST.
[0363] 21. It is proposed to signal a syntax element that is signaled before the residual (which can be transformed) to indicate information related to RST in the current block.
[0364] a. In one example, the signaling of information related to RST may not depend on the non - zero or zero coefficients calculated when parsing the residual.
[0365] b. In one example, non - zero or zero coefficients may not be counted when parsing the residual.
[0366] c. In one example, the coding block flag (cbf) flag for a sub - block that is set to all zeros by RST may not be signaled and is inferred as 0.
[0367] d. In one example, the valid flag of the coefficients set to zero by RST may not be signaled and is inferred as 0.
[0368] e. The scanning order of the parsing residual block may depend on whether and how the RST is applied.
[0369] i. In one example, coefficients set to zero by the RST may not be scanned.
[0370] f. The arithmetic coding and decoding context of the parsing residual block may depend on whether and how the RST is applied.
[0371] 22. It is recommended that whether and how to apply the quantization matrix may depend on whether and how the RST is applied.
[0372] a. In one example, different quantization matrices may be applied regardless of whether the RST is applied.
[0373] b. Alternatively, whether and how to apply the RST may depend on whether and how the quantization matrix is applied.
[0374] i. In one example, when applying the quantization matrix to a block, the RST may not be applied.
[0375] 23. It is recommended to apply the RST to the quantized coefficients / residuals.
[0376] a. In one example, when using transform skipping, the RST may be applied to the residuals.
[0377] b. In one example, the RST may be applied to the quantized transform coefficients of a block.
[0378] 24. It is recommended to apply the RST to the sub-block transform block.
[0379] a. In one example, the RST may be applicable to the top-left coefficient generated by the sub-block transform.
[0380] Figure 16 is a block diagram of the video processing apparatus 1600. The apparatus 1600 may be used to implement one or more methods described herein. The apparatus 1600 may be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1600 may include one or more processors 1602, one or more memories 1604, and video processing hardware 1606. The processor 1602 may be configured to implement one or more methods described herein. The (one or more) memories 1604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1606 may be used to implement some of the techniques described in this document in hardware circuits.
[0381] Figure 17FIG. 1700 is a flow diagram of an example method for video processing. Method 1700 includes determining (1702) a constraint rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between a bitstream representation of a current video block and pixels of the current video block. Method 1700 includes performing (1704) the conversion by applying the dimensionality-reduced quadratic transform according to the constraint rule. The dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block. The dimensionality-reduced quadratic transform and the primary transform are applied together in a specific order during the conversion.
[0382] Other embodiments and techniques are described in the following examples.
[0383] 1. A video processing method includes: determining a constraint rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between a bitstream representation of a current video block and pixels of the current video block, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the constraint rule; wherein the dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block, and wherein the dimensionality-reduced quadratic transform and the primary transform are applied together in a specific order during the conversion.
[0384] 2. The method according to example 1, wherein the conversion includes encoding the current video block into a bitstream representation, and wherein the specific order includes: first applying the primary transform in the forward direction, then selectively applying the dimensionality-reduced quadratic transform in the forward direction, and then quantizing the output of the dimensionality-reduced quadratic transform in the forward direction.
[0385] 3. The method according to example 1, wherein the conversion includes: decoding the current video block from the bitstream representation, and wherein the specific order includes: first dequantizing the bitstream representation, then selectively applying the dimensionality-reduced quadratic transform in the reverse direction, and then applying the primary transform in the reverse direction to the output of the dimensionality-reduced quadratic transform in the reverse direction.
[0386] 4. The method according to any one of examples 1-3, wherein the constraint rule specifies clipping the range of the output of the dimensionality-reduced quadratic transform in the reverse direction to the range [MinCoef, MaxCoef], including MinCoef and MaxCoef, where MinCoef and / or MaxCoef are two integer values that are functions of the conditions of the current video block.
[0387] 5. The method according to example 4, wherein the condition of the current video block is the type of color or luminance component represented by the current video block.
[0388] 6. The method according to Example 1, wherein the constraint rule stipulates that a dimensionality-reduced quadratic transform is applied to one or more M×N sub-blocks of the current video block, and the remaining sub-blocks of the current video block are set to zero.
[0389] 7. The method according to Example 1, wherein the constraint rule stipulates that the dimensionality-reduced quadratic transform is applied differently to different sub-blocks of the current video block.
[0390] 8. The method according to any one of Examples 1-5, wherein the constraint rule stipulates that since the size of the current video block is 4×H or W×4, the dimensionality-reduced quadratic transform is applied exactly to one M×N sub-block of the current video block, where H is the height in integer pixels and W is the width in integer pixels.
[0391] 9. The method according to Example 8, wherein H > 8 or W > 8.
[0392] 10. The method according to any one of Examples 1-9, wherein the current video block is a non-square region of the video.
[0393] 11. The method according to Example 2 or 3, wherein the constraint rule stipulates that the transform coefficients of the main transform in the forward direction are set to zero, or zero coefficients are filled into the output of the quadratic transform in the reverse direction.
[0394] Other embodiments of Examples 1-5 are described in Item 1 of Section 4. Other embodiments of Examples 6-7 are described in Item 2 of Section 4. Other embodiments of Examples 8-9 are described in Item 3 of Section 4. Other embodiments of Examples 10-11 are described in Item 4 of Section 4.
[0395] 12. A video processing method, comprising: determining a constraint rule for selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of the current video block and the pixels of the current video block and between the bitstream representation of an adjacent video region and the pixels of the adjacent region, and performing the conversion by applying the dimensionality-reduced quadratic transform according to the constraint rule; wherein the dimensionality-reduced quadratic transform has a reduced dimension from the dimensions of the current video block and the adjacent video region, and wherein the dimensionality-reduced quadratic transform and the main transform are applied together in a specific order during the conversion.
[0396] 13. The method according to Example 12, wherein the adjacent video region includes the upper left block of the current video block.
[0397] 14. The method according to Example 12, wherein the current video block and the adjacent video region correspond to sub-blocks of a parent video block.
[0398] Other embodiments of Examples 12-14 are described in Item 5 of Section 4.
[0399] 15. A video processing method, comprising: determining a zeroing rule to selectively apply a dimensionality-reducing quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reducing quadratic transform according to the zeroing rule; wherein the dimensionality-reducing quadratic transform has a dimension reduced from the dimension of the current video block; and wherein the zeroing rule specifies the maximum number of coefficients used by the dimensionality-reducing quadratic transform.
[0400] 16. The method according to example 15, wherein the maximum number of the coefficients is a function of the component identification of the current video block.
[0401] 17. The method according to example 16, wherein the maximum number of coefficients for a luminance video block and a chrominance video block is different.
[0402] 18. The method according to any one of examples 15-17, wherein the zeroing rule specifies a zeroing range, which is a function of the codec information of the current video block.
[0403] 19. The method according to any one of examples 15-17, wherein the zeroing rule specifies a zeroing range, which is a function of the quantization parameter of the current video block.
[0404] 20. The method according to any one of examples 15-19, wherein the zeroing range is indicated in the bitstream representation by a field included at the sequence parameter set level, or the picture parameter set level, or the picture header, or the slice header, or the slice group header, or the codec tree unit row, or the codec tree unit, or the codec unit, or the video data unit level.
[0405] Other embodiments of examples 15-17 are described in item 6 of section 4. Other embodiments of example 18 are described in item 7 of section 4. Other embodiments of example 19 are described in item 8 of section 4. Other embodiments of example 20 are described in item 9 of section 4.
[0406] 21. A video processing method, comprising: determining a condition for selectively applying a dimensionality-reducing quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by applying the dimensionality-reducing quadratic transform according to the condition; wherein the dimensionality-reducing quadratic transform has a dimension reduced from the dimension of the current video block; and wherein the condition is signaled in the bitstream representation.
[0407] 22. The method according to example 21, wherein the condition is the color format or the use of separate plane coding or based on the color identification of the current video block.
[0408] Other embodiments of Examples 21-22 are described in item 10 of Section 4.
[0409] 23. The method according to any one of Examples 21-22, wherein for the chrominance and luminance components, the condition is signaled in the bitstream representation, respectively.
[0410] Other embodiments of Example 23 are described in item 11 of Section 4.
[0411] 24. The method according to any one of Examples 21-23, wherein the condition depends on the coding and decoding information of the current video block and adjacent video regions.
[0412] 25. The method according to Example 24, wherein the condition excludes the application to the current video block coded using a specific intra prediction mode.
[0413] Other embodiments of Examples 24-25 are described in item 12 of Section 4.
[0414] 26. The method according to Example 24, wherein the condition specifies the application to the current video block coded inter.
[0415] 27. The method according to Example 24, wherein the condition specifies the application to the current video block coded using the intra block copy mode.
[0416] Other embodiments of Examples 25-26 are described in item 13 of Section 4.
[0417] 28. The method according to Example 21, wherein the condition is signaled in the bitstream representation at a certain level such that all blocks within that level meet the condition, wherein the level is a sequence parameter set level or a picture parameter set level, or a picture header, or a slice header, or a slice group header, or a coding tree unit row, or a coding tree unit, or a coding unit, or at the video data unit level.
[0418] Other embodiments of Example 28 are described in item 14 of Section 4.
[0419] 29. The method according to Example 21, wherein the condition is to code the current video block using the transform skip mode.
[0420] Other embodiments of Example 29 are described in item 17 of Section 4.
[0421] 30. A video processing method, comprising: selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by conditionally applying the dimensionality-reduced quadratic transform; wherein the dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block; and wherein the conversion includes selectively applying a position-dependent intra prediction combination (PDPC) based on a coexistence rule.
[0422] 31. The method according to example 30, wherein the coexistence rule excludes applying the PDPC to the current video block due to applying the quadratic transform.
[0423] 32. The method according to example 30, wherein the coexistence rule stipulates applying the PDPC to the current video block due to applying the quadratic transform.
[0424] 33. The method according to example 30, wherein selectively applying the quadratic transform is performed on the current video block using the PDPC.
[0425] Other embodiments of examples 30-33 are described in item 15 of section 4.
[0426] 34. A video processing method, comprising: applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by conditionally applying the dimensionality-reduced quadratic transform; wherein the dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block; and wherein the applying controls using neighboring samples for intra prediction during the conversion.
[0427] Other embodiments of example 34 are described in item 16 of section 4.
[0428] 35. A video processing method, comprising: selectively applying a dimensionality-reduced quadratic transform during the conversion between the bitstream representation of a current video block and the pixels of the current video block, and performing the conversion by conditionally applying the dimensionality-reduced quadratic transform; wherein the dimensionality-reduced quadratic transform has a dimension reduced from that of the current video block; and wherein the selective applying controls the use of a quantization matrix during the conversion.
[0429] 36. The method according to example 35, wherein the use of the quantization matrix occurs only due to applying the quadratic transform.
[0430] Other embodiments of examples 35-36 are described in item 22 of section 4.
[0431] 37. The method according to any one of examples 1-36, wherein the primary transform and the quadratic transform are stored as transform matrices with a bit width less than 8.
[0432] 38. The method according to any one of Examples 1-36, wherein the primary transform and the secondary transform are stored as a prediction transform matrix.
[0433] 39. The method according to any one of Examples 1-36, wherein the primary transform can be derived from the secondary transform using a first rule, or wherein the secondary transform can be derived from the primary transform using a second rule.
[0434] 40. The method according to any one of Examples 1-36, wherein the bitstream representation includes information about the secondary transform or the primary transform before the residual information for the current video block.
[0435] Other embodiments of Examples 37-40 are described in Items 18, 19, 20, and 21 of Section 4.
[0436] 41. A video processing apparatus, comprising a processor configured to implement one or more of the methods described in Examples 1-40.
[0437] 42. A computer-readable medium having stored thereon code that, when executed by a processor, causes the processor to implement any one or more of the methods described in Examples 1-40.
[0438] It will be appreciated that the disclosed techniques may be embodied in a video encoder or decoder to improve compression efficiency using techniques including the use of a dimensionality-reduced secondary transform.
[0439] Figure 18 is a block diagram showing an example video processing system 1800 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 1800. System 1800 may include an input 1802 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1802 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0440] System 1800 may include a codec component 1804 that may implement various codec or encoding methods described herein. The codec component 1804 may reduce the average bitrate of a video from the input 1802 to the output of the codec component 1804 to generate a codec representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. As represented by component 1806, the output of the codec component 1804 may be stored or transmitted via a connected communication. The stored or transmitted bitstream (or codec) representation of the video received at the input 1802 may be used by component 1808 to generate pixel values or a displayable video that is sent to the display interface 1810. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "codec" operations or tools, it should be understood that encoding tools or operations are used at the encoder and the corresponding decoding tools or operations of the inverse encoding result will be performed by the decoder.
[0441] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), or a Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described herein may be implemented in various electronic devices, such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0442] Figure 19 is a flowchart of an example method 1900 of video processing in accordance with the present technology. Method 1900 includes, at operation 1910, determining that output values from a reduced-dimension inverse quadratic transform are constrained within a range [min, max] for a conversion between a block of a video and a bitstream representation of the video, including min and max. The reduced dimension is a dimension reduced from the dimension of the block, and min and max are integer values. Method 1900 includes, at operation 1920, performing the conversion based on the determination. In some embodiments, the reduced-dimension inverse quadratic transform includes a low-frequency non-separable inverse transform, where the low frequency corresponds to the reduced dimension.
[0443] In some embodiments, the coefficients after the inverse quantization step are constrained to the range [qmin, qmax], including qmin and qmax, where qmin and qmax are positive integers. At least one of (1) min equals qmin or (2) max equals qmax is satisfied. In some embodiments, the range is based on the color component of the block. In some embodiments, at least one of min or max is based on the bit depth of the color component. In some embodiments, the range is based on the shape of the block. In some embodiments, the range is based on whether the block is square or non-square. In some embodiments, the range is based on the dimension of the block. In some embodiments, one of min or max is signaled in the bitstream representation. In some embodiments, the range is signaled by a sequence parameter set, a picture parameter set, a slice header, a slice group header, a coding tree unit, or a coding unit.
[0444] In some embodiments, for the luma component of the block, min is -(1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)), and max is (1 << (extended_precision_processing_flag? Max(15, BitDepthY + 6) : 15)). BitDepthY is the bit depth of the luma component, and where extended_precision_processing_flag is a variable signaled in the bitstream representation.
[0445] In some embodiments, for the chroma component of the block, min is -(1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)), and max is (1 << (extended_precision_processing_flag? Max(15, BitDepthC + 6) : 15)). BitDepthC is the bit depth of the luma component, and where extended_precision_processing_flag is a variable signaled in the bitstream representation.
[0446] In some embodiments, min equals -(1 << 15), and max equals (1 << 15) - 1.
[0447] In some embodiments, extended_precision_processing_flag is signaled in the sequence parameter set. In some embodiments, the coefficients of the block after a secondary transform applied between the forward main transform and the quantization step are limited to a certain range.
[0448] Figure 20 It is a flowchart of an exemplary method 2000 for video processing according to the present technology. The method 2000 includes, at operation 2010, determining a manner of applying a reduced-dimension quadratic transform to sub-blocks of a block for conversion between the block of the video and the bitstream representation of the video, based on the number of sub-blocks to which a quadratic transform can be applied. The quadratic transform is applied to the block between the forward main transform and the quantization step, or between the inverse quantization step and the inverse main transform. The reduced dimension is a dimension reduced from the dimension of the block. The method 2000 further includes, at operation 2020, performing the conversion based on the determination.
[0449] In some embodiments, the reduced-dimension quadratic transform includes a low-frequency non-separable transform, where the low frequency corresponds to the reduced dimension. In some embodiments, the reduced dimension corresponds to the dimension of the sub-block.
[0450] In some embodiments, the sub-block has a dimension of 4×4. In some embodiments, the sub-block is associated with a sub-block index. Coefficients of sub-blocks outside a non-zero range are set to zero, and the zeroing range is determined based on the sub-block index. In some embodiments, coefficients of sub-blocks outside a non-zero range are set to zero. The non-zero range is determined based on the number of sub-blocks to which a quadratic transform can be applied.
[0451] In some embodiments, the number of sub-blocks to which a quadratic transform can be applied is greater than 1. The quadratic transform is applied to a first sub-block in a first manner, and the quadratic transform is applied to a second sub-block in a second manner different from the first manner. In some embodiments, coefficients of the first sub-block outside a first non-zero range are set to zero. Coefficients of the second sub-block outside a second non-zero range are set to zero, and the first non-zero range is different from the second non-zero range. In some embodiments, the first non-zero range is greater than the second non-zero range. In some embodiments, the first non-zero range is represented as 16, and the second non-zero range is represented as 8.
[0452] In some embodiments, in the case where the quadratic transform is applied to only one sub-block, coefficients of the only sub-block outside a first non-zero range are set to zero. In the case where the quadratic transform is applied to multiple sub-blocks, coefficients of the multiple sub-blocks outside a second non-zero range are set to zero. In some embodiments, the first non-zero range is different from the second non-zero range. In some embodiments, the second non-zero range is represented as 8.
[0453] Figure 21It is a flowchart of another exemplary method for video processing according to the present technology. Method 2100 includes, at operation 2110, for the conversion between a block of a video and a bitstream representation of the video, determining that a dimensionality reduction quadratic transform is applied to a single sub-block of the block when the dimensions of the block satisfy a condition. The quadratic transform is performed between the forward main transform and the quantization step, or between the inverse quantization step and the inverse main transform. The reduced dimension is the dimension reduced from the dimension of the block. Method 2100 further includes, at operation 2120, performing the conversion based on the determination.
[0454] In some embodiments, the reduced dimension corresponds to the dimension of the sub-block. In some embodiments, the single sub-block to which the quadratic transform can be applied is the upper left sub-block of the current block. In some embodiments, the dimension of the single sub-block is M×N, where M and N are positive integers. In some embodiments, M = N = 4. In some embodiments, the condition specifies that the dimension of the block is 4×H or W×4, and where H > 8 and W > 8. In some embodiments, at least one of (1) H > T1 or (2) W > T2 is satisfied, where T1 and T2 are greater than 8. In some embodiments, T1 = T2 = 16. In some embodiments, at least one of (1) H < T1 or (2) W < T2 is satisfied, where T1 and T2 are greater than 8. In some embodiments, T1 = T2 = 32. In some embodiments, the condition specifies that the dimension of the block is M×H or W×N, and where H ≥ N and W ≥ M.
[0455] Figure 22 It is a flowchart of another exemplary method for video processing according to the present technology. Method 2200 includes, at operation 2210, for the conversion between a block of a video and a bitstream representation of the video, determining that a dimensionality reduction quadratic transform can be applied to a region in the block having dimensions K×L. K and L are positive integers and K is not equal to L. The quadratic transform is performed between the forward main transform and the quantization step, or between the inverse quantization step and the inverse main transform. The reduced dimension is the dimension reduced from the dimension of the block. Method 2200 further includes, at operation 2220, performing the conversion based on the determination.
[0456] In some embodiments, the reduced dimension corresponds to the dimension of the region. In some embodiments, the coefficients of the region outside the non-zero range are set to zero. In some embodiments, the non-zero range is represented as the upper left region in the block, and the size of the upper left region is M×M, where M is less than or equal to K and L.
[0457] Figure 23FIG. 2300 is a flow chart of another exemplary method of video processing in accordance with the present technology. Method 2300 includes, at operation 2310, determining a non-zero range based on characteristics of a block for conversion between the block of the video and a bitstream representation of the video. The non-zero range corresponds to a range outside of which coefficients associated with a dimensionality-reduced secondary transform are set to zero. The secondary transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimensionality is a dimensionality reduced from the dimensionality of the block. Method 2300 further includes, at operation 2320, performing the conversion based on the determination.
[0458] In some embodiments, characteristics of the block include color components of the block. In some embodiments, a first non-zero range for a luminance component of the block is different from a second non-zero range for a chrominance component of the block. In some embodiments, characteristics of the block include codec information of the block. In some embodiments, the codec information includes information indicating whether the block is coded in an intra mode or a non-intra mode. In some embodiments, the codec information includes information indicating whether the block is coded in an intra mode, an inter mode, or an intra block copy mode. In some embodiments, the codec information includes a reference picture of motion information. In some embodiments, characteristics of the block include a quantization parameter of the block. In some embodiments, the first non-zero range corresponds to a first quantization parameter, the second non-zero range corresponds to a second quantization parameter, and wherein the first non-zero range is less than or equal to the second non-zero range in the case where the first quantization parameter is greater than or equal to the second quantization parameter.
[0459] In some embodiments, different non-zero ranges are associated with different transform matrices for the secondary transform. In some embodiments, the non-zero range is signaled in a bitstream representation in a sequence parameter set, a picture parameter set, a picture header, a slice header, a slice group header, a codec tree unit (CTU) row, a CTU, or a coding unit. In some embodiments, multiple non-zero ranges may be applied to the secondary transform, and a value indicating selection of one of the multiple non-zero ranges is signaled in a bitstream representation in a sequence parameter set, a picture parameter set, a picture header, a slice header, a slice group header, a codec tree unit (CTU) row, a CTU, or a coding unit.
[0460] In some embodiments, performing the conversion includes generating a bitstream representation based on the block of the video. In some embodiments, performing the conversion includes generating the block of the video from the bitstream representation.
[0461] Figure 24AFIG. 2400 is a flowchart of an example method of video encoding according to the present technology. Method 2400 includes, at operation 2410, determining that a dimensionality-reduced quadratic transform is applicable to two adjacent sub-blocks of a video block. Each of the two adjacent sub-blocks has dimensions of M×N, where M and N are positive integers. The quadratic transform is performed between the forward primary transform and the quantization step. The reduced dimension is a dimension reduced from the dimension of the block. Method 2400 further includes, at operation 2420, generating an encoded / decoded representation of the video based on the determination.
[0462] In some embodiments, the reduced dimension corresponds to the dimensions of the two adjacent blocks. In some embodiments, the method includes arranging the coefficients of the two adjacent sub-blocks into a one-dimensional vector having 2×M×N elements. In some embodiments, the method includes obtaining M×N transform elements by applying the quadratic transform to the one-dimensional vector using a transform matrix. The transform matrix has a first dimension of 2×M×N elements and a second dimension of M×N elements. In some embodiments, the method includes rearranging the M×N transform elements into a first sub-block of the two adjacent sub-blocks. In some embodiments, the method includes setting the elements in a second sub-block of the two adjacent sub-blocks to zero. In some embodiments, both the forward primary transform and the quadratic transform are performed at the sub-block level.
[0463] Figure 24B FIG. 2450 is a flowchart of an example method of video decoding according to the present technology. Method 2450 includes, at operation 2460, determining that a dimensionality-reduced quadratic transform is applicable to two adjacent sub-blocks of a video block. Each of the two adjacent sub-blocks has dimensions of M×N, where M and N are positive integers. The quadratic transform is performed between the inverse quantization step and the inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. Method 2450 further includes, at operation 2470, generating blocks of the video by parsing the encoded / decoded representation of the video according to the determination.
[0464] In some embodiments, the reduced dimension corresponds to the dimensions of the two adjacent blocks. In some embodiments, the method includes arranging the coefficients of a first sub-block of the two adjacent sub-blocks into a one-dimensional vector having M×N elements. In some embodiments, the method includes obtaining 2×M×N transform elements by applying the quadratic transform to the one-dimensional vector using a transform matrix. The transform matrix has a first dimension of M×N elements and a second dimension of 2×M×N elements. In some embodiments, the method includes rearranging the 2×M×N transform elements into the two adjacent sub-blocks. In some embodiments, M = N = 4.
[0465] In some embodiments, the dimensionality-reduced quadratic transform includes a low-frequency non-separable transform, and the low frequency corresponds to the reduced dimension.
[0466] Figure 25A flowchart of another example method for video processing according to the present technology. Method 2500 includes, at operation 2510, determining whether to apply a dimensionality-reducing secondary transform to a block of a video according to a rule based on characteristics associated with the block for conversion between the block of the video and a bitstream representation of the video. The secondary transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is a dimension reduced from the dimension of the block. Method 2500 includes, at operation 2520, performing the conversion based on the determination.
[0467] In some embodiments, the characteristics associated with the block include encoding and decoding information of the block or encoding and decoding information of adjacent blocks. In some embodiments, the rule specifies that the secondary transform is not applied to the block if the encoding and decoding information indicates that the block or an adjacent block is encoded and decoded in one or more specific encoding and decoding modes. In some embodiments, one or more specific encoding and decoding modes include at least one of the following: linear mode (LM) mode, LM-T mode, LM-A mode, one or more wide-angle intra prediction modes, block differential pulse codec modulation (BDPCM) mode, differential pulse codec modulation (DPCM) mode, residual domain block differential pulse codec modulation (RBDPCM) mode, matrix-based intra prediction (MIP) mode, or one or more angular intra prediction modes. In some embodiments, the rule specifies that the secondary transform is not applied to the chrominance component of the block if the block is encoded and decoded in a joint chrominance residual encoding and decoding mode. Encoding and decoding the block in a joint chrominance residual encoding and decoding mode includes determining a joint residual, which is an average of the residuals associated with the chrominance components of the block. In some embodiments, the rule specifies that the secondary transform is applied to the luminance component of the block and not to the chrominance component of the block encoded in the LM mode, LM-T mode, or LM-A mode.
[0468] In some embodiments, the characteristics associated with the block include the coefficients or residuals of the block after a quantization or inverse quantization step. In some embodiments, the rule specifies that the secondary transform is applied to the residuals if a transform skip mode is used to encode and decode the block. The transform skip mode is a mode that skips a forward or inverse primary transform. In some embodiments, the rule specifies that the secondary transform is applied to the quantized transform coefficients of the block.
[0469] In some embodiments, the characteristics associated with the block include whether an intra encoding tool is used to encode and decode the block. In some embodiments, the rule specifies that the secondary transform may be applied to the block if an inter encoding tool is used to encode and decode the block. In some embodiments, the rule specifies that the secondary transform may be applied to the block if an intra block copy encoding tool is used to encode and decode the block. In some embodiments, the rule specifies that the secondary transform may be applied to the block if a combined inter-intra prediction encoding tool is used to encode and decode the block.
[0470] In some embodiments, the characteristics associated with a block include information associated with the chroma format of the block. In some embodiments, a rule specifies that the secondary transform is not applied to the chroma component of the block. In some embodiments, the rule specifies that the secondary transform is not applied to the chroma component of the block when the chroma format of the block is 4:0:0. In some embodiments, the rule specifies that the secondary transform is not applied to the chroma component of the block when the chroma components of the chroma format are separately encoded and decoded. In some embodiments, a rule specifies that the secondary transform is applied to the block. The non-zero range of the secondary transform associated with the dimensions of the block is determined based on the color components of the block, and the non-zero range is the range outside of which the coefficients of the block are set to zero. In some embodiments, for the same dimension of the block, the first non-zero range of the chroma component of the block is less than the second non-zero range of the luminance component of the block.
[0471] In some embodiments, it is determined whether the position-dependent intra prediction combination (PDPC) encoding / decoding step is applied to the block based on whether the secondary transform is applicable. In some embodiments, the PDPC encoding / decoding step is not applied when the secondary transform is applied to the block. In some embodiments, the PDPC encoding / decoding step is applied when the secondary transform is applied to the block.
[0472] In some embodiments, the characteristics associated with a block include whether the position-dependent intra prediction combination (PDPC) encoding / decoding step is applied to the block. In some embodiments, the rule specifies that the secondary transform is not applied to the block when the PDPC encoding / decoding step is applicable. In some embodiments, it is determined whether to filter the neighboring samples of the block for the intra prediction encoding / decoding step based on whether the secondary transform is applied to the block. In some embodiments, the neighboring samples are not filtered when the secondary transform is applied to the block. In some embodiments, the neighboring samples are filtered when the secondary transform is applied to the block.
[0473] In some embodiments, the characteristics associated with a block include whether to filter the neighboring samples of the block for the intra prediction encoding / decoding step applied to the block. In some embodiments, the rule specifies that the secondary transform is not applicable when the neighboring samples are filtered. In some embodiments, the rule specifies that the secondary transform is not applicable when the neighboring samples are not filtered.
[0474] In some embodiments, the characteristics associated with a block include whether the block is encoded or decoded in a transform skip mode, in which the forward or inverse main transform is skipped. In some embodiments, the block is encoded or decoded using the transform skip mode, and a secondary transform may be applied to the block. In some embodiments, a first transform matrix for the secondary transform when the transform skip mode is enabled is different from a second transform matrix for the secondary transform when the transform skip mode is disabled. In some embodiments, it is determined whether a quantization matrix is applied to the block based on whether the secondary transform is applied. In some embodiments, a first quantization matrix is applied when the secondary transform is applicable, and a second quantization matrix different from the first quantization matrix is applied when the secondary transform is not applicable.
[0475] In some embodiments, the characteristics associated with a block include whether a quantization matrix is applied to the block. In some embodiments, the rule stipulates that the secondary transform is not applicable when the quantization matrix is applied. In some embodiments, the characteristics associated with a block include whether a sub-block level transform is applied to the block. In some embodiments, the rule stipulates that the secondary transform may be applied to the coefficients of the top-left sub-block of the block generated by the sub-block level transform. In some embodiments, it is determined based on whether the secondary transform is applied to the block the scan order for parsing the residual block after the quantization or inverse quantization step. In some embodiments, the coefficients set to zero by the secondary transform are not scanned. In some embodiments, it is determined based on whether the secondary transform is applied to the block the arithmetic coding and decoding context for parsing the residual block after the quantization or inverse quantization step.
[0476] In some embodiments, information related to the secondary transform is signaled at one or more levels in the bitstream representation, the one or more levels including a picture parameter set, a slice, a header, a picture header, a slice group header, a slice, a coding tree unit row, or a coding tree unit. In some embodiments, whether the secondary transform is applicable is based on the one or more levels at which the information is signaled. In some embodiments, the information is signaled separately for the luminance component and the chrominance component encoded within the coding tree unit. In some embodiments, when the secondary transform is not applied to the block, one or more syntax elements related to the secondary transform are not included in the bitstream representation of the block. In some embodiments, one or more syntax elements related to the secondary transform are signaled before the quantized transform residuals in the bitstream representation. In some embodiments, the one or more syntax elements are signaled independently of the number of coefficients determined when parsing the quantized residuals. In some embodiments, the number of coefficients is not calculated when parsing the quantized residuals. In some embodiments, a syntax flag indicating that all sub-blocks are set to zero by the secondary transform is not included in the bitstream representation, where the implied value of the syntax flag is 0. In some embodiments, a syntax flag indicating that coefficients are set to zero by the secondary transform is not included in the bitstream representation, where the implied value of the syntax flag is 0.
[0477] Figure 26 It is a flowchart of another exemplary method for video processing according to the present technology. Method 2600 includes, at operation 2610, determining bit-precision constraints for coefficients of one or more transformation matrices of a reduced-dimensional quadratic transform applied to a block for conversion between a block of a video and a bitstream representation of the video. The quadratic transform is performed between a forward primary transform and a quantization step, or between an inverse quantization step and an inverse primary transform. The reduced dimension is the dimension reduced from the dimension of the block. Method 2600 further includes, at operation 2620, performing the conversion based on the determination.
[0478] In some embodiments, the bit-precision constraints include storing coefficients of one or more transformation matrices with a bit width less than 8. In some embodiments, the bit-precision constraints include storing coefficients of one or more transformation matrices based on a correlation between the one or more transformation matrices. In some embodiments, a difference between a first element and a second element in a transformation matrix is stored, wherein the first element is derived based on the second element. In some embodiments, a difference between a first element in a first transformation matrix and a second element in a second transformation matrix is stored, wherein the first element is derived based on the second element. In some embodiments, the difference is represented by a bit width less than 8. In some embodiments, the bit width is 6 or 4.
[0479] In some embodiments, the reduced-dimensional quadratic transform includes a low-frequency non-separable transform, and the low frequency corresponds to the reduced dimension.
[0480] Figure 27 It is a flowchart of another exemplary method for video processing according to the present technology. Method 2700 includes, at operation 2710, determining to use a reduced-dimensional quadratic transform (RST) tool for a codec unit of a video for conversion between the codec unit of the video and a bitstream representation of the video, based on rules associated with one or more transform units in the codec unit. The RST tool includes, during encoding, applying a forward quadratic transform between a forward primary transform and a quantization step, or during decoding, applying an inverse quadratic transform between an inverse quantization step and an inverse primary transform. Sizes of the forward quadratic transform and the inverse quadratic transform are smaller than a size of the codec unit. Method 2700 further includes, at operation 2720, performing the conversion based on the determination.
[0481] In some embodiments, the rules are related to a dimension of the codec unit and a maximum transform size. In some embodiments, the rules specify disabling the quadratic transform when a width or a height of the codec unit is greater than the maximum transform size.
[0482] In some embodiments, the rule is related to the dimension of the transform unit relative to the dimension of the codec unit. In some embodiments, the rule specifies that the second transform is disabled when the dimension of the codec unit is greater than the dimension of the transform unit.
[0483] In some embodiments, the rule is related to the number of one or more transform units in the codec unit. In some embodiments, the rule specifies that the second transform is disabled for the codec unit when the number of one or more transform units in the codec unit is greater than 1. In some embodiments, the rule specifies that when the number of one or more transform units in the codec unit is greater than 1, the second transform is applied to a single transform unit among the one or more transform units. In some embodiments, the second transform is applied to the first transform unit in the codec unit. In some embodiments, the second transform is applied to the last transform unit in the codec unit. In some embodiments, the rule specifies that when the number of one or more transform units in the codec unit is greater than 1, the second transform is applied to multiple transform units. In some embodiments, the second transform is independently applied to each of the one or more transform units. In some embodiments, the rule specifies that whether the second transform is applied to the second transform unit in the codec unit is independent of whether the second transform is applied to the first transform unit in the codec unit. In some embodiments, for one transform unit among the multiple transform units, the rule is related to the number of non-zero coefficients of the transform unit, regardless of the number of non-zero coefficients of any other transform unit in the codec unit. In some embodiments, the rule specifies that the second transform is disabled for the transform unit when the number of non-zero coefficients is less than a threshold. In some embodiments, the threshold is 2. In some embodiments, a syntax element indicating that the second transform is disabled for the transform unit is omitted in the bitstream representation.
[0484] In some embodiments, the rule is related to a single transform unit among one or more transform units in the codec unit. In some embodiments, when the number of non-zero coefficients in the single transform unit is less than a threshold, the rule stipulates disabling the second transform for the codec unit. In some embodiments, when the number of non-zero coefficients in a sub-region of the single transform unit is less than a threshold, the rule stipulates disabling the second transform for the codec unit. In some embodiments, the sub-region includes the upper left region of the single transform unit. In some embodiments, the dimension of the sub-region is 4×4. In some embodiments, the threshold is 2. In some embodiments, the single transform unit includes the first transform unit or the last transform unit among the one or more transform units.
[0485] In some embodiments, the rule is related to a syntax flag, and the syntax flag is associated with a transform unit among the one or more transform units. In some embodiments, when the syntax flag is omitted in the bitstream representation, the rule stipulates disabling the second transform for the transform unit. In some embodiments, when the syntax flag is omitted in the bitstream representation, the rule stipulates that the second transform is applied to the transform unit. In some embodiments, the rule stipulates that when the codec unit includes a single transform unit, the syntax flag is the same as another syntax flag associated with the codec unit, and the another syntax flag is determined based on information about the coefficients of the codec unit. In some embodiments, the rule stipulates that when the codec unit includes multiple transform units, the syntax flag associated with the last transform unit among the multiple transform units is the same as another syntax flag associated with the codec unit, and the another syntax flag is determined based on information about the coefficients of the codec unit. In some embodiments, the syntax flag associated with other transform units among the multiple transform units indicates disabling the second transform for the other transform units. In some embodiments, the syntax flag associated with other transform units among the multiple transform units indicates that the second transform can be applied to the other transform units.
[0486] Figure 28A flowchart of another example method for video processing according to the present technology. Method 2800 includes, at operation 2810, determining, according to rules, a way to apply a reduced-dimensionality second transform (RST) tool to one or more color components of a block for conversion between a video block including one or more color component blocks and a bitstream representation of the video. The RST tool includes, during encoding, applying a forward second transform between a forward primary transform and a quantization step, or during decoding, applying an inverse second transform between an inverse quantization step and an inverse primary transform. The sizes of the forward second transform and the inverse second transform are smaller than the size of the codec unit. Method 2800 further includes, at operation 2810, performing the conversion based on the determination.
[0487] In some embodiments, the rules specify that, in the case of splitting the block using a single codec tree, whether the second transform is applied to the first component of the block and whether the second transform is applied to the second component of the block are determined independently. In some embodiments, the first component includes a luminance component, and wherein the second component includes a chrominance component. In some embodiments, the rules specify that whether the second transform is applied to the first component of the block is determined based on the number of non-zero coefficients of the first component without considering the number of non-zero coefficients of the second component.
[0488] In some embodiments, the rules specify that, in the case of splitting the block using a single codec tree, whether the second transform is applied to the first component of the block is determined based on whether the second transform is applied to the second component of the block. In some embodiments, the rules specify that whether the second transform is applied to the first component of the block is determined based on the number of non-zero coefficients of the second component. In some embodiments, the first component includes a chrominance component, and wherein the second component includes a luminance component. In some embodiments, the first component includes a red component or a blue component, and wherein the second component includes a green component. In some embodiments, the rules specify that the second transform is disabled for a component of the block if the number of non-zero coefficients of the component or a sub-region of the component is less than a threshold. In some embodiments, the sub-region is the upper left region of the block. In some embodiments, the dimension of the sub-region is 4×4. In some embodiments, the threshold is 2.
[0489] In some embodiments, a syntax flag indicating disabling the secondary transformation for the components of the block is omitted in the bitstream representation. In some embodiments, whether the rule applies is based on the dimension W×H of the coding unit, transformation unit, or the block. In some embodiments, the rule applies when W>T, where T is a positive integer. In some embodiments, the rule applies when H>T, where T is a positive integer. In some embodiments, the rule applies when W>T and H>T, where T is a positive integer. In some embodiments, T is 64. In some embodiments, T is equal to the maximum transformation size of the coding unit, transformation unit, or the block.
[0490] In some embodiments, performing the transformation includes generating the bitstream representation based on the block of the video. In some embodiments, performing the transformation includes generating the block of the video from the bitstream representation.
[0491] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the transformation from the video block to the bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the transformation from the bitstream representation of the video to the video block will be performed using the video processing tool or mode enabled based on the decision or determination.
[0492] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the transformation from the video block to the bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.
[0493] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the transformation from the pixel representation of the video to the corresponding bitstream representation, and vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits juxtaposed or scattered at different positions within the bitstream. For example, a macroblock may be encoded according to the transformed and encoded error residual values and also using bits in the header and other fields in the bitstream.
[0494] The disclosures, other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing unit” or “data processing apparatus” includes all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computer clusters. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to an appropriate receiver device.
[0495] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including a compiled or interpreted language, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0496] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0497] For example, a processor suitable for executing a computer program includes general and special-purpose microprocessors, and any one or more of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0498] Although this patent document contains many details, it should not be construed as limiting any subject matter or the scope of any claims, but rather as a description of the characteristics of particular embodiments of a particular technology. Certain features described in the context of separate embodiments of this patent document may also be implemented in combination in a single embodiment. Conversely, the various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0499] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to mean that such operations must be performed in the particular order shown or in sequential order to achieve a desired result, or that all illustrated operations must be performed. Additionally, the separation of various system components described in the embodiments of this patent document should not be understood to be required in all embodiments.
[0500] Only some implementations and examples are described herein, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Determining whether to use a reduced-dimensional secondary transform RST tool for a codec unit of a video based on rules associated with one or more transform units in the codec unit for conversion between a codec unit of the video and a bitstream of the video; wherein the RST tool includes, during encoding, applying a forward secondary transform between a forward primary transform and a quantization step, or during decoding, applying an inverse secondary transform between an inverse quantization step and an inverse primary transform, and wherein dimensions of the forward secondary transform and the inverse secondary transform are smaller than dimensions of the codec unit; and Performing the conversion based on the determination; wherein the rules are related to the number of one or more transform units in the codec unit; wherein the rules specify that, when the number of one or more transform units in the codec unit is greater than 1, the secondary transform is applied to a first transform unit or a last transform unit among the one or more transform units; 2. The method according to claim 1, wherein The rules are related to the dimension of the codec unit and a maximum transform size; 3. The method according to claim 2, wherein, The rules specify that, when the width or height of the codec unit is greater than the maximum transform size, the secondary transform is disabled; 4. The method according to claim 1, wherein The rules are related to the dimension of a transform unit relative to the dimension of the codec unit; 5. The method according to claim 4, wherein, The rules specify that, when the dimension of the codec unit is greater than the dimension of the transform unit, the secondary transform is disabled; 6. The method according to claim 1, wherein, The rules specify that, when the number of one or more transform units in the codec unit is greater than 1, the secondary transform is applied to multiple transform units; 7. The method according to claim 6, wherein, The secondary transform is independently applied to each of the one or more transform units; 8. The method according to claim 6, wherein, The rules specify that whether the secondary transform is applied to a second transform unit in the codec unit is independent of whether the secondary transform is applied to a first transform unit in the codec unit; 9. The method according to claim 6, wherein, For one transform unit among the multiple transform units, the rules are related to the number of non-zero coefficients of the transform unit, regardless of the number of non-zero coefficients of any other transform unit in the codec unit; 10. The method according to claim 9, wherein, When the number of non-zero coefficients is less than a threshold, the rules specify that the secondary transform is disabled for the transform unit; 11. The method according to claim 10, wherein, The threshold is 2; 12. The method according to claim 10 or 11, wherein a syntax element indicating that the secondary transform is disabled for the transform unit is omitted in the bitstream; 13. The method according to claim 1, wherein The rules are related to a single transform unit among one or more transform units in the codec unit; 14. The method according to claim 13, wherein, The rules specify that, when the number of non-zero coefficients of the single transform unit is less than a threshold, the secondary transform is disabled for the codec unit; 15. The method according to claim 13, wherein, The rules specify that, when the number of non-zero coefficients in a sub-region of the single transform unit is less than a threshold, the secondary transform is disabled for the codec unit; 16. The method according to claim 15, wherein The sub-region includes an upper left region of the single transform unit; 17. The method according to claim 15 or 16, wherein The size of the sub-region is 4×4; 18. The method according to any one of claims 14 to 16, wherein, The threshold is 2; 19. The method according to any one of claims 13 to 16, wherein The single transform unit includes a first transform unit or a last transform unit among the one or more transform units.
20. The method according to any one of claims 1 to 11, wherein The rule is related to a syntax flag, and the syntax flag is associated with one transform unit among the one or more transform units.
21. The method according to claim 20, wherein, The rule stipulates that when the syntax flag is omitted in the bitstream, the secondary transform is disabled for the transform unit.
22. The method according to claim 20, wherein The rule stipulates that when the syntax flag is omitted in the bitstream, the secondary transform is applied to the transform unit.
23. The method according to claim 20, wherein, The rule stipulates that when the codec unit includes a single transform unit, the syntax flag is the same as another syntax flag associated with the codec unit, and the other syntax flag is determined based on information about the coefficients of the codec unit.
24. The method according to claim 20, wherein, The rule stipulates that when the codec unit includes multiple transform units, the syntax flag associated with the last transform unit among the multiple transform units is the same as another syntax flag associated with the codec unit, and the other syntax flag is determined based on information related to the coefficients of the codec unit.
25. The method according to claim 24, wherein, The syntax flag associated with other transform units among the multiple transform units indicates that the secondary transform is disabled for the other transform units.
26. The method according to claim 24, wherein, The syntax flag associated with other transform units among the multiple transform units indicates that the secondary transform is applied to the other transform units.
27. The method according to claim 25 or 26, wherein The syntax flag indicating that the secondary transform is disabled for a component of the codec unit is omitted in the bitstream.
28. The method according to any one of claims 1 to 11, wherein Whether to apply the rule is based on the size W×H of the codec unit or transform unit.
29. The method according to any one of claims 1 to 11, wherein, Performing the conversion includes generating the bitstream based on the codec unit of the video.
30. The method according to any one of claims 1 to 11, wherein Performing the conversion includes generating the codec unit of the video from the bitstream.
31. A video processing method, comprising: For the conversion between blocks of a video including one or more color component blocks and the bitstream of the video, according to a rule, determining a manner of applying a dimensionality reduction secondary transform RST tool to one or more color components of the block, wherein the RST tool includes, during encoding, applying a forward secondary transform between a forward main transform and a quantization step, or during decoding, applying a reverse secondary transform between an inverse quantization step and a reverse main transform, and wherein the sizes of the forward secondary transform and the reverse secondary transform are smaller than the size of the block; and Performing the conversion based on the determination, wherein the rule stipulates that when the block is segmented using a single codec tree, whether to apply the secondary transform to the first component of the block is determined based on whether the secondary transform is applied to the second component of the block.
32. The method according to claim 31, wherein, The rule stipulates that whether to apply the secondary transform to the first component of the block is determined based on the number of non-zero coefficients of the second component.
33. The method according to claim 31 or 32, wherein, The first component includes a chrominance component, and wherein the second component includes a luminance component.
34. The method according to claim 31 or 32, wherein The first component includes a red component or a blue component, and wherein the second component includes a green component.
35. The method according to claim 31, wherein, The rule stipulates that when the number of non-zero coefficients of the component or a sub-region of the component is less than a threshold, the secondary transform is disabled for the component of the block.
36. The method according to claim 35, wherein, The sub-region is the upper left region of the block.
37. The method according to claim 35, wherein, The size of the sub-region is 4×4.
38. The method according to any one of claims 35 to 37, wherein The threshold is 2.
39. The method according to any one of claims 31 to 32, wherein, Omit in the bitstream a syntax flag indicating disabling of the secondary transformation for components of the block.
40. The method according to any one of claims 31 to 32, wherein Whether to apply the rule is based on the size W×H of the transform unit or the block.
41. The method according to claim 40, wherein, Apply the rule when W>T, where T is a positive integer.
42. The method according to claim 40, wherein, Apply the rule when H>T, where T is a positive integer.
43. The method according to claim 40, wherein, Apply the rule when W>T and H>T, where T is a positive integer.
44. The method according to any one of claims 41 to 43, wherein, T is 64.
45. The method according to any one of claims 41 to 43, wherein, T is equal to the maximum transform size of the transform unit or the block.
46. The method according to any one of claims 31 to 32, wherein Performing the conversion includes generating the bitstream based on the block of the video.
47. The method according to any one of claims 31 to 32, wherein, Performing the conversion includes generating the block of the video from the bitstream.
48. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, The instruction, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 47.
49. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for causing a processor to execute the method according to any one of claims 1 to 47.
Citation Information
Patent Citations
Method and apparatus of enhanced multiple transforms and non-separable secondary transform for video coding
WO2018166429A1