Video processing method, apparatus and medium

CN115668923BActive Publication Date: 2026-09-08DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180038623.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-27
Filing Date
2021-05-24
Publication Date
2026-09-08
Estimated Expiration
2041-05-24

AI Technical Summary

Technical Problem

[0004]尽管在视频压缩方面取得了进步,但数字视频仍占互联网和其他数字通信网络的最大带宽使用

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668923B_ABST
    Figure CN115668923B_ABST
Patent Text Reader

Abstract

Devices, systems, and methods, including methods for transform design, for digital video coding are described. In one representative aspect, a video processing method includes performing a conversion between a current video block of a video and a bitstream of the video based on a rule that specifies that a selection of a set of transform matrices for performing a transform operation during the conversion is based on a low-frequency non-separable transform index indicated in the bitstream, that the rule specifies that the transform operation includes coding the current video block into the bitstream by applying a forward transform to residual values of the current video block during an encoding operation, or that the rule specifies that the transform operation includes generating the current video block from the bitstream by applying an inverse transform to scaling coefficients indicated in the bitstream during a decoding operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference of related applications

[0002] This application is based on International Patent Application No. PCT / CN2021 / 095390, filed on May 24, 2021, which claims priority and interest in International Patent Application No. PCT / CN2020 / 092592, filed on May 27, 2020. All of the aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology

[0004] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This paper describes devices, systems, and methods related to digital video coding and decoding, particularly multi-transformation methods for video coding and decoding. The described methods can be applied to both existing video coding and decoding standards (e.g., High Efficiency Video Codec (HEVC)) and future video coding and decoding standards (e.g., Multi-Functional Video Codec (VVC)) or codecs.

[0006] In one representative aspect, the disclosed techniques can be used to provide example methods for video processing. These methods include performing a rule-based transformation between a current video block and a bitstream of video, wherein the rule specifies that the set of transform matrices selected during the transformation to perform the transform operation is based on low-frequency inseparable transform indices indicated in the bitstream, and the rule specifies that the transform operation includes, during encoding, encoding / decoding the current video block into the bitstream by applying a forward transform to the residual values ​​of the current video block, or the rule specifies that the transform operation includes, during decoding, generating the current video block from the bitstream by applying an inverse transform to scaling factors indicated in the bitstream.

[0007] In another representative aspect, the disclosed technique can be used to provide an example method for video processing. This method includes: determining a zeroing range for the current video block using rules for a conversion between a current video block and a bitstream of the video; and performing the conversion based on the determination, wherein the rules specify the zeroing range based on the size of the current video block, and during a secondary transformation operation in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0008] In another representative aspect, the disclosed technique can be used to provide an example method for video processing. This method includes: a conversion between a current video block and a bitstream of the video; determining, according to rules, whether to disable a zeroing operation for the current video block; and performing the conversion based on the determination, wherein the rules specify that the zeroing operation includes using a master transform, in which coefficients within a range are considered to have zero values.

[0009] In another representative aspect, the disclosed technique can be used to provide an example method for video processing. This method includes: determining, for a conversion between a current video block and a bitstream of video, whether a zeroing range has been expanded in the main transform of the current video block; and performing the conversion based on the determination, wherein, during the operation of the main transform in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0010] In another representative aspect, the disclosed techniques can be used to provide example methods for video processing. These methods include: determining one or more master transform matrices from a master transform set according to rules for a transformation between a current video block and a bitstream of the video; and performing the transformation according to the determination, wherein the one or more master transform matrices include one or more of Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VII, transform skip mode, identity transform, and transforms based on a training process.

[0011] In another representative aspect, the disclosed techniques can be used to provide example methods for video processing. These methods include: converting one or more video blocks of a video to a bitstream of the video; determining a transform set for a primary transform or a secondary transform tool according to rules, wherein the rules specify that the transform set is determined from two or more transform sets having specific characteristics; and performing the transformation based on the determination.

[0012] In another representative aspect, the disclosed technique can be used to provide an example method for video processing. This method includes: performing a conversion between a current video block and a bitstream of the video; determining a zeroing range for the main transform of the current video block based on a prediction mode of the current video block; and performing the conversion based on the determination, wherein, during the operation of the main transform in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0013] In another representative aspect, the disclosed technique can be used to provide an example method for video processing. This method includes: for a conversion between a current video block and a bitstream of video, determining a zeroing range for a secondary transform of the current video block based on the number of samples in the current video block; and performing the transformation based on the determination, wherein, during the operation of the secondary transform in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0014] In another representative aspect, the disclosed techniques can be used to provide example methods for video processing. These methods include performing a conversion between a current block of video and a bitstream representation of the video based on a selection of a master transform included in a transform set, wherein the conversion includes using a master transform, and the master transform includes one or more of Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VII, transform skip mode, identity transform, or a transform based on a training process.

[0015] In another representative aspect, the above methods are embodied in processor-executable code and stored in a computer-readable program medium.

[0016] In another representative aspect, a device configured or operable to perform the methods described above is disclosed. This device may include a processor programmed to implement the methods.

[0017] In another representative aspect, video decoder devices can implement the methods described herein.

[0018] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings and description. Attached Figure Description

[0019] Figure 1 An example of an encoder block diagram is shown;

[0020] Figure 2 Examples of 67 intra-frame prediction modes are shown;

[0021] Figure 3A and 3B An example of a reference sample for a wide-angle intra-prediction mode with a non-square block is shown;

[0022] Figure 4 An example of discontinuities is shown when using wide-angle intra-frame prediction;

[0023] Figures 5A-5D An example of samples used by the Location-Related Intra-Prediction Combination (PDPC) method is shown;

[0024] Figure 6 An example of four reference lines adjacent to the prediction block is shown;

[0025] Figure 7A Examples of 4×8 and 8×4 block partitioning are shown;

[0026] Figure 7B Examples of block partitioning are shown, except for 4×8, 8×4, and 4×4.

[0027] Figure 8 An example is shown where a 4×8 sample block is divided into two independent decodeable regions;

[0028] Figure 9 An example of the processing order for pixel rows that maximizes throughput of a 4×N block using a vertical predictor is shown;

[0029] Figure 10 An example of a quadratic transformation in JEM is shown;

[0030] Figure 11 An example of the proposed simplified quadratic transformation (RST) is shown;

[0031] Figure 12 Examples of forward and inverse simplified transformations are shown;

[0032] Figure 13 An example of a positive RST 8×8 process with a 16×48 matrix is ​​shown;

[0033] Figure 14 An example of scanning positions 17 to 64 in an 8×8 block for non-zero elements is shown;

[0034] Figure 15 Examples of subblock transformation modes SBT-V and SBT-H are shown.

[0035] Figure 16 A flowchart illustrating an example of a method with multiple variations based on the disclosed technique is shown;

[0036] Figure 17 This is a block diagram of an example of a video processing device;

[0037] Figure 18 A block diagram of an example video codec system is shown;

[0038] Figure 19 A block diagram of an example encoder is shown;

[0039] Figure 20 A block diagram of an example decoder is shown;

[0040] Figure 21 This is a block diagram of an example video processing system that can implement the publicly available technology; and

[0041] Figures 22 to 29 This is a flowchart illustrating an example of a video processing method. Detailed Implementation

[0042] 1 Introduction

[0043] Due to the ever-increasing demand for higher resolution video, video encoding and decoding methods and technologies are ubiquitous in modern technology. Video codecs typically consist of electronic circuitry or software that compresses or decompresses digital video and are constantly being improved to provide higher encoding and decoding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end latency (delay). Compression formats typically conform to standard video compression specifications, such as the High Efficiency Video Codec (HEVC) standard (also known as H.265 or MPEG-H Part 2), the yet-to-be-finalized Multi-Functional Video Codec (VVC) standard, or other current and / or future video codec standards.

[0044] Implementations of the disclosed techniques can be applied to existing video codec standards (e.g., HEVC, H.265) and future standards to improve runtime performance. Section headings are used in this document to improve readability and do not in any way limit the discussion or embodiments (and / or implementations) to the respective sections.

[0045] 2. Implementation examples and examples of the multi-transformation method

[0046] 2.1 Color Space and Chromaticity Subsampling

[0047] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes the range of colors as tuples of numbers, typically 3 or 4 values ​​or color components (e.g., RGB). Essentially, a color space is a refinement of a coordinate system and its subspaces.

[0048] For video compression, the most commonly used color spaces are YCbCr and RGB.

[0049] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also known as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with primary color) is different from Y (luminance), meaning that light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.

[0050] Chromatic subsampling is a practice of encoding images by achieving a lower resolution for chromatic information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences.

[0051] 4:4:4 format.Each of the three Y'CbCr components has the same sampling rate, therefore there is no chromaticity subsampling. This scheme is sometimes used in high-end cinema scanners and film post-production.

[0052] 4:2:2 format. The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved. This reduces the bandwidth of the uncompressed video signal by one-third, with almost no visible difference.

[0053] 4:2:0 format. In the 4:2:0 scheme, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line. Therefore, the data rate is the same. Cb and Cr are subsampled in both the horizontal and vertical directions with a factor of 2. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.

[0054] In MPEG-2, Cb and Cr are located at the same position in the horizontal direction. Cb and Cr are located between pixels in the vertical direction (in the gaps).

[0055] In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in the gaps, in the middle of the alternating luminance samples.

[0056] In a 4:2:0 DV, Cb and Cr are located in the same position in the horizontal direction. In the vertical direction, they are located in the same position on alternating lines.

[0057] 2.2 Encoding and decoding process of typical video codecs

[0058] Figure 1 An example of a VVC codec block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image to reduce the mean square error between the original and reconstructed samples, respectively, by adding an offset and by applying a Finite Impulse Response (FIR) filter. Codec-side information signaling informs the offset and filter coefficients. ALF is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts created in the previous stages.

[0059] 2.3 Intra-mode encoding and decoding with 67 intra-prediction modes

[0060] To capture arbitrary edge directions presented in natural video, the number of intra-frame directional modes has been expanded from the 33 used by HEVC to 65. Additional directional modes are... Figure 2The arrows are depicted as red dots, and the planar and DC modes remain the same. These dense directional intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction.

[0061] like Figure 1 As shown, the conventional angular intra-prediction direction is defined as ranging from 45 degrees to -135 degrees clockwise. In VTM2, for non-square blocks, several conventional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are signaled using the original method and remapped to the wide-angle mode index after parsing. The total number of intra-prediction modes remains unchanged (e.g., 67), and the intra-mode encoding / decoding remains unchanged.

[0062] In HEVC, each intra-codec block has a square shape, with each side's length being a power of 2. Therefore, generating the intra-predictor using DC mode does not require division. In VTV2, blocks can have rectangular shapes, and in general, division must be performed for each block. To avoid division in DC prediction, only the longer sides are used to calculate the average of non-square blocks.

[0063] 2.4 Wide-angle intra-frame prediction for non-square blocks

[0064] In some embodiments, the conventional angular intra-prediction direction is defined as ranging from 45 degrees to -135 degrees clockwise. In VTM2, for non-square blocks, several conventional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are signaled using the original method and remapped to the wide-angle mode index after resolution. The total number of intra-prediction modes for a given block remains unchanged (e.g., 67), and the intra-mode encoding / decoding remains unchanged.

[0065] To support these predicted directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined as follows: Figure 3A and 3B The example is shown in the image.

[0066] In some embodiments, the number of modes replaced in the wide-angle orientation mode depends on the aspect ratio of the block. Table 1 shows the replaced intra-prediction modes.

[0067] Table 1: Intra-prediction modes replaced by wide-angle mode

[0068]

[0069] like Figure 4As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the increased gap Δp. α The negative impact.

[0070] 2.5 Example of Location-Related Intra-Prediction Combination (PDPC)

[0071] In VTM2, the intra-prediction results for planar modes are further modified using the Position-Related Intra-Prediction Combination (PDPC) method. PDPC is an intra-prediction method that combines unfiltered boundary reference samples with HEVC-type intra-prediction samples with filtered boundary reference samples. PDPC is applicable to the following intra-modal modes that do not require signaling notification: planar, DC, horizontal, vertical, lower-left corner mode and its eight adjacent corner modes, and upper-right corner mode and its eight adjacent corner modes.

[0072] The predicted sample point pred(x,y) is predicted using a linear combination of the intra-frame prediction mode (DC, plane, angle) and the reference sample point according to the following equation:

[0073] pred(x,y)=(wL×R -1,y + wT×R x,-1 – wTL ×R -1,-1 +(64 – wL – wT+wTL)×pred(x,y) + 32 ) >> shift

[0074] In this article, R x,-1 R -1,y R represents the reference sample points located at the top and left of the current sample point (x, y), respectively. -1,-1 This represents the reference sample point located at the top left corner of the current block.

[0075] In some embodiments, and if PDPC is applied to DC, planar, horizontal, and vertical intra-frame modes, no additional boundary filters are required, as in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters.

[0076] Figures 5A-5D Reference samples (R) of PDPC applied to various prediction modes are shown. x,-1 R -1,y and R -1,-1 The definition of ). The predicted sample point pred(x',y') is located at (x',y') within the prediction block. The reference sample point R. x,-1 The coordinates x are given by the following formula: x = x' + y' + 1, and the reference point R is... -1,yThe coordinates y are similarly given by the following formula: y = x' + y' + 1.

[0077] In some embodiments, the PDPC weights depend on the prediction mode and are shown in Table 2, where S = shift.

[0078] Table 2: Examples of PDPC weights based on prediction patterns

[0079]

[0080] 2.6 Multiple Reference Lines (MRL)

[0081] Multi-reference line (MRL) intra-prediction uses more reference lines for intra-prediction. Figure 6 The example shown illustrates four reference lines, where the samples for segments A and F are not taken from reconstructed neighboring samples, but are instead filled with the nearest samples from segments B and E, respectively. HEVC intra-prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used.

[0082] The index (mrl_idx) of the selected reference line is signaled and used to generate the intra-predictor. For reference line indices greater than 0, only the additional reference line modes are included in the MPM list, and only the MPM index is signaled, with no remaining modes. The reference line index is signaled before the intra-predictor modes, and in the case of signaling a non-zero reference line index, the planar mode and DC mode are excluded from the intra-predictor modes.

[0083] 2.7 Intra-frame sub-block segmentation (ISP)

[0084] In JVET-M0102, an ISP was proposed, which divides the intra-prediction block of luminance into two or four sub-segments vertically or horizontally based on the block size dimension, as shown in Table 3. Figure 7A and Figure 7B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples.

[0085] Table 3: Number of sub-segments depending on block size

[0086]

[0087] For each of these sub-segments, a residual signal is generated by entropy decoding of the coefficients sent by the encoder, followed by inverse quantization and inverse transform. The sub-segment is then intra-predicted, and the corresponding reconstructed samples are finally obtained by adding the residual signal to the predicted signal. Therefore, the reconstructed values ​​of each sub-segment can be used to generate the prediction for the next sub-segment, and the process is repeated for the next sub-segment, and so on. All sub-segments share the same intra-frame mode.

[0088] Based on the intra-frame mode and segmentation used, two different types of processing orders (referred to as normal and reverse orders) are employed. In the normal order, the first sub-segment to be processed is the one containing the top-left sample of the CU, and then it continues downwards (horizontal segmentation) or to the right (vertical segmentation). As a result, the reference samples used to generate the sub-segment prediction signal are located only to the left and top of the line. On the other hand, the reverse processing order either starts with the sub-segment containing the bottom-left sample of the CU and continues upwards, or starts with the sub-segment containing the top-right sample of the CU and continues to the left.

[0089] 2.8 Block Differential Pulse Codec Modulation Codec (BDPCM)

[0090] BDPCM was proposed in JVET-M0057. Since the shape of the horizontal (or vertical) predictor uses the left (A) (or top (B)) pixels for the prediction of the current pixel, the most throughput-efficient way to process a block is to process all pixels in a column (or row) in parallel and process these columns (or rows) sequentially. To increase throughput, we introduce the following processing: when the predictor selected on the block is vertical, a block with a width of 4 is divided into two halves with horizontal boundaries, and when the predictor selected on the block is horizontal, a block with a height of 4 is divided into two halves with vertical boundaries.

[0091] When dividing a block, samples from one region are not allowed to use pixels from another region to compute predictions: if this happens, the predicted pixel will be replaced by a reference pixel in the prediction direction. This is in... Figure 8 The image shows the different positions of the current pixel X in a vertically predicted 4×8 block.

[0092] Because of this property, it is now possible to process 4×4 blocks in 2 cycles, and 4×8 or 8×4 blocks in 4 cycles, and so on. Figure 9 As shown.

[0093] Table 4 summarizes the number of cycles required to process a data block, depending on the block size. It shows that it is common for any block with two dimensions greater than or equal to 8 to be processed at 8 pixels or more per cycle.

[0094] Table 4: Throughput of blocks with sizes of 4×N and N×4

[0095]

[0096] 2.9 Quantization Residual Domain BDPCM

[0097] In JVET-N0413, Quantization Residual Domain BDPCM (hereinafter referred to as RBDPCM) was proposed. Similar to intra-frame prediction, intra-frame prediction is performed on the entire block by copying samples in the prediction direction (horizontal or vertical prediction). The residual is quantized, and the deviation between the quantization residual and its predictor (horizontal or vertical) quantization value is encoded and decoded.

[0098] For a block of size M (rows) × N (columns), assume It is the prediction residual after performing intra-frame prediction using unfiltered samples from the top or left block boundaries, either horizontally (copying left neighboring pixel values ​​row by row across the prediction block) or vertically (copying the top neighboring row to each row in the prediction block). Assuming... Residual The quantized version is then used, where the residual is the difference between the original block and the predicted block values. The block DPCM is then applied to the quantized residual samples to obtain a modified version with element-wise... M×N array When the vertical BDPCM is signaled:

[0099]

[0100] For horizontal prediction, similar rules are applied, and residual quantization samples are obtained in the following manner:

[0101]

[0102] Residual Quantization Samples It is sent to the decoder.

[0103] On the decoder side, the above calculation is reversed to produce For the vertical prediction case,

[0104]

[0105] Regarding the horizontal situation

[0106]

[0107] Inverse quantization residual It is added to the intra-block prediction value to produce reconstructed sample values.

[0108] One advantage of this approach is that the DPCM inverse operation can be performed dynamically during coefficient resolution, simply by adding a predictor during coefficient resolution, or the DPCM inverse operation can be performed after resolution.

[0109] Transform skipping is always used in the quantization residual domain BDPCM.

[0110] 2.10 Multiple Transform Set (MTS) in VVC

[0111] VTM4 allows for large block sizes up to 64×64, primarily suitable for higher resolution video such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed out, leaving only low-frequency coefficients. For example, for an M×N transform block, with M as the block width and N as the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of transform coefficients are retained. When using transform skip mode on a large block, the entire block is used without zeroing out any values.

[0112] In addition to DCT-II already used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of both inter-frame and intra-frame codec blocks. It uses multi-select transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 4 below shows the basis functions of the selected DST / DCT.

[0113] Table 4: Basis functions of the transformation matrix used in VVC

[0114]

[0115] To maintain the orthogonality of the transformation matrices, the transformation matrices are quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformation coefficients within a 16-bit range, all coefficients are 10 bits after both the horizontal and vertical transformations.

[0116] To control the MTS scheme, separate enable flags are specified for intra-frame and inter-frame use at the SPS level. When MTS is enabled in SPS, signaling is sent to the CU level flag to indicate whether MTS has been applied. Here, MTS applies only to luminance. Signaling is sent to the MTS CU level flag when the following conditions are met.

[0117] ○ Both width and height are less than or equal to 32

[0118] ○ The CBF mark equals one

[0119] If the MTS CU flag is zero, DCT2 is applied bidirectionally. However, if the MTS CU flag is 1, two additional signaling flags are used to indicate the transform type in the horizontal and vertical directions, respectively. The transform and signaling notification mapping table is shown in Table 5. For transform matrix precision, an 8-bit master transform core is used. Therefore, all transform cores used by HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Furthermore, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, use an 8-bit master transform core.

[0120]

[0121] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed out for DST-7 and DCT-8 blocks with a size (width or height, or both) equal to 32. Only the coefficients in the 16×16 low-frequency region are retained.

[0122] As in HEVC, transform skip mode can be used to encode and decode block residuals. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU level MTS_CU_flag is not equal to 0. The block size limit for transform skip is the same as the MTS in JEM4, which indicates that transform skip applies to the CU when both the block width and height are equal to or less than 32.

[0123] 2.11 The Simplified Quadratic Transform (RST) proposed in JVET-N0193

[0124] 2.11.1 Indivisible Quadratic Transformation (NSST) of JEM

[0125] In JEM, a quadratic transform is applied between the forward master transform and quantization (on the encoder side) and between the inverse quantization and the inverse master transform (on the decoder side). For example... Figure 10 As shown, a 4×4 (or 8×8) quadratic transformation is performed based on the block size. For example, for each 8×8 block, the 4×4 quadratic transformation is applied to the smaller block (i.e., min(width, height) < 8), and the 8×8 quadratic transformation is applied to the larger block (i.e., min(width, height) > 4).

[0126] The following describes the application of inseparable transformations using an input as an example. To apply inseparable transformations, a 4×4 input block X...

[0127]

[0128] First, it is represented as a vector.

[0129]

[0130] Inseparable transformations are calculated as ,in The vector indicates the transformation coefficients, and T is a 16×16 transformation matrix. (16×1 coefficient vector) The scan order (horizontal, vertical, or diagonal) of this block is then reorganized into 4×4 blocks. Coefficients with smaller indices are placed in the 4×4 coefficient blocks along with their smaller scan indices. There are a total of 35 transform sets, and each transform set uses 3 inseparable transform matrices (kernels). The mapping from intra-frame prediction modes to transform sets is predefined. For each transform set, the selected inseparable quadratic transform (NSST) candidate is further specified by a quadratic transform index explicitly signaled. This index is signaled once per intra-frame CU in the bitstream after the transform coefficients.

[0131] 2.11.2 Simplified Quadratic Transform (RST) in JVET-N0193

[0132] JVET-K0099 introduced the RST (also known as the Low-Frequency Inseparable Transform (LFNST)), and JVET-L0133 introduced four transform sets (instead of 35) mappings. In JVET-N0193, 16×64 (further simplified to 16×48) and 16×16 matrices were used. For ease of notation, the 16×64 (simplified to 16×48) transform is represented as RST8×8, and the 16×16 transform is represented as RST4×4. Figure 11 An example of RST is shown.

[0133] 2.11.3 RST Calculation

[0134] The main idea of ​​the Reduction Transformation (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N (R < N) is the reduction factor.

[0135] The RT matrix is ​​an R×N matrix, as shown below:

[0136]

[0137] The R rows of the transformation form R bases in N-dimensional space. The inverse transformation matrix of RT is the transpose of its forward transformation. The forward and inverse RT are as follows: Figure 12 As shown.

[0138] In this proposal, a simplification factor of 4 (1 / 4 size) for RST8×8 is applied. Therefore, a 16×64 direct matrix is ​​used instead of 64×64, which is the traditional size of an 8×8 inseparable transform matrix. In other words, a 64×16 inverse RST matrix is ​​used on the decoder side to generate the core (master) transform coefficients in the top-left region of the 8×8. Forward RST8×8 uses a 16×64 (or 8×64 for an 8×8 block) matrix, so it only produces non-zero coefficients in the top-left 4×4 region of a given 8×8 region. In other words, if RST is applied, then the 8×8 region outside the top-left 4×4 region will only have zero coefficients. For RST4×4, a 16×16 (or 8×16 for a 4×4 block) direct matrix multiplication is applied.

[0139] The inverse RST is conditionally applied when the following two conditions are met:

[0140] ○ Block size is greater than or equal to a given threshold (W>=4 && H>=4)

[0141] ○ The switch skip mode flag is equal to zero.

[0142] If both the width (W) and height (H) of the transform coefficient block are greater than 4, then RST8×8 is applied to the top-left 8×8 region of the transform coefficient block. Otherwise, RST4×4 is applied to the top-left min(8,W)×min(8,H) region of the transform coefficient block.

[0143] If the RST index is equal to 0, then RST is not applied. Otherwise, RST is applied, and its kernel is selected along with the RST index. The RST selection method and encoding / decoding for the RST index will be explained later.

[0144] Furthermore, RST is applied to intra-CUs in both intra-frame and inter-frame stripes, as well as to both luma and chroma. If dual-tree is enabled, the luma and chroma RST indices are signaled separately. For inter-frame stripes (where dual-tree is disabled), a single RST index is signaled and used for both luma and chroma.

[0145] At the 13th JVET conference, Intra-Segmentation (ISP) was adopted as a new intra-prediction mode. When ISP mode is selected, Restricted Sub-Segmentation (RST) is disabled, and no signaling is used to notify the RST index, because the performance improvement is negligible even if RST is applied to every feasible segment block. Furthermore, disabling RST on the residuals of ISP predictions reduces coding complexity.

[0146] 2.11.4 RST Selection

[0147] The RST matrix is ​​selected from four transform sets, each consisting of two transforms. Which transform set to apply is determined based on the intra-frame prediction mode as follows:

[0148] (1) If one of the three CCLM modes is indicated, then transform set 0 is selected.

[0149] (2) Otherwise, select the transformation set according to the following table:

[0150] Transform set selection table

[0151]

[0152] The range of the index in the table above (represented as IntraPredMode) is [-14, 83], which is the transform mode index used for wide-angle intra-prediction.

[0153] 2.11.5 Simplified Dimensional RST Matrix

[0154] As a further simplification, a 16×48 matrix is ​​used instead of a 16×64 matrix with the same transformation set configuration. Each 16×48 matrix obtains 48 input data points (e.g., ...) from three 4×4 blocks (excluding the bottom right 4×4 block) in the top left 8×8 block. Figure 13 (As shown).

[0155] 2.11.6 RST Signaling Notification

[0156] A positive RST 8×8 with R=16 uses a 16×64 matrix, therefore it produces non-zero coefficients only in the top-left 4×4 region of a given 8×8 region. In other words, if RST is applied, then the 8×8 region other than the top-left 4×4 region produces only zero coefficients. As a result, when any non-zero element is detected in the 8×8 block region other than the top-left 4×4, the RST index is not encoded or decoded (it is in...). Figure 14 (as shown in the image), because this means that no RST was applied. In this case, the RST index is inferred to be zero.

[0157] 2.11.7 Zeroing Range

[0158] Normally, any coefficient in a 4×4 subblock can be non-zero before applying the inverse RST to the subblock. However, in some cases, certain coefficients in the 4×4 subblock must be zero before the inverse RST can be applied to the subblock.

[0159] Let nonZeroSize be a variable. Any coefficient with an index not less than nonZeroSize must be zero when rearranged into a one-dimensional array before the inverse RST.

[0160] When nonZeroSize equals 16, there is no zeroing constraint on the coefficients in the top left 4×4 sub-block.

[0161] In JVET-N0193, when the current block size is 4×4 or 8×8, nonZeroSize is set to equal to 8 (i.e., a coefficient with a scan index in the range [8, 15]). Figure 14 (As shown) should be 0). For other block dimensions, nonZeroSize is set to equal to 16.

[0162] 2.12 Affine Linear Weighted Intra-Prediction (ALWIP, or Matrix-Based Intra-Prediction)

[0163] JVET-N0217 proposed the Affine Linear Weighted Intra Prediction (ALWIP, or Matrix-Based Intra Prediction (MIP)).

[0164] Two tests were performed in JVET-N0217. In Test 1, ALWIP was designed with a memory limit of 8KB and a maximum of 4 multiplications per sample. Test 2 was similar to Test 1, but with further simplification in terms of memory requirements and model architecture.

[0165] ○ A single set of matrices and offset vectors for all block shapes.

[0166] ○ Reduce the number of patterns used for all block shapes to 19.

[0167] ○ Reduce memory requirements to 5760 10-bit values, or 7.20 kilobytes.

[0168] ○ Linear interpolation of the predicted samples is performed in a single step in each direction, instead of iterative interpolation in the first test.

[0169] 2.13 Sub-block Transformation

[0170] For inter-frame predicted CUs with cu_cbf equal to 1, signaling can be used to notify cu_sbt_flag to indicate whether to decode the entire residual block or a sub-part of the residual block. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is encoded / decoded using the inferred adaptive transform, and the other portion of the residual block is zeroed out. SBT is not applicable to combining inter-frame and intra-frame modes.

[0171] In the sub-block transformation, position-dependent transformations are applied to the luminance transform blocks in SBT-V and SBT-H (always using the chroma TB of DCT-2). The two positions in SBT-H and SBT-V are associated with different core transformations. More specifically, Figure 15The code specifies the horizontal and vertical transformations for each SBT location. For example, the horizontal and vertical transformations for SBT-V location 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the corresponding transformation is set to DCT-2. Therefore, the sub-block transformations jointly specify the TU tiling, cbf, and horizontal and vertical transformations of the residual block, which can be considered a syntax shortcut for cases where the main residual of the block is on one side of the block.

[0172] 3. Examples of the shortcomings of existing implementation methods

[0173] The current design has the following problems:

[0174] ○ The main transforms (including DCT-2, DST-7 and DCT-8) may not be applicable to different prediction modes.

[0175] ○ The main transformation may not be applicable to certain block sizes.

[0176] ○ The main transformation may not be applicable to different color components.

[0177] 4 Example methods of multiple transformations

[0178] The embodiments of the currently disclosed technology overcome the shortcomings of existing implementations, thereby providing video codecs with higher encoding / decoding efficiency but lower computational complexity. As described in this document, the methods for multi-transformation can enhance both existing and future video codec standards, as illustrated in the examples below describing various implementations. The examples of the disclosed technology provided below illustrate general concepts and are not intended to be limiting. In the examples, unless the contrary is explicitly stated, the various features described in these examples can be combined.

[0179] In the following examples, the CU may include information associated with all three color components having a single-tree codec structure. Alternatively, the CU may include information associated only with the luma color component of the monochrome codec. Alternatively, the CU may include information associated only with the luma color component having a dual-tree codec structure (e.g., the Y component in YCbCr format or the G component in GBR format). Alternatively, the CU may include information associated only with the two chroma components having a dual-tree codec structure (e.g., the Cb and Cr components in YCbCr format or the B and R components in GBR format).

[0180] In the following examples, "block" can refer to a codec unit (CU) or transform unit (TU) of video data, or any rectangular area. "Current block" can refer to the codec unit (CU) currently being decoded / encoded, the transform unit (TU) currently being decoded / encoded, or any codec rectangular area of ​​video data currently being decoded / encoded. "CU" or "TU" can also be referred to as "codec block" and "transform block".

[0181] In the following examples, the term "transformation" can refer to the transformation process or the transformation matrix.

[0182] In the following examples, "current block" can refer to the codec unit (CU) currently being decoded / encoded or the transform unit (TU) currently being decoded / encoded.

[0183] In the following example, the encoding / decoding information may include prediction mode (e.g., intra / inter / IBC mode), motion vector, reference picture, inter-frame prediction direction, intra-frame prediction mode, CIIP (Combined Intra-Inter-Frame Prediction) mode, ISP mode, affine intra-frame mode, the transform kernel used, transform skip flag, etc., i.e., the information required for encoding.

[0184] In the following examples, the master transform represents the transform applied to the prediction error prior to quantization and / or secondary transform (if necessary). On the decoder side, the master transform indicates the transform used to generate a temporary block, which is used to derive the final reconstructed block of the block.

[0185] In the following examples, the quadratic transform (e.g., LFSNT) represents a transform applied to the prediction error (if needed) after the main transform and before quantization. On the decoder side, the quadratic transform indicates the transform used to generate a temporary block, which serves as input to the main transform process.

[0186] 1. The main transformation (matrix) included in the transformation set can be, but is not limited to, one of DCT-II, DST-VII, DCT-VIII, transformation skip mode, identity transformation, or transformation from training (such as KLT-based transformation).

[0187] 2. The choice of the main transformation matrix can depend on the color components.

[0188] a. In one example, the first set of transformation matrices can be used for the luminance (or G) component, while the second set of transformation matrices is used for the chrominance (or B / R) component.

[0189] b. In one example, each color component may correspond to a separate set of transformations.

[0190] 3. The choice of the main transformation matrix can depend on the size of the block.

[0191] a. In one example, two or more sets of master transformation matrices can be defined for different block sizes.

[0192] b. In one example, if both the width and height of the block are less than M, the first set of principal transformation matrices can be used, for example, M=8, 16, 32. Otherwise, the second set of principal transformation matrices is used.

[0193] c. In one example, if the width or height of the block is less than M, the first set of principal transformation matrices can be used, for example, M=8, 16, 32. Otherwise, the second set of principal transformation matrices is used.

[0194] 4. The choice of the main transformation matrix can depend on the prediction mode.

[0195] a. In one example, two different sets of master transform matrices can be specified for intra-frame mode and inter-frame mode, respectively.

[0196] b. In one example, two or more different sets of master transform matrices can be specified for intra-frame mode.

[0197] c. In one example, two or more different sets of master transform matrices can be specified for inter-frame modes.

[0198] d. In one example, the transMatrix with transformation dimensions of 4, 8, 16, and 32 is included in the transformation set.

[0199] i. In one example, DCT-8 is not used in intra-frame mode when using transMatrix.

[0200] If the transformation dimension is 4, then the following applies:

[0201] transMatrix[m][n] =

[0202] {

[0203] { 46 80 77 44}

[0204] { 71 52 -49 -79}

[0205] { 82 -42 -48 75}

[0206] { -51 75 -75 50}

[0207] }

[0208] Otherwise, if the transformation dimension is 8, the following applies:

[0209] transMatrix[m][n] =

[0210] {

[0211] { -33 -60 -80 -90 -85 -66 -41 -21}

[0212] { -48 -76 -71 -23 43 84 85 56}

[0213] { -68 -75 -10 69 76 -5 -77 -77}

[0214] { -87 -40 74 60 -54 -67 34 77}

[0215] { 76 -16 -76 51 55 -84 -24 88}

[0216] { 79 -69 -14 78 -70 3 72 -72}

[0217] { 62 -87 65 -22 -27 69 -88 58}

[0218] { 35 -59 75 -82 80 -71 58 -31}

[0219] },

[0220] Otherwise, if the transformation dimension is 16, the following applies:

[0221] transMatrix[m][n] =

[0222] {

[0223] { -24 -40 -53 -65 -75 -82 -87 -89 -87 -82 -74 -63 -50 -37 -25 -14}

[0224] { -36 -60 -75 -80 -77 -64 -41 -9 26 58 81 91 89 75 56 35}

[0225] { -51 -80 -85 -64 -27 17 58 83 84 58 13 -36 -73 -87 -78 -55}

[0226] { -61 -87 -69 -16 47 85 81 33 -31 -76 -81 -41 21 71 86 66}

[0227] { 66 78 28 -47 -88 -61 12 77 81 19 -58 -89 -49 29 82 80}

[0228] { -78 -68 17 88 67 -26 -86 -51 40 84 33 -56 -85 -20 63 84}

[0229] { 83 45 -61 -82 12 87 36 -68 -72 30 88 16 -79 -60 36 84}

[0230] { -95 -14 95 30 -87 -40 76 55 -63 -63 50 70 -34 -76 10 77}

[0231] { 77 -17 -77 32 70 -48 -62 63 52 -77 -40 86 29 -94 -26 96}

[0232] { 83 -45 -60 77 19 -91 31 72 -69 -34 87 -16 -78 58 54 -80}

[0233] { 77 -69 -19 86 -58 -36 90 -47 -46 89 -32 -58 83 -15 -74 72}

[0234] { -76 91 -33 -47 90 -64 -9 74 -78 19 54 -86 48 29 -83 61}

[0235] { -56 85 -67 17 41 -79 76 -34 -26 75 -87 53 13 -76 97 -58}

[0236] { 49 -82 89 -73 36 13 -56 82 -84 59 -15 -33 70 -88 77 -38}

[0237] { -30 58 -79 89 -87 75 -51 22 10 -41 67 -83 86 -78 59 -28}

[0238] { -14 28 -42 54 -63 74 -84 91 -92 87 -80 72 -61 50 -35 15}

[0239] },

[0240] Otherwise, if the transformation dimension is 32, the following applies:

[0241] transMatrix[m][n] = transMatrixCol0to15[m][n], where m = 0...15, n = 0...15

[0242] transMatrixCol0to15 =

[0243] {

[0244] { 25 33 40 47 54 61 67 73 78 83 86 87 88 88 88 87}

[0245] { -43 -56 -67 -75 -81 -84 -85 -82 -77 -70 -61 -49 -34 -18 1 19}

[0246] { 61 80 90 92 87 75 58 36 11 -15 -40 -60 -75 -84 -87 -84}

[0247] { 57 73 79 75 59 35 5 -27 -54 -74 -82 -81 -68 -45 -13 22}

[0248] { -63 -78 -77 -57 -24 15 52 79 88 77 50 14 -25 -59 -80 -84}

[0249] { -72 -84 -68 -33 12 53 79 83 65 31 -14 -55 -81 -84 -59 -14}

[0250] { 78 82 54 1 -53 -87 -88 -55 0 56 92 89 48 -11 -66 -97}

[0251] { 83 77 29 -32 -73 -79 -45 11 58 74 53 4 -50 -81 -66 -15}

[0252] { 101 81 8 -71 -105 -72 6 76 95 53 -22 -79 -82 -33 35 80}

[0253] { 76 50 -21 -78 -70 -2 66 82 32 -44 -87 -63 8 75 85 25}

[0254] { -96 -43 52 97 47 -46 -92 -48 40 88 54 -27 -81 -58 16 72}

[0255] {-106 -27 85 90 -4 -87 -72 22 92 58 -40 -91 -42 52 87 22}

[0256] { -55 -8 54 45 -26 -68 -20 63 69 -19 -88 -43 62 90 -1 -94}

[0257] {-105 17 106 31 -84 -72 46 96 6 -90 -54 58 86 -11 -88 -37}

[0258] {-101 32 102 -9 -91 -15 81 40 -67 -59 40 73 -9 -77 -17 70}

[0259] { -85 51 82 -51 -85 46 88 -34 -88 20 94 -6 -94 -10 88 21}

[0260] },

[0261] transMatrix[ m ][ n ]=transMatrixCol16to31[ m − 16 ][ n ], wherein m =16..31, n = 0..15

[0262] transMatrixCol16to31 =

[0263] {

[0264] { 84 81 78 74 70 66 62 57 52 47 40 33 26 19 14 9}

[0265] { 36 51 63 72 80 85 87 87 84 78 71 62 52 41 30 20}

[0266] { -74 -58 -39 -17 4 25 44 60 72 79 81 78 72 62 49 34}

[0267] { 53 75 87 87 77 59 33 2 -31 -59 -80 -92 -93 -84 -69 -49}

[0268] { -74 -50 -16 22 56 80 87 79 54 17 -23 -59 -84 -94 -89 -71}

[0269] { 36 76 92 80 41 -10 -57 -87 -90 -67 -26 21 60 84 87 72}

[0270] { -86 -37 27 76 94 75 28 -26 -67 -80 -62 -22 22 56 68 60}

[0271] { 44 81 78 33 -31 -83 -93 -54 13 72 98 74 13 -55 -96 -98}

[0272] { 71 14 -52 -79 -51 7 60 71 38 -21 -71 -76 -34 29 76 84}

[0273] { -55 -91 -53 31 88 75 5 -68 -90 -45 32 87 77 7 -66 -100}

[0274] { 65 -5 -73 -72 -3 69 77 11 -70 -86 -19 67 91 34 -51 -99}

[0275] { -66 -79 -3 77 71 -12 -80 -59 24 78 48 -31 -76 -46 28 80}

[0276] { -68 52 109 20 -89 -83 26 99 46 -62 -90 -6 79 65 -19 -90}

[0277] { 65 69 -26 -82 -17 72 56 -41 -77 -4 73 47 -45 -72 0 74}

[0278] { 40 -60 -55 46 68 -25 -83 -4 86 44 -72 -79 36 101 17 -106}

[0279] { -78 -34 72 46 -65 -60 51 72 -31 -78 12 82 9 -85 -31 84}

[0280] },

[0281] 5. The choice of the main transform matrix can depend on the intra-frame prediction method (including but not limited to ISP, MIP, MRL, CCLM, BDPCM).

[0282] a. In one example, two or more sets of master transform matrices can be used for different intra-frame modes.

[0283] b. In one example, the first set of transform matrices can be used for CCLM codec blocks, and the second set of transform matrices can be used for non-CCLM codec blocks.

[0284] c. In one example, the first set of transform matrices can be used for normal intra-frame prediction codec blocks, and the second set of transform matrices can be used for multiple reference line (MRL) blocks.

[0285] 6. The choice of the main transformation matrix can depend on the inter-frame prediction method, including but not limited to merge mode, AMVP mode, and affine mode.

[0286] a. Two or more sets of master transform matrices can be used for different inter-frame modes.

[0287] b. In one example, the first set of transform matrices can be used for merge mode blocks and normal inter-frame prediction codec blocks, and the second set of transform matrices can be used for non-merge codec blocks.

[0288] c. In one example, the first set of transform matrices can be used for AMVP mode blocks and normal inter-frame prediction codec blocks, and the second set of transform matrices can be used for non-AMVP codec blocks.

[0289] 7. In one example, when the codec tool X is applied to a block, the first set of transformation matrices can be used, while when the codec tool X is not applied to a block, the second set of transformation matrices can be used.

[0290] a. The encoding / decoding tool X can be: ISP, MIP, LMCS, ALF, SAO, DQ, affine, SMVD, SBT, AMVR, BDOF, PROF, DMVR, LFNST, LIC, OBMC, JCCR, IBC.

[0291] 8. The choice of the master transformation matrix can depend on the transformation mode, including but not limited to implicit MTS, explicit MTS, SBT mode, JCCR, and different LFNST indices.

[0292] a. In one example, two or more sets of master transformation matrices can be used for different transformation modes.

[0293] b. In one example, the first set of transformation matrices can be used for implicit MTS, and the second set of transformation matrices can be used for explicit MTS.

[0294] c. In one example, the first set of transformation matrices can be used for the SBT mode, and the second set of transformation matrices can be used for the non-SBT mode.

[0295] i. In one example, transMatrix with transformation dimensions of 4, 8, 16, and 32 can be used in SBT.

[0296] 1) In one example, DST-7 is not used for SBT when using transMatrix.

[0297] If the transformation dimension is 4, then the following applies:

[0298] transMatrix[m][n] =

[0299] {

[0300] { 19 50 80 84}

[0301] { 61 84 11 -75}

[0302] { 88 -5 -75 55}

[0303] { 68 -83 65 -28}

[0304] },

[0305] Otherwise, if the transformation dimension is 8, the following applies:

[0306] transMatrix[m][n] =

[0307] {

[0308] { 13 27 48 71 86 86 77 62}

[0309] { -28 -58 -84 -79 -31 36 79 82}

[0310] { -64 -94 -59 31 83 45 -34 -73}

[0311] { -83 -62 40 79 -21 -85 -8 80}

[0312] { -96 3 88 -27 -69 58 50 -67}

[0313] { -84 69 16 -78 65 13 -82 59}

[0314] { -62 91 -72 25 29 -72 83 -43}

[0315] { -26 54 -72 82 -85 78 -58 24}

[0316] },

[0317] Otherwise, if the transformation dimension is 16, the following applies:

[0318] transMatrix[m][n] =

[0319] {

[0320] { 9 14 21 30 40 52 64 75 82 85 86 84 81 76 71 63}

[0321] { 17 27 41 58 75 85 87 78 55 23 -11 -44 -70 -86 -91 -83}

[0322] { -39 -63 -85 -97 -89 -57 -7 46 80 86 67 32 -8 -44 -65 -67}

[0323] { -52 -76 -84 -64 -11 52 88 77 22 -47 -87 -82 -38 21 67 79}

[0324] { 81 97 66 -6 -76 -90 -32 50 85 45 -29 -78 -69 -15 45 70}

[0325] { -83 -75 -9 67 74 2 -72 -66 18 88 57 -38 -94 -58 31 88}

[0326] { 92 56 -42 -90 -20 76 58 -47 -81 8 83 42 -60 -83 2 80}

[0327] { -90 -18 77 54 -63 -65 48 75 -38 -81 18 89 0 -93 -28 80}

[0328] { 95 -20 -92 14 88 -23 -83 37 72 -54 -61 65 51 -71 -48 70}

[0329] { 86 -56 -64 73 29 -88 19 78 -64 -43 85 2 -84 36 70 -64}

[0330] { 75 -74 -23 88 -55 -38 87 -38 -51 84 -23 -70 81 6 -87 59}

[0331] { -64 93 -37 -46 83 -46 -33 94 -89 22 52 -77 32 45 -88 51}

[0332] { -49 87 -72 14 49 -84 76 -29 -32 80 -85 39 28 -81 91 -45}

[0333] { 41 -80 85 -57 11 36 -72 91 -90 69 -28 -24 67 -85 72 -31}

[0334] { 14 -40 66 -83 89 -86 76 -65 62 -69 76 -76 65 -46 23 -5}

[0335] { -25 57 -76 81 -76 62 -39 9 26 -58 82 -93 93 -80 57 -23}

[0336] },

[0337] Otherwise, if the transformation dimension is 32, the following condition applies:

[0338] transMatrix[ m ][ n ] = transMatrixCol0to15[ m ][ n ], where m = 0...15,n = 0...15

[0339] transMatrixCol0to15 =

[0340] {

[0341] { -9 -11 -14 -16 -20 -24 -28 -33 -38 -43 -48 -54 -59 -65 -70 -74}

[0342] { -19 -24 -30 -37 -45 -53 -61 -68 -75 -79 -84 -86 -87 -84 -79 -72}

[0343] { -35 -44 -55 -65 -75 -84 -89 -90 -88 -80 -65 -45 -19 9 35 58}

[0344] { 43 55 65 74 79 76 67 51 30 4 -25 -52 -76 -90 -89 -78}

[0345] { -71 -87 -96 -93 -79 -56 -23 14 49 78 92 89 69 33 -10 -50}

[0346] { -70 -83 -82 -67 -38 -1 38 69 83 77 49 4 -44 -79 -89 -70}

[0347] { -86 -95 -80 -41 10 60 91 93 61 6 -52 -89 -92 -55 8 64}

[0348] { -76 -82 -53 0 56 90 81 35 -24 -71 -83 -52 8 65 86 58}

[0349] { -89 -78 -29 35 82 87 43 -24 -79 -86 -37 39 88 77 7 -64}

[0350] { 78 63 7 -56 -85 -55 17 76 80 23 -53 -88 -50 29 84 67}

[0351] { 91 58 -25 -90 -82 -8 70 90 31 -56 -91 -39 52 90 40 -52}

[0352] {-101 -46 52 96 47 -45 -87 -39 53 85 22 -64 -76 2 77 68}

[0353] { 92 27 -74 -90 -3 88 70 -33 -94 -35 67 82 -15 -85 -38 58}

[0354] { -82 -6 82 60 -43 -86 -4 88 50 -58 -80 13 85 32 -72 -72}

[0355] { -90 10 98 33 -88 -67 64 93 -32 -99 -3 92 28 -77 -50 65}

[0356] { 68 -15 -74 -6 72 18 -66 -35 69 42 -69 -50 66 52 -57 -63}

[0357] },

[0358] transMatrix[ m ][ n ] = transMatrixCol16to31[ m − 16 ][ n ], wherein m =16...31, n = 0...15

[0359] transMatrixCol16to31 =

[0360] {

[0361] { -78 -80 -82 -83 -84 -85 -85 -85 -83 -82 -80 -78 -75 -73 -70 -66}

[0362] { -62 -49 -32 -15 2 19 35 48 60 71 78 82 86 88 87 83}

[0363] { 76 86 89 87 79 65 47 27 5 -17 -36 -52 -65 -73 -75 -72}

[0364] { -52 -19 18 51 75 91 94 85 66 37 4 -30 -56 -75 -85 -81}

[0365] { -76 -83 -73 -49 -12 29 62 80 81 67 39 6 -26 -53 -70 -71}

[0366] { -25 29 73 94 84 50 1 -48 -83 -93 -78 -42 4 45 72 78}

[0367] { 90 75 29 -29 -70 -81 -62 -18 32 71 80 61 22 -21 -56 -69}

[0368] { -3 -61 -86 -66 -9 54 89 79 24 -45 -90 -92 -51 10 65 88}

[0369] { -87 -47 21 76 77 29 -36 -79 -74 -17 51 92 78 17 -53 -88}

[0370] { -11 -78 -83 -20 62 95 51 -32 -90 -76 1 78 93 40 -40 -88}

[0371] { -93 -34 57 89 29 -56 -81 -34 44 83 43 -38 -84 -57 22 79}

[0372] { -29 -86 -49 45 87 32 -57 -84 -18 71 79 -3 -83 -75 13 82}

[0373] { 73 -12 -79 -45 51 83 7 -76 -66 26 88 39 -62 -86 -3 78}

[0374] { 44 92 2 -95 -46 69 76 -22 -87 -23 72 63 -42 -90 -11 80}

[0375] { 64 -47 -80 20 84 8 -74 -38 60 65 -35 -82 7 82 24 -65}

[0376] { 56 68 -47 -85 43 96 -29 -106 12 107 8 -105 -24 97 45 -80}

[0377] },

[0378] ii. In one example, a transMatrix with a transformation dimension of 4, 8, 16, or 32 can be used for SBT.

[0379] 1) In one example, DCT-8 is not used for SBT when using transMatrix.

[0380] If the transformation dimension is 4, then the following applies:

[0381] transMatrix[m][n] =

[0382] {

[0383] { -90 -80 -43 -13}

[0384] { -75 31 88 44}

[0385] { -49 85 -34 -75}

[0386] { -16 43 -75 93}

[0387] },

[0388] Otherwise, if the transformation dimension is 8, the following applies:

[0389] transMatrix[m][n] =

[0390] {

[0391] { 71 84 87 81 64 41 21 9}

[0392] { 85 74 21 -46 -88 -83 -50 -20}

[0393] { 77 23 -62 -83 -8 75 86 49}

[0394] { 73 -19 -82 -2 85 16 -82 -79}

[0395] { 67 -63 -46 78 6 -84 21 92}

[0396] { 53 -81 25 53 -86 45 48 -92}

[0397] { 39 -82 84 -51 -1 53 -86 69}

[0398] { -19 47 -66 74 -78 81 -75 46}

[0399] },

[0400] Otherwise, if the transformation dimension is 16, the following applies:

[0401] transMatrix[m][n] =

[0402] {

[0403] { 75 81 85 87 86 85 81 76 67 56 44 33 24 16 10 6}

[0404] { -85 -89 -79 -58 -29 2 37 69 88 93 86 72 54 36 21 12}

[0405] { -76 -70 -38 7 52 84 89 68 24 -28 -69 -91 -90 -74 -49 -26}

[0406] { 80 62 4 -59 -89 -74 -18 49 86 75 23 -42 -82 -87 -66 -38}

[0407] { 71 37 -35 -85 -65 5 72 80 18 -62 -91 -49 27 82 91 63}

[0408] { -83 -16 73 86 4 -78 -69 16 80 50 -38 -87 -45 40 88 72}

[0409] { 71 -8 -84 -44 61 76 -25 -90 -17 81 56 -50 -84 -13 75 88}

[0410] { 69 -30 -80 12 83 0 -83 -16 84 29 -77 -52 71 74 -42 -102}

[0411] { 81 -64 -72 69 53 -74 -41 76 23 -81 -6 83 -11 -84 5 89}

[0412] { 61 -71 -26 80 -22 -71 63 42 -90 8 88 -55 -60 80 37 -91}

[0413] { -57 86 -16 -69 80 3 -80 68 16 -85 59 38 -91 36 66 -79}

[0414] { 46 -83 51 21 -77 64 6 -77 94 -43 -35 84 -65 -14 89 -79}

[0415] { 37 -80 76 -30 -35 78 -79 34 30 -82 97 -65 2 61 -89 60}

[0416] { 34 -77 91 -73 31 21 -66 88 -87 68 -33 -12 55 -82 79 -44}

[0417] { 4 -17 35 -54 70 -78 78 -72 70 -76 81 -82 81 -72 50 -22}

[0418] { -24 59 -81 94 -94 81 -57 29 2 -30 51 -68 81 -81 65 -33}

[0419] },

[0420] Otherwise, if the transformation dimension is 32, the following applies:

[0421] transMatrix[m][n] = transMatrixCol0to15[m][n], where m = 0...15, n = 0...15

[0422] transMatrixCol0to15 =

[0423] {

[0424] { -73 -76 -78 -81 -83 -85 -86 -86 -87 -87 -87 -85 -83 -81 -77 -74}

[0425] { -94 -96 -96 -92 -85 -74 -60 -47 -34 -16 0 18 35 50 61 72}

[0426] { 76 81 76 63 43 21 -1 -25 -46 -64 -80 -91 -93 -85 -69 -48}

[0427] { 69 70 59 41 15 -12 -41 -70 -89 -92 -82 -61 -31 7 42 73}

[0428] { -80 -77 -56 -19 24 62 91 98 81 43 0 -44 -78 -91 -81 -60}

[0429] { 80 69 38 -11 -54 -82 -92 -70 -24 29 66 82 70 35 -9 -55}

[0430] { -88 -62 -3 49 80 85 54 5 -50 -87 -84 -49 16 71 92 71}

[0431] { -84 -49 17 64 74 55 14 -34 -72 -67 -27 33 78 75 27 -34}

[0432] { -99 -49 40 93 86 32 -47 -88 -65 -3 59 83 52 -12 -71 -90}

[0433] { -91 -38 52 100 58 -24 -83 -74 1 76 85 22 -62 -89 -37 40}

[0434] { 67 9 -54 -70 -18 51 80 35 -58 -93 -29 58 88 29 -63 -90}

[0435] { 70 1 -95 -73 43 117 52 -72 -113 -27 76 90 4 -74 -64 14}

[0436] { -69 11 82 45 -53 -72 -1 64 56 -32 -80 -14 74 58 -37 -93}

[0437] { 71 -28 -85 -8 71 39 -50 -58 23 64 15 -67 -63 37 102 19}

[0438] { 81 -41 -104 11 101 28 -89 -54 70 60 -48 -72 28 77 -1 -82}

[0439] { 67 -45 -81 42 78 -23 -78 8 71 -5 -64 -10 68 23 -75 -41}

[0440] },

[0441] transMatrix[ m ][ n ] = transMatrixCol16to31[ m − 16 ][ n ], where m = 16...31, n = 0...15

[0442] transMatrixCol16to31 =

[0443] {

[0444] { -69 -64 -58 -52 -46 -42 -38 -33 -27 -23 -18 -15 -13 -10 -8 -6}

[0445] { 80 85 85 84 81 76 72 67 60 54 46 40 33 28 23 17}

[0446] { -24 -2 19 38 55 72 84 90 91 87 80 71 60 52 43 31}

[0447] { 91 97 92 72 43 9 -23 -51 -70 -82 -86 -81 -71 -61 -48 -35}

[0448] { -23 17 50 76 86 81 62 35 2 -36 -60 -77 -81 -76 -66 -49}

[0449] { -80 -73 -47 -10 35 71 92 90 62 19 -28 -71 -91 -90 -81 -63}

[0450] { 25 -33 -72 -83 -66 -22 29 70 90 79 39 -15 -58 -81 -85 -68}

[0451] { -81 -89 -45 25 82 101 60 -6 -66 -94 -82 -40 21 75 98 86}

[0452] { -48 28 95 91 24 -54 -88 -63 -2 55 82 57 13 -38 -73 -64}

[0453] { 78 60 5 -64 -88 -39 41 82 58 -6 -75 -93 -41 30 76 77}

[0454] { -22 59 82 35 -58 -103 -39 61 92 47 -38 -99 -68 9 79 86}

[0455] { 74 41 -30 -71 -39 35 77 36 -44 -78 -26 62 79 21 -55 -75}

[0456] { -11 88 64 -43 -96 -25 77 86 -20 -94 -62 47 99 47 -54 -92}

[0457] {-102 -73 70 97 -22 -97 -27 84 65 -44 -92 -14 86 70 -30 -67}

[0458] { -20 72 47 -54 -64 34 75 -8 -77 -23 71 58 -58 -97 12 106}

[0459] { 86 45 -76 -54 72 64 -70 -70 66 75 -49 -94 38 109 -4 -101}

[0460] },

[0461] d. In one example, the first set of transformation matrices can be used in JCCR mode, and the second set of transformation matrices can be used in non-JCCR mode.

[0462] e. In one example, the set of transformation matrices can be selected based on the LFNST index.

[0463] i. In one example, signaling notifications for the MTS index can be conditional on the LFNST index.

[0464] 1) In one example, the MTS index may not be signaled when the LFNST index is equal to certain specified values ​​(e.g., 1 and 2).

[0465] 9. In the example above, for two sets of transformations (e.g., represented by first and second sets, for the principal transformation and / or the second transformation), the following may apply:

[0466] a. In one example, the second group may include all the transformations in the first group and have additional transformations.

[0467] i. Alternatively, the first set may include all transformations in the second set, and have additional transformations.

[0468] b. In one example, if the first transformation set and the second transformation set are different, then at least one matrix of the first transformation set is not included in the second transformation set, or at least one matrix of the second transformation set is not included in the first transformation set.

[0469] c. In one example, two sets can have the same number of transformations.

[0470] d. In one example, the two sets can have different numbers of transformations.

[0471] Zeroing range in primary / secondary conversion

[0472] 10. The zeroing range in the master transform (e.g., including MTS and DCT2) may depend on the prediction mode.

[0473] a. In one example, the zeroing range in the main transform may differ for intra-frame mode and inter-frame mode.

[0474] b. In one example, the zeroing range used in the main transform for intra-frame and inter-frame modes may differ depending on certain block sizes.

[0475] c. In one example, whether to apply the zeroing operation depends on the prediction mode.

[0476] 11. The zeroing range in a second transformation may depend on whether the block size is smaller than (or not smaller than, or larger than, or not larger than) the specified size.

[0477] a. In one example, different zeroing ranges can be specified for different block sizes.

[0478] i. In one example, when the block size is less than M×N, the zeroing range can be set to X. (In other words, any coefficient with a scan order index greater than or equal to X should be equal to zero.)

[0479] 1) In one example, when the size of the block (width or height, or both width and height) is less than or equal to 8, the zeroing range of the block is set to X.

[0480] ii. In one example, when the block size is greater than M×N, the zeroing range can be set to Y. (In other words, any coefficient with a scan order index greater than or equal to Y should be equal to 0.)

[0481] 1) In one example, when the size of the block (width or height, or both width and height) is greater than or equal to 8, the zeroing range of the block can be set to Y.

[0482] 2) In one example, when the width of the block is greater than 8 and the height is greater than or equal to 8, or when the width of the block is greater than or equal to 8 and the height is greater than 8, the zeroing range of the block can be set to Y.

[0483] iii. For example, X is different from Y.

[0484] iv. For example, X equals 16.

[0485] v. For example, Y equals 32.

[0486] 12. The zeroing range in a quadratic transformation may depend on the number of samples in the block.

[0487] a. In one example, different zeroing ranges can be specified based on the number of samples in the transform block.

[0488] i. In one example, when the number of samples in the transform block is less than N, the zeroing range can be set to X. (In other words, any coefficient with a scan order index greater than or equal to X should be equal to 0.)

[0489] 1) In one example, when the number of samples in a block is less than 256, the zeroing range of the block can be set to X.

[0490] ii. In one example, when the number of samples in the transform block is greater than N, the zeroing range can be set to Y. (In other words, any coefficient with a scan order index greater than or equal to Y should be equal to 0.)

[0491] 1) In one example, when the number of samples in a block is greater than 256, the zeroing range of the block can be set to Y.

[0492] iii. For example, X is different from Y.

[0493] iv. For example, X equals 16.

[0494] v. For example, Y equals 32.

[0495] 13. Zeroing can be disabled in the primary or secondary transform (also known as LFNST).

[0496] a. In one example, zeroing can be disabled for some block sizes in the main transform.

[0497] b. In one example, zeroing can be disabled for some transformation modes in the primary or secondary transformation, including but not limited to implicit MTS, explicit MTS, SBT mode, JCCR, and different LFNST indices.

[0498] i. In one example, zeroing can be disabled for MTS.

[0499] ii. In one example, zeroing can be disabled for DCT2.

[0500] c. In one example, zeroing can be disabled for certain dimensions in a quadratic transformation.

[0501] i. In one example, zeroing can be disabled for LFNST 4×4.

[0502] ii. In one example, zeroing can be disabled for LFNST 8×8.

[0503] iii. In one example, zeroing can be disabled for blocks larger than M×N, such as M=16, N=16.

[0504] iv. In one example, zeroing can be disabled when the number of samples in a block is greater than N, for example, N=64.

[0505] 14. The zeroing range can be expanded in the main transformation or the secondary transformation (also known as LFNST).

[0506] a. In one example, the zeroing range can be expanded for some block sizes in the main transform.

[0507] b. In one example, the zeroing range can be expanded for some transformation modes in the primary or secondary transformation, including but not limited to implicit MTS, explicit MTS, SBT mode, JCCR, and different LFNST indices.

[0508] i. In one example, the zeroing range for MTS can be expanded.

[0509] ii. In one example, the zeroing range for DCT2 can be expanded.

[0510] c. In one example, the zeroing range can be expanded for certain block sizes in a quadratic transformation.

[0511] i. In one example, the zeroing range can be expanded for LFNST 4×4.

[0512] ii. In one example, the zeroing range can be expanded for LFNST 8×8.

[0513] iii. In one example, for blocks larger than M×N, such as M=16, N=16, X=32, the zeroing range can be extended to X.

[0514] iv. In one example, when the number of samples in a block is greater than N, for example N=64 and X=32, the zeroing range can be expanded to X.

[0515] 15. The transformation process can be implemented using a butterfly calculation method.

[0516] a. In one example, some principal transformations can be implemented using butterfly computation.

[0517] b. In one example, some quadratic transformations can be achieved using butterfly computation.

[0518] c. In one example, the transformation matrix can be adjusted to facilitate butterfly computation.

[0519] The above example can be incorporated into the context of methods such as method 1600, which can be implemented in a video encoder or decoder.

[0520] Figure 16 A flowchart of an exemplary method for video processing is shown. Method 1600 includes, in step 1602, performing a conversion between the current block of the video and the bitstream representation of the video based on the selection of a master transform included in a transform set. In some embodiments, the conversion includes using a master transform, and the master transform includes one or more of Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VII, transform skip mode, identity transform, or a transform based on a training process.

[0521] In some embodiments, the following technical solutions can be implemented:

[0522] 1. A video processing method comprising performing a conversion between a current block of video and a bitstream representation of video based on a selection of a master transform included in a transform set, wherein the conversion includes using a master transform, and the master transform includes one or more of Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VII, transform skip mode, identity transform, or a transformation based on a training process.

[0523] 2. The method described in Solution 1, wherein the transformation based on the training process is a transformation based on the Karhunen-Loeve transformation (KLT).

[0524] 3. The method described in Solution 1, wherein the selection of the primary transform is based on the chroma components of the video.

[0525] 4. The method as described in Solution 3, wherein the transform set includes a first transform matrix set for the luminance component of the video and a second transform matrix set for the chrominance component.

[0526] 5. As described in Solution 1, wherein the selection of the primary transformation is based on the size of the current block.

[0527] 6. The method described in Solution 5, wherein a first principal transformation matrix is ​​selected when the height or width of the current block is less than M, and a second principal transformation matrix is ​​selected otherwise, and M is a positive integer.

[0528] 7. The method described in Solution 1, wherein the selection of the primary transform is based on the prediction mode of the current block.

[0529] 8. The method described in Solution 7, wherein the prediction mode is an inter-frame mode or an intra-frame mode.

[0530] 9. The method as described in Solution 1, wherein the selection of the primary transform is based on the intra-prediction method applied to the current block.

[0531] 10. The method of Solution 9, wherein the intra-prediction mode includes one or more of the following: intra-block segmentation (ISP), matrix-based intra-prediction (MIP), multiple reference lines (MRL) method, cross-component linear model (CCLM), or block differential pulse coding-decoding modulation-decoding (BDPCM).

[0532] 11. The method as described in Solution 1, wherein the selection of the primary transform is based on the inter-frame prediction method applied to the current block.

[0533] 12. The method as described in Solution 1, wherein when the encoding / decoding tool is applied to the current block, the transform set includes a first transform set, and when the encoding / decoding tool is not applied to the current block, the transform set includes a second transform set.

[0534] 13. The method of solution 12, wherein the encoding / decoding tools include one or more of the following: Intra-block Segmentation (ISP), Matrix-based Intra-frame Prediction (MIP), Luminance Mapping with Chroma Scaling (LMCS), Adaptive Loop Filtering (ALF), Sample Adaptive Offset (SAO), Dependent Quantization (DQ), Affine, Symmetric Motion Vector Differentiation (SMVD), Sub-block Transform (SBT), Adaptive Motion Vector Resolution (AMVR), Bidirectional Optical Flow (BDOF), Predictive Refinement with Optical Flow (PROF), Decoder-Side Motion Vector Refinement (DMVR), Low-Frequency Inseparable Transform (LFNST), Local Illumination Compensation (LIC), Overlapping Block Motion Compensation (OBMC), Joint Encoding / Decoding or Chroma Residual (JCCR) or Intra-block Copying (IBC).

[0535] 14. The method described in Solution 1, wherein the selection of the primary transform is based on the transform mode of the current block.

[0536] 15. The method as described in Solution 14, wherein the transformation mode includes one or more of implicit MTS, explicit MTS, SBT mode, JCCR, and different LFNST indices.

[0537] 16. The method as described in Solution 1, wherein the transform set includes a first transform set and a second transform set.

[0538] 17. The method as described in Solution 16, wherein the number of first transformations in the first transformation set is the same as the number of second transformations in the second transformation set.

[0539] 18. The method described in Solution 1, wherein the zeroing range of the main transform is based on the prediction mode of the current block.

[0540] 19. The method as described in Solution 1, wherein the transformation further includes using a quadratic transformation included in the transformation set.

[0541] 20. The method as described in Solution 19, wherein the zeroing range of the secondary transformation is based on a comparison of at least one dimension of the current block with a predetermined threshold.

[0542] 21. The method as described in Solution 19, wherein the zeroing range of the second transformation is based on the number of samples in the current block.

[0543] 22. The method as described in Solution 19, wherein the zeroing operation in the primary transformation and / or secondary transformation is disabled.

[0544] 23. The method as described in Solution 19, wherein the zeroing range of the primary transformation and / or secondary transformation is expanded.

[0545] 24. The method of any one of solutions 1 to 23, wherein the main transformation is implemented using butterfly computation.

[0546] 25. The method of any one of solutions 1 to 24, wherein the transformation generates the current block from the bitstream representation.

[0547] 26. The method of any one of solutions 1 to 24, wherein the transformation generates a bitstream representation from the current block.

[0548] 27. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method described in any one of solutions 1 to 26.

[0549] 28. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method of any one of solutions 1 to 26.

[0550] Figure 17 This is a block diagram of a video processing apparatus 1700. Apparatus 1700 can be used to implement one or more of the methods described herein. Apparatus 1700 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 1700 may include one or more processors 1702, one or more memories 1704, and video processing hardware 1706. The processors 1702(s) may be configured to implement one or more methods described herein (including, but not limited to, method 1600). The memories 1704(s) may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 1706 may be used to implement some of the techniques described herein in hardware circuitry.

[0551] In some embodiments, the video encoding and decoding methods and technical solutions described in this document can be implemented using devices implemented on a hardware platform, such as those related to... Figure 17 As stated above.

[0552] Figure 18 This is a block diagram illustrating an example video codec system 300 that can utilize the techniques disclosed herein.

[0553] like Figure 18 As shown, the video encoding / decoding system 300 may include a source device 310 and a destination device 320. The source device 310 generates encoded video data and may be referred to as a video encoding device. The destination device 320 can decode the encoded video data generated by the source device 310 and may be referred to as a video decoding device.

[0554] The source device 310 may include a video source 312, a video encoder 314, and an input / output (I / O) interface 316.

[0555] Video source 312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 314 encodes the video data from video source 312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax elements. I / O interface 316 includes a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 320 via network 330a through I / O interface 316. Encoded video data may also be stored on storage medium / server 330b for access by destination device 320.

[0556] Destination device 320 may include I / O interface 326, video decoder 324 and display device 322.

[0557] I / O interface 326 may include a receiver and / or a modem. I / O interface 326 may acquire encoded video data from source device 310 or storage medium / server 330b. Video decoder 324 may decode the encoded video data. Display device 322 may display the decoded video data to a user. Display device 322 may be integrated with destination device 320 or may be external to destination device 320 configured to connect to an external display device.

[0558] The video encoder 314 and the video decoder 324 can operate according to video compression standards such as High Efficiency Video Coding (HEVC), Multi-Functional Video Coding (VVC), and other current and / or other standards.

[0559] Figure 19 This is a block diagram illustrating an example of a video encoder 400, which can be... Figure 18 The video encoder 314 in the system 300 shown in the figure.

[0560] The video encoder 400 can be configured to perform any or all of the techniques disclosed herein. Figure 19 In the example, the video encoder 400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 400. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0561] The functional components of the video encoder 400 may include a segmentation unit 401, a prediction unit 402 (which may include a mode selection unit 403, a motion estimation unit 404, a motion compensation unit 405, and an intra-frame prediction unit 406), a residual generation unit 407, a transform unit 408, a quantization unit 409, an inverse quantization unit 410, an inverse transform unit 411, a reconstruction unit 412, a buffer 413, and an entropy coding unit 414.

[0562] In other examples, the video encoder 400 may include more, fewer, or different functional components. In one example, the prediction unit 402 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0563] Furthermore, some components, such as the motion estimation unit 404 and the motion compensation unit 405, can be highly integrated, but for interpretive purposes... Figure 19 The examples are shown separately.

[0564] The segmentation unit 401 can segment an image into one or more video blocks. The video encoder 400 and the video decoder 500 can support various video block sizes.

[0565] The mode selection unit 403 can, for example, select one of the intra-frame or inter-frame encoding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 407 to generate residual block data and to the reconstruction unit 412 to reconstruct the encoded / decoded blocks for use as reference images. In some examples, the mode selection unit 403 can select a combined intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 403 can also select the resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for the blocks in the inter-frame prediction case.

[0566] To perform inter-frame prediction for the current video block, motion estimation unit 404 can generate motion information for the current video block by comparing one or more reference frames from buffer 413 with the current video block. Motion compensation unit 405 can determine the predicted video block for the current video block based on the motion information of the image from buffer 413 (rather than the image associated with the current video block) and decoded samples.

[0567] The motion estimation unit 404 and the motion compensation unit 405 can perform different operations on the current video block, for example, the different operations performed depend on whether the current video block is in an I-band, a P-band, or a B-band.

[0568] In some examples, motion estimation unit 404 can perform unidirectional prediction of the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 404 can then generate a reference index indicating that the reference video block is present in the reference images of list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0569] In other examples, motion estimation unit 404 can perform bidirectional prediction of the current video block. Motion estimation unit 404 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 404 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 404 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 405 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0570] In some examples, the motion estimation unit 404 can output the complete set of motion information for the decoder's decoding process.

[0571] In some examples, the motion estimation unit 404 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 404 may signal the motion information of the current video block by referencing the motion information of another video block. For example, the motion estimation unit 404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0572] In one example, the motion estimation unit 404 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.

[0573] In another example, motion estimation unit 404 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicating video block. Video decoder 500 can use the motion vector of the indicating video block and the motion vector difference to determine the motion vector of the current video block.

[0574] As discussed above, the video encoder 400 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 400 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.

[0575] Intra-prediction unit 406 can perform intra-prediction on the current video block. When intra-prediction unit 406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0576] The residual generation unit 407 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0577] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 407 may not perform a subtraction operation.

[0578] The transform processing unit 408 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0579] After the transform processing unit 408 generates a transform coefficient video block associated with the current video block, the quantization unit 409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0580] The inverse quantization unit 410 and the inverse transform unit 411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 402 to produce a reconstructed video block associated with the current block for storage in the buffer 413.

[0581] After the video block is reconstructed in reconstruction unit 412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0582] The entropy encoding unit 414 can receive data from other functional components of the video encoder 400. When the entropy encoding unit 414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0583] Figure 20 This is a block diagram illustrating an example of a video decoder 500, which can be... Figure 18 The video decoder 314 in the system 300 shown in the figure.

[0584] The video decoder 500 can be configured to perform any or all of the techniques disclosed herein. Figure 20 In the example, the video decoder 500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 500. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0585] exist Figure 20 In the example, the video decoder 500 includes an entropy decoding unit 501, a motion compensation unit 502, an intra-frame prediction unit 503, an inverse quantization unit 504, an inverse transform unit 505, a reconstruction unit 506, and a buffer 507. In some examples, the video decoder 500 can perform operations related to the video encoder 400 ( Figure 19 The decoding process is the overall inversion of the encoding process described.

[0586] Entropy decoding unit 501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 501 can decode the entropy-encoded video, and based on the entropy-encoded video data, motion compensation unit 502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 502 can determine such information, for example, by performing AMVP and merge modes.

[0587] The motion compensation unit 502 can generate motion compensation blocks, possibly based on interpolation filters. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0588] The motion compensation unit 502 can use the interpolation filter used by the video encoder 400 during the encoding of the video block to calculate the interpolation values ​​of a sub-integer number of pixels of the reference block. The motion compensation unit 502 can determine the interpolation filter used by the video encoder 400 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0589] The motion compensation unit 502 can use some syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0590] Intra-prediction unit 503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 503 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 501. Inverse transform unit 503 applies an inverse transform.

[0591] The reconstruction unit 506 can sum the residual block using the corresponding prediction block generated by the motion compensation unit 402 or the intra-frame prediction unit 503 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on the display device.

[0592] Figure 21 This is a block diagram of an example video processing system 2100 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 2100. System 2100 may include an input 2102 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 2102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0593] System 2100 may include an encoding / decoding component 2104 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 2104 can reduce the average bit rate of the video from input 2102 to the output of encoding / decoding component 2104 to produce an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 2104 may be stored or transmitted via connected communication, as represented by component 2106. The stored or communicated bitstream (or encoded / decoded) representation of the video received at input 2102 may be used by component 2108 to generate pixel values ​​or displayable video that is sent to display interface 2110. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that the encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations will be inverted by the decoder to retrieve the encoded / decoded results.

[0594] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0595] Figure 22 This is a flowchart of example method 2200 for video processing. Operation 2202 includes: performing a transformation between a current video block and a bitstream of video based on rules, wherein the rules specify that the set of transformation matrices selected during the transformation to perform the transformation operation is based on a low-frequency inseparable transformation index indicated in the bitstream, and the rules specify that the transformation operation includes, during the encoding operation, encoding / decoding the current video block into the bitstream by applying a forward transform to the residual value of the current video block, or the rules specify that the transformation operation includes, during the decoding operation, generating the current video block from the bitstream by applying an inverse transform to a scaling factor indicated in the bitstream.

[0596] In some embodiments, the rule specifies whether the bitstream includes an index indicating the transform matrix set based on the low-frequency inseparable transform index. In some embodiments, the rule specifies that, in response to the low-frequency inseparable transform index being equal to a specific value, the index indicating the transform matrix set is not included in the bitstream. In some embodiments, the specific value is equal to 1 or 2.

[0597] Figure 23This is a flowchart of example method 2300 for video processing. Operation 2302 includes a conversion between the current video block and the video bitstream, using rules to determine the zeroing range of the current video block. Operation 2304 includes performing the conversion according to the determination, wherein the rules specify the zeroing range based on the size of the current video block, and during the secondary transformation operation in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0598] In some embodiments, the rule specifies that in response to the current video block having a first size, the zeroing range is a first zeroing range; the rule specifies that in response to the current video block having a second size, the zeroing range is a second zeroing range, the second zeroing range being different from the first zeroing range, and the second size being different from the first size. In some embodiments, the rule specifies that in response to the current video block having a first size less than or equal to M×N, the zeroing range is a first zeroing range, where M and N are integers. In some embodiments, one or more transform coefficients of the current video block having a scan order index greater than or equal to the first zeroing range are equal to zero. In some embodiments, the rule specifies that in response to the current video block having a first size less than or equal to 8, the zeroing range is a first zeroing range. In some embodiments, the first size includes the width of the current video block, or the height of the current video block, or the width and height of the current video block.

[0599] In some embodiments, the rule specifies that the zeroing range is a second zeroing range in response to the current video block having a second size greater than M×N, where M and N are integers. In some embodiments, one or more transform coefficients of the current video block having a scan order index greater than or equal to the second zeroing range are equal to zero. In some embodiments, the rule specifies that the zeroing range is a second zeroing range in response to the current video block having a second size greater than 8. In some embodiments, the second size includes the width of the current video block, or the height of the current video block, or the width and height of the current video block. In some embodiments, the rule specifies that the zeroing range is a second zeroing range in response to: (1) the width of the current video block is greater than 8 and the height of the current video block is greater than or equal to 8, or (2) the width of the current video block is greater than or equal to 8 and the height of the current video block is greater than 8. In some embodiments, the second zeroing range is 16, and the first zeroing range is 8.

[0600] Figure 24 This is a flowchart of example method 2400 for video processing. Operation 2402 includes a conversion between the current video block and the video bitstream, determining whether to disable the zeroing operation for the current video block based on rules. Operation 2404 includes performing the conversion based on the determination, where the rules specify that the zeroing operation includes using a master transform, in which coefficients within a range are treated as having zero values.

[0601] In some embodiments, the master transform includes Discrete Cosine Transform (DCT)-II.

[0602] Figure 25 This is a flowchart of example method 2500 for video processing. Operation 2502 includes a conversion between the current video block and the video bitstream, determining whether the zeroing range has been expanded in the main transform of the current video block. Operation 2504 includes performing the conversion based on the determination, wherein, during the operation of the main transform in the conversion, the transform coefficients of the current video block in the zeroing range are treated as having zero values.

[0603] In some embodiments, the principal transform includes a multiple transform set (MTS). In some embodiments, the principal transform includes a discrete cosine transform (DCT)-II.

[0604] Figure 26 This is a flowchart of example method 2600 for video processing. Operation 2602 includes determining one or more master transform matrices from a master transform set according to rules for a transformation between the current video block and the video bitstream. Operation 2604 includes performing the transformation based on the determination, wherein the one or more master transform matrices include one or more of Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VII, transform skip mode, identity transform, and transformations based on the training process.

[0605] In some embodiments, the transformation based on the training process includes a transformation based on the Karhunen-Loeve transform (KLT). In some embodiments, a rule specifies that the determination of one or more master transform matrices is based on the chroma components of the video. In some embodiments, the transform set includes a first set of transform matrices for the luma components of the video and a second set of transform matrices for the chroma components of the video. In some embodiments, each color component of the video corresponds to a separate transform set. In some embodiments, a rule specifies that the determination of one or more master transform matrices is based on the size of the current video block. In some embodiments, two or more sets of master transform matrices are defined for different video block sizes. In some embodiments, a rule specifies that one or more master transform matrices from a first set are determined in response to the height and width of the current block being less than M, or a rule specifies that one or more master transform matrices from a second set are determined in response to the height and width of the current block being greater than or equal to M, and M is a positive integer. In some embodiments, a rule specifies that one or more master transform matrices from a first set are determined in response to the height or width of the current block being less than M, or a rule specifies that one or more master transform matrices from a second set are determined in response to the height or width of the current block being greater than or equal to M, and M is a positive integer. In some embodiments, M is equal to 8, 16, or 32.

[0606] In some embodiments, the rule specifies that one or more master transform matrices are determined based on a prediction mode for the current video block. In some embodiments, two different sets of master transform matrices are specified for intra-frame and inter-frame modes, respectively, where the prediction mode includes either an intra-frame or inter-frame mode. In some embodiments, two or more different sets of master transform matrices are specified for the intra-frame mode, where the prediction mode includes the intra-frame mode. In some embodiments, two or more different sets of master transform matrices are specified for the inter-frame mode, where the prediction mode includes the inter-frame mode. In some embodiments, the transform set includes transform matrices with a transform dimension of 4, 8, 16, or 32. In some embodiments, the rule specifies that one or more master transform matrices are determined based on an intra-frame prediction mode applied to the current video block. In some embodiments, the intra-frame prediction mode includes Intra-Segmentation Block (ISP), Matrix-Based Intra-Frame Prediction (MIP), Multiple Reference Line (MRL) methods, Cross-Component Linear Model (CCLM), or Block Differential Pulse Codec Modulation-Codec (BDPCM).

[0607] In some embodiments, two or more sets of master transform matrices are used for different intra-prediction modes. In some embodiments, a first set of transform matrices is used for a first set of video blocks on which CCLM is applied, and a second set of transform matrices is used for a second set of video blocks on which an intra-prediction mode different from CCLM is applied. In some embodiments, a first set of transform matrices is used for a first set of video blocks on which an intra-prediction mode different from MRL and MIP is applied, and a second set of transform matrices is used for a second set of video blocks on which MRL is applied. In some embodiments, a rule specifies that one or more master transform matrices are determined based on the inter-prediction mode applied to the current block. In some embodiments, the inter-prediction mode includes merge mode, advanced motion vector prediction (AMVP) mode, or affine mode. In some embodiments, two or more sets of master transform matrices are used for different inter-prediction modes. In some embodiments, a first set of transform matrices is used for a first set of video blocks on which merge mode or inter-prediction mode is applied, and a second set of transform matrices is used for a second set of video blocks on which inter-prediction mode but not merge mode is applied.

[0608] In some embodiments, a first set of transform matrices is used on a first set of video blocks to which AMVP mode or inter-frame prediction mode is applied, and a second set of transform matrices is used on a second set of video blocks to which inter-frame prediction mode is applied instead of AMVP mode. In some embodiments, the transform set includes a first transform set in response to the application of encoding / decoding tools to the current video block, and a second transform set in response to the application of encoding / decoding tools not to the current video block. In some embodiments, the encoding / decoding tools include Intra-Segmentation in Sub-Block (ISP), Matrix-Based Intra-Frame Prediction (MIP), Luminance Mapping with Chroma Scaling (LMCS), Adaptive Loop Filtering (ALF), Sample Adaptive Offset (SAO), Dependent Quantization (DQ), Affine, Symmetric Motion Vector Difference (SMVD), Sub-Block Transform (SBT), Adaptive Motion Vector Resolution (AMVR), Bidirectional Optical Flow (BDOF), Predictive Refinement with Optical Flow (PROF), Decoder-Side Motion Vector Refinement (DMVR), Low-Frequency Inseparable Transform (LFNST), Local Illumination Compensation (LIC), Overlapping Block Motion Compensation (OBMC), Joint Encoding / Decoding or Chroma Residual (JCCR) or Intra-Block Copy (IBC). In some embodiments, the rule specifies that the determination of one or more master transform matrices is based on the transform mode of the current video block. In some embodiments, the transform mode includes implicit multiple transform set (MTS), explicit MTS, sub-block transform (SBT) mode, joint codec or chroma residual (JCCR) or different low-frequency inseparable transform (LFNST) indices.

[0609] In some embodiments, two or more master transformation matrix sets are used for different transformation modes. In some embodiments, a first set of transformation matrices is used for implicit MTS, and a second set of transformation matrices is used for explicit MTS. In some embodiments, a first set of transformation matrices is used for SBT, and a second set of transformation matrices is used for transformation modes different from SBT. In some embodiments, SBT uses transformation matrices with a transformation dimension equal to 4, 8, 16, or 32. In some embodiments, a first set of transformation matrices is used for JCCR, and a second set of transformation matrices is used for transformation modes different from JCCR.

[0610] Figure 27 This is a flowchart of example method 2700 for video processing. Operation 2702 includes conversion between one or more video blocks and a video bitstream, determining a transform set for a primary or secondary transform tool according to rules, wherein the rules specify that the transform set is determined from two or more transform sets having specific characteristics. Operation 2704 includes performing the conversion based on the determination.

[0611] In some embodiments, two or more transform sets include a first transform set and a second transform set, and the second transform set includes one or more transforms as well as all transforms from the first transform set.

[0612] Figure 28 This is a flowchart of an example method 2800 for video processing. Operation 2802 includes a conversion between a current video block and the video bitstream, determining the zeroing range of the main transform for the current video block based on the prediction mode of the current video block. Operation 2804 includes performing the conversion based on the determination, wherein, during the operation of the main transform in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0613] In some embodiments, the zeroing range is a first range in response to the prediction mode being an intra-frame prediction mode, and the zeroing range is a second range in response to the prediction mode being an inter-frame prediction mode, and the first range is different from the second range.

[0614] Figure 29 This is a flowchart of example method 2900 for video processing. Operation 2902 includes a conversion between the current video block and the video bitstream, determining a zeroing range for the secondary transform of the current video block based on the number of samples in the current video block. Operation 2904 includes performing the conversion based on the determination, wherein, during the operation of the secondary transform in the conversion, the transform coefficients of the current video block within the zeroing range are treated as having zero values.

[0615] In some embodiments, in response to the current video block having a first number of samples, the zeroing range is a first zeroing range; in response to the current video block having a second number of samples, the zeroing range is a second zeroing range; the first zeroing range is different from the second zeroing range; the first number of samples is different from the second number of samples; and the current video block is a transform block. In some embodiments, the primary transform or secondary transform is implemented using butterfly computation. In some embodiments, a temporary video block for deriving the final reconstructed block of the current video block is generated by applying a forward primary transform to the prediction error before performing the quantization operation or secondary transform, using the primary transform during the encoding / decoding operation during the conversion, and applying an inverse primary transform during the decoding operation during the conversion.

[0616] In some embodiments, the quadratic transform is used during the encoding operation during conversion by applying a positive quadratic transform to the output of the forward master transform, the forward master transform being applied to the residual of the current video block before quantization, and the quadratic transform is used during the decoding operation during conversion by applying an inverse quadratic transform to the output of the dequantized current video block before applying the inverse master transform. In some embodiments, performing the conversion includes encoding the video into a bitstream. In some embodiments, performing the conversion includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments, performing the conversion includes decoding the video from the video.

[0617] In some embodiments, a video decoding apparatus includes a processor configured to implement one or more of methods 2200 to 2900. In some embodiments, a video encoding apparatus includes a processor configured to implement one or more of methods 2200 to 2900. In some embodiments, a computer program product having computer instructions stored thereon, which, when executed by a processor, cause the processor to implement one or more of methods 2200 to 2900. In some embodiments, a non-transitory computer-readable storage medium storing a bitstream generated according to one or more of methods 2200 to 2900. In some embodiments, a non-transitory computer-readable storage medium storing instructions that cause a processor to implement one or more of methods 2200 to 2900. In some embodiments, a method for generating a bitstream includes: generating a bitstream of video according to one or more of methods 2200 to 2900, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, apparatus, or bitstream generated according to the disclosed method or system described in this document.

[0618] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-occurring or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residual values ​​of the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0619] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on the decision or determination, the conversion from video blocks to a bitstream representation of the video will be performed using that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.

[0620] Some embodiments of the technology disclosed herein include deciding or determining to disable video processing tools or modes. In one example, when video processing tools or modes are disabled, the encoder will not use the tools or modes in the conversion of video blocks to a bitstream representation of the video. In another example, when video processing tools or modes are disabled, the decoder will process the bitstream using knowledge that the bitstream has not yet been modified based on the video processing tools or modes that have been decided or determined to be disabled.

[0621] Based on the foregoing, it will be understood that specific embodiments of the technology disclosed herein have been described for illustrative purposes, but various modifications may be made without departing from the scope of the invention. Accordingly, the technology disclosed herein is not limited to what is claimed in the appended claims.

[0622] The implementations and functional operations of the subject matter described in this patent document can be implemented in various systems, digital electronic circuits, or in computer software, firmware, or hardware, containing the structures disclosed in this specification and their equivalents, or combinations thereof. Implementations of the subject matter described in this specification can be implemented as one or more computer program products encoded on a tangible and non-volatile computer-readable medium, i.e., one or more computer program instruction modules, for execution by a data processing device or for controlling the operation of a data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable and propagable signals, or combinations thereof. The terms "data processing unit" or "data processing device" encompass all devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof.

[0623] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0624] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be performed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).

[0625] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data contain all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices). The processor and memory may be supplemented by dedicated logic circuitry or integrated into dedicated logic circuitry.

[0626] Intended to include the instruction manual and accessories Figure 1 The above is to be considered as exemplary only, where exemplary means example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.

[0627] While this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as descriptions of features specified in particular embodiments of a particular invention. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0628] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0629] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.

Claims

1. A video processing method, comprising: For the conversion between the current video block and the bitstream of the video, rules are used to determine the zeroing range of the current video block; as well as The conversion shall be performed based on the determination. The rule specifies that the zeroing range is based on the size of the current video block, and During the secondary transformation operation in the transformation, the transform coefficients of the current video block within the zeroing range are considered to have zero values. The rule stipulates that when the current video block has a first size, the zeroing range is a first zeroing range. The rule stipulates that when the current video block has a second size, the zeroing range is the second zeroing range. Wherein, the second zeroing range is different from the first zeroing range, and The second dimension is different from the first dimension. The rule specifies that the selection of the transform kernel used to perform the main transform operation during the conversion is based on the first index of the secondary transform indicated in the bitstream. Wherein, the second transformation is the Low Frequency Inseparable Transform (LFNST), and the first index is the index of the LFNST; The rule specifies whether the bitstream includes a second index, whereby the second index indicates that the selection of the transform kernel is based on the first index of the quadratic transform, and the second index is an index used for the Multi-Transform Selection System (MTS). The rule stipulates that when the first index is equal to a specific value, the second index is not included in the bitstream.

2. The method according to claim 1, wherein, When the zeroing range is the first zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the first zeroing range are equal to zero.

3. The method according to claim 1, wherein, When the zeroing range is the second zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the second zeroing range are equal to zero.

4. The method according to claim 1, wherein, The rule stipulates that when the current video block has a first size equal to 4×4 or equal to 8×8, the zeroing range is the first zeroing range.

5. The method according to claim 1, wherein, The rule specifies that the zeroing range is the second zeroing range when the current video block has the following second size: (1) The width of the current video block is greater than 8, and the height of the current video block is greater than or equal to 8, or (2) The width of the current video block is greater than or equal to 8, and the height of the current video block is greater than 8.

6. The method according to claim 1, wherein, The second zeroing range is 16, and the first zeroing range is 8.

7. The method according to claim 1, wherein, The rule specifies that the main transform operation includes, during the encoding operation, applying a forward main transform to the residual value of the current video block to encode and decode the current video block into the bitstream, or The rule specifies that the main transform operation includes generating the residual value of the current video block during the decoding operation by applying the inverse main transform to a scaling factor indicated in the bitstream.

8. The method according to claim 1, wherein, The specific value is equal to 1 or 2.

9. The method according to claim 1, wherein, The primary transformation or the secondary transformation is implemented using butterfly computation.

10. The method according to claim 1, in, The quadratic transform is applied to the output of the forward master transform, and then used during the encoding operation during the transformation. The forward master transform is applied to the residual of the current video block before quantization. The quadratic transform is used during the decoding operation during the conversion by applying the inverse quadratic transform to the inverse quantized output of the current video block before applying the inverse master transform.

11. The method according to claim 1, wherein, Performing the conversion includes encoding the video into the bitstream.

12. The method according to claim 1, wherein, Performing the conversion includes decoding the video from the bitstream.

13. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: For the conversion between the current video block and the bitstream of the video, rules are used to determine the zeroing range of the current video block; and The conversion shall be performed based on the determination. The rule specifies that the zeroing range is based on the size of the current video block. During the secondary transformation operation in the conversion, the transform coefficients of the current video block within the zeroing range are considered to have zero values. The rule stipulates that when the current video block has a first size, the zeroing range is a first zeroing range. The rule stipulates that when the current video block has a second size, the zeroing range is the second zeroing range. Wherein, the second zeroing range differs from the first zeroing range, and the second size differs from the first size. The rule specifies that the selection of the transform kernel used to perform the main transform operation during the conversion is based on the first index of the secondary transform indicated in the bitstream. Wherein, the second transformation is the Low Frequency Inseparable Transform (LFNST), and the first index is the index of the LFNST; The rule specifies whether the bitstream includes a second index, whereby the second index indicates that the selection of the transform kernel is based on the first index of the quadratic transform, and the second index is an index used for the Multi-Transform Selection System (MTS). The rule stipulates that when the first index is equal to a specific value, the second index is not included in the bitstream.

14. The apparatus according to claim 13, wherein, When the zeroing range is the first zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the first zeroing range are equal to zero. When the zeroing range is the second zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the second zeroing range are equal to zero. The rule stipulates that when the current video block has a first size equal to 4×4 or equal to 8×8, the zeroing range is the first zeroing range, and The rule stipulates that the zeroing range is the second zeroing range when the current video block has the following second dimensions: (1) The width of the current video block is greater than 8, and the height of the current video block is greater than or equal to 8, or (2) The width of the current video block is greater than or equal to 8, and the height of the current video block is greater than 8. The second zeroing range is 16, and the first zeroing range is 8.

15. The apparatus according to claim 13, wherein, The rule specifies that the main transform operation includes, during the encoding operation, encoding and decoding the current video block into the bitstream by applying a forward main transform to the residual value of the current video block, or the rule specifies that the main transform operation includes, during the decoding operation, generating the residual value of the current video block by applying an inverse main transform to a scaling factor indicated in the bitstream.

16. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: For the conversion between the current video block and the bitstream of the video, rules are used to determine the zeroing range of the current video block; and The conversion shall be performed based on the determination. in, The rule specifies that the zeroing range is based on the size of the current video block. During the secondary transformation operation in the conversion, the transform coefficients of the current video block within the zeroing range are considered to have zero values. The rule stipulates that when the current video block has a first size, the zeroing range is a first zeroing range. The rule stipulates that when the current video block has a second size, the zeroing range is the second zeroing range. Wherein, the second zeroing range differs from the first zeroing range, and the second size differs from the first size. The rule specifies that the selection of the transform kernel used to perform the main transform operation during the conversion is based on the first index of the secondary transform indicated in the bitstream. Wherein, the second transformation is the Low Frequency Inseparable Transform (LFNST), and the first index is the index of the LFNST; The rule specifies whether the bitstream includes a second index, whereby the second index indicates that the selection of the transform kernel is based on the first index of the quadratic transform, and the second index is an index used for the Multi-Transform Selection System (MTS). The rule stipulates that when the first index is equal to a specific value, the second index is not included in the bitstream.

17. The non-transitory computer-readable storage medium according to claim 16, wherein, When the zeroing range is the first zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the first zeroing range are equal to zero. When the zeroing range is the second zeroing range, one or more transform coefficients of the scan sequence index of the current video block that are greater than or equal to the second zeroing range are equal to zero. The rule stipulates that when the current video block has a first size equal to 4×4 or equal to 8×8, the zeroing range is the first zeroing range, and The rule stipulates that the zeroing range is the second zeroing range when the current video block has the following second dimensions: (1) The width of the current video block is greater than 8, and the height of the current video block is greater than or equal to 8, or (2) The width of the current video block is greater than or equal to 8, and the height of the current video block is greater than 8. The second zeroing range is 16, and the first zeroing range is 8.

18. A non-transitory computer-readable recording medium storing a bitstream of video, wherein a computer program is also stored thereon, When the computer program is executed by a processor, it performs the following steps to generate the bitstream: For the current video block of the video, use rules to determine the zeroing range of the current video block; and The bit stream is generated based on the determination. The rule specifies that the zeroing range is based on the size of the current video block. During the secondary transformation operation in the generation process, the transform coefficients of the current video block within the zeroing range are considered to have zero values. The rule stipulates that when the current video block has a first size, the zeroing range is a first zeroing range. The rule stipulates that when the current video block has a second size, the zeroing range is the second zeroing range. Wherein, the second zeroing range differs from the first zeroing range, and the second size differs from the first size. The rule specifies that the selection of the transform kernel used to perform the main transform operation during the generation period is based on the first index of the secondary transform indicated in the bitstream. Wherein, the second transformation is the Low Frequency Inseparable Transform (LFNST), and the first index is the index of the LFNST; The rule specifies whether the bitstream includes a second index, whereby the second index indicates that the selection of the transform kernel is based on the first index of the quadratic transform, and the second index is an index used for the Multi-Transform Selection System (MTS). The rule stipulates that when the first index is equal to a specific value, the second index is not included in the bitstream.

19. The non-transitory computer-readable recording medium according to claim 18, wherein, When the zeroing range is the first zeroing range, one or more transform coefficients of the current video block having a scan order index greater than or equal to the first zeroing range are equal to zero. When the zeroing range is the second zeroing range, one or more transform coefficients of the scan sequence index of the current video block that are greater than or equal to the second zeroing range are equal to zero. The rule stipulates that when the current video block has a first size equal to 4×4 or equal to 8×8, the zeroing range is the first zeroing range, and The rule stipulates that the zeroing range is the second zeroing range when the current video block has the following second dimensions: (1) The width of the current video block is greater than 8, and the height of the current video block is greater than or equal to 8, or (2) The width of the current video block is greater than or equal to 8, and the height of the current video block is greater than 8. The second zeroing range is 16, and the first zeroing range is 8.

20. A method for storing a bitstream of video, comprising: For the current video block of the video, use rules to determine the zeroing range of the current video block; and The bit stream is generated based on the determination. The bitstream is stored in a non-transitory computer-readable recording medium. in, The rule specifies that the zeroing range is based on the size of the current video block. During the secondary transformation operation in the generation process, the transform coefficients of the current video block within the zeroing range are considered to have zero values. The rule stipulates that when the current video block has a first size, the zeroing range is a first zeroing range. The rule stipulates that when the current video block has a second size, the zeroing range is the second zeroing range. Wherein, the second zeroing range differs from the first zeroing range, and the second size differs from the first size. The rule specifies that the selection of the transform kernel used to perform the main transform operation during the generation period is based on the first index of the secondary transform indicated in the bitstream. Wherein, the second transformation is the Low Frequency Inseparable Transform (LFNST), and the first index is the index of the LFNST; The rule specifies whether the bitstream includes a second index, whereby the second index indicates that the selection of the transform kernel is based on the first index of the quadratic transform, and the second index is an index used for the Multi-Transform Selection System (MTS). The rule stipulates that when the first index is equal to a specific value, the second index is not included in the bitstream.

21. A video decoding apparatus comprising a processor configured to implement the method of any one of claims 1 to 10 and 12.

22. A video encoding apparatus comprising a processor configured to implement the method of any one of claims 1 to 11.

23. A non-transitory computer-readable storage medium storing a bitstream of video, wherein a computer program is also stored thereon, When the computer program is executed by a processor, it implements the method according to any one of claims 2 to 11 to generate the bit stream.

24. A non-transitory computer-readable storage medium for storing instructions that cause a processor to implement the method of any one of claims 2 to 12.

25. A bitstream generation method, comprising: The method according to any one of claims 1 to 11 generates a bitstream of video, and The bit stream is stored on a non-transitory computer-readable recording medium.

Citation Information

Patent Citations

  • Transformation and quadratic transformation matrix training method, encoder and related device

    CN110636313A