Implicit Multiple Transform Set Signaling Notification in Video Coding and Decoding
By determining the use of the identity transformation mode based on the decoding coefficient relationship between the video block and the representative block, or enabling the zeroing operation at the video area level, the problem of inefficient encoding and decoding of video blocks in the prior art is solved, and more efficient video processing is achieved.
Patent Information
- Application Number
- CN202180020117.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-07
- Filing Date
- 2021-03-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-03-08
AI Technical Summary
When existing video encoding and decoding technologies process video blocks, it is difficult to effectively determine the use of the identity transformation mode, resulting in low encoding and decoding efficiency.
Optimize the encoding and decoding process of the video block by determining whether to apply horizontal or vertical identity transformation based on the decoding coefficient relationship between the video block and the representative block, or enable the zeroing operation at the video region level.
It improves the encoding and decoding efficiency of video blocks, reduces the computational complexity and storage requirements, and enhances the performance of video processing.
Smart Images

Figure CN115315944B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority and benefit of International Patent Application No. PCT / CN2020 / 078334, filed on March 7, 2020. The entire disclosure of the above application is incorporated by reference into a part of the disclosure of this application. Technical field
[0003] This patent document relates to image encoding and decoding, as well as video encoding and decoding. Background art
[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. With the increase in the number of connected user devices capable of receiving and displaying video, it is expected that the bandwidth demand for digital video usage will continue to grow. Summary of the invention
[0005] This document discloses techniques that can be used by a video encoder and a decoder, which use control information useful for decoding the encoded - decoded representation to process the encoded - decoded representation of video.
[0006] In one exemplary aspect, a video processing method is disclosed. For the conversion between a current video block of a video and the bitstream of the video, an identity transform mode for the conversion of the current video block is determined according to a rule, and the rule stipulates that the use is based on representative coefficients of one or more representative blocks of the video. The method further includes performing the conversion based on the determination.
[0007] In another exemplary aspect, a video processing method is disclosed. The method includes the conversion between a current video block of a video and the bitstream of the video, and determines a default transform applicable to the current video block according to a rule, and the rule stipulates that the identity transform is not used for the conversion of the current video block. The method further includes performing the conversion based on the determination.
[0008] In another exemplary aspect, a video processing method is disclosed. The method includes performing the conversion between a video and the bitstream of the video according to a rule. The rule stipulates that an indication is included at the video region level. The indication indicates whether a zeroing operation that sets some residual coefficients to zero is applied to the transform blocks of video blocks in the video region.
[0009] In another exemplary aspect, a video processing method is disclosed. The method includes performing the conversion between a current video block of a video and the bitstream of the video according to a rule. An identity transform mode is applied to the current video block during the conversion, and the rule stipulates that a zeroing operation is enabled, and during the zeroing operation, non - zero coefficients are restricted within a sub - region of the current video block.
[0010] In another exemplary aspect, a video processing method is disclosed. The method includes: converting a current video block of a video with a bitstream of the video to determine a zeroing type of the current video block for a zeroing operation. The method further includes performing the conversion based on the determination. The current video block is encoded and decoded by applying an identity transform to the current video block. The zeroing type of the video block defines a sub-region of the video block within which non-zero coefficients are restricted for the zeroing operation.
[0011] In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a current video block of a video and a bitstream of the video according to a rule. The rule stipulates that in a case where at least one non-zero coefficient is outside a zeroing region determined by an identity transform mode, using the identity transform mode to convert the current video block is prohibited. The zeroing region includes a region where non-zero coefficients are restricted for the zeroing operation.
[0012] In another exemplary aspect, a video processing method is disclosed. The method includes: converting a video block of a video with an encoded and decoded representation of the video, determining whether a horizontal identity transform or a vertical identity transform is applied to the video block based on a rule; and performing the conversion based on the determination. The rule stipulates a relationship between the determination and representative coefficients of decoded coefficients of one or more representative blocks from the video.
[0013] In another exemplary aspect, another video processing method is disclosed. The method includes: converting a video block of a video with an encoded and decoded representation of the video, determining whether a horizontal identity transform or a vertical identity transform is applied to the video block based on a rule; and performing the conversion based on the determination. The rule stipulates a relationship between the determination and decoded luminance coefficients of the video block.
[0014] In another exemplary aspect, another video processing method is disclosed. The method includes: converting a video block of a video with an encoded and decoded representation of the video, determining whether a horizontal identity transform or a vertical identity transform is applied to the video block based on a rule; and performing the conversion based on the determination. The rule stipulates a relationship between the determination and a value V, where the value V is associated with decoded coefficients or representative coefficients of a representative block.
[0015] In another exemplary aspect, another video processing method is disclosed. The method includes determining that one or more syntax fields are present in an encoded and decoded representation of a video, where the video includes one or more video blocks; and based on the one or more syntax fields, determining whether a horizontal identity transform or a vertical identity transform is enabled for the video blocks in the video.
[0016] In another example aspect, another video processing method is disclosed. The method includes making a first determination as to whether to enable the use of an identity transform for the conversion between a video block of a video and an encoded / decoded representation of the video; making a second determination as to whether to enable a zeroing operation during the conversion; and performing the conversion based on the first determination and the second determination.
[0017] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video block of a video and an encoded / decoded representation of the video; wherein the video block is represented as an encoded block in the encoded / decoded representation, wherein non-zero coefficients of the encoded block are restricted within one or more sub-regions; and wherein an identity transform is applied to generate the encoded block.
[0018] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0019] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0020] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0021] These features and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 An example video encoder block diagram is shown.
[0023] Figure 2 An example of 67 intra prediction modes is shown.
[0024] Figure 3A An example of reference samples for wide-angle intra prediction is shown.
[0025] Figure 3B Another example of reference samples for wide-angle intra prediction is shown.
[0026] Figure 4 The discontinuity problem when the direction exceeds 45 degrees is shown.
[0027] Figure 5A An example definition of samples used by PDPC applied to diagonal intra mode and adjacent angle intra mode is shown.
[0028] Figure 5B Another example definition of samples used by PDPC applied to diagonal intra mode and adjacent angle intra mode is shown.
[0029] Figure 5C Another example definition of samples used by PDPC applied to the intra diagonal mode and the adjacent angular intra mode is shown.
[0030] Figure 5D Yet another example definition of samples used by PDPC applied to the intra diagonal mode and the adjacent angular intra mode is shown.
[0031] Figure 6 An example of the partitioning of 4×8 blocks and 8×4 blocks is shown.
[0032] Figure 7 An example of the partitioning of all blocks other than 4×8, 8×4, and 4×4 is shown.
[0033] Figure 8 An example of the secondary transform in JEM is shown.
[0034] Figure 9 An example of the simplified secondary transform LFNST is shown.
[0035] Figure 10A An example of the positive simplified transform is shown.
[0036] Figure 10B An example of the inverse simplified transform is shown.
[0037] Figure 11 An example of the positive LFNST8×8 process with a 16×48 matrix is shown.
[0038] Figure 12 An example of the scan positions 17 to 64 of non-zero elements is shown.
[0039] Figure 13 Examples of the sub-block transform modes SBT-V and SBT-H are shown.
[0040] Figure 14A An example of scan region based coefficient coding (SRCC) is shown.
[0041] Figure 14B Another example of scan region based coefficient coding (SRCC) is shown.
[0042] Figure 15A An example of the IST limitation according to the non-zero coefficient position is shown.
[0043] Figure 15B Another example of the IST limitation according to the non-zero coefficient position is shown.
[0044] Figure 16A An example of the TS coding / decoding block of the zeroing type is shown.
[0045] Figure 16B Shows another example of a TS encoding / decoding block of the zeroing type.
[0046] Figure 16C Shows another example of a TS encoding / decoding block of the zeroing type.
[0047] Figure 16D Shows another TS encoding / decoding block of the zeroing type.
[0048] Figure 17 Is a block diagram of an example video processing system.
[0049] Figure 18 Is a block diagram showing a video encoding / decoding system according to some embodiments of the present disclosure.
[0050] Figure 19 Is a block diagram showing an encoder according to some embodiments of the present disclosure.
[0051] Figure 20 Is a block diagram showing a decoder according to some embodiments of the present disclosure.
[0052] Figure 21 Is a block diagram of a video processing apparatus.
[0053] Figure 22 Is a flowchart of an example method of video processing.
[0054] Figure 23 Is a flowchart representation of a video processing method according to the present technology.
[0055] Figure 24 Is a flowchart representation of another video processing method according to the present technology.
[0056] Figure 25 Is a flowchart representation of another video processing method according to the present technology.
[0057] Figure 26 Is a flowchart representation of another video processing method according to the present technology.
[0058] Figure 27 Is a flowchart representation of another video processing method according to the present technology.
[0059] Figure 28 Is a flowchart representation of another video processing method according to the present technology.
[0060] Figure 29 Is a flowchart representation of another video processing method according to the present technology. Detailed implementation
[0061] The use of section headings in this document is for ease of understanding and does not limit the technologies and embodiments disclosed in each section to only that section. Additionally, the use of the H.266 term in some of the descriptions is for ease of understanding and not to limit the scope of the disclosed technologies. Thus, the technologies described herein also apply to other video codec protocols and designs.
[0062] 1. Overview
[0063] This document relates to video coding and decoding technologies. Specifically, it relates to transform skip modes and transform types (e.g., including the identity transform) in video coding and decoding. It can be applied to existing video coding standards (e.g., HEVC), or standards that are near completion (Versatile Video Coding). It can also be applied to future video coding standards or video codecs.
[0064] 2. Preliminary Discussion
[0065] Video coding and decoding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding and decoding standards have been based on a hybrid video coding and decoding structure that utilizes temporal prediction plus transform coding. To explore future video coding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the Versatile Video Coding (VVC) standard, with the goal of reducing the bitrate by 50% compared to HEVC.
[0066] 2.1. Coding and Decoding Processes of Typical Video Codecs
[0067] Figure 1An example of the encoder block diagram of VVC is shown, which includes three in-loop filtering blocks: the Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture, signal the side information for encoding and decoding of the offset and filter coefficients, and reduce the mean squared error between the original samples and the reconstructed samples by adding the offset and by applying a Finite Impulse Response (FIR) filter, respectively. ALF is located in the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts created in the previous stages.
[0068] 2.2. Intra Mode Coding and Decoding with 67 Intra Prediction Modes
[0069] To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The additional directional modes are depicted by the dashed arrows in Figure 2 and the planar mode and DC mode remain unchanged. These dense directional intra prediction modes are applicable to all block sizes as well as for luma and chroma intra prediction.
[0070] The traditional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction, as shown in Figure 2 . In VTM2, for non-square blocks, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled using the original method and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0071] In HEVC, each intra-coded block has a square shape with a side length that is a power of 2. Therefore, no division operation is required to generate the intra prediction value using the DC mode. In VVV2, the blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.
[0072] 2.3. Wide-Angle Intra Prediction for Non-Square Blocks
[0073] The traditional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VTM2, for non-square blocks, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled using the original method and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes for a specific block remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0074] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined as shown in Figures 3A to 3B as follows.
[0075] The number of modes of the replacement mode in the wide-angle direction mode depends on the aspect ratio of the block. The intra prediction modes of the replacement are shown in Table 1.
[0076] Table 1: Intra Prediction Modes Replaced by Wide-Angle Mode
[0077]
[0078]
[0079] As Figure 4 shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, low-pass reference sample filtering and side smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap Δp α .
[0080] 2.4. Position-Dependent Intra Prediction Combination
[0081] In VTM2, the result of intra prediction of the planar mode is further modified by the position dependent intraprediction combination (PDPC) method. PDPC is an intra prediction method that calls a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, bottom-left angular mode and its eight adjacent angular modes, and top-right angular mode and its eight adjacent angular modes.
[0082] Using a linear combination of the intra prediction modes (DC, planar, angular) and reference samples, the prediction sample pred(x,y) is predicted according to the following equation:
[0083] pred(x,y) = (wL × R -1,y + wT × R x,-1 – wTL × R -1,-1 +(64 – wL – wT + wTL) × pred(x,y)+32) >> 6
[0084] where R x,-1 , R -1,y represent the reference samples located above and to the left of the current sample (x,y), respectively, and R -1,-1 represents the reference sample located at the upper left corner of the current block.
[0085] If PDPC is applied to DC intra mode, planar intra mode, horizontal intra mode, and vertical intra mode, no additional boundary filtering is required, while it is required in the case of HEVC DC mode boundary filtering or horizontal / vertical mode edge filtering.
[0086] Figures 5A to 5D The reference samples (R x,-1 , R -1,y and R -1,-1 ) to which PDPC is applied for various prediction modes are defined. The predicted sample pred(x’, y’) is located at (x’, y’) within the prediction block. The coordinate x of the reference sample R x,-1 is given by: x = x’ + y’ + 1, and the coordinate y of the reference sample R -1,y is similarly given by: y = x’ + y’ + 1. Figure 5A The diagonal upper right mode is shown. Figure 5B The diagonal lower left mode is shown. Figure 5C The adjacent diagonal upper right mode is shown. Figure 5D An adjacent diagonal lower left mode is shown.
[0087] The PDPC weights depend on the prediction mode, as shown in Table 2.
[0088] Table 2: Examples of PDPC weights according to prediction mode
[0089]
[0090]
[0091] 2.5. Intra sub-block partitioning (ISP)
[0092] In some embodiments, ISP is proposed, and ISP vertically or horizontally divides the luma intra prediction block into 2 sub-partitions or 4 sub-partitions according to the block size dimension, as shown in Table 3. Figure 6 and Figure 7 show examples of two possibilities. All sub-partitions satisfy the condition of having at least 16 samples.
[0093] Table 3: The number of sub-partitions depends on the block size.
[0094] Block size Number of sub - divisions 4×4 Not divided 4×8 and 8×4 2 All other cases 4
[0095] For each of these sub - partitions, a residual signal is generated by entropy - decoding the coefficients transmitted by the encoder and then inverse - quantizing and inverse - transforming them. Then, the sub - partition is intra - predicted, and finally, the corresponding reconstructed samples are obtained by adding the residual signal to the predicted signal. Thus, the reconstructed values of each sub - partition will be available for generating the prediction of the next one, and this process will be repeated and so on. All sub - partitions share the same intra - mode.
[0096] Based on the intra - mode and the utilized partitioning, two different classes of processing orders are used, which are referred to as the normal order and the reverse order. In the normal order, the first sub - partition to be processed is the one containing the top - left sample of the CU, and then it continues down (for horizontal partitioning) or to the right (for vertical partitioning). As a result, the reference samples used to generate the sub - partition prediction signals are only located to the left and above these lines. On the other hand, the reverse processing order starts from the sub - partition containing the bottom - left sample of the CU and continues up, or starts from the sub - partition containing the top - right sample of the CU and continues to the left.
[0097] 2.6. Multiple Transform Set (MTS)
[0098] In addition to DCT - II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding of both inter - coded and intra - coded blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST - VII and DCT - VIII. Table 4 shows the basis functions of the selected DST / DCT.
[0099] Table 4: Transform Types and Basis Functions
[0100]
[0101] There are two ways to enable MTS, one is explicit MTS; the other is implicit MTS.
[0102] 2.6.1. Implicit MTS
[0103] Implicit MTS is a new tool in VVC. The derivation of the variable implicitMtsEnabled is as follows:
[0104] Whether to enable implicit MTS depends on the value of the variable implicitMtsEnabled. The derivation of the variable implicitMtsEnabled is as follows:
[0105] – If sps_mts_enabled_flag is equal to 1 and one or more of the following conditions are true, then implicitMtsEnabled is set to be equal to 1:
[0106] – IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT (i.e., ISP is enabled).
[0107] – cu_sbt_flag is equal to 1 (i.e., ISP is enabled), and Max(nTbW, nTbH) is less than or equal to 32
[0108] – sps_explicit_mts_intra_enabled_flag is equal to 0 (i.e., explicit MTS is disabled), CuPredMode[0][xTbY][yTbY] is equal to MODE_INTRA, and lfnst_idx[x0][y0] is equal to 0, and intra_mip_flag[x0][y0] is equal to 0
[0109] – Otherwise, implicitMtsEnabled is set to be equal to 0.
[0110] The derivation of the variable trTypeHor that specifies the horizontal transform kernel and the variable trTypeVer that specifies the vertical transform kernel is as follows:
[0111] – If one or more of the following conditions are true, then set trTypeHor and trTypeVer to be equal to 0 (e.g., DCT2).
[0112] – cIdx is greater than 0 (i.e., for chrominance components)
[0113] – IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT and lfnst_idx is not equal to 0
[0114] – Otherwise, if implicitMtsEnabled is equal to 1, then the following applies:
[0115] – If cu_sbt_flag is equal to 1, then trTypeHor and trTypeVer are specified in Table 40 according to cu_sbt_horizontal_flag and cu_sbt_pos_flag.
[0116] – Otherwise (cu_sbt_flag is equal to 0), the derivation of trTypeHor and trTypeVer is as follows:
[0117] trTypeHor = (nTbW >= 4 && nTbW <= 16)? 1 : 0 (1188)
[0118] trTypeVer = (nTbH >= 4 && nTbH <= 16)? 1 : 0 (1189)
[0119] – Otherwise, trTypeHor and trTypeVer are specified in Table 39 according to mts_idx.
[0120] The derivation of variables nonZeroW and nonZeroH is as follows:
[0121] – If ApplyLfnstFlag equals 1, nTbW is greater than or equal to 4, and nTbH is greater than or equal to 4, the following conditions apply:
[0122] nonZeroW = (nTbW == 4 || nTbH == 4)? 4 : 8 (1190)
[0123] nonZeroH = (nTbW == 4 || nTbH == 4)? 4 : 8 (1191)
[0124] – Otherwise, the following applies:
[0125] nonZeroW = Min(nTbW, (trTypeHor > 0)? 16 : 32) (1192)
[0126] nonZeroH = Min(nTbH, (trTypeVer > 0)? 16 : 32) (1193)
[0127] 2.6.2. Explicit MTS
[0128] To control the MTS scheme, a flag is used to specify whether there is an explicit MTS for intra / inter in the bitstream. In addition, two separate enable flags are specified at the SPS level for intra and inter respectively to indicate whether the explicit MTS is enabled. When MTS is enabled at the SPS, the CU-level transform index can be signaled to indicate whether MTS is applied. Here, MTS only applies to luminance. When the following conditions are met, the MTS CU-level index (represented by mts_idx) is signaled.
[0129] - Both width and height are less than or equal to 32
[0130] - CBF luminance flag equals one
[0131] - Non-TS
[0132] - Non-ISP
[0133] - Non-SBT
[0134] - LFNST is disabled
[0135] - There are non-zero coefficients that are not at the DC position (the upper-left position of the block).
[0136] - There are no non-zero coefficients outside the upper-left 16×16 region.
[0137] If the first binary bit of mts_idx is equal to zero, DCT2 is applicable in both directions. However, if the first binary bit of mts_idx is equal to one, two other binary bits are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 5. For the transform matrix precision, an 8-bit primary transform kernel is used. Therefore, all the transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point DCT-2, 16-point DCT-2, and 32-point DCT-2. In addition, other transform kernels including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8 use an 8-bit primary transform kernel.
[0138] Table 5: Signaling of MTS
[0139]
[0140] To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 blocks and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16×16 low-frequency region are retained.
[0141] As in HEVC, the residual of a block can be coded and decoded in the transform skip mode. To avoid redundancy in syntax coding and decoding, when the MTS_CU_flag at the CU level is not equal to zero, the transform skip flag is not signaled. The block size limit for transform skip is the same as that of MTS in JEM4, which indicates that transform skip is applicable to a CU when both the block width and block height are equal to or less than 32.
[0142] 2.6.3. Zeroing in MTS
[0143] In VTM8, large block size transforms with a size up to 64×64 are enabled, which is mainly applicable to higher-resolution videos, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) not less than 64, the high-frequency transform coefficients of the blocks to which the DCT2 transform is applied are set to zero, so that only the low-frequency coefficients are retained, and all other coefficients are forced to zero without being signaled. For example, for an M×N transform block, where M is the block width and N is the block height, when M is not less than 64, only the left 32 columns of the transform coefficients are retained. Similarly, when N is not less than 64, only the first 32 rows of the transform coefficients are retained.
[0144] For transform blocks with dimensions (width or height, or both width and height) not less than 32, the high-frequency transform coefficients of the blocks to which DCT8 or DST7 transforms are applied are set to zero, so that only the low-frequency coefficients are retained, and all other coefficients are forced to zero without being notified. For example, for a transform block of M×N, where M is the block width and N is the block height, when M is not less than 32, only the left 16 columns of the transform coefficients are retained. Similarly, when N is not less than 32, only the first 16 rows of the transform coefficients are retained.
[0145] 2.7. Low frequency non-separable secondary transform (LFNST)
[0146] 2.7.1. JEM non-separable secondary transform (NSST)
[0147] In JEM, a secondary transform is applied between the forward primary transform and quantization (at the encoder) and between the inverse quantization and the inverse primary transform (at the decoder side). As Figure 8 shown, 4×4 (or 8×8) secondary transforms are performed according to the block size. For example, for each 8×8 block, a 4×4 secondary transform is applied to small blocks (e.g., min(width,height)<8), and an 8×8 secondary transform is applied to larger blocks (e.g., min(width,height)>4).
[0148] The application of the non-separable transform is described below using the input as an example. To apply the non-separable transform, the 4×4 input block X
[0149]
[0150] is first represented as a vector
[0151]
[0152] The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. Subsequently, the 16x1 coefficient vector Reorganized into 4x4 blocks. Coefficients with smaller indices are placed in the 4x4 coefficient block together with smaller scan indices. There are a total of 35 transform sets, and each transform set uses 3 non-separable transform matrices (kernels). The mapping from the intra prediction mode to the transform set is predefined. For each transform set, the selected non-separable quadratic transform candidate is further specified by a quadrature transform index signaled explicitly. After the transform coefficients, this index is signaled once per frame per CU in the bitstream.
[0153] 2.7.2. Reduced Secondary Transform (LFNST)
[0154] In some embodiments, LFNST is introduced and a mapping using 4 transform sets (instead of 35 transform sets) is used. In some implementations, 16×64 (which can be further reduced to 16×48) matrices and 16×16 matrices are used for 8×8 blocks and 4×4 blocks respectively. For ease of annotation, the 16×64 (which can be further reduced to 16×48) transform is denoted as LFNST8×8, and the 16×16 transform is denoted as LFNST4×4. Figure 9 An example of LFNST is shown.
[0155] LFNST calculation
[0156] The main idea of the reduced transform (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N (R < N) is the reduction factor.
[0157] The RT matrix is an R×N matrix as follows:
[0158]
[0159] where the R rows of the transform are the R bases of the N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. The forward RT and inverse RT are as Figure 10A and Figure 10B depicted.
[0160] In this proposal, the LFNST 8×8 with a reduction factor of 4 (1 / 4 size) is applied. Thus, instead of 64×64, a 16×64 direct matrix is used, which is the size of the traditional 8×8 non-separable transform matrix. In other words, a 64×16 inverse LFNST matrix is used on the decoder side to generate the core (primary) transform coefficients in the upper-left 8×8 region. The positive LFNST 8×8 uses a 16×64 (or 8×64 for an 8×8 block) matrix, such that it generates non-zero coefficients only in the upper-left 4×4 region within a given 8×8 region. In other words, if LFNST is applied, the 8×8 region except for the upper-left 4×4 region will have only zero coefficients. For LFNST 4×4, a 16×16 (or 8×16 for a 4×4 block) direct matrix multiplication is applied.
[0161] The inverse LFNST is conditionally applied when the following two conditions are met:
[0162] a. The block size is greater than or equal to a given threshold (W >= 4 && H >= 4)
[0163] b. The transform skip mode flag is equal to zero
[0164] If both the width (W) and height (H) of the transform coefficient block are greater than 4, the LFNST 8x8 is applied to the upper-left 8×8 region of the transform coefficient block. Otherwise, the LFNST 4x4 is applied to the upper-left min(8, W)×min(8, H) region of the transform coefficient block.
[0165] If the LFNST index is equal to 0, the LFNST is not applied. Otherwise, the LFNST is applied, and its kernel is selected together with the LFNST index. The LFNST selection method and the encoding / decoding of the LFNST index will be explained later.
[0166] In addition, the LFNST is applied to the intra CUs in both intra and inter strips, as well as to luminance and chrominance. If the dual-tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter strips (dual-tree is disabled), a single LFNST index is signaled and used for both luminance and chrominance.
[0167] At the 13th JVET conference, intra sub-partitioning (ISP) was adopted as a new intra prediction mode. When the ISP mode is selected, the LFNST is disabled, and the LFNST index is not signaled, because even if the LFNST is applied to each feasible partition block, the performance improvement is limited. In addition, disabling the LFNST for the residuals of ISP prediction can reduce the encoding complexity.
[0168] LFNST selection
[0169] The LFNST matrix is selected from four transform sets, each transform set consisting of two transforms. Which transform set to apply is determined by the intra prediction mode as follows:
[0170] 1) If one of the three CCLM modes is indicated, transform set 0 is selected.
[0171] 2) Otherwise, the transform set selection is performed according to Table 6.
[0172] Table 6: Transform Set Selection Table
[0173] IntraPredMode Tr. set index IntraPredMode<0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode 1
[0174] The index for accessing the table, denoted as IntraPredMode, ranges from [-14, 83], which is the transform mode index for wide-angle intra prediction.
[0175] Reduction LFNST matrix of dimensions
[0176] As a further simplification, a 16×48 matrix is applied instead of a 16×64 with the same transform set configuration. Each matrix obtains 48 input data from three 4×4 blocks excluding the right-bottom 4×4 block from the left-top 8×8 block ( Figure 11 ).
[0177] LFNST signaling
[0178] The positive LFNST8×8 with R = 16 uses a 16×64 matrix, so it only produces non-zero coefficients in the upper-left 4×4 area within a given 8×8 area. In other words, if LFNST is applied, the 8×8 area only produces zero coefficients except for the upper-left 4×4 area. Therefore, when any non-zero element is detected in the 8×8 block area other than the upper-left 4×4 (as shown in Figure 12 ), the LFNST index is not encoded or decoded because this means that LFNST is not applied. In this case, the LFNST index is inferred to be zero.
[0179] Zeroing range
[0180]
[0181]
[0182] Let nonZeroSize be a variable. It is required that any coefficient with an index not less than nonZeroSize must be zero when rearranging any coefficient to a 1-D array before the inverse LFNST.
[0182] When nonZeroSize is equal to 16, the coefficients in the left-top 4×4 sub-block have no zeroing constraint.
[0183] In some examples, when the current block size is 4×4 or 8×8, nonZeroSize is set to be equal to 8. For other block sizes, nonZeroSize is set to be equal to 16.
[0184] 2.8. Affine linear weighted intra prediction (ALWIP, also known as matrix-based intra prediction)
[0185] In some embodiments, affine linear weighted intra prediction (ALWIP, also known as matrix-based intra prediction (MIP)) is used.
[0186] In some embodiments, two tests are conducted. In Test 1, ALWIP is designed to have a memory limit of 8K bytes and at most 4 multiplications per sample point. Test 2 is similar to Test 1, but further simplifies the design in terms of memory requirements and model architecture.
[0187] · A single set of matrices and offset vectors for all block shapes.
[0188] · The number of modes for all block shapes is reduced to 19.
[0189] · The memory requirement is reduced to 5760 10-bit values, i.e., 7.20 kilobytes.
[0190] · The linear interpolation of the predicted sample points is performed in a single step in each direction, instead of the iterative interpolation in the first test.
[0191] 2.9. Sub-block transform
[0192] For an inter-predicted CU with cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether to decode the entire residual block or a sub-part of the residual block. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded with an inferred adaptive transform, and another part of the residual block is set to zero. SBT is not applied to combined inter-intra modes.
[0193] In the sub-block transform, position-dependent transforms are applied to the luminance transform blocks in SBT-V and SBT-H (chrominance TBs always use DCT-2). Two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are in Figure 13It is stipulated in [reference]. For example, the horizontal and vertical transforms at position 0 of SBT-V are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the corresponding transform is set to DCT-2. Therefore, the sub-block transform jointly stipulates TU tiling, cbf, and the horizontal and vertical transforms of the residual block, which can be regarded as a syntax shortcut for the case where the main residual of the block is on one side of the block.
[0194] 2.10. Coefficient Coding and Decoding Based on Scanned Region (SRCC)
[0195] SRCC has been adopted by AVS-3. For SRCC, as Figures 14A to 14B shown in [reference], the lower-right position (SRx, SRy) is signaled, and only the coefficients within the rectangle with four corners (0, 0), (SRx, 0), (0, SRy), (SRx, SRy) are scanned and signaled. All coefficients outside the rectangle are zero.
[0196] 2.11. Implicit Selection of Transform (IST)
[0197] As disclosed in PCT / CN2019 / 090261 (incorporated herein by reference), an implicit selection of the transform solution is given, where the selection of the transform matrix (DCT2 for horizontal and vertical transforms, or DST7 for both) is determined by the parity of the non-zero coefficients in the transform block.
[0198] The proposed method is applied to the luminance component of the intra-coded and decoded blocks, excluding those blocks coded with DT, and the allowed block sizes range from 4×4 to 32×32. The transform type is hidden in the transform coefficients. Specifically, the parity of the number of valid coefficients (e.g., non-zero coefficients) in a block is used to represent the transform type. An odd number indicates the application of DST-VII, and an even number indicates the application of DCT-II.
[0199] To remove the 32-point DST-7 introduced by IST, it is proposed to limit the use of IST according to the range of the remaining scanned region when using SRCC. As Figures 15A to 15B shown in [reference], when the x coordinate or y coordinate of the lower-right position in the remaining scanned region is not less than 16, IST is not allowed. That is, for this case, DCT-II is directly applied.
[0200] For another case, when using run-length coefficient coding and decoding, each non-zero coefficient needs to be checked. When the x coordinate or y coordinate of a non-zero coefficient position is not less than 16, IST is not allowed.
[0201] The corresponding syntax changes are represented by bold italic and underlined text as follows:
[0202]
[0203]
[0204]
[0205]
[0206] 3. Examples of technical problems solved by the disclosed technical solutions
[0207] The current designs of IST and MTS have the following problems:
[0208] 1. The TS mode in VVC is signaled at the block level. However, DCT2 and DST7 work well for residual blocks in sequences captured by cameras, while for videos with screen content, the transform skip (TS) mode is used more frequently compared to DST7. It is necessary to study how to determine the use of the TS mode in a more efficient way.
[0209] 2. In VVC, the maximum allowed TS block size is set to 32×32. How to support large TS blocks still requires further study.
[0210] 4. Example techniques and embodiments
[0211] The items listed below should be regarded as examples for explaining general concepts. These items should not be interpreted in a narrow way. In addition, these items can be combined in any way.
[0212] min(x, y) yields the smaller one of x and y.
[0213] Implicit determination of transform skip mode / identity transform
[0214] It is proposed to determine whether to apply a horizontal and / or vertical identity transform (IT) (e.g., transform skip mode) to the current first block based on the decoding coefficients of one or more representative blocks. This method is called "implicit determination of IT". When both the horizontal transform and the vertical transform are IT, the transform skip (TS) mode is used for the current first block.
[0215] A "block" may be a transform unit (TU) / prediction unit (PU) / coding / decoding unit (CU) / transform block (TB) / prediction block (PB) / coding / decoding block (CB). A TU / PU / CU may include one or more color components, such as only the luminance component for a dual-tree partition, and the currently coded color component is luminance; and for two chrominance components of a dual-tree partition, the currently coded color component is chrominance; or for three color components in a single-tree case.
[0216] 1. Decoded coefficients may be associated with one or more representative blocks in the same or different color components of the current block.
[0217] a. In one example, the representative block is the first block, and the decoded coefficients associated with the first block are used to determine the use of IT on the first block.
[0218] b. In one example, the determination of using IT for the first block may depend on the decoded coefficients of multiple blocks, which include at least one block different from the first block.
[0219] i. In one example, the multiple blocks may include the first block.
[0220] ii. In one example, the multiple blocks may include one or more blocks adjacent to the first block.
[0221] iii. In one example, the multiple blocks may include one or more blocks having the same block dimension as the first block.
[0222] iv. In one example, the multiple blocks may include the last N decoded blocks that are before the first block in decoding order and satisfy specific conditions (such as having the same prediction mode as the current block, for example, all intra-coded or IBC-coded, or having the same dimension as the current block). N is an integer greater than 1.
[0223] v. In one example, the multiple blocks may include one or more blocks with different color components from those of the first block.
[0224] 1) In one example, the first block may be in the luminance component. The multiple blocks may include blocks in the chrominance component (for example, the second block in the Cb / B component and the third block in the Cr / R component).
[0225] a) In one example, the three blocks are in the same coding / decoding unit.
[0226] b) Additionally, optionally, implicit MTS is only applied to luminance blocks and not to chrominance blocks.
[0227] 2) In one example, the first block in the first color component and the multiple blocks included in the multiple blocks that are not in the first component color component can be at corresponding positions or juxtaposed positions in the picture.
[0228] 2. The decoding coefficient used to determine the use of IT is called the representative coefficient.
[0229] a. In one example, the representative coefficient only includes coefficients that are not equal to zero (referred to as valid coefficients).
[0230] b. In one example, the representative coefficient can be modified before being used to determine the use of IT.
[0231] i. For example, the representative coefficient can be clipped before being used to derive the transform.
[0232] ii. For example, the representative coefficient can be scaled before being used to derive the transform.
[0233] iii. For example, an offset can be added to the representative coefficient before being used to derive the transform.
[0234] iv. For example, the representative coefficient can be filtered before being used to derive the transform.
[0235] v. For example, before being used to derive the transform, the coefficient or representative coefficient can be mapped to other values (e.g., through a look-up table or dequantization).
[0236] c. In one example, the representative coefficient is all the significant coefficients in the representative block.
[0237] d. Optionally, the representative coefficient is a part of the significant coefficients in the representative block.
[0238] i. In one example, the representative coefficient is those odd decoding significant coefficients.
[0239] 1) Optionally, the representative coefficient is those even decoding significant coefficients.
[0240] ii. In one example, the representative coefficient is those decoding significant coefficients that are greater than or not less than a threshold.
[0241] 1) Optionally, the representative coefficient is those decoding significant coefficients whose magnitude is greater than or not less than a threshold.
[0242] iii. In one example, the representative coefficient is those decoding significant coefficients that are less than or not greater than a threshold.
[0243] 1) Optionally, the representative coefficient is those decoding significant coefficients whose magnitude is less than or not greater than a threshold.
[0244] iv. In one example, the representative coefficients are the first K (K >= 1) decoded valid coefficients in the decoding order.
[0245] v. In one example, the representative coefficients are the last K (K >= 1) decoded valid coefficients in the decoding order.
[0246] vi. In one example, the representative coefficients can be the coefficients at predefined positions in the block.
[0247] 1) In one example, the representative coefficients can include only one coefficient located at the (xPos, yPos) coordinates relative to the representative block. For example, xPos = yPos = 0.
[0248] 2) In one example, the representative coefficients can include only one coefficient located at the (xPos, yPos) coordinates relative to the representative block. And xpo and / or ypo satisfy the following conditions:
[0249] a) In one example, xPos is not greater than the threshold Tx (e.g., 31) and / or yPos is not greater than the threshold Ty (e.g., 31).
[0250] b) In one example, xPos is not less than the threshold Tx (e.g., 32) and / or yPos is not less than the threshold Ty (e.g., 32).
[0251] 3) For example, the position can depend on the dimension of the block.
[0252] vii. In one example, the representative coefficients can be those coefficients at predefined positions in the coefficient scan order.
[0253] e. Optionally, the representative coefficients can also include those zero coefficients.
[0254] f. Optionally, the representative coefficients can be coefficients derived from the decoded coefficients, such as by clipping to a range, by quantization.
[0255] g. In one example, the representative coefficients can be the coefficients before the last valid coefficient (which can include the last valid coefficient).
[0256] 3. The determination of using IT for the first block can depend on the decoded luminance coefficients of the first block.
[0257] a. Additionally, optionally, the determined use of IT is only applied to the luminance component of the first block, while DCT2 is always used for the chrominance component of the first block.
[0258] b. Additionally, optionally, the determined use of IT is applied to all color components of the first block. That is, the same transformation matrix is applied to all color components of the first block.
[0259] 4. The determination of the use of IT can depend on a function of representative coefficients, such as a function that takes representative coefficients as input and outputs a value V.
[0260] a. In one example, V is derived as the number of representative coefficients.
[0261] i. Optionally, V is derived as the sum of representative coefficients.
[0262] 1) Optionally, V is derived as the sum of the levels (or their absolute values) of representative coefficients.
[0263] 2) Optionally, V can be derived as the level (or its absolute value) of one representative coefficient (such as the last one).
[0264] 3) Optionally, V can be derived as the number of representative coefficients with even levels.
[0265] 4) Optionally, V can be derived as the number of representative coefficients with odd levels.
[0266] 5) Additionally, optionally, the sum can be clipped to derive V.
[0267] ii. Optionally, V is derived as the output of a function that defines the residual energy distribution.
[0268] 1) In one example, the function returns the ratio of the sum of the absolute values of some representative coefficients to the sum of the absolute values of all representative coefficients.
[0269] 2) In one example, the function returns the ratio of the sum of the squares of the absolute values of some representative coefficients to the sum of the squares of the absolute values of all representative coefficients.
[0270] iii. Optionally, V is derived as whether at least one representative coefficient is outside a sub-region of the representative block.
[0271] 1) In one example, the sub-region is defined as the upper-left sub-region of the representative block, for example, the upper-left quarter of the representative block.
[0272] b. In one example, the determination of the use of IT can depend on the parity of V.
[0273] i. For example, if V is even, IT is used; while if V is odd, IT is not used.
[0274] 1) Optionally, if V is even, IT is used; if V is odd, IT is not used.
[0275] ii. In one example, if V is less than threshold T1, IT is used; if V is greater than threshold T2, IT is not used.
[0276] 1) Optionally, if V is greater than threshold T1, IT is used; if V is less than threshold T2, IT is not used.
[0277] iii. For example, the threshold can depend on encoding / decoding information such as block dimension, prediction mode.
[0278] iv. For example, the threshold can depend on QP.
[0279] c. In one example, the determination of the use of IT can depend on the combination of V and other encoding / decoding information (e.g., prediction mode, slice type / picture type, block dimension).
[0280] 5. The determination of the use of IT can further depend on the encoding / decoding information of the current block.
[0281] a. In one example, the determination can also depend on mode information (e.g., inter, intra, or IBC).
[0282] b. In one example, the transform determination can depend on the scan region that is the smallest rectangle covering all valid coefficients (e.g., as depicted in FIG. 14).
[0283] i. In one example, if the size (e.g., width times height) of the scan region associated with the current block is greater than a given threshold, a default transform (such as DCT-2), including horizontal and vertical transforms, can be utilized. Otherwise, rules defined in bullet 3 can be used (e.g., IT when V is even, DCT-2 when V is odd).
[0284] ii. In one example, if the width of the scan region associated with the current block is greater than (or less than) a given maximum width (e.g., 16), a default horizontal transform (such as DCT-2) can be utilized. Otherwise, rules defined in bullet 3 can be used.
[0285] iii. In one example, if the height of the scan region associated with the current block is greater than (or less than) a given maximum height (e.g., 16), a default vertical transform (such as DCT-2) can be utilized. Otherwise, rules defined in bullet 3 can be used.
[0286] iv. In one example, the given size is L×K, where L and K are integers such as 16.
[0287] v. In one example, the default transform matrix can be DCT-2 or DST-7.
[0288] 6. One or more of the methods disclosed in bullet points 1 to 5 can only be applied to specific blocks.
[0289] a. For example, one or more of the methods disclosed in bullet points 1 to 5 can only be applied to blocks of IBC encoding / decoding other than DT and / or blocks of intra-frame encoding / decoding.
[0290] b. For example, one or more of the methods disclosed in bullet points 1 to 5 can only be applied to blocks with specific constraints on coefficients.
[0291] i. A rectangle with four corners (0, 0), (CRx, 0), (0, CRy), (CRx, CRy) is defined as a constrained rectangle, for example, in the SRCC method. In one example, one or more of the methods disclosed in bullet points 1 to 5 can be applied only when all coefficients outside the constrained rectangle are zero. For example, CRx = CRy = 16.
[0292] 1) For example, CRx = SRx and CRy = SRy, where (SRx, SRy) is defined as described in Section 2.14 in SRCC.
[0293] 2) Additionally, optionally, the above method is applied only when the block width or block height is greater than K.
[0294] a) In one example, K is equal to 16.
[0295] b) In one example, the above method is applied only when the block width is greater than K1 and K1 is equal to CRx; or when the block height is greater than K2 and K2 is equal to CRy.
[0296] ii. One or more of the methods can be applied only when the last non-zero coefficient (in the forward scan order) satisfies certain conditions. For example, when the horizontal coordinate / vertical coordinate is not greater than a threshold (e.g., 16 / 32).
[0297] 7. When it is determined not to use IT, a default transform such as DCT-2 or DST-7 can be used instead.
[0298] a. Optionally, when it is determined not to use IT, a selection can be made from multiple default transforms such as DCT-2 or DST-7.
[0299] 8. Whether and / or how to apply the methods disclosed above can be signaled at the video region level (such as sequence level / picture level / strip level / slice group level / slice level).
[0300] a. In one example, it can be signaled (e.g., a flag) in the sequence header / picture header / SPS / VPS / DCI / DPS / PPS / APS / strip header / slice group header.
[0301] i. Additionally, optionally, one or more syntax elements (e.g., one or more flags) can be signaled to specify whether to enable the implicitly determined method of IT.
[0302] 1) In one example, a first flag can be signaled to control the use of the implicitly determined method of IT for IBC-coded blocks at the video region level.
[0303] a) Additionally, optionally, the flag can be signaled under the condition of checking whether IBC is enabled.
[0304] 2) In one example, a second flag can be signaled to control the use of the implicitly determined method of IT for intra-coded blocks at the video region level (e.g., blocks with DT mode can be excluded).
[0305] 3) In one example, a second flag can be signaled to control the use of the implicitly determined method of IT for inter-coded blocks at the video region level (e.g., blocks with DT mode can be excluded).
[0306] 4) In one example, a second flag can be signaled to control the use of the implicitly determined method of IT for intra-coded blocks and inter-coded blocks at the video region level (e.g., blocks with DT mode can be excluded).
[0307] 5) In one example, a second flag can be signaled to control the use of the implicitly determined method of IT for IBC-coded blocks and inter-coded blocks at the video region level (e.g., blocks with DT mode can be excluded).
[0308] ii. Additionally, optionally, when the implicitly determined method of IT is enabled for the video region, the following can be further applied:
[0309] 1) In one example, for IBC-coded blocks, if IT is used for the block, apply the TS mode; otherwise, use DCT2.
[0310] 2) In one example, for intra-coded blocks (e.g., blocks with DT mode can be excluded), if IT is used for the block, apply the TS mode; otherwise, use DCT2.
[0311] iii. Additionally, optionally, when the implicit determination method of IT is disabled for the video region, the following can be further applied;
[0312] 1) In one example, for IBC encoded / decoded blocks, DCT-2 is used.
[0313] 2) In one example, for intra-encoded / decoded blocks (e.g., excluding blocks with DT mode), DCT-2 or DST-7 can be determined on-the-fly, such as determined by IST.
[0314] 9. At the video region level, such as sequence level / picture level / strip level / slice group level / slice level, signaling indicates whether to apply zeroing to transform blocks (including identity transform).
[0315] a. In one example, this indication (e.g., a flag) can be signaled in the sequence header / picture header / SPS / VPS / DCI / DPS / PPS / APS / strip header / slice group header.
[0316] b. In one example, when this indication specifies enabling zeroing, only IT transform is allowed.
[0317] c. In one example, when this indication specifies disabling zeroing, only non-IT transform is allowed.
[0318] d. Additionally, optionally, the binarization / context modeling / allowed range of the last significant coefficient / lower right position in SRCC (e.g., the maximum X / Y coordinates relative to the upper left position of the block) can depend on this indication.
[0319] 10. The first rule (e.g., in the above bullets 1 to 7) can be used to determine the use of IT for the first block, and the second rule can be used to determine the transform type excluding IT.
[0320] a. In one example, the first rule can be defined as the residual energy distribution.
[0321] b. In one example, the second rule can be defined as the parity of representative coefficients.
[0322] Transform skip
[0323] 11. Apply zeroing to IT (e.g., TS) encoded / decoded blocks, where non-zero coefficients are restricted to a specific sub-region of the block.
[0324] a. In one example, the zeroing range of the IT (e.g., TS) encoding / decoding block is set to the upper-right K*L sub-region of the block, where K is set to min(T1, W) and L is set to min(T2, H), where W and H are the block width / block height respectively, and T1 / T2 are two thresholds.
[0325] i. In one example, T1 and / or T2 can be set to 32 or 16.
[0326] ii. Additionally, optionally, the last non-zero coefficient should be within the K*L sub-region.
[0327] iii. Additionally, optionally, the lower-right position (SRx, SRy) in the SRCC method should be within the K*L sub-region.
[0328] 12. Multiple zeroing types of the IT (e.g., TS) encoding / decoding block are defined, where each type corresponds to a sub-region of the block and non-zero coefficients only exist in that sub-region.
[0329] a. In one example, non-zero coefficients only exist in the upper-left K0*L0 sub-region of the block.
[0330] b. In one example, non-zero coefficients only exist in the upper-right K1*L1 sub-region of the block.
[0331] i. Additionally, optionally, an indication of the lower-left position of the sub-region with non-zero coefficients can be signaled.
[0332] c. In one example, non-zero coefficients only exist in the lower-left K2*L2 sub-region of the block.
[0333] i. Additionally, optionally, an indication of the upper-right position of the sub-region with non-zero coefficients can be signaled.
[0334] d. In one example, non-zero coefficients only exist in the lower-right K3*L3 sub-region of the block.
[0335] i. Additionally, optionally, an indication of the upper-left position of the sub-region with non-zero coefficients can be signaled.
[0336] e. Additionally, optionally, an indication of the zeroing type of IT can be further explicitly signaled or immediately derived.
[0337] 13. When at least one valid coefficient is outside the zeroing region defined by IT (e.g., TS), for example, outside the upper-left K0*L0 sub-region of the block, IT (e.g., TS) is not used in that block.
[0338] a. Additionally, optionally, for this case, a default transform is used.
[0339] 14. When there is at least one valid coefficient outside the zeroing region defined by another transform matrix (e.g., DST7 / DCT2 / DCT8), for example, outside the upper-left K0*L0 sub-region of the block, IT (e.g., TS) is used in the block.
[0340] a. Additionally, optionally, for this case, it is inferred to use the TS mode.
[0341] Figures 16A to 16D Multiple zeroing types of the TS encoding / decoding block are shown. Figure 16A The upper-left K0*L0 sub-region is shown. Figure 16B The upper-right K1*L1 sub-region is shown. Figure 16C The lower-left K2*L2 sub-region is shown. Figure 16D The lower-right K3*L3 sub-region is shown.
[0342] General
[0343] 15. The determination of the transform matrix can be done at the CU / CB level or the TU level.
[0344] a. In one example, the determination is made at the CU level, where all TUs share the same transform matrix.
[0345] i. Additionally, optionally, when a CU is partitioned into multiple TUs, the coefficients in one TU (e.g., the first TU or the last TU) or part of the TUs or all TUs can be used to determine the transform matrix.
[0346] b. Whether to use the CU-level solution or the TU-level solution can depend on the block size of a block and / or the VPDU size and / or the maximum CTU size and / or the encoding / decoding information.
[0347] i. In one example, when the block size is greater than the VPDU size, the CU-level determination method can be applied.
[0348] 16. Whether and / or how to apply the above-disclosed method can depend on the encoding / decoding information, which can include:
[0349] a. Block dimensions.
[0350] i. In one example, for a block whose width and height are both not greater than a threshold (e.g., 32), the above implicit MTS method can be applied.
[0351] b. QP
[0352] c. Picture or slice type (such as I-frame or P / B-frame, I-slice or P / B-slice)
[0353] i. In one example, the proposed method can be enabled for I-frames, but disabled for P / B-frames.
[0354] d. Structure splitting method (single-tree or dual-tree)
[0355] i. In one example, for a slice / picture / tile / slice to which single-tree splitting is applied, the above implicit MTS method can be applied.
[0356] e. Coding / decoding mode (such as inter-frame mode / intra-frame mode / IBC mode, etc.).
[0357] i. In one example, for a block coded / decoded intra-frame, the above implicit MTS method can be applied.
[0358] f. Coding / decoding method (such as intra-sub-block splitting, Derived Tree (DT) method, etc.).
[0359] i. In one example, for an intra-frame coded / decoded block to which DT is applied, the above implicit MTS method can be disabled.
[0360] ii. In one example, for an intra-frame coded / decoded block to which ISP is applied, the above implicit MTS method can be disabled.
[0361] g. Color component
[0362] i. In one example, for a luminance block, the above implicit MTS method can be applied, while for a chrominance block, this method is not applied.
[0363] h. Intra-frame prediction mode (such as DC, vertical, horizontal, etc.).
[0364] i. Motion information (such as MV and reference index).
[0365] j. Standard profile / level / hierarchy
[0366] Figure 17 FIG. 1700 is a block diagram showing an example video processing system 1700 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 1700. System 1700 can include an input 1702 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or can be received in a compressed or coded format. Input 1702 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, Passive Optical Network (PON), etc.), and wireless interfaces (e.g., Wi-Fi or cellular interfaces).
[0367] System 1700 may include an encoding / decoding component 1704 that may implement various encoding / decoding or encoding methods described in this document. The encoding / decoding component 1704 may reduce the average bit rate of a video from the input 1702 to the output of the encoding / decoding component 1704 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 1704 may be stored or transmitted via a communication of a connection represented by component 1706. The stored or transmitted bitstream (or encoded / decoded) representation of the video received at the input 1702 may be used by component 1708 to generate pixel values or a viewable video to be sent to the display interface 1710. The process of generating a user-viewable video from the bitstream is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding / decoding" operations or tools, it will be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding / decoding results will be performed by the decoder.
[0368] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0369] Figure 21 is a block diagram of a video processing apparatus 2100. The apparatus 2100 may be used to implement one or more methods described herein. The apparatus 2100 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 2100 may include one or more processors 2102, one or more memories 2104, and video processing hardware 2106. The (multiple) processors 2102 may be configured to implement one or more methods described in this document. The one or more memories 2104 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2106 may be used to implement some of the techniques described in this document in hardware circuitry.
[0370] Figure 18 is a block diagram showing an example video encoding / decoding system 100 that may utilize the techniques of the present disclosure.
[0371] As Figure 18As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data that may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the source device 110 may be referred to as a video decoding device.
[0372] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0373] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0374] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0375] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be located external to the destination device 120, which is configured to interface with an external display device.
[0376] The video encoder 114 and the video decoder 124 may operate according to video compression standards (e.g., the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards).
[0377] Figure 19 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may be Figure 18 the video encoder 114 in the system 100 shown.
[0378] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 19 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, the processor may be configured to perform any or all of the techniques described in the present disclosure.
[0379] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0380] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0381] In addition, some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be highly integrated, but are shown separately in Figure 11 the example for purposes of explanation.
[0382] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0383] The mode selection unit 203 may select, for example, one of the coding / decoding modes (intra or inter) based on an error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra prediction and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel accuracy or integer pixel accuracy).
[0384] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures other than the picture associated with the current video block from buffer 213.
[0385] For example, the motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0386] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0387] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0, and may also search for another reference video block of the current video block in the reference pictures of list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector that indicates the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0388] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.
[0389] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0390] In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0391] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0392] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0393] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0394] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0395] In other examples, for the current video block, there may be no residual data for the current video block. For example, in the skip mode, the residual generation unit 207 may not perform a subtraction operation.
[0396] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0397] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0398] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block based on the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0399] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0400] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0401] Figure 20 is a block diagram showing an example of the video decoder 300, and the video decoder 300 can be Figure 18 the video decoder 114 in the system 100 shown.
[0402] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 20 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0403] In Figure 20 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding pass that is generally opposite to the encoding pass ( Figure 19 ) described with respect to the video encoder 200.
[0404] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and the motion compensation unit 302 can determine motion information based on the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge mode.
[0405] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. The syntax elements may include an identifier for the interpolation filter to be used with sub-pixel precision.
[0406] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during video block encoding to compute the interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and use the interpolation filter to generate a prediction block.
[0407] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0408] The intra prediction unit 303 may form a prediction block based on spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (e.g., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0409] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0410] A list of preferred solutions for some embodiments is provided below.
[0411] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0412] 1. A video processing method (e.g., Figure 22 the method 2200 described in) includes: converting between a video block of a video and an encoded / decoded representation of the video, determining whether to apply a horizontal identity transform or a vertical identity transform to the video block based on a rule (2202); and performing the conversion based on the determination (2204), where the rule specifies the relationship between the determination and representative coefficients of decoded coefficients of one or more representative blocks from the video.
[0413] 2. The method of Solution 1, wherein one or more representative blocks belong to the color component to which the video block belongs.
[0414] 3. The method of Solution 1, wherein one or more representative blocks belong to a color component different from the color component of the video block.
[0415] 4. The method according to any one of Solutions 1-3, wherein the one or more representative blocks correspond to the video block.
[0416] 5. The method according to any one of Solutions 1-3, wherein the one or more representative blocks do not include the video block.
[0417] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., bullet points 1 and 2).
[0418] 6. The method according to any one of Solutions 1-5, wherein the representative coefficient includes a decoded coefficient having a non-zero value.
[0419] 7. The method according to any one of Solutions 1-6, wherein the relationship specifies using the representative coefficient based on a modified coefficient determined by modifying the representative coefficient.
[0420] 8. The method according to any one of Solutions 1-7, wherein the representative coefficient corresponds to a significant coefficient of the decoded coefficient.
[0421] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0422] 9. A video processing method, comprising: converting between a video block of a video and an encoded / decoded representation of the video, determining based on a rule whether to apply a horizontal identity transform or a vertical identity transform to the video block; and performing the conversion based on the determination, wherein the rule specifies a relationship between the determination and the decoded luminance coefficient of the video block.
[0423] 10. The method of Solution 1, wherein performing the conversion includes applying a horizontal identity transform luminance component or a vertical identity transform luminance component of the video block and DCT2 to the chrominance component of the video block.
[0424] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1 and item 4).
[0425] 11. A video processing method, comprising: converting between a video block of a video and an encoded / decoded representation of the video, determining based on a rule whether to apply a horizontal identity transform or a vertical identity transform to the video block; and performing the conversion based on the determination, wherein the rule specifies a relationship between the determination and a value V associated with the decoded coefficient or the representative coefficient of the representative block.
[0426] 12. The method of Solution 11, wherein V is equal to the number of representative coefficients.
[0427] 13. The method of Solution 11, wherein V is equal to the sum of the values of the representative coefficients.
[0428] 14. The method of Solution 11, wherein V is a function of the residual energy distribution of the representative coefficients.
[0429] 15. The method of any one of Solutions 11 - 14, wherein the relationship is defined relative to the parity of the value V
[0430] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 5).
[0431] 16. The method of any one of the above solutions, wherein the rule states that the relationship further depends on the codec information of the video block.
[0432] 17. The method of Solution 16, wherein the codec information is the codec mode of the video block.
[0433] 18. The method of Solution 16, wherein the codec information includes the smallest rectangular region covering all valid coefficients of the video block.
[0434] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 6).
[0435] 19. The method of any one of the above solutions, wherein the determination is performed because the video block has a mode or a constraint on the coefficients.
[0436] 20. The method of Solution 19, wherein the type corresponds to the intra block copy (IBC) mode.
[0437] 21. The method of Solution 19, wherein the constraint on the coefficients causes the coefficients outside the rectangle of the current block to be zero.
[0438] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 7).
[0439] 22. The method of any one of Solutions 1 - 21, wherein, in the case where the horizontal identity transform and the vertical identity transform are not used in the determination, the DCT - 2 transform or the DST - 7 transform is used to perform the transform.
[0440] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 9).
[0441] 23. The method of any one of Solutions 1-22, wherein one or more syntax fields in the codec representation indicate whether the method is enabled for a video block.
[0442] 24. The method of Solution 23, wherein the one or more syntax fields are included at the sequence level or picture level or slice group level or slice level or sub-picture level.
[0443] 25. The method of any one of Solutions 23-24, wherein the one or more syntax fields are included in a slice header or a picture header.
[0444] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., Items 1 and 8).
[0445] 26. A video processing method, comprising: determining that one or more syntax fields are present in a codec representation of a video, wherein the video includes one or more video blocks; and determining, based on the one or more syntax fields, whether to enable a horizontal identity transform or a vertical identity transform for the video blocks in the video.
[0446] 27. The method of Solution 1, wherein, in response to an implicit determination that a transform skip mode indicated by the one or more syntax fields is enabled, for the conversion between a first video block of a video and the codec representation of the video, determining, based on a rule, whether to apply a horizontal identity transform or a vertical identity transform to the video block; and performing the conversion based on the determination, wherein the rule specifies a relationship between the determination and representative coefficients of decoded coefficients of one or more representative blocks from the video.
[0447] 28. The method of Solution 27, wherein the first video block is coded and decoded in an intra block copy mode.
[0448] 29. The method of Solution 27, wherein the first video block is coded and decoded in an intra mode.
[0449] 30. The method of Solution 27, wherein the first video block is coded and decoded in an intra mode instead of a derived tree (DT) mode.
[0450] 31. The method of Solution 27, wherein the determination is based on the parity of the number of non-zero coefficients in the first video block.
[0451] 32. The method of Solution 27, wherein when the parity of the number of non-zero coefficients in the first video block is even, applying a horizontal identity transform and a vertical identity transform to the first video block.
[0452] 33. The method of Solution 27, wherein when the parity of the number of non-zero coefficients in the first video block is even, a horizontal identity transform and a vertical identity transform are not applied to the first video block.
[0453] 34. The method of Solution 33 applies DCT-2 to the first video block.
[0454] 35. The method of Solution 32 further includes: in response to the implicit determination that the transform skip mode indicated by one or more syntax fields is disabled, the horizontal identity transform and the vertical identity transform are not applied to the first video block.
[0455] 36. The method of Solution 32, wherein DCT-2 is applied to the first video block.
[0456] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 9, 10).
[0457] 37. A video processing method includes: making a first determination on whether to enable the use of the identity transform for the conversion between a video block of a video and the codec representation of the video; making a second determination on whether to enable the zeroing operation during the conversion; and performing the conversion based on the first determination and the second determination.
[0458] 38. The method according to Solution 37, wherein one or more syntax fields at a first level in the codec representation indicate the first determination.
[0459] 39. The method according to any one of Solutions 37-38, wherein one or more syntax fields at a second level in the codec representation indicate the second determination.
[0460] 40. The method according to any one of Solutions 38-39, wherein the first level and the second level correspond to header fields at the sequence or picture level or parameter sets at the sequence level or picture level or adaptive parameter sets.
[0461] 41. The method according to any one of Solutions 37-40, wherein the conversion uses the identity transform or the zeroing operation, but not both.
[0462] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 12 and 13).
[0463] 42. A video processing method includes: performing a conversion between a video block of a video and the codec representation of the video; wherein the video block is represented as a codec block in the codec representation, wherein the non-zero coefficients of the codec block are restricted to one or more sub-regions; and wherein the identity transform is applied to generate the codec block.
[0464] 43. The method of Solution 1, wherein the one or more sub-regions include an upper-right sub-region of a video block having dimensions of K×L, where K and L are integers, K is min(T1, W), L is min(T2, H), where W and H are the width and height of the video block respectively, and T1 and T2 are thresholds.
[0465] 44. The method according to any one of Solutions 42-43, wherein the codec representation indicates the one or more sub-regions.
[0466] The following solutions illustrate example embodiments of the techniques discussed in the previous section (Items 16 and 17).
[0467] 45. The method according to any one of Solutions 1-44, wherein the video region includes a video codec unit.
[0468] 46. The method according to Solutions 1-45, wherein the video region is a prediction unit or a transform unit.
[0469] 47. The method according to any one of Solutions 1-46, wherein the video block satisfies specific dimensional conditions.
[0470] 48. The method according to any one of Solutions 1-47, wherein the video block is coded and decoded using a pre-specified quantization parameter range.
[0471] 49. The method according to any one of Solutions 1-48, wherein the video region includes a video picture.
[0472] 50. The method according to any one of Solutions 1 to 49, wherein the conversion includes encoding the video into a codec representation.
[0473] 51. The method according to any one of Solutions 1 to 49, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0474] 52. A video decoding device, including a processor configured to implement the method described in one or more of Solutions 1 to 51.
[0475] 53. A video codec device, including a processor configured to implement the method described in one or more of Solutions 1 to 51.
[0476] 54. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of Solutions 1 to 51.
[0477] 55. The method, device or system described in this document.
[0478] Figure 23is a flowchart illustration of a video processing method according to the present technology. Method 2300 includes, at operation 2310, determining, according to a rule, to use an identity transform mode for a transform between a current video block of a video and a bitstream of the video, where the rule specifies that the use is based on representative coefficients of one or more representative blocks of the video. Method 2300 further includes, at operation 2320, performing the transform based on the determination.
[0479] In some embodiments, the identity transform mode includes a transform skip mode. In the transform skip mode, a residual representing a prediction error between the current video block and a reference video block is represented in the bitstream without applying a transform. In some embodiments, in response to the transform skip mode being applied to the transform of the current video block, the transform skip mode includes a horizontal transform mode and / or a vertical transform mode.
[0480] In some embodiments, determining the use of the identity transform mode includes an implicit determination of the identity transform. In some embodiments, one or more representative blocks belong to the same color component. In some embodiments, the color component includes a luminance component. In some embodiments, one or more representative blocks belong to different color components. In some embodiments, the video block belongs to a luminance component of the video and one or more representative blocks belong to a chrominance component of the video. In some embodiments, one or more representative blocks and the video block are in the same codec unit. In some embodiments, one or more representative blocks are located at a collocated position of a picture of the video block.
[0481] In some embodiments, one or more representative blocks include the current video block and the use of the identity transform mode for the current video block is based on representative coefficients associated with the current video block. In some embodiments, the use of the identity transform mode for the current video block is based on representative coefficients of one or more representative blocks, where at least one representative block is different from the video block. In some embodiments, one or more representative blocks include the current video block. In some embodiments, one or more representative blocks include neighboring blocks of the current video block. In some embodiments, one or more representative blocks include at least N blocks that satisfy a condition regarding the video block, where N is an integer greater than 1. In some embodiments, the condition is satisfied in a case where at least N blocks are coded using the same prediction mode as the video block. In some embodiments, the condition is satisfied in a case where at least N blocks have the same dimension as the video block. In some embodiments, the representative coefficients used to determine the use of the identity transform mode for the current video block include decoded coefficients.
[0482] In some embodiments, the representative coefficients include only non-zero coefficients. In some embodiments, the non-zero coefficients are represented as significant coefficients. In some embodiments, the representative coefficients are modified before being used to determine the use of an identity transform mode for the current video block. In some embodiments, at least one of the representative coefficients is modified based on: (1) clipping at least one of the representative coefficients, (2) scaling at least one of the representative coefficients, (3) adding an offset to at least one of the representative coefficients, (4) filtering at least one of the representative coefficients, or (5) mapping at least one of the representative coefficients to another value.
[0483] In some embodiments, the representative coefficients include all non-zero coefficients in one or more representative blocks. In some embodiments, the representative coefficients include a portion of the non-zero coefficients in one or more representative blocks. In some embodiments, the representative coefficients include the even non-zero coefficients in one or more representative blocks. In some embodiments, the representative coefficients include the odd non-zero coefficients in one or more representative blocks. In some embodiments, the representative coefficients include the portion of the non-zero coefficients whose absolute value is greater than or equal to a threshold. In some embodiments, the representative coefficients include the portion of the non-zero coefficients whose absolute value is less than or equal to a threshold. In some embodiments, the representative coefficients include the first or the last K non-zero coefficients in the decoding order, where K is greater than or equal to 1. In some embodiments, the representative coefficients include the coefficients at predefined positions in one or more representative blocks. In some embodiments, the representative coefficients include only one coefficient located at a position (xPos, yPos) relative to the representative block, where xPos and yPos satisfy a condition. In some embodiments, the condition specifies that xPos is less than or equal to a first threshold. In some embodiments, the condition specifies that yPos is greater than a second threshold. In some embodiments, xPos = 0 and yPos = 0. In some embodiments, the position (xPos, yPos) is based on the dimensions of the video block. In some embodiments, the representative coefficients include the coefficients before the last non-zero coefficient. In some embodiments, the representative coefficients include the coefficients before the last non-zero coefficient and the last non-zero coefficient.
[0484] In some embodiments, the representative coefficients include zero coefficients and non-zero coefficients. In some embodiments, the representative coefficients are derived based on modified decoded coefficients. In some embodiments, the representative coefficients include the representative coefficients associated with the luminance component of the current video block. In some embodiments, the use of the identity transform mode for the current video block is applied only to the luminance component of the current video block. In some embodiments, a discrete cosine transform 2 (DCT-2) is applied to one or more chrominance components of the current video block. In some embodiments, the use of the identity transform mode for the current video block is applied to all color components of the current video block.
[0485] In some embodiments, the use of the identity transform mode for the current video block is determined based on a function of representative coefficients of the output value V. In some embodiments, the value V is derived based on the number of representative coefficients. In some embodiments, the value V is derived based on the number of representative coefficients whose levels are even. In some embodiments, the value V is derived based on the following: (1) the sum of the levels of the representative coefficients, (2) one level of the representative coefficients, or (3) the number of representative coefficients whose levels are odd. In some embodiments, the function of the representative coefficients defines a residual energy distribution. In some embodiments, the function returns the ratio of (1) the sum of the absolute values of some of the representative coefficients to (2) the sum of the absolute values of all the representative coefficients. In some embodiments, the function returns the ratio of (1) the sum of the squares of the absolute values of some of the representative coefficients to (2) the sum of the squares of the absolute values of all the representative coefficients. In some embodiments, the value V is determined based on whether at least one representative coefficient is outside a sub-region of the representative block. In some embodiments, the use of the identity transform mode for the current video block is based on the parity of the value V. In some embodiments, the identity transform mode is used when the value V is an even value, and the identity transform mode is not used when the value V is an odd value. In some embodiments, the identity transform mode is used when the value V is less than a first threshold, and the identity transform mode is not used when the value V is greater than a second threshold. In some embodiments, the identity transform mode is used when the value V is less than a third threshold, and the identity transform mode is not used when the value V is greater than a fourth threshold.
[0486] In some embodiments, the use of the identity transform mode is further based on the codec information of the current video block. In some embodiments, the codec information includes at least one of a prediction mode, a slice type, a picture type, a block dimension, a flag indicating whether the identity transform mode is enabled at the sequence level, or a flag indicating whether the identity transform mode is enabled in the picture header. In some embodiments, the codec information includes information about the codec mode of the current video block. In some embodiments, the codec information includes information about a scan region, which is the smallest rectangular region covering all the representative coefficients. In some embodiments, a default transform is used when the dimension of the scan region is greater than a threshold. In some embodiments, the dimension includes a width, a height, or a size equal to the width multiplied by the height.
[0487] In some embodiments, whether a rule applies to a current video block is based on the codec characteristics of the current video block. In some embodiments, the codec characteristics of the current video block include the codec mode of the block, and the codec mode includes at least an intra block copy codec mode or an intra codec mode. In some embodiments, the codec characteristics of the current video block include constraints on the coefficients of the block. In some embodiments, the constraint is satisfied when all coefficients in the rectangular region of the current video block are zero. In some embodiments, the constraint is satisfied when the last non-zero coefficient is less than or equal to a threshold. In some embodiments, the codec characteristics of the current video block include the dimensions of the codec block.
[0488] Figure 24 is a flowchart representation of a video processing method according to the present technology. Method 2400 includes, at operation 2410, converting between a current video block of a video and a bitstream of the video, and applying a default transform to the current video block according to a rule. The rule specifies not to use an identity transform for the conversion of the current video block. Method 2400 includes, at operation 2420, performing the conversion based on the determination.
[0489] In some embodiments, the identity transform mode includes a transform skip mode. In the transform skip mode, the residual of the prediction error between the current video block and a reference video block is represented in the bitstream without applying a transform. In some embodiments, the default transform includes a discrete cosine transform 2 (DCT-2) or a discrete sine transform 7 (DST-7). In some embodiments, the default transform is selected from a plurality of default transform candidates. In some embodiments, the determination of applicability is indicated at a video region level. In some embodiments, the video region includes a sequence, a picture, a slice, a slice group, or a tile. In some embodiments, the determination of applicability to a video block is indicated in a sequence header, a picture header, a sequence parameter set, a video parameter set, a decoder parameter set, a picture parameter set, an adaptive parameter set, a slice header, or a slice group header.
[0490] In some embodiments, one or more syntax elements are used to indicate whether the determination applies to a video block. In some embodiments, for a block coded in an intra block copy codec mode, a first syntax element is used at the video region level. In some embodiments, for a block coded in an intra codec mode, a second syntax element is used at the video region level. In some embodiments, for a block coded in an inter codec mode, a second syntax element is used at the video region level. In some embodiments, for a block coded in an intra codec mode and a block coded in an inter codec mode, a second syntax element is used at the video region level. In some embodiments, for a block coded in an intra block copy codec mode and a block coded in an inter codec mode, a second syntax element is used at the video region level.
[0491] In some embodiments, when an identity transform is used for a block encoded or decoded using an intra block copy codec mode or an intra codec mode, a transform skip mode is applied to the block. In some embodiments, when the identity transform is disabled for a block, DCT-2 or DST-7 is determined for transformation.
[0492] Figure 25 is a flowchart representation of a video processing method according to the present technology. Method 2500 includes, at operation 2510, performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies an indication at a video region level. The indication indicates whether a zeroing operation that sets some residual coefficients to zero is applied to a transform block of a video block in a video region.
[0493] In some embodiments, a video region includes a sequence, a picture, a slice, a slice group, or a tile. In some embodiments, a video region level includes a sequence header, a picture header, a sequence parameter set, a video parameter set, a decoder parameter set, a picture parameter set, an adaptive parameter set, a slice header, or a slice group header. In some embodiments, only an identity transform is allowed when the zeroing operation is enabled, while only a non-identity transform is allowed when the zeroing operation is disabled. In some embodiments, information on a coefficient codec tool based on a scan region is based on the indication.
[0494] In some embodiments, the conversion is performed according to a second rule that specifies a transform type of a video block, where the transform type does not include an identity transform. In some embodiments, the rule is defined as a residual energy distribution rule, and the second rule is defined as the parity of representative coefficients.
[0495] In some embodiments, a transform matrix is determined at a codec unit level, a codec block level, or a transform unit level. In some embodiments, when all transform units share the same transform matrix, the determination is made at the codec unit level. In some embodiments, whether the determination is made at the codec unit level or at the transform unit level is based on codec information of a video block.
[0496] In some embodiments, the applicability of one of the above methods is based on the codec information of the current video block. In some embodiments, the codec information includes the dimensions of the video block. In some embodiments, the method is applicable to the case where the width and / or height of the current video block is less than or equal to a threshold. In some embodiments, the method is applicable to the case where the width and / or height of the current video block is less than a threshold. In some embodiments, the threshold is equal to 32. In some embodiments, the codec information includes the segmentation method applied to the current video block. In some embodiments, the segmentation method includes a single tree and / or a double tree. In some embodiments, the codec information includes the codec mode of the video block. In some embodiments, the codec mode includes an inter prediction mode, an intra prediction mode, or an intra block copy prediction mode. In some embodiments, the codec information includes the quantization parameter, picture or slice type, codec method, color component, intra prediction mode, or motion information associated with the video block. In some embodiments, the codec information includes the profile, level, or tier of the video codec standard.
[0497] Figure 26 is a flowchart representation of a video processing method 2600 according to the present technology. Method 2600 includes, at operation 2610, performing a conversion between a current video block of a video and a bitstream of the video according to a rule. An identity transform mode is applied to the current video block during the conversion, and the rule specifies enabling a zeroing operation during which non-zero coefficients are restricted within a sub-region of the current video block.
[0498] In some embodiments, the identity transform mode includes a transform skip mode. In the transform skip mode, the residual of the prediction error between the current video block and a reference video block is represented in the bitstream without applying a transform. In some embodiments, the current video block is a prediction residual block.
[0499] In some embodiments, during the zeroing operation, the sub-region is set to the upper right region having a size of K×L. K is equal to min(T1, 2), L is equal to min(T2, H), W represents the width of the video block, H represents the height of the video block, and T1 and T2 represent two thresholds. In some embodiments, T1 is equal to 16 or 32, and T2 is equal to 16 or 32.
[0500] In some embodiments, the last non-zero coefficient of the current video block is within the sub-region. In some embodiments, the lower right position represented as (SRx, SRy) used in a coefficient codec tool based on a scan region is within the sub-region.
[0501] Figure 27is a flowchart representation of a video processing method 2700 according to the present technology. Method 2700 includes, at operation 2710, determining a zeroing type of a current video block of a video for a conversion between the current video block of the video and a bitstream of the video. Method 2700 further includes, at operation 2720, performing the conversion according to the determination. The current video block is encoded and decoded by applying an identity transform to the current video block. The zeroing type of the video block defines a sub-region of the video block within which non-zero coefficients are restricted for zeroing operations.
[0502] In some embodiments, the zeroing type includes a first type of video block that includes an upper left sub-region of size K0×L0. In some embodiments, the zeroing type includes a second type of video block that includes an upper right sub-region of size K1×L1. In some embodiments, the zeroing type includes a third type of video block that includes a lower left sub-region of size K2×L2. In some embodiments, the zeroing type includes a fourth type of video block that includes a lower right sub-region of size K3×L3. In some embodiments, the position of the sub-region is indicated in the bitstream. In some embodiments, the zeroing type of the video block is determined during the conversion.
[0503] Figure 28 is a flowchart representation of a video processing method 2800 according to the present technology. Method 2800 includes, at operation 2810, performing a conversion between a current video block of a video and a bitstream of the video according to a rule. The rule stipulates that, in a case where at least one non-zero coefficient is outside a zeroing region determined by an identity transform mode, using the identity transform mode to convert the current video block is prohibited. The zeroing region includes a region where non-zero coefficients are restricted for zeroing operations. In some embodiments, a default transform is used in the video block.
[0504] Figure 29 is a flowchart representation of a video processing method 2900 according to the present technology. Method 2900 includes, at operation 2910, performing a conversion between a video block of a video and a bitstream of the video according to a rule. The rule stipulates that, in a case where at least one non-zero coefficient is outside a zeroing region determined by a transform matrix that is not an identity transform, using the identity transform mode during the conversion of the video block is enabled. The zeroing region includes a region where non-zero coefficients are restricted for zeroing operations.
[0505] In some embodiments, the transform matrix includes a discrete sine transform 7 (DST7), a discrete cosine transform 2 (DCT2), or a discrete cosine transform 8 (DCT8). In some embodiments, a transform skip mode is used in the video block.
[0506] In some embodiments, the conversion includes encoding the video into a bitstream. In some embodiments, the conversion includes decoding the bitstream to generate a video.
[0507] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from the pixel representation of a video to the corresponding bitstream, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream of the current video block may, for example, correspond to bits juxtaposed or scattered at different positions within the bitstream. For example, a macroblock may be encoded based on the transformed and coded error residual values and also using bits in the headers and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream knowing that some fields may or may not be present based on the determination as described in the above solutions. Similarly, the encoder may determine to include or exclude certain syntax fields and generate the coded representation accordingly by including or excluding the syntax fields from the coded representation.
[0508] The disclosed solutions and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in: digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of the foregoing. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution or control of the operation by a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, e.g., including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs being discussed, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device.
[0509] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can also be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0510] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0511] For example, processors suitable for executing computer programs include general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer does not necessarily require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0512] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features that are specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although the above features may be described as acting in a particular combination and even initially claimed as such, in some cases, one or more features may be deleted from the claimed combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0513] Similarly, although these operations are described in a particular order in the drawings, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the operations shown be performed to obtain a desirable result. Additionally, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0514] Only some implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A method for video data processing, comprising: For the conversion between the current video block of a video and the bitstream of the video, determining the use of an implicitly selected transform skip mode for the conversion of the current video block according to a rule, wherein the rule stipulates that the use is based on the number of representative coefficients in the current video block; and Performing the conversion based on the determination, wherein, in the implicitly selected transform skip mode, the prediction residual between the current video block and a reference video block exists in the bitstream and no transform needs to be applied.
2. The method according to claim 1, wherein, the use of the implicitly selected transform skip mode for the current video block is based on the parity of the number of representative coefficients.
3. The method according to claim 2, wherein, the implicitly selected transform skip mode is used at least based on the number of representative coefficients being an even value, and in response to the number of representative coefficients being an odd value, the implicitly selected transform skip mode is not used.
4. The method according to claim 1, wherein, the use of the implicitly selected transform skip mode is further based on the codec information of the current video block.
5. The method according to claim 4, wherein, the codec information includes at least one of a prediction mode, a block dimension, or a flag indicating whether to enable the implicitly selected transform skip mode at a picture header.
6. The method according to claim 5, wherein, the implicitly selected transform skip mode is used at least based on the block dimension satisfying that both the width and height of the current video block are less than or equal to a threshold.
7. The method according to claim 6, wherein, the threshold is equal to 32.
8. The method according to claim 5, wherein, the implicitly selected transform skip mode is used at least based on the prediction mode of the current video block being an intra block copy codec mode, an intra codec mode, or an inter codec mode.
9. The method according to claim 1, wherein, the implicitly selected transform skip mode is only applied to luminance video blocks and not to chrominance video blocks.
10. The method according to claim 1, wherein, the representative coefficient is an even coefficient in the current video block.
11. The method according to claim 1, wherein, the use of the implicitly selected transform skip mode for the current video block is also based on whether a zeroing operation is applied.
12. The method according to claim 11, wherein, the zeroing range of the zeroing operation is set outside the upper left K*L sub-region of the current video block, where K is set to min(T1, W), L is set to min(T2, H), where W and H are the block width and block height of the current video block respectively, and T1 and T2 are equal to 16.
13. The method according to claim 12, wherein, for the current video block, based on the lower right position (SRx, SRy) of the scanning region of the coefficient codec for the scanning region (SRCC) tool being within the K*L sub-region.
14. The method according to claim 1, wherein, the conversion includes encoding the video into the bitstream.
15. The method according to claim 1, wherein, the conversion includes decoding the video from the bitstream.
16. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: for a conversion between a current video block of a video and the bitstream of the video, determine a use of an implicitly selected transform skip mode for the conversion of the current video block according to a rule, wherein the rule specifies that the use is based on the number of representative coefficients in the current video block; and perform the conversion based on the determination, wherein, in the implicitly selected transform skip mode, a prediction residual between the current video block and a reference video block exists in the bitstream and no transform is applied.
17. The apparatus according to claim 16, wherein, the use of the implicitly selected transform skip mode for the current video block is determined based on the parity of the number of representative coefficients, and wherein the implicitly selected transform skip mode is used at least based on the number of representative coefficients being an even value, and in response to the number of representative coefficients being an odd value, the implicitly selected transform skip mode is not used.
18. A non-transitory computer-readable storage medium storing instructions that cause a processor to: for a conversion between a current video block of a video and the bitstream of the video, determine a use of an implicitly selected transform skip mode for the conversion of the current video block according to a rule, wherein, the rule specifies that the use is based on the number of representative coefficients in the current video block; and perform the conversion based on the determination, wherein, in the implicitly selected transform skip mode, a prediction residual between the current video block and a reference video block exists in the bitstream and no transform is applied.
19. A non-transitory computer-readable storage medium storing a bitstream of a video, the bitstream being generated by a method executed by a video processing apparatus, wherein the method comprises: for a current video block of a video, determine a use of an implicitly selected transform skip mode for the current video block according to a rule, wherein the rule specifies that the use is based on the number of representative coefficients in the current video block; and generate the bitstream based on the determination, wherein, in the implicitly selected transform skip mode, a prediction residual between the current video block and a reference video block exists in the bitstream and no transform is applied.
20. The non-transitory computer-readable storage medium according to claim 19, wherein, the use of the implicitly selected transform skip mode for the current video block is determined based on the parity of the number of representative coefficients, and wherein the implicitly selected transform skip mode is used at least based on the number of representative coefficients being an even value, and in response to the number of representative coefficients being an odd value, the implicitly selected transform skip mode is not used.
21. A method for storing a bitstream of a video, comprising: For a current video block of the video, determining the use of an implicitly selected transform skip mode for the current video block according to a rule, wherein the rule stipulates that the use is based on the number of representative coefficients in the current video block; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable storage medium, wherein, in the implicitly selected transform skip mode, a prediction residual between the current video block and a reference video block exists in the bitstream and no transform needs to be applied.
22. A video decoding apparatus, comprising a processor configured to implement the method according to any one of claims 2 and 4 to 15.
23. A video encoding apparatus, comprising a processor configured to implement the method according to any one of claims 2 and 4 to 15.
24. A computer-readable storage medium having computer instructions stored thereon, which when executed by a processor cause the processor to implement the method according to any one of claims 2 and 4 to 15.
Citation Information
Patent Citations
Coefficient dependent encoding of transform matrix selection
CN110839158A
Group flag in transform coefficient coding for video coding
US20130272414A1