Method, apparatus, and storage medium for processing video data
By introducing a secondary conversion tool in video encoding, the application of the video block is determined based on the size of the video block, the problem of low video encoding efficiency in the prior art is solved, and higher encoding efficiency and lower bandwidth occupation are achieved.
Patent Information
- Application Number
- JP2024119547
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-07
- Filing Date
- 2024-07-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-06-08
AI Technical Summary
When existing video encoding technology processes high-resolution video, bandwidth occupies a large amount of bandwidth and low encoding efficiency, making it difficult to meet future video transmission needs.
A video processing method is proposed. By introducing a secondary conversion tool in the video encoding process, it determines whether to apply the tool based on the width and height of the video block, and performs corresponding conversions during the encoding and decoding process.
Improves the compression performance of video encoding, reduces the bandwidth required for video transmission, and improves video quality and encoding efficiency.
Smart Images

Figure 0007684495000049 
Figure 0007684495000050 
Figure 0007684495000051
Abstract
Description
Background Art
[0001] Cross - reference to Related Applications Based on the patent laws and / or regulations applicable with respect to the Paris Convention, this application has been duly made to claim the priority and benefit of International Patent Application No. PCT / CN2019 / 090446, filed on July 7, 2019. For all purposes under the law, the entire disclosure of the foregoing application is incorporated by reference as part of the disclosure of this application.
[0002] Technical Field This patent document relates to video processing technologies, devices, and systems.
[0003] Background Despite progress in video compression, digital video still occupies the largest bandwidth in the Internet and other digital communication networks. It is expected that as the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth for the use of digital video will continue to grow.
Summary of the Invention
[0004] Devices, systems, and methods are related to digital video processing. The methods described may be applicable to both existing video - coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video - coding standards or video codecs.
[0005] In a representative aspect, the disclosed technology can be used to provide a video processing method. The method includes performing a conversion between a current video block of a video and a coded representation of the video, and the step of performing the conversion includes determining the applicability of a secondary conversion (or a secondary transform) tool for the current video block based on the width (W) and / or height (H) of the current video block. The secondary conversion tool includes applying a forward secondary conversion to the output of a forward primary conversion (or a forward primary transform) applied to the residual of the video block before quantization during encoding, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding.
[0006] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining whether a current video block of a coding unit of a video meets a condition according to a rule, and performing a conversion between the current video block and a coded representation of the video according to the determination. The condition is related to the characteristics of one or more color components of the video, the size of the current video block, or the coefficients in a part of the residual block of the current video block. The rule stipulates that the presence of side information regarding the secondary conversion tool in the coded representation is controlled by the condition. The secondary conversion tool includes applying a forward secondary conversion to the output of a forward primary conversion applied to the residual of the video block before quantization during encoding, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding.
[0007] In yet another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes performing a conversion between a current video block of a video and a coded representation of the video, and performing the conversion includes determining how to use a secondary conversion tool and / or signaling information related to the secondary conversion tool according to rules independent of the partition tree type applied to the current video block, where the secondary conversion tool applies a forward secondary conversion to the output of a forward primary conversion applied to the residual of the video block during encoding, before quantization, or applies an inverse secondary conversion to the output of the inverse quantization of the video block during decoding, before applying an inverse primary conversion.
[0008] In yet another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining the applicability of a secondary conversion tool to a current video block of a coding unit of a video, where the coding unit includes a plurality of conversion units and the determination is based on a single conversion unit of the coding unit, and performing a conversion between the current video block and a coded representation of the video based on the determination, where the secondary conversion tool applies a forward secondary conversion to the output of a forward primary conversion applied to the residual of the video block during encoding, before quantization, or applies an inverse secondary conversion to the output of the inverse quantization of the video block during decoding, before applying an inverse primary conversion.
[0009] In yet another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes determining, for a current video block of a coding unit of a video, the applicability of a secondary conversion tool and / or the presence of side information related to the secondary conversion tool, where the coding unit includes a plurality of conversion units and the determination is made at a conversion unit level or a prediction unit level; and performing a conversion between current video blocks of the coded representation of the video based on the determination. The secondary conversion tool includes applying a forward secondary conversion to an output of a forward primary conversion applied to a residual of a video block during encoding before quantization, or applying an inverse secondary conversion to an output of inverse quantization of a video block before applying an inverse primary conversion during decoding.
[0010] In yet another representative aspect, the above method is embodied in a form of code executable by a processor and stored in a computer-readable program medium.
[0011] In yet another representative aspect, a device configured or operable to perform the above method is disclosed. The device may include a processor programmed to implement this method.
[0012] In yet another representative aspect, a video decoder device can implement the method described in this application.
[0013] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims.
Brief Description of the Drawings
[0014]
Figure 1
[0015]
Figure 2
[0016]
Figure 3
[0017]
Figure 4
[0018]
Figure 5
[0019]
Figure 6
[0020]
Figure 7
[0021]
Figure 8
[0022]
Figure 9
[0023]
Figure 10
[0024]
Figure 11
[0025]
Figure 12
[0026]
Figure 13
[0027]
Figure 14
[0028]
Figure 15
[0029]
Figure 16
[0030]
Figure 17
[0031]
Figure 18
[0032]
Figure 19
[0033]
Figure 20
[0034]
Figure 21
[0035]
Figure 22A
Figure 22B
Figure 22C
Figure 22D
[0036]
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0037] Embodiments of the disclosed technology may be applicable to existing video coding standards (e.g., HEVC, H.265) and future standards in order to improve compression performance. The section headings are used in this case to improve the readability of the description and are not intended to limit the description or embodiments (and / or implementations) to only the individual sections in any way.
[0038] 1. Video Coding Introduction Due to the increasing demand for higher-resolution videos, video coding methods and technologies are present everywhere in modern technology. A video codec typically includes electronic circuitry or software that compresses or decompresses digital video and is constantly improved to provide higher coding efficiency. A video codec converts uncompressed video into a compressed format or vice versa. There is a complex relationship between video quality, the amount of data (determined by the bit rate) used to represent the video, the complexity of the encoding and decoding algorithms, the sensitivity to data loss and errors, the ease of editing, random access, and end-to-end latency. Compressed formats usually conform to standard video compression standards such as the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2), the Versatile Video Coding (VVC) standard that is expected to be finalized, or other current and / or future video coding standards.
[0039] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction and transform coding are used. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into a reference software named Joint Exploration Model (JEM)[3][4]. In April 2018, the Joint Video Expert Team (JVET) was launched between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard aiming for a 50% bitrate reduction compared to HEVC.
[0040] 2.1 Coding Flow of a Typical Video Codec Figure 1 shows an example of an encoder block diagram of VVC, which includes three in-loop filtering blocks, namely: a deblocking filter (DF), a sample adaptive offset (SAO), and an ALF. Different from the DF that uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean squared error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter (the coded side information signals the offset and the filter coefficients respectively). ALF is located at the final processing stage of each picture and can be regarded as a tool that attempts to capture and repair the artifacts generated in the previous stage.
[0041] 2.2 Intra Coding in VVC 2.2.1 Intra Mode Coding with 67 Intra Prediction Modes To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 as used in HEVC to 65. The additional direction modes are depicted as dotted arrows in Figure 2, and the planar and DC modes remain the same. These more densely packed directional intra prediction modes are applied to all block sizes and to both luma and chroma intra prediction.
[0042] The conventional angular intra prediction directions are defined from 45 degrees to -135 degrees in the clockwise direction as shown in Figure 2. In VTM2, some of the conventional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original method and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode coding does not change.
[0043] In HEVC, all intra-coded blocks are square, and the length of each side is a power of 2. Therefore, no splitting process is required to generate an intra predictor using the DC mode. In VVV2, blocks can have a rectangular shape, and in the general case, it is necessary to use a splitting process for each block. To avoid the splitting process for DC prediction, only the long side is used to calculate the average for non-square blocks.
[0044] In addition to 67 intra prediction modes, wide-angle intra prediction of non-square blocks (WAIP) and position-dependent intra prediction combination (PDPC) methods are further enabled for specific blocks. PDPC is applied without signaling for the following intra-modes: planar, DC, horizontal, vertical, the lower left angular mode and its eight adjacent angular modes, and the upper right angular mode and its eight adjacent angular modes.
[0045] 2.2.2 Affine Linear Weighted Intra Prediction (ALWIP or Matrix-Based Intra Prediction) Affine Linear Weighted Intra Prediction (ALWIP, also known as Matrix-based Intra Prediction (MIP)) was proposed in JVET-N0217.
[0046] 2.2.2.1 Generation of Reduced Prediction Signal by Matrix-Vector Multiplication First, adjacent reference samples are downsampled by averaging to generate a reduced reference signal bdry red . Then, a reduced prediction signal pred red is calculated by computing a matrix-vector product and adding an offset:
[0047]
Equation
[0048] Here, A is a matrix having 4 columns when W = H = 4 and 8 columns in all other cases, and W red ·H red rows. b is a vector of size W red ·H red .
[0049] 2.2.2.2 Description of the entire ALWIP process The overall processes of averaging, matrix-vector multiplication, and linear interpolation are shown for various shapes in FIGS. 3-6. Note that the remaining shapes are treated in the same way as one of the cases shown.
[0050] Assuming a 1.4×4 block, ALWIP takes two averages along each axis of the boundary. The resulting four input samples enter a matrix-vector multiplication. The matrix is taken from the set S 0 . After adding an offset, this results in 16 final prediction samples. Linear interpolation is not required to generate the prediction signal. Thus, a total of (4·16) / (4·4) = 4 multiplications are performed per sample.
[0051] Assuming a 2.8×8 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is taken from the set S 1 . This results in 16 samples at the odd positions of the prediction block. Thus, a total of (8·16) / (8·8) = 2 multiplications are performed per sample. After adding an offset, these samples are interpolated vertically by using the reduced upper and lower boundaries. Horizontal interpolation follows by using the original left boundary.
[0052] Assuming a 3.8×4 block, ALWIP takes four averages along the horizontal axis of the boundary and four original boundary values at the left boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is taken from the set S 1It is taken from. This results in 16 samples at the odd horizontal positions and each vertical position of the prediction block. Thus, a total of (8·16) / (8·4) = 4 multiplications per sample are performed. After adding the offset, these samples are interpolated horizontally by using the original left boundary.
[0053] Assuming a 4.16×16 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is taken from set S 2 It is taken from. This results in 64 samples at the odd positions of the prediction block. Thus, a total of (8·64) / (16·16) = 2 multiplications per sample are performed. After adding the offset, these samples are interpolated vertically by using the eight averages of the upper boundary. Horizontal interpolation continues by using the original left boundary. In this case, the interpolation process does not add any multiplications. Thus, overall, two multiplications per sample are required to compute the ALWIP prediction.
[0054] For larger shapes, the procedure is essentially the same, and it is easy to verify that the number of multiplications per sample is less than 4.
[0055] For a W×8 (W>8) block, only horizontal interpolation is required because the samples are at odd horizontal positions and given at each vertical position.
[0056] Finally, for a W×4 (W>8) block, let A_kbe be the matrix resulting from excluding all rows corresponding to the odd entries along the horizontal axis of the downsampled block. Thus, the output size is 32, and again, only horizontal interpolation remains to be performed. The transposed case is handled accordingly. 2.2.2.3 Syntax and Semantics 7.3.6.5 Coding Unit Syntax
[0057]
Number
[0058] 2.2.3 Multiple reference line (MRL) Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 7, an example of four reference lines is depicted, where the samples of segments A and F are not fetched from the reconstructed adjacent samples but are padded with the nearest samples from segments B and E respectively. HEVC intra-picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used.
[0059] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line indices greater than 0, only include the additional reference line mode in the MPM list and only signal the MPM index without the remaining modes. The reference line index is signaled before the intra prediction mode, and the Planar and DC modes are excluded from the intra prediction mode if a non-zero reference line index is signaled.
[0060] MRL is disabled for the first line of the block inside the CTU and prohibits using the extended reference samples outside the current CTU line. Also, PDPC is disabled if additional lines are used.
[0061]
Number
[0062]
Number
[0063]
Number
[0064]
Number
[0065]
Number
[0066] To maintain the orthogonality of the transformation matrix, the transformation matrix is quantized more precisely than the transformation matrix of HEVC. After horizontal transformation and after vertical transformation, all coefficients should have 10 bits in order to keep the intermediate values of the transformed coefficients within the 16-bit range.
[0067] To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter respectively. When MTS is enabled in the SPS, a CU-level flag is signaled to specify whether MTS is applied or not. Here, MTS is applied only to luma. The MTS CU-level flag is signaled when the following conditions are met.
[0068] ○ Both the width and height are 32 or less
[0069] ○ The CBF flag is equal to 1
[0070]
Number
[0071] To reduce the complexity of the large-sized DST-7 and DCT-8, for DST-7 and DCT-8 blocks having a size equal to 32 (width or height, or both width and height), the high-frequency transform coefficients are zeroed out. Only the coefficients within the 16×16 low-frequency region are retained.
[0072] In addition to the different transforms being applied, VVC also supports a mode called transform skip (TS) that is similar to the concept of TS in HEVC. TS is treated as a special case of MTS.
[0073] 2.4.2 Reduced secondary transform (RST) proposed in JVET-N0193 2.4.2.1 Non-separable secondary transform (NSST) in JEM In JEM, secondary transform is applied between the forward primary transform and quantization (encoder side), and between inverse quantization and inverse primary transform (decoder side). As shown in Figure 10, the 4×4 (or 8×8) secondary transform is executed depending on the block size. For example, the 4×4 secondary transform is applied to small blocks (i.e., min (width, height) < 8), and for each 8×8 block, the 8×8 secondary transform is applied to larger blocks (i.e., min (width, height) > 4).
[0074] The application of non-separable transform is described using the following input as an example. To apply non-separable transform, first,
[0075]
Number
[0076]
Number
[0077]
Number
[0078]
Number
[0079] 2.4.2.2 Reduced Secondary Transform (RST) in JVET-N0193 RST (Low Frequency Non-Separable Transform, LFNST) was introduced in JVET-K0099, and the mapping of four transform sets (instead of 35 transform sets) was introduced in JVET-L0133. In this JVET-N0193, 16×64 matrices (further reduced to 16×48 matrices) and 16×16 matrices are used. For notation convenience, the 16×64 (reduced to 16×48) transform is denoted as RST8×8, and the 16×16 transform is denoted as RST4×4. Figure 11 shows an example of RST.
[0080] 2.4.2.2.1 RST Calculation The main idea of the reduced transform (RT) is to map an N-dimensional vector to an R-dimensional vector in another space, where R / N (R < N) is the reduction factor.
[0081]
Number
[0082] Here, the R rows of the transform matrix are the R bases in the N-dimensional space. The inverse transform matrix for RT is the transpose of its forward transform matrix. The forward and inverse RTs are depicted in Figure 12.
[0083] In this contribution, RST8x8 with a reduction factor of 4 (1 / 4 size) is applied. Therefore, instead of the conventional 64×64 non-separable transform matrix size, a 16×64 direct matrix is used. In other words, a 64×16 inverse RST matrix is used on the decoder side to generate core (primary) transform coefficients in the upper left 8×8 region. The forward RST8x8 uses a 16×64 (or 8×64 for an 8×8 block) matrix, and as a result, generates non-zero coefficients only in the upper left 4×4 region within a given 8×8 region. In other words, when RST is applied, the 8×8 region excluding the upper left 4×4 region will have only zero coefficients. In the case of RST4×4, 16×16 (8×16 for a 4×4 block), direct matrix multiplication is applied
[0084] Inverse RST is conditionally applied when the following two conditions are met:
[0085] ○ The block size is greater than or equal to a given threshold (W >= 4 && H >= 4)
[0086] ○ The transform skip mode flag is equal to zero
[0087] When both the width (W) and height (H) of the transform coefficient block are greater than 4, RST8x8 is applied to the upper left 8×8 region of the transform coefficient block. Otherwise, RST4x4 is applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block
[0088] When the RST index is equal to 0, RST is not applied. Otherwise, RST is applied and its kernel is selected using the RST index. The method of selecting RST and the coding of the RST index will be described later
[0089] Furthermore, RST is applied to intra CUs in both intra and inter slices, and to both luma and chroma. When dual-tree is enabled, the RST indices for luma and chroma are signaled separately. In the case of inter slice (dual-tree is disabled), a single RST index is signaled and used for both luma and chroma.
[0090] 2.4.2.2.2 Constraints of RST When the ISP mode is selected, RST is disabled and the RST index is not signaled because, even if RST is applied to all appropriate partition blocks, the performance improvement is only marginal. Furthermore, disabling RST for the residuals predicted by ISP may reduce the encoding complexity.
[0091] 2.4.2.2.3 RST Selection The RST matrix is selected from four sets of transforms, each of which consists of two transforms. Which set of transforms is applied is determined from the intra prediction mode as follows:
[0092] (1) If one of the three CCLM modes is specified, transform set 0 is selected.
[0093]
Number
[0094] The index for accessing the above table denoted as IntraPredMode has a range of [-14, 83], which is the transform mode index used for wide-angle intra prediction.
[0095] 2.4.2.2.4 RST Matrix for Reduced Dimensions For further simplification, using the same set of transform settings, a 16×48 matrix is applied instead of a 16×64 matrix, each of which takes in 48 data from three 4×4 blocks within the upper left 8×8 block (excluding the lower right 4×4 block) (shown in FIG. 13).
[0096] 2.4.2.2.5 RST Signaling Forward RST8x8 uses a 16×48 matrix, and as a result, generates non-zero coefficients only in the upper left 4×4 area among the first three 4×4 areas. In other words, when RST8x8 is applied, only the upper left 4×4 (by RST8x8) and lower right 4×4 (by primary transform) areas may have non-zero coefficients. As a result, if any non-zero elements are detected in the upper right 4×4 and lower left 4×4 block areas (shown in FIG. 14 and called the "zero-out" areas), the RST index is not coded because that indicates that RST was not applied in that case. In such cases, the RST index is presumed to be zero.
[0097] 2.4.2.2.6 Zero-Out Areas in One CG Normally, before applying the inverse RST to a 4×4 sub-block, some coefficients within the 4×4 sub-block may be non-zero. However, in some cases, before the inverse RST is applied to the sub-block, some coefficients of the 4×4 sub-block are restricted to be zero.
[0098] Let nonZeroSize be a variable. Before the inverse RST, any coefficient having an index not smaller than nonZeroSize when rearranged into a 1-D array is required to be zero.
[0099] When nonZeroSize is equal to 16, there is no zero-out constraint on the coefficients of the upper left 4×4 sub-block.
[0100]
Number
[0101] 2.4.3 Sub - block Transformation In the case of an inter - predicted CU having cu_cbf equal to 1, cu_sbt_flag may be signaled to indicate whether the entire residual block or a sub - part of the residual block is decoded. In the former case, inter - MTS information is further analyzed to determine the transformation type of the CU. In the latter case, a part of the residual block is coded with an estimated adaptive transformation and the other part of the residual block is zeroed out. SBT is not applied to the combined inter - intra mode.
[0102]
Number
[0103] 2.4.4 Quantized Residual Domain Block Differential Pulse Code Modulation Coding (QR - BDPCM) In JVET - N0413, quantized residual domain BDPCM (hereinafter referred to as RBDPCM) has been proposed. Similar to intra prediction, intra prediction is performed for the entire block by performing sample copy in the prediction direction (horizontal or vertical prediction). The residual is quantized, and the delta between the quantized residual and the quantized value of its predictor (horizontal or vertical) is coded.
[0104] When the block has a size of M (rows)×N (columns), r i,j, where \(0\leq i\leq M - 1\) and \(0\leq j\leq N - 1\), is the prediction residual after performing intra prediction horizontally (copying the left adjacent pixel values across the prediction block line by line using non-filtered samples from the upper or left block boundary samples) or vertically (copying the upper adjacent line to each line of the predicted block). Let \(Q(r i,j ), where \(0\leq i\leq M - 1\) and \(0\leq j\leq N - 1\), be the quantized version of the residual \(r i,j ), where the residual is the difference between the values of the original block and the predicted block. Then, block DPCM is applied to the quantized residual samples, resulting in a modified \(M\times N\) array \(R ~ i,j having elements \(r ~ ). When vertical BDPCM is signaled, it is as follows:
[0105]
Number
[0106] For horizontal prediction, similar rules apply and the residual quantization samples are obtained as follows
[0107]
Number
[0108]
Number
[0109]
Number
[0110]
Number
[0111] For the horizontal case, it is as follows.
[0112] [Number]
[0113] Inverted quantization residue Q -1 (Q(r i,j )) is added to the intra-block prediction value to generate the reconstructed sample value.
[0114] When QR-BDPCM is selected, there is no applied transform.
[0115] 2.5 Coefficient entropy coding 2.5.1 Coefficient coding of transform application blocks In HEVC, the transform coefficients of a coding block are coded using non-overlapping coefficient groups (or sub-blocks), and each CG contains the coefficients of a 4×4 block of the coding block. The CGs within a coding block and the transform coefficients within a CG are coded according to a predefined scan order.
[0116] The CGs within a coding block and the transform coefficients within a CG are coded according to a predefined scan order. Both the CGs and the coefficients within a CG follow the diagonal upward right scan order. Examples of the 4×4 block and 8×8 scan orders are depicted in FIGS. 16 and 17, respectively.
[0117] Note that the coding order is the reverse scan order (i.e., decoding from CG3 to CG0 in FIG. 17), and when decoding a block, the coordinates of the last non-zero coefficient are decoded first.
[0118] Coding of the transform coefficient levels of a CG having at least one non-zero transform coefficient may be separated into a plurality of scan paths. In the first path, a first bin (denoted by bin0 and also referred to as significant_coeff_flag, which indicates that the magnitude of the coefficient is greater than 0) is coded. Next, two scan paths may be applied to the context of coding the second / third bins (designated by bin1 and bin2 respectively and also referred to as coeff_abs_greater1_flag and coeff_abs_greater2_flag). Finally, two or more scan paths for coding the sign information and the remaining values of the coefficient levels (also called coeff_abs_level_remaining) are invoked as needed. Note that only the bins of the first three scan paths are coded in regular mode, and those bins are referred to as regular bins in the following description.
[0119] In VVC 3, for each CG, the regularly coded bins and the bypass-coded bins are separated in the coding order; first, all the regularly coded bins for the sub-block are sent, and then the bypass-coded bins are sent. The transform coefficient levels of the sub-block are coded as follows in five paths at the scan position:
[0120] ○ Path 1: Coding of significance (sig_flag), greater than 1 flag (gt1_flag), parity (par_level_flag), greater than 2 flag (gt2_flag) is processed in the coding order. If sig_flag is equal to 1, first, gt1_flag is coded (which specifies whether the absolute level is greater than 1). If gt1_flag is equal to 1, par_flag is additionally coded (which specifies the parity of the absolute level minus 2).
[0121] ○ Path 2: The coding of the remaining absolute level (remainder) is processed for all scan positions having gt2_flag equal to 1 or gt1_flag equal to 1. Non-binary syntax elements are binaryized with Golomb-Rice code, and the resulting bins are coded in the bypass mode of the arithmetic coding engine.
[0122] ○ Path 3: The absolute level (absLevel) of coefficients for which sig_flag was not coded in the first pass (due to reaching the limit of regularly coded bins) is fully coded in the bypass mode of the arithmetic coding engine using Golomb-Rice code.
[0123] ○ Path 4: Code the sign (sign_flag) for all scan positions having sig_coeff_flag equal to 1.
[0124] For 4x4 sub-blocks, it is guaranteed that no more than 32 regularly coded bins (sig_flag, par_flag, gt1_flag and gt2_flag) are encoded or decoded. In the case of 2×2 chroma sub-blocks, the number of regularly coded bins is limited to 8.
[0125] (Path 3) The Rice parameter (ricePar) for coding the non-binary syntax element remainder is derived in the same way as in HEVC. At the start of each sub-block, ricePar is set equal to 0. After coding the syntax element remainder, the Rice parameter is modified according to a predefined formula. (Path 4) To code the non-binary syntax element absLevel, the sum of absolute values sumAbs in the local template is determined. The variables ricePar and posZero are determined by table look-up based on the dependent quantization and sumAb. The intermediate variable codeValue is derived as follows:
[0126] ○ If absLevel[k] is equal to 0, codeValue is set equal to posZero;
[0127] ○ Otherwise, if absLevel[k] is less than or equal to posZero, codeValue is set equal to absLevel[k] - 1;
[0128] ○ Otherwise (absLevel[k] is greater than posZero), codeValue is set equal to absLevel[k].
[0129] The value of codeValue is coded using a Golomb-Rice code together with the Rice parameter ricePar.
[0130] 2.5.1.1 Context Modeling for Coefficient Coding The selection of the probability model for the syntax elements related to the absolute value of the transform coefficient levels depends on the values of the absolute value levels in the local neighborhood or the values of the partially reconstructed absolute value levels. The templates used are shown in Figure 18.
[0131] The selected probability model depends on the sum of the absolute value levels (or the partially reconstructed absolute value levels) in the local neighborhood and the number of absolute value levels greater than 0 in the local neighborhood (given by the number of sig_coeff_flags equal to 1). Context modeling and binarization depend on the following metrics for the local neighborhood:
[0132] ○ numSig: The number of non-zero levels in the local neighborhood;
[0133] ○ sumAbs1: The sum of the partially reconstructed absolute value levels (absLevel1) after the first pass in the local neighborhood;
[0134] ○ sumAbs: The sum of the reconstructed absolute value levels in the local neighborhood;
[0135] ○ Diagonal position (d): The sum of the horizontal and vertical coordinates of the current scan position within the transform block.
[0136] Based on the values of numSig, sumAbs1, and d, a probability model for coding sig_flag, par_flag, gt1_flag, and gt2_flag is selected. The Rice parameter for binarizing abs_remainder is selected based on the values of sumAbs and numSig.
[0137] 2.5.1.2 Dependent Quantization (DQ) Furthermore, the same HEVC scalar quantization is used together with a new concept called dependent scale quantization. Dependent scale quantization refers to an approach where the set of allowable reconstructed values of the transform coefficients depends on the values of the transform coefficient levels that precede the current transform coefficient level in the reconstruction order. The main effect of this approach is that, compared to the conventional independent scalar quantization used in HEVC, the allowable reconstruction vectors are packed more densely in the N - dimensional vector space (where N represents the number of transform coefficients in the transform block). This means that for a given average number of allowable reconstruction vectors per unit volume in the N - dimensional space, the average distortion between the input vector and the closest reconstruction vector is reduced. The approach of dependent scale quantization is realized by: (a) defining two scalar quantizers with different reconstruction levels, and (b) defining a process for switching between the two scalar quantizers.
[0138] Two scalar quantizers, denoted as Q0 and Q1, are shown in FIG. 19. The positions of the available reconstruction levels are uniquely specified by the quantization step size Δ. The scalar quantizer (Q0 or Q1) used is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the transform coefficient levels preceding the current transform coefficient in the coding / reconstruction order.
[0139]
Number
[0140] 2.5.2 Coefficient Coding of TS-Coded Blocks and QR-BDPCM-Coded Blocks QR-BDPCM follows the context modeling method of TS-coded blocks.
[0141] Modified Transform Coefficient Level Coding for TS Residuals Compared with the case of regular residual coding, the residual coding for TS includes the following changes:
[0142] (1) There is no signaling of the last x / y position.
[0143] (2) coded_sub_block_flag is coded for all sub-blocks except the last sub-block when all preceding flags are equal to 0.
[0144] (3) sig_coeff_flag context modeling using a reduced template
[0145] (4) Single context model for abs_level_gt1_flag and par_level_flag
[0146] (5) Context modeling for sign flag, additional flags greater than 5, 7, 9
[0147] (6) Modified Rice parameter for remainder binary evolution
[0148] [Number] TIFF0007684495000036.tif255164TIFF0007684495000037.tif170170
[0149] 3 Defects of existing implementations The current design has the following problems:
[0150] (1) The four predefined transform sets for the chroma component are the same as those for the luma component. Furthermore, luma and chroma blocks with the same intra prediction mode use the same transform set. However, typically the chroma signal is smoother compared to the luma component. Using the same set may not be optimal.
[0151] (2) RST is applied only to specific CGs, not all CGs. However, the decision regarding signaling the RST index depends on the number of non-zero coefficients in all blocks. If all coefficients of the CG to which RST is applied are zero, there is no need to signal the RST index. However, the current design may still signal the index, wasting unnecessary bits.
[0152] (3) The RST index is signaled after residual coding because it is necessary to record whether non-zero coefficients exist at a predetermined location and how many there are (e.g., numZeroOutSigCoeff, numSigCoeff in Section 2.3.2.2.7). Such a design makes the analysis process more complex.
[0153] (4) The RST index is context-coded, and context modeling depends on the coded luma / chroma intra prediction mode and the MTS index. Such a design introduces a delay in analysis regarding the reconstruction of the intra prediction mode. Also, eight contexts are introduced, which may impose a burden on the hardware implementation.
[0154] (a) DM and CCLM share the same context index offset, but since they are two different chroma intra prediction methods, it doesn't make sense.
[0155] (5) The current design of non-TS residual coding first codes the coefficient information, followed by coding of the RST index (i.e., whether to use RST or not, and if so, which matrix is selected). In such a design, in the entropy coding of the residual, the on / off information of RST cannot be considered.
[0156] (6) RST is always applied to the upper left region of the transform block to which the primary transform is applied. However, when based on different primary transforms, it is not always correct that energy concentrates in the upper left region of the transform block.
[0157] (7) The decision on whether to signal RST-related information is made in different ways for the dual-tree and single-tree coding structures.
[0158] If there are more than one TU in a CU (for example, the CU size is 128×128), whether to analyze the RST-related information can only be determined after decoding all the TUs. For example, in the case of a 128×128 CU, the first PB cannot process without waiting for the LFNST index that comes after the last PB. This does not necessarily break the overall 64×64-based decoder pipeline (if CABAC is separable), but increases the data buffering fourfold for a specific number of decoder pipeline stages. This is costly.
[0159] 4 Example methods for context modeling of residual coding Embodiments of the technology disclosed herein overcome the drawbacks of existing implementations, thereby providing video coding with higher coding efficiency. The method for context modeling for residual coding based on the disclosed technology may enhance both existing and future video coding standards and will be elucidated in the following examples described for various implementations. The examples of the disclosed technology provided below illustrate general concepts and are not intended to be construed as limiting. In the examples, various features described in these examples can be combined unless otherwise specified explicitly.
[0160] In the following description, "block" may refer to a coding unit (CU), a transform unit (TU), or some rectangular region of video data. "Current block" may refer to a decoded / encoded coding unit (CU), a transform unit (TU) currently being decoded / encoded, or some coding rectangular region of video data currently being decoded / encoded. "CU" or "TU" may also be known as "coding block" and "transform block".
[0161] In these embodiments, RST may be any variation of the design in JVET-N0193. RST may apply the secondary transformation to one block or may be any technique (e.g., the RST proposed in JVET-N0193 applied to a transform skip (TS)-coded block) capable of applying a certain transformation to a transform skip (TS)-coded block.
[0162] Furthermore, the "zero-out region" or "zero-out CG" may always specify a region / CG having zero coefficients due to the reduced transform size used in the secondary transform process. For example, if the secondary transform size is 16×32 and the CG size is 4×4, it is applied to the first two CGs, but only the first CG may have non-zero coefficients, and the second 4×4 CG may also be referred to as a zero-out CG. Selection of the transformation matrix in RST 1. The sub-region to which RST is applied may be a sub-region that is not the upper left part of the block. a. In one example, RST may be applied to the upper right, lower right, lower left, or central sub-region of the block. b. The sub-region to which RST is applied may depend on the intra prediction mode and / or the primary transform matrix (e.g., DCT-II, DST-VII, identity transform). 2. The selection of the transform set and / or transform matrix used in RST may depend on the color component. a. In one example, one set of transform matrices may be used for the luma (or G) component and one set may be used for the chroma components (or B / R components). b. In one example, each color component may correspond to one set. c. In one example, at least one matrix differs in any of two or more sets for different color components. 3. The selection of the transform set and / or transform matrix used in RST may depend on the intra prediction method (e.g., CCLM, multiple reference line-based intra prediction method, matrix-based intra prediction method). a. In one example, one set of transformation matrices may be used for CCLM-coded blocks and others may be used for non-CCLM-coded blocks. b. In one example, one set of transformation matrices may be used for normally intra-prediction-coded blocks and others may be used for blocks with multiple reference lines enabled (i.e., those that do not use adjacent lines for intra prediction). c. In one example, one set of transformation matrices may be used for blocks using joint chroma residual coding and others may be used for blocks to which joint chroma residual coding is not applied. d. In one example, at least one matrix differs in any of two or more sets for different intra prediction methods. e. Alternatively, RST may be disabled for blocks coded with a specific intra prediction direction and / or a specific coding tool, e.g., CCLM, and / or joint chroma residual coding, and / or a specific color component (e.g., chroma). 4. The selection of the transform set and / or transformation matrix used in RST may depend on the primary transform. a. In one example, when the primary transform applied to one block is the identity transform (e.g., the TS mode is applied to one block), the transform set and / or transformation matrix used in RST may be different from those for other types of primary transforms. b. In one example, when the horizontal and vertical 1-D primary transforms applied to one block are in the same basis (e.g., both are DCT-II), the transform set and / or transformation matrix used in RST may be different from those for its primary transform with different bases in different directions (vertical or horizontal). Signaling of RST side information and residual coding 5. Whether and / or how to signal the side information (e.g., st_idx) of RST may depend on the last non-zero coefficient within the block (in scanning order). a. In one example, the RST may be enabled and the index of the RST may be signaled only if the last non-zero coefficient is located within the CG to which the RST is applied. b. In one example, if the last non-zero coefficient is not located within the CG to which the RST is applied, the RST is disabled and the signaling of the RST is skipped. 6. Whether and how to signal the side information of the RST (e.g., st_idx) may depend on the coefficients of a specific color component rather than all available color components within the CU. a. In one example, only the side information may be used to determine whether and how to signal the side information of the RST. i. Alternatively, further, the above method is applied only when the block size meets certain conditions. 1) The condition is W < T1 or H < T2. 2) For example, T1 = T2 = 4. Therefore, for a 4X4 CU, the luma block size is 4x4, and the two chroma blocks in the 4:2:0 format are 2x2. In this case, only the luma information may be used. ii. Alternatively, further, the above method is applied only when the current partition type tree is a single tree. b. Whether to use the information of one color component or the information of all color components may depend on the block size / coded information. 7. Whether and how to signal the side information of the RST (e.g., st_idx) may depend on the coefficients within a partial region of one block rather than the entire block. a. In one example, the partial region may be defined as the CG to which the RST is applied. b. In one example, the partial region may be defined as the first or last M (e.g., M = 1 or 2) CGs in the scanning order or the reverse scanning order of the block. i. In one example, M may depend on the block dimension. ii. In one example, when the block size is 4xN and / or Nx4 (N > 8), M is set to 2. iii. In one example, when the block size is 4x8 and / or 8x4 and / or WxH (W ≥ 8, H ≥ 8), M is set to 1. c. In one example, information about a block having dimensions W×H (e.g., the number of non-zero coefficients of the block) may be allowed to be taken into account to determine the use of RST and / or the signaling of RST-related information. i. For example, when W < T1 or H < T2, the number of non-zero coefficients of the block may not be counted. For example, T1 = T2 = 4. d. In one example, a partial region may be defined as the upper left M×N region of the current block having dimensions W×H. i. In one example, M may be smaller than W, and / or N may be smaller than H. ii. In one example, M and N may be fixed numbers, for example, M = N = 4. iii. In one example, M and / or N may depend on W and / or H. iv. In one example, M and / or N may depend on the maximum allowable transform size. 1) For example, when W is greater than 8 and H is equal to 4, M = 8 and N = 4. 2) For example, when H is greater than 8 and W is equal to 4, M = 4 and N = 8. 3) For example, when neither of the above two conditions is satisfied, M = 4 and N = 4. v. Alternatively, furthermore, these methods may be applied only for specific block dimensions where the condition of 7.c is not satisfied. e. In one example, the partial region may be the same for all blocks. i. Alternatively, this may be changed based on the block dimensions and / or the coded information. f. In one example, the partial region may depend on a given range of the scanning order index. i. In one example, the partial region may cover coefficients located within a specific range with those scanning order indices within [dxS, IdxE] including both ends, based on the coefficient scan order (e.g., reverse decoding order) of the current block having dimensions W×H. 1) In one example, IdxS is equal to 0. 2) In one example, IdxE may be smaller than W×H - 1. 3) In one example, IdxE may be a fixed number. For example, IdxE = 15. 4) In one example, IdxE may depend on W and / or H. a. For example, when W is greater than 8 and H is equal to 4, IdxE = 31. b. For example, when H is greater than 8 and W is equal to 4, IdxE = 31. c. For example, when W is equal to 8 and H is equal to 8, IdxE = 7. d. For example, when W is equal to 4 and H is equal to 4, IdxE = 7. e. For example, when neither of the above two conditions a) and b) is satisfied, IdxE = 15. f. For example, when none of the above two conditions a), b), c) and d) is satisfied, IdxE = 15. g. For example, when neither of the above two conditions c) and d) is satisfied, IdxE = 15. ii. Alternatively, furthermore, these methods may be applied only for specific block dimensions such that the condition of 7.c is not satisfied. g. In one example, this may depend on the positions of non - zero coefficients within the partial region. h. In one example, this may depend on the energy (such as the sum of squares or the sum of absolute values) of non - zero coefficients within the partial region. i. In one example, this may depend on the number of non-zero coefficients within a partial region of one block rather than the entire block. i. Alternatively, this may depend on the number of non-zero coefficients within a partial region of one or more blocks within a CU. ii. If the number of non-zero coefficients within a partial region of one block is less than a threshold, signaling of side information for RST may be skipped. iii. In one example, the threshold is fixed to be N (e.g., N = 1 or 2). iv. In one example, the threshold may depend on slice type / picture type / partition tree type (dual or single) / video content (screen content or content captured by a camera). v. In one example, the threshold may depend on a color format such as 4:2:0 or 4:4:4, and / or color components such as Y or Cb / Cr. 8. If there are no non-zero coefficients in the CG where RST may be applied, RST shall be disabled. a. In one example, when RST is applied to one block, at least one CG to which RST is applied shall contain at least one non-zero coefficient. b. In one example, for 4×N and / or N×4 (N > 8), when RST is applied, the first two 4×4 CGs shall contain at least one non-zero coefficient. c. In one example, for 4×8 and / or 8×4, when RST is applied, the top-left 4×4 shall contain at least one non-zero coefficient. d. In one example, for WxH (W >= 8 and H >= 8), when RST is applied, the top-left 4×4 shall contain at least one non-zero coefficient. e. The compliant bitstream shall satisfy one or more of the above conditions. 9. Syntax elements related to RST may be signaled before coding the residual (e.g., transform coefficients / things directly quantized). a. In one example, the counting of the number of non-zero coefficients in the zero-out region (e.g., numZeroOutSigCoeff) and the number of non-zero coefficients in the entire block (e.g., numSigCoeff) is excluded in the coefficient analysis process. b. In one example, RST-related syntax elements (e.g., st_idx) may be coded before residual_coding. c. RST-related syntax elements may be signaled conditionally (e.g., according to the coded block flag, the usage of the TS mode). vi. In one example, RST-related syntax elements (e.g., st_idx) may be coded after the signaling of the coded block flag or after the signaling of the TS / MTS-related syntax elements. vii. In one example, when the TS mode is enabled (e.g., when the decoded transform_skip_flag is equal to 1), the signaling of the RST-related syntax elements is skipped. d. Residual-related syntax may not be signaled for zero-out CG. e. How to code the residual (e.g., scanning order, binary evolution, syntax to be decoded, context modeling) may depend on RST. i. In one example, instead of the diagonal right scanning order, the raster scanning order may be applied. 1) The raster scanning order is from left to right, from top to bottom, or the reverse order. 2) Alternatively, instead of the diagonal right scanning order, the vertical scanning order (from top to bottom, from left to right, or the reverse order) may be applied. 3) Alternatively, furthermore, the context modeling may be modified. a. In one example, context modeling may depend on the previously coded information of the nearest N neighbors in scan order within the template, rather than using the right, bottom, and bottom - right neighbors. b. In one example, context modeling may depend on the previously coded information (e.g., assuming the current index is equal to 0, -1, -2,...) within the template according to the scanned index. ii. In one example, different binarization methods (e.g., Rice parameter derivation) may be applied to code the residuals associated with RST - coded blocks and non - RST - coded blocks. iii. In one example, the signaling of certain syntax elements may be skipped for RST - coded blocks. 1) The signaling of the coded_sub_block_flag for the CG - coded block with respect to the CG to which RST is applied may be skipped. a. In one example, when RST8x8 is applied to the first 3 CGs in diagonal scan order, the signaling of the coded_sub_block_flag is skipped for the 2nd and 3rd CGs, e.g., the upper - right 4×4 CG and the lower - left 4×4 CG of the upper - left 8x8 region of the block. i. Alternatively, further, the corresponding coded_sub_block_flag is assumed to be 0, i.e., all coefficients are zero. b. In one example, when RST is applied to 1 block, the signaling of the coded_sub_block_flag is skipped for the first CG in scan order (or the last CG in reverse scan order). ii. Alternatively, further, the coded_sub_block_flag for the upper - left CG within the block is assumed to be 1, i.e., it contains at least one non - zero coefficient. c. An example of an 8×8 block is shown in FIG. 21. When RST8x8 or RST4x4 is applied to an 8x8 block, the coded_sub_block_flag of CG0 is presumed to be 1, and the coded_sub_block_flag of CG1 and CG2 is presumed to be 0. 2) The signaling of the magnitude and / or sign flag of the coefficient for a specific coordinate may be skipped. a. In one example, when the index for one CG in the scan order is not less than the maximum allowable index where non-zero coefficients may exist (e.g., nonZeroSize in section 0), the signaling of the coefficient may be skipped. b. In one example, the signaling of syntax elements such as sig_coeff_flag, abs_level_gtX_flag, par_level_flag, abs_remainder, coeff_sign_flag, and dec_abs_level may be skipped. 3) Alternatively, the signaling of the residual (e.g., CG coded block flag, coefficient magnitude, and / or sign flag for a specific coordinate) may be maintained, but the context modeling may be modified to be different from other CGs. iv. In one example, the coding of the residual in the CG to which RST is applied and other CGs may be different. 1) Regarding the above clauses, they may be applied only to the CG to which RST is applied. 10. RST-related syntax elements may be signaled before other transformation instructions such as transform skip and / or MTS index. a. In one example, the signaling of the transform skip may depend on the RST information. i. In one example, when RST is applied within a block, the transform skip instruction is not signaled and is presumed to be 0 for the block. b. In one example, the signaling of the MTS index may depend on the RST information. i. In one example, when RST is applied within a block, one or more MTS conversion instructions are not signaled and are presumed not to be used for the block. 11. In arithmetic coding for different parts within one block, it is proposed to use different context modeling methods. a. In one example, the block is treated as having two parts: the first M CGs in the scanning order and the remaining CGs. i. In one example, M is set to 1. ii. In one example, M is set to 2 for 4xN and Nx4 (N > 8) blocks, and set to 1 for all other cases. b. In one example, the block is treated as having two parts: a sub-region where RST is applied and a sub-region where RST is not applied. i. When RST4x4 is applied, the sub-region where RST is applied is the first one or two CGs of the current block. ii. When RST4x4 is applied, the sub-region where RST is applied is the first three CGs of the current block. c. In one example, in the context modeling process for the first part within one block, using the previously coded information is disabled, but it is proposed to enable it for the second part. d. In one example, when decoding the first CG, the information of one or more remaining CGs may not be allowed to be used. i. In one example, when coding the CG coding block flag for the first CG, the value of the second CG (e.g., right or down) is not considered. ii. In one example, when coding the CG coding block flag for the first CG, the values of the second and third CGs (e.g., right and down CGs for WxH (W ≥ 8 and H ≥ 8)) are not considered. iii. In one example, when coding the current coefficient, if the neighbors in the context template are in different CGs, the information from these neighbors is prohibited from being used. e. In one example, when decoding the coefficients in the region where RST is applied, the information in the remaining regions where RST is not applied may be prohibited from being used. f. Alternatively, furthermore, the above method may be applied under specific conditions. i. The conditions may include whether RST is enabled. ii. The conditions may include the block size. Context modeling in arithmetic coding of RST side information 12. When coding the RST index, context modeling may depend on whether explicit or implicit Multiple-Transform Selection (MTS) is enabled. a. In one example, when implicit MTS is enabled, different contexts may be selected for blocks coded in the same intra prediction mode. i. In one example, block dimensions such as shape (square or non-square) are used to select the context. b. In one example, instead of checking the transform index (e.g., tu_mts_idx) coded for explicit MTS, the basis of the transform matrix may be used. i. In one example, for the transform matrix basis using DCT-II for both horizontal and vertical 1-D transforms, the corresponding context may be different from other types of transform matrices. 13. When coding the RST index, context modeling may depend on whether CCLM is enabled (e.g., sps_cclm_enabled_flag). a. Alternatively, whether or not to enable the selection of a context for RST index coding, or how to select it, may depend on whether CCLM is applied to one block. b. In one example, context modeling may depend on whether CCLM is enabled for the current block. i. The following is an example. intraModeCtx = sps_cclm_enabled_flag? ( intra_chroma_pred_mode[ x0 ][ y0 ] is CCLM: intra_chroma_pred_mode[ x0 ][ y0 ] is DM)? 1 : 0. c. Alternatively, whether or not to enable the selection of a context for RST index coding, or how to select it, may depend on whether the current chroma block is coded in DM mode. i. The following is an example. intraModeCtx = ( intra_chroma_pred_mode[ x0 ][ y0 ] == (sps_cclm_enabled_flag? 7:4) )? 1 : 0. 14. When coding the RST index, context modeling may depend on the block dimension / split depth (e.g., quadtree depth and / or BT / TT depth). 15. When coding the RST index, context modeling may depend on the color format and / or color component. 16. When coding the RST index, context modeling may be independent of the intra prediction mode and / or MTS index. 17. When coding the RST index, the first and / or second bin may be context-coded with only one context or may be bypass-coded. Starting the RST process under conditions 18. Whether to start the inverse RST process may depend on the CG coded block flag. a. In one example, if the top - left CG coded block flag is zero, there is no need to start the process. i. In one example, if the top - left CG coded block flag is zero and the block size is not equal to 4xN / Nx4 (N > 8), there is no need to start the process. b. In one example, if both of the first two CG coded block flags in the scanning order are equal to zero, there is no need to start the process. i. In one example, if both of the first two CG coded block flags in the scanning order are equal to zero and the block size is equal to 4xN / Nx4 (N > 8), there is no need to start the process. 19. Whether to start the inverse RST process may depend on the block size. a. In one example, for a specific block size such as 4×8 / 8×4, RST may be disabled. Alternatively, further, the signaling of RST - related syntax elements may be skipped. Unification of dual-tree and single-tree coding 20. The usage of RST and / or the signaling of RST - related information may be determined in the same way in dual - tree and single - tree coding. a. For example, if the number of non - zero coefficients to be counted (e.g., numSigCoeff defined in JVET - N0193) is not greater than T1 in the case of dual - tree coding or not greater than T2 in the case of single - tree coding, RST should not be applied and the related information should not be signaled, where T1 is equal to T2. b. In one example, both T1 and T2 are set to N, for example, N = 1 or 2. Consider multiple TUs within a CU. Whether to apply RST and / or how to apply it may depend on the block dimensions W×H. a. In one example, if W>T1 or H>T2, RST may not be applied. b. In one example, if W>T1 and H>T2, RST may not be applied. c. In one example, if W*H>=T, RST may not be applied. d. Regarding the above examples, the following apply: i. In one example, the block is a CU. ii. In one example, T1=T2=64. iii. In one example, T1 and / or T2 may depend on the maximum allowable transformation size. For example, T1=T2=maximum allowable transformation size. iv. In one example, T is set to 4096. e. Alternatively, further, if it is determined that RST is not to be applied, the relevant information may not be signaled. 22. When there are N (N>1) TUs within a CU, in order to determine the usage of RST and / or the signaling of RST-related information, the coded information of only one of the N TUs is used. a. In one example, the first TU of the CU in the decoding order may be used for the determination. b. In one example, the top-left TU of the CU in the decoding order may be used for the determination. c. In one example, the determination using a specific TU may be performed in the same way as in the case where there is only one TU in the CU. 23. The usage of RST and / or the signaling of RST-related information may be performed at the TU level or the PU level instead of the CU level. a. Alternatively, further, different TUs / PUs within a CU may select different secondary transformation matrices or enable / disable control flags. b. Alternatively, further, for the dual-tree case, chroma blocks may be coded and different color components may select different secondary transformation matrices or enable / disable control flags. c. Alternatively, whether to signal RST-related information at which video unit level may depend on the partition tree type (dual or single). d. Alternatively, whether to signal RST-related information at which video unit level may depend on the relationship between the CU / PU / TU and the maximum allowable transform block size, e.g., larger or smaller.
[0163] 5 Implementation Examples of the Disclosed Technology In the following exemplary embodiments, changes additional to JVET-N0193 are highlighted in gray. Text to be deleted is marked with double brackets (e.g., [[a]] indicates the deletion of the character "a").
[0164]
Number
[0165]
Number
[0166]
Number
[0167] The above example can be incorporated in the methods described below, e.g., in the context of methods 2200, 2210, 2220, 2230, 2240, and 2250, which may be implemented in a video decoder or a video encoder.
[0168] FIG. 22A shows a flowchart of an exemplary method for video processing. Method 2210 includes, at step 2212, performing a conversion between a current video block of a video and a coded representation of the video. In some implementations, the step of performing the conversion includes determining the applicability of a secondary conversion tool to the current video block based on the width (W) and / or height (H) of the current video block. In some implementations, the step of performing the conversion includes determining the usage of the secondary conversion tool and / or the signaling of information related to the secondary conversion tool according to rules independent of the partition tree type applied to the current video block.
[0169] FIG. 22B shows a flowchart of an exemplary method for video processing. Method 2220 includes, at step 2222, determining whether a current video block of a coding unit of a video satisfies a condition according to a rule. Method 2220 further includes, at step 2224, performing a conversion between the current video block and a coded representation of the video according to the determination. In some implementations, the condition is related to characteristics of one or more color components of the video, the size of the current video block, or coefficients in a portion of a residual block of the current video block. In some implementations, the rule defines that the presence of side information regarding a secondary conversion tool in the coded representation is controlled by the condition.
[0170] FIG. 22C shows a flowchart of an exemplary method for video processing. Method 2230 includes, at step 2232, determining the applicability of a secondary transformation tool to a current video block of a coding unit of video, where the coding unit includes a plurality of transformation units and the determination is based on a single transformation unit of the coding unit. Method 2230 further includes, at step 2234, performing a transformation between the current video block and the coded representation of the video.
[0171] In the operations shown in FIGS. 22A through 22C, the secondary transformation tool includes applying a forward secondary transformation to the output of the forward primary transformation applied to the residual of the video block during encoding, prior to quantization, or applying an inverse secondary transformation to the output of the inverse quantization of the video block during decoding, prior to applying the inverse primary transformation.
[0172] In yet another representative aspect, the disclosed techniques can be used to provide a video processing method. The method includes determining the applicability of a secondary transformation tool and / or the presence of side information associated with the secondary transformation tool for a current video block of a coding unit of video, where the coding unit includes a plurality of transformation units and the determination is made at the transformation unit level or the prediction unit level, and performing a transformation between the current video block of the coded representation of the video based on the determination, where the secondary transformation tool includes applying a forward secondary transformation to the output of the forward primary transformation applied to the residual of the video block during encoding, prior to quantization, or applying an inverse secondary transformation to the output of the inverse quantization of the video block during decoding, prior to applying the inverse primary transformation.
[0173] In some embodiments, the video coding method may be implemented using an apparatus implemented on a hardware platform as described with respect to FIG. 23 or FIG. 24.
[0174] Figure 23 is a block diagram of a video processing apparatus 2300. The apparatus 2310 may be used to implement one or more of the methods described in the present application. The apparatus 2300 may be embodied in a smartphone, a tablet, a computer, a monolithic Internet of Things (IoT) receiver, or the like. The apparatus 2300 may include one or more processors 2302, one or more memories 2304, and video processing hardware 2306. The processor 2302 may be configured to implement one or more of the methods described in this document (including, but not limited to, methods 2200, 2210, 2220, 2230, 2240, and 2250). The memories 2304 may be used to store data and code used to implement the methods and techniques described in the present application. The video processing hardware 2306 may be used to implement some of the techniques described in this document in a hardware circuit.
[0175] Figure 24 is a block diagram of another example of a video processing system in which the disclosed technology may be implemented. Figure 24 is a block diagram showing an exemplary video processing system 2400 in which various technologies disclosed in the present application may be implemented. Various implementations may include some or all of the components of system 4100. The system 2400 may include an input 2402 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or in a compressed or encoded format. The input 2402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0176] System 2400 may include a coding component 2404 capable of implementing various coding or encoding methods described in this document. The coding component 2404 can reduce the average bitrate of the video from the input 2402 to the output of the coding component 2404 to generate a coded representation of the video. Thus, coding techniques are sometimes called video compression or video transcoding techniques. The output of the coding component 2404 may be stored or transmitted via a communication connected as represented by component 2406. The stored or communicated bitstream (or coded) representation of the video received at the input 2402 may be used by component 2408 to generate pixel values or viewable video to be transmitted to the display interface 2410. The process of generating a video viewable by the user from the bitstream representation is sometimes called video decompression. Further, certain video processing operations are referred to as "coding" operations or tools, but it will be understood that coding tools or operations are used in an encoder and corresponding decoding tools or operations that process the result of coding in reverse will be performed by a decoder.
[0177] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)), a Displayport, etc. Examples of a storage interface may include a serial advanced technology attachment (SATA), PCI, an IDE interface, etc. The techniques described in this document can be embodied in various electronic devices such as a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0178] Some embodiments of the disclosed technology involve making a determination or decision to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in processing video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the conversion from a video block to a video bitstream representation will use the video processing tool or mode when enabled based on the determination or decision. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream along with information indicating that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a video bitstream representation to a video block will be performed using the video processing tool or mode enabled based on the determination or decision.
[0179] Some embodiments of the disclosed technology involve making a determination or decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode when converting a video block to a video bitstream representation. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream along with information indicating that the bitstream has not been modified using the video processing tool or mode disabled based on that determination or decision.
[0180] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may correspond to bits that are located at the same position or are spread to different locations within the bitstream, as defined by the syntax, for example. For example, a macroblock may be encoded from the perspective of the transformed and coded error residual values and also using the bits of the header and other fields within the bitstream.
[0181] Various techniques and embodiments can be described using the following clause format. The first set of clauses describes specific features and aspects of the techniques disclosed in the above sections.
[0182] 1. A video processing method, comprising: selecting a set of transforms or a transform matrix for applying a reduced second-order transform to a current video block based on the characteristics of the current video block; and applying the selected set of transforms or transform matrix to a portion of the current video block as part of the conversion between the current video block and the bitstream representation of the video including the current video block. A method comprising the above.
[0183] 2. The method according to clause 1, wherein a portion of the current video block is the upper-right sub-region, the lower-right sub-region, the lower-left sub-region, or the central sub-region of the current video block.
[0184] 3. The method according to clause 1 or 2, wherein the characteristics of the current video block are the intra prediction mode or the first-order transform matrix of the current video block.
[0185] 4. The method according to clause 1, wherein the above characteristics are the color components of the current video block.
[0186] 5. The method according to clause 4, wherein the first transformation set is selected for the luma component of the current video block, and a second transformation set different from the first transformation set is selected for one or more chroma components of the current video block.
[0187] 6. The method according to clause 1, wherein the above feature is the intra prediction mode or intra coding method of the current video block.
[0188] 7. The method according to clause 6, wherein the intra prediction method includes a multiple reference line (MRL)-based prediction method or a matrix-based intra prediction method.
[0189] 8. The method according to clause 6, wherein the first transformation set is selected when the current video block is a cross-component linear model (CCLM) coded block, and a second transformation set different from the first transformation set is selected when the current video block is a non-CCLM coded block.
[0190] 9. The method according to clause 6, wherein the first transformation set is selected when the current video block is coded by a joint chroma residual coding method, and a second transformation set different from the first transformation set is selected when the current video block is not coded by a joint chroma residual coding method.
[0191] 10. The method according to clause 1, wherein the above feature is the primary transformation of the current video block.
[0192] 11. A video processing method comprising: determining, based on one or more coefficients associated with the current video block, a selective inclusion of signaling of side information for application of reduced secondary transformation (RST) in the bitstream representation of the current video block; and Performing a conversion between a current video block and a video including the bitstream representation of the current video block based on a decision; A method comprising.
[0193] 12. The method according to clause 11, wherein one or more coefficients include the last non-zero coefficient in the scanning order of the current video block.
[0194] 13. The method according to clause 11, wherein one or more coefficients include a plurality of coefficients within a partial region of the current video block.
[0195] 14. The method according to clause 13, wherein the partial region includes one or more coding groups to which RST can be applied.
[0196] 15. The method according to clause 13, wherein the partial region includes the first M coding groups or the last M coding groups in the scanning order of the current video block.
[0197] 16. The method according to clause 13, wherein the partial region includes the first M coding groups or the last M coding groups in the reverse scanning order of the current video block.
[0198] 17. The method according to clause 13, wherein making the decision is further based on the energy of one or more non-zero coefficients of a plurality of coefficients.
[0199] 18. A video processing method comprising: For applying a reduced secondary transform (RST) to a current video block, a step of constructing a bitstream representation of the current video block, wherein syntax elements related to the RST are signaled in the bitstream representation before coding residual information; and Performing a conversion between the current video block and the bitstream representation of the current video block based on the construction; A method comprising.
[0200] 19. The method according to clause 18, wherein signaling the syntax elements related to RST is based on the usage of at least one coded block flag or conversion selection mode.
[0201] 20. The method according to clause 18, wherein the bitstream representation excludes the coding residual information corresponding to the coding group with all zero coefficients.
[0202] 21. The method according to clause 18, wherein the coding residual information is based on the application of RST.
[0203] 22. A video processing method comprising: constructing a bitstream representation of a current video block for applying a reduced second transformation (RST) to the current video block, wherein the syntax elements related to RST are signaled in the bitstream representation before either a transform skip indication or a multiple transform set (MTS) index; and performing a transformation between the current video block and the bitstream representation of the current video block based on the construction. The method as described above.
[0204] 23. The method according to clause 22, wherein the transform skip indication or MTS index is based on the syntax elements related to RST.
[0205] 24. A video processing method comprising: constructing a context model for coding an index of a reduced second transformation (RST) based on the characteristics of a current video block; and performing a transformation between the current video block and the bitstream representation of the video including the current video block based on the construction. The method as described above.
[0206] 25. The method according to clause 24, wherein the above feature is the explicit or implicit feasibility of a multiple transform selection (MTS) process.
[0207] 26. The method according to clause 24, wherein the above feature is the feasibility of a cross-component linear model (CCLM) coding mode in the current video block.
[0208] 27. The method according to clause 24, wherein the above feature is the size of the current video block.
[0209] 28. The method according to clause 24, wherein the above feature is the split depth of the partitioning process applied to the current video block.
[0210] 29. The method according to clause 28, wherein the partitioning process is a quadtree (QT) partitioning process, a binary tree (BT) partitioning process, or a ternary tree (TT) partitioning process.
[0211] 30. The method according to clause 24, wherein the above feature is the color format or color component of the current video block.
[0212] 31. The method according to clause 24, wherein the above feature excludes the intra prediction mode of the current video block and the index of the multiple transform selection (MTS) process.
[0213] 32. A video processing method comprising: determining, based on the features of the current video block, a decision regarding the selective application of an inverse scaled second transform (RST) process to the current video block; and performing a conversion between the current video block and the bitstream representation of the video including the current video block based on the decision. The method includes.
[0214] 33. The method according to clause 32, wherein the feature is the coded block flag of the current video block's coding group.
[0215] 34. The method according to clause 33, wherein the inverse RST process is not applied and the coded block flag of the upper left coding group is zero.
[0216] 35. The method according to clause 33, wherein the inverse RST process is not applied and the coded block flags for the first and second coding groups in the scanning order of the current video block are zero.
[0217] 36. The method according to clause 32, wherein the feature is the height (M) or width (N) of the current video block.
[0218] 37. The method according to clause 36, wherein the inverse RST process is not applied and (i) M = 8 and N = 4, or (ii) M = 4 and N = 8.
[0219] 38. A video processing method comprising: determining whether to selectively apply an inverse reduction second transformation (RST) process to a current video block based on a feature of the current video block; and performing a conversion between the current video block and a bitstream representation of a video including the current video block based on the determination, wherein the bitstream representation includes side information regarding RST, and the side information is included based on a luma component or a single color coefficient of the current video block.
[0220] 39. The method according to clause 38, wherein the side information is further included based on the dimensions of the current video block.
[0221] 40. The side information is the method according to clause 38 or 39, which is included without considering the block information for the current video block.
[0222] 41. The conversion is the method according to any one of clauses 1 - 40, which includes generating a bitstream representation from the current video block.
[0223] 42. The conversion is the method according to any one of clauses 1 - 40, which includes generating the current video block from a bitstream representation.
[0224] 43. An apparatus in a video system including a processor and a non - transient memory with instructions, where the instructions, when executed by the processor, cause the processor to execute the method according to any one of clauses 1 - 42.
[0225] 44. A computer program product stored on a non - transient computer - readable medium, including program code for executing the method according to any one of clauses 1 - 42.
[0226] The second set of clauses describes specific features and aspects of the technology disclosed in the above sections, such as exemplary implementations 6, 7, 20 - 23.
[0227] 1. A video processing method, comprising: executing a conversion between a current video block of a video and a coded representation of the video, the step of executing the conversion including determining the applicability of a secondary conversion tool for the current video block based on the width (W) and / or height (H) of the current video block. The secondary transformation tool includes, during encoding, applying a forward secondary transformation to the output of the forward primary transformation applied to the residual of a video block before quantization, or during decoding, applying an inverse secondary transformation to the output of the inverse quantization of a video block before applying the inverse primary transformation.
[0228] 2. The method according to clause 1, wherein the secondary transformation tool corresponds to a low-frequency non-separable transform (LFNST) tool.
[0229] 3. The method according to clause 1, wherein the secondary transformation tool is not applied when W>T1 or H>T2, and T1 and T2 are integers.
[0230] 4. The method according to clause 1, wherein the secondary transformation tool is not applied when W>T1 and H>T2, and T1 and T2 are integers.
[0231] 5. The method according to clause 1, wherein the secondary transformation tool is not applied when W*H>=T, and T is an integer.
[0232] 6. The method according to any one of clauses 1-5, wherein the block is a coding unit.
[0233] 7. The method according to clause 3 or 4, wherein T1=T2=64.
[0234] 8. The method according to clause 3 or 4, wherein T1 and / or T2 depend on the maximum allowable transform size.
[0235] 9. The method according to clause 5, wherein T is 4096.
[0236] 10. The method according to clause 1, wherein the determining step determines not to apply the secondary transformation tool, and the information related to the secondary transformation tool is not signaled.
[0237] 11. A video processing method, A step of determining whether the current video block of the video coding unit meets the conditions according to the rules, and A step of performing a conversion between the current video block and the coded representation of the video according to the determination, including The conditions are related to the characteristics of one or more color components of the video, the size of the current video block, or the coefficients in a part of the residual block of the current video block. The rule stipulates that the presence of side information regarding the secondary conversion tool in the coded representation is controlled by the conditions. The secondary conversion tool includes applying a forward secondary conversion to the output of the forward primary conversion applied to the residual of the video block during encoding before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying the inverse primary conversion during decoding.
[0238] 12. The method according to clause 11, wherein the secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool.
[0239] 13. The method according to clause 11 or 12, wherein the characteristics of one or more color components correspond only to the luma information of the coding unit including the current video block.
[0240] 14. The method according to clause 13, wherein the condition is satisfied only when the video coding unit has a height (H) smaller than T1 and a width (W) smaller than T2, and T1 and T2 are integers.
[0241] 15. The method according to clause 14, wherein T1 = T2 = 64.
[0242] 16. The method according to clause 12, wherein the condition is satisfied only when the partition type tree applied to the coding unit is a single tree.
[0243] 17. The method according to clause 11, wherein the rule makes a determination using one color component or all color components based on the dimensions of the current video block and / or the coded information.
[0244] 18. The method according to clause 11, wherein the determination is made according to the rule without using the information of the current video block having width (W) and height (H).
[0245] 19. The method according to clause 18, wherein the information of the current video block includes the number of non-zero coefficients of the current video block when W < T1 or H < T2, and T1 and T2 are integers.
[0246] 20. The method according to clause 11, wherein the determination is made according to the rule based on the coefficients within a part of the current video block, and the part is defined as the upper left MxN region of the current video block having width (W) and height (H), and M, N, W, and H are integers.
[0247] 21. The method according to clause 20, wherein M is smaller than W and / or N is smaller than H.
[0248] 22. The method according to clause 20, wherein M and N are fixed numbers.
[0249] 23. The method according to clause 20, wherein M and / or N depend on W and / or H.
[0250] 24. The method according to clause 20, wherein M and / or N depend on the maximum transform size.
[0251] 25. The method according to clause 11, wherein the determination is made according to the rule based on the coefficients within a part of the current video block, and the part is defined to be the same for all video blocks of the video.
[0252] 26. The method according to clause 11, wherein the determination is made according to a rule based on coefficients within a part of the current video block, and the part is determined depending on the dimensions of the current video block and / or the coded information.
[0253] 27. The method according to clause 11, wherein the determination is made according to a rule based on coefficients within a part of the current video block, and the part is determined depending on a given range of the scanning order index of the current video block.
[0254] 28. The method according to clause 27, wherein the rule defines a part having a scanning order index within the range [IdxS, IdxE] including both ends satisfying at least one of 1) IdxS being equal to 0, 2) IdxE being smaller than the product of W and (H - 1), 3) IdxE being a fixed number, or 4) IdxE depending on W and / or H, where W and H correspond to the width and height of the current video block, respectively.
[0255] 29. The method according to clause 11, wherein the condition is satisfied when the current video block has specific dimensions.
[0256] The method according to clause 11, wherein the determination is made according to a rule based on non - zero coefficients within a part of the current video block, and the part is determined depending on the number of non - zero coefficients in the current video block and / or other blocks within the coding unit.
[0257] 31. A video processing method, executing a conversion between a current video block of a video and a coded representation of the video wherein the step of executing the conversion includes determining, according to a rule independent of the partition tree type applied to the current video block, the usage of a secondary conversion tool and / or the signaling of information related to the secondary conversion tool. The secondary transformation tool includes applying a forward secondary transformation to the output of the forward primary transformation applied to the residual of the video block during encoding, before quantization, or applying an inverse secondary transformation to the output of the inverse quantization of the video block during decoding, before applying the inverse primary transformation.
[0258] 32. The method according to clause 31, wherein the secondary transformation tool corresponds to a low-frequency non-separable transform (LFNST) tool.
[0259] 33. The method according to clause 31, wherein the partition tree type applied to the current video block is a dual tree type or a single tree type.
[0260] 34. The rule stipulates not to use the secondary transformation tool when the number of counted non-zero coefficients is not greater than T, and the value of T is determined independently of the partition tree type.
[0261] 35. T is equal to 1 or 2 in the method according to clause 34.
[0262] 36. A video processing method, a step of determining the applicability of a secondary transformation tool to the current video block of the current video block of the coding unit of the video, where the coding unit includes a plurality of transformation units and the determination is based on a single transformation unit of the coding unit, a step of performing a conversion between the current video block and the coded representation of the video based on the determination, and The secondary transformation tool includes applying a forward secondary transformation to the output of the forward primary transformation applied to the residual of the video block during encoding, before quantization, or applying an inverse secondary transformation to the output of the inverse quantization of the video block during decoding, before applying the inverse primary transformation.
[0263] 37. The secondary conversion tool is the method described in clause 36 that is compatible with the low-frequency non-separable transform (LFNST) tool.
[0264] 38. The single conversion unit is the method described in clause 37 that corresponds to the first conversion unit of the coding unit in the decoding order.
[0265] 39. The single conversion unit is the method described in clause 37 that corresponds to the top-left conversion unit of the coding unit in the decoding order.
[0266] 40. The single conversion unit is determined using a similar rule applied when there is only one conversion unit in the coding unit, according to any one of clauses 36 - 39.
[0267] 41. A video processing method, For the current video block of the coding unit of the video, a step of determining the applicability of the secondary conversion tool and / or the existence of side information related to the secondary conversion tool, where the coding unit includes a plurality of conversion units and the determination is made at the conversion unit level or the prediction unit level, And a step of performing a conversion between the current video blocks of the coded representation of the video based on the determination, The secondary conversion tool includes applying a forward secondary conversion to the output of the forward primary conversion applied to the residual of the video block during encoding, before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block during decoding, before applying the inverse primary conversion.
[0268] 42. The secondary conversion tool is the method described in clause 41 that is compatible with the low-frequency non-separable transform (LFNST) tool.
[0269] 43. The coding unit is the method according to clause 41, including a flag indicating the applicability of the secondary conversion tool or different prediction units or different conversion units using different secondary conversion matrices.
[0270] 44. The different color components are the method according to clause 41, using a flag indicating the applicability of the secondary conversion tool or different secondary conversion matrices when the dual tree is enabled and the chroma block is coded.
[0271] 45. The determining step is the method according to clause 41, determining the existence of side information based on the partition tree type applied to the current video block.
[0272] 46. The determining step is the method according to clause 41, determining the existence of side information based on whether the coding unit, prediction unit, or conversion unit is larger or smaller than the maximum allowable conversion block size.
[0273] 47. Performing the conversion includes generating a coded representation from the video or generating the video from the coded representation, and is the method according to any one of clauses 1 - 46.
[0274] 48. An apparatus in a video system including a processor and a non - transient memory with instructions, where the instructions, when executed by the processor, cause the processor to execute the method according to any one of clauses 1 - 47.
[0275] 49. A computer program product stored on a computer - readable medium, including program code for executing the method according to any one of clauses 1 - 47.
[0276] From the above, although specific embodiments of the technology disclosed herein have been described for purposes of illustration in this application, it will be understood that various modifications may be made without departing from the scope of the invention. Accordingly, the technology disclosed herein is not limited except as by the appended claims.
[0277] The implementation of the subject matter and the functional operations described in this patent document can be realized in various systems, digital electronic circuits, or computer software, firmware, or hardware, or combinations of one or more of them, including the structures disclosed herein and their structural equivalents. The implementation of the subject matter described herein can be realized as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or combinations of one or more of them. The term "data processing unit" or "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code for creating an execution environment for the computer program in question, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations of one or more of them.
[0278] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored within a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), within a single file dedicated to the program in question, or within multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, where the multiple computers can be located at one site or distributed across multiple sites and interconnected by a communication network.
[0279] The processes and logic flows described herein can be executed by one or more programmable processors executing one or more computer programs, by acting on input data to generate output, thereby performing functions. The processes and logic flows can also be executed by special-purpose logic circuits, such as, for example, FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and can also be implemented as such devices.
[0280] Processors suitable for the execution of a computer program include, for example, any one or more processors of both general and special purpose microprocessors, and any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or will receive data from or transfer data to, or be operatively coupled to both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, for example, all forms of non-volatile memory, media and memory devices including semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0281] The specification is intended to be considered only as exemplary along with the drawings, where exemplary means illustrative. As used in this application, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.
[0282] Although this patent document contains many details, these should not be construed as limitations regarding the scope of any invention or what may be claimed. Instead, they should be construed as descriptions of features that may be specific to particular embodiments of a particular invention. Specific features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, although a feature may have been described or even initially claimed as acting in a particular combination, one or more features in the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0283] Similarly, in the figures, operations are depicted in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or sequentially in order to achieve the desired result, or that all of the illustrated operations be performed. Further, the way in which various system components are divided in the embodiments described in this patent document should not be understood as requiring such division in all embodiments.
[0284] Only a few implementation examples and embodiments are described, and other implementations, extensions, and modifications are possible based on what is described and illustrated in this patent document.
[0285] (Appendix 1) A video processing method, comprising: executing a conversion between a current video block of a video and a coded representation of the video including, the step of performing the conversion includes determining the applicability of a secondary conversion tool to the current video block based on the width (W) and / or height (H) of the current video block, The secondary conversion tool includes applying a forward secondary conversion to the output of a forward primary conversion applied to the residual of a video block during encoding, before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block during decoding, before applying an inverse primary conversion. (Appendix 2) The method according to Appendix 1, wherein the secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool. (Appendix 3) The method according to Appendix 1, wherein the secondary conversion tool is not applied when W>T1 or H>T2, and T1 and T2 are integers. (Appendix 4) The method according to Appendix 1, wherein the secondary conversion tool is not applied when W>T1 and H>T2, and T1 and T2 are integers. (Appendix 5) The method according to Appendix 1, wherein the secondary conversion tool is not applied when W*H>=T, and T is an integer. (Appendix 6) The method according to any one of Appendices 1-5, wherein the block is a coding unit. (Appendix 7) The method according to Appendix 3 or 4, wherein T1=T2=64. (Appendix 8) The method according to Appendix 3 or 4, wherein T1 and / or T2 depend on the maximum allowable conversion size. (Appendix 9) The method according to Appendix 5, wherein T is 4096. (Appendix 10) The method according to Appendix 1, wherein the determining step determines not to apply the secondary conversion tool, and information related to the secondary conversion tool is not signaled. (Appendix 11) A video processing method, A step of determining whether a current video block of a video coding unit meets a condition according to a rule, A step of performing a conversion between the current video block and a coded representation of the video according to the determination, including, the condition is related to features of one or more color components of the video, the size of the current video block, or coefficients in a part of a residual block of the current video block, the rule defines that the presence of side information regarding a secondary conversion tool in the coded representation is controlled by the condition, the secondary conversion tool includes applying a forward secondary conversion to an output of a forward primary conversion applied to a residual of a video block during encoding before quantization, or applying an inverse secondary conversion to an output of inverse quantization of the video block before applying an inverse primary conversion during decoding, a method. (Appendix 12) The secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool, the method described in Appendix 11. (Appendix 13) The feature of the one or more color components corresponds only to the luma information of the coding unit including the current video block, the method described in Appendix 11 or 12. (Appendix 14) The condition is satisfied only when the coding unit of the video has a height (H) smaller than T1 and a width (W) smaller than T2, where T1 and T2 are integers, the method described in Appendix 13. (Appendix 15) T1 = T2 = 64, the method described in Appendix 14. (Appendix 16) The condition is satisfied only when the partition type tree applied to the coding unit is a single tree, the method described in Appendix 12. (Appendix 17) The method according to Appendix 11, wherein the rule performs the determination using one color component or all color components based on the dimensions of the current video block and / or the coded information. (Appendix 18) The method according to Appendix 11, wherein the determination is performed according to the rule without using the information of the current video block having a width (W) and a height (H). (Appendix 19) The information of the current video block is W <t1又はh<t2である場合に前記現在のビデオ·ブロックの非ゼロ係数の数を含み、t1及びt2は整数である、付記18に記載の方法。(Appendix 20) The determination is made according to the rule based on coefficients within a part of the current video block, the part being defined as the upper left MxN region of the current video block having a width (W) and a height (H), where M, N, W, and H are integers, the method according to Appendix 11. (Appendix 21) M is smaller than W, and / or N is smaller than H, the method according to Appendix 20. (Appendix 22) M and N are fixed numbers, the method according to Appendix 20. (Appendix 23) M and / or N depend on W and / or H, the method according to Appendix 20. (Appendix 24) M and / or N depend on the maximum transformation size, the method according to Appendix 20. (Appendix 25) The determination is made according to the rule based on coefficients within a part of the current video block, the part being defined to be the same for all video blocks of the video, the method according to Appendix 11. (Appendix 26) The determination is made according to the rule based on coefficients within a part of the current video block, the part being defined depending on the dimensions and / or coded information of the current video block, the method according to Appendix 11. (Appendix 27) The determination is made according to the rule based on coefficients within a part of the current video block, the part being defined depending on a given range of the scanning order index of the current video block, the method according to Appendix 11. (Appendix 28) The method according to Appendix 27, which defines a portion having a scanning order index within the range of [IdxS, IdxE] including both ends that satisfy at least one of the following: 1) IdxS is equal to 0, 2) IdxE is smaller than the product of W and (H-1), 3) IdxE is a fixed number, or 4) IdxE depends on W and / or H, where W and H correspond to the width and height of the current video block, respectively. (Appendix 29) The method according to Appendix 11, wherein the condition is satisfied when the current video block has specific dimensions. (Appendix 30) The method according to Appendix 11, wherein the determination is made according to the rule based on non-zero coefficients within a part of the current video block, and the part is determined depending on the number of non-zero coefficients in the current video block and / or other blocks within the coding unit. (Appendix 31) A video processing method, executing a conversion between a current video block of a video and the coded representation of the video including, in the step of executing the conversion, determining the usage of a secondary conversion tool and / or the signaling of information related to the secondary conversion tool according to a rule independent of the partition tree type applied to the current video block, wherein the secondary conversion tool includes applying a forward secondary conversion to the output of a forward primary conversion applied to the residual of a video block during encoding before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding. (Appendix 32) The method according to Appendix 31, wherein the secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool. (Appendix 33) The method described in Appendix 31, wherein the partition tree type applied to the current video block is a dual tree type or a single tree type. (Appendix 34) The method described in Appendix 31, wherein the rule stipulates not to use the secondary conversion tool when the number of counted non-zero coefficients is not greater than T, and the value of T is determined independently of the partition tree type. (Appendix 35) The method described in Appendix 34, wherein T is equal to 1 or 2. (Appendix 36) A video processing method, A step of determining the applicability of a secondary conversion tool for a current video block of a coding unit of a video, wherein the coding unit includes a plurality of conversion units, and the determination is based on a single conversion unit of the coding unit. A step of performing a conversion between the current video block and the coded representation of the video based on the determination. The method includes applying a forward secondary conversion to the output of a forward primary conversion applied to the residual of a video block during encoding before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding. The secondary conversion tool includes applying a forward secondary conversion to the output of a forward primary conversion applied to the residual of a video block during encoding before quantization, or applying an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding. (Appendix 37) The method described in Appendix 36, wherein the secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool. (Appendix 38) The method described in Appendix 37, wherein the single conversion unit corresponds to the first conversion unit of the coding unit in the decoding order. (Appendix 39) The method described in Appendix 37, wherein the single conversion unit corresponds to the upper left conversion unit of the coding unit in the decoding order. (Appendix 40) The method according to any one of Appendices 36 to 39, wherein the single conversion unit is determined using the same rules applicable when there is only one conversion unit in the coding unit. (Appendix 41) A video processing method, a step of determining the applicability of a secondary conversion tool and / or the presence of side information related to the secondary conversion tool for a current video block of a coding unit of a video, wherein the coding unit includes a plurality of conversion units, and the determination is made at a conversion unit level or a prediction unit level; a step of performing conversion between current video blocks of the coded representation of the video based on the determination; The method includes that the secondary conversion tool applies a forward secondary conversion to the output of a forward primary conversion applied to the residual of a video block during encoding before quantization, or applies an inverse secondary conversion to the output of the inverse quantization of the video block before applying an inverse primary conversion during decoding. (Appendix 42) The method according to Appendix 41, wherein the secondary conversion tool corresponds to a low-frequency non-separable transform (LFNST) tool. (Appendix 43) The method according to Appendix 41, wherein the coding unit includes a flag indicating the applicability of the secondary conversion tool or different prediction units or different conversion units using different secondary conversion matrices. (Appendix 44) For different color components, a flag indicating the applicability of the secondary conversion tool or different secondary conversion matrices are used when dual trees are enabled and chroma blocks are coded, according to the method described in Appendix 41. (Appendix 45) The method according to Appendix 41, wherein the step of determining determines the presence of the side information based on the partition tree type applied to the current video block. (Appendix 46) The determining step is the method according to Supplementary Note 41, in which the coding unit, the prediction unit, or the conversion unit determines the presence of the side information based on whether it is larger or smaller than the maximum allowable conversion block size. (Supplementary Note 47) Executing the conversion includes the method according to any one of Supplementary Notes 1-46, which includes generating the coded representation from the video or generating the video from the coded representation. (Supplementary Note 48) An apparatus in a video system including a processor and a non-transitory memory with instructions, where the instructions, when executed by the processor, cause the processor to execute the method according to any one of Supplementary Notes 1-47. (Supplementary Note 49) A computer program stored in a computer-readable medium, including program code for executing the method according to any one of Supplementary Notes 1-47.
Claims
1. 1. A method for processing video data, comprising: determining, during conversion between a current video block of a video and a bitstream of the video, whether a secondary conversion tool is applied to the current video block based on a relationship between at least one of a width (W) and a height (H) of the current video block and a maximum allowed conversion size (T); and performing the conversion based on the determination; wherein the secondary transformation tool is not applied to the current video block if W>T and / or H>T; and using the secondary transform tool includes applying a forward secondary transform to an output of a forward primary transform applied to a residual of the current video block during encoding, prior to quantization; or The method, wherein using the secondary transform tool includes applying an inverse secondary transform to an output of an inverse quantization of the current video block before applying an inverse primary transform during decoding.
2. The method of claim 1 , wherein the secondary transformation tool corresponds to a Low Frequency Non-Separable Transform (LFNST) tool.
3. 2. The method of claim 1, further comprising determining whether side information for the current video block associated with the secondary conversion tool is included in the bitstream, and the side information is excluded from the bitstream if W>T and / or H>T.
4. 2. The method of claim 1, wherein each of the width (W) and height (H) is a size characteristic of the current video block that corresponds to luma information only.
5. The method of claim 1 , wherein the current video block is a coding unit.
6. 10. The method of claim 1, wherein whether side information is included in the bitstream is determined further based on a position of a last non-zero coefficient in a residual of the current video block.
7. 7. The method of claim 6, wherein if the side information is included in the bitstream, the last non-zero coefficient is located within a region of the current video block to which the secondary transform tool is applied.
8. 8. The method of claim 7, wherein when the size of the current video block is 4x4 or 8x8, the region corresponds to the first eight coefficients in a top-left 4x4 coding group.
9. 8. The method of claim 7, wherein if the width or the height of the current video block is greater than or equal to 4 and the size of the current video block is not 4x4 or 8x8, then the region corresponds to an upper left 4x4 coding group.
10. 2. The method of claim 1, wherein in response to side information indicating that the secondary transformation tool is enabled, a first syntax element indicating that a transformation kernel is to be used with a primary transformation tool is not present in the bitstream.
11. 2. The method of claim 1, wherein the transforming includes encoding the current video block into the bitstream.
12. 2. The method of claim 1, wherein the converting includes decoding the current video block from the bitstream.
13. 1. An apparatus for processing video data comprising a processor and a non-transitory memory having instructions that, when executed by the processor, cause the processor to: determining, during conversion between a current video block of a video and a bitstream of the video, whether a secondary conversion tool is applied to the current video block based on a relationship between at least one of a width (W) and a height (H) of the current video block and a maximum allowed conversion size (T); and performing the conversion based on the determination; wherein the secondary transformation tool is not applied to the current video block if W>T and / or H>T; and using the secondary transform tool includes applying a forward secondary transform to an output of a forward primary transform applied to a residual of the current video block during encoding, prior to quantization; or The apparatus, wherein using the secondary transform tool includes applying an inverse secondary transform to an output of an inverse quantization of the current video block before applying an inverse primary transform during decoding.
14. A non-transitory computer-readable storage medium storing instructions that cause a processor to: determining, during conversion between a current video block of a video and a bitstream of the video, whether a secondary conversion tool is applied to the current video block based on a relationship between at least one of a width (W) and a height (H) of the current video block and a maximum allowed conversion size (T); and performing the conversion based on the determination; wherein the secondary transformation tool is not applied to the current video block if W>T and / or H>T; and using the secondary transform tool includes applying a forward secondary transform to an output of a forward primary transform applied to a residual of the current video block during encoding, prior to quantization; or The storage medium, wherein using the secondary transform tool includes applying an inverse secondary transform to an output of inverse quantization of the current video block before applying an inverse primary transform during decoding.
15. 1. A method of storing a video bitstream, comprising: For a current video block of a video, determining whether a secondary transformation tool is to be applied to the current video block based on a relationship between at least one of a width (W) and a height (H) of the current video block and a maximum allowable transformation size (T); and generating the bitstream based on the determination; storing the bitstream on a non-transitory computer readable storage medium; wherein the secondary transformation tool is not applied to the current video block if W>T and / or H>T; and using the secondary transform tool includes applying a forward secondary transform to an output of a forward primary transform applied to a residual of the current video block during encoding, prior to quantization; or The method of storing, wherein using the secondary transform tool includes applying an inverse secondary transform to an output of inverse quantization of the current video block before applying an inverse primary transform during decoding.
Citation Information
Patent Citations
Image processing device and method
WO2017195666A1