Video signal processing method and apparatus using secondary transform
The use of quadratic transforms in video signal processing enhances coding efficiency by applying a secondary low-frequency non-separable transform, addressing inefficiencies in existing methods and improving compression and decoding processes.
Patent Information
- Application Number
- JP2025120179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-07-07
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-29
- Estimated Expiration
- 2040-06-25
AI Technical Summary
Existing video signal processing methods lack efficiency in coding, particularly in handling spatial and temporal correlations in video compression.
Implementing a video signal processing method that utilizes quadratic transforms, including a secondary low-frequency non-separable transform, to enhance coding efficiency by parsing syntax elements and applying specific conditions for transform blocks.
Improves coding efficiency by optimizing transform processes, leading to more effective video signal compression and decoding.
Smart Images

Figure 2025142070000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video signal processing method and apparatus, and more particularly to a video signal processing method and apparatus for encoding or decoding a video signal. [Background technology]
[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over a communication line or storing it in a form suitable for a storage medium. Compression coding can be used to encode audio, video, text, and other data, but video compression is the technology that specifically targets video. Video signal compression is performed by removing redundant information by taking into account spatial correlation, temporal correlation, and stochastic correlation. However, with the recent development of various media and data transmission media, more efficient video signal processing methods and devices are needed. Summary of the Invention [Problem to be solved by the invention]
[0003] An object of the present invention is to increase the coding efficiency of video signals.
[0004] The present invention has the objective of increasing coding efficiency via quadratic transforms. [Means for solving the problem]
[0005] This specification provides a video signal processing method that utilizes a quadratic transform.
[0006] In more detail, the video signal decoding apparatus includes a processor, and the processor parses syntax elements related to secondary transform of a coding unit from a bitstream of a video signal if one or more predetermined conditions are satisfied, and determines whether the secondary transform is applied to a transform block included in the coding unit based on the parsed syntax elements. If the secondary transform is applied to the transform block, the processor performs a secondary inverse transform based on one or more coefficients of a first sub-block, which is one of one or more sub-blocks constituting the transform block, to obtain one or more inverse transform coefficients for the first sub-block, and performs a primary inverse transform based on the one or more inverse transform coefficients to obtain residual samples for the transform block. The secondary transform is a low frequency non-separable transform. The transform block is a block to which a first-order transform that can be performed separately for vertical transform and horizontal transform is applied, and a first condition among the one or more preset conditions is that an index value indicating a position of a first coefficient among the one or more coefficients of the first sub-block is greater than a preset threshold value.
[0007] Also, in this specification, the syntax element is characterized in that it includes information indicating whether the secondary transform is applied to the coding unit, and information indicating a transform kernel to be used for the secondary transform.
[0008] In addition, in this specification, the first coefficient is a last significant coefficient according to a preset scan order, and the significant coefficient is a non-zero coefficient.
[0009] In addition, in this specification, the first sub-block is characterized as being the first sub-block in a preset scan order.
[0010] In addition, in this specification, a second condition among the one or more preset conditions is that the width and height of the transformation block are 4 pixels or more.
[0011] In addition, in this specification, the preset critical value is 0.
[0012] In addition, in this specification, the preset scan order is characterized by being an up-right diagonal scan order.
[0013] In addition, in this specification, the third condition among the one or more pre-set conditions is when a transform skip flag value included in the bitstream is not a specific value, and if the transform skip flag value has the specific value, the transform skip flag indicates that the first transform and the second transform are not applied to the transform block.
[0014] In addition, in this specification, a fourth condition among the one or more preset conditions is characterized in that at least one of the one or more coefficients of the first sub-block is not 0, and the at least one or more coefficients are present at a position other than the first position according to a preset scan order.
[0015] In addition, in this specification, the coding unit is composed of a plurality of coding blocks, and if at least one of the transform blocks corresponding to each of the plurality of coding blocks satisfies one or more of the predetermined conditions, the syntax elements related to the secondary transform are parsed.
[0016] Also, in this specification, a video signal decoding apparatus includes a processor, wherein the processor performs a primary transform on residual samples of a block included in a coding unit to obtain a plurality of primary transform coefficients for the block, performs a secondary transform based on one or more coefficients of the plurality of primary transforms to obtain one or more secondary transform coefficients for a first sub-block that is one of sub-blocks constituting the block, and encodes information about the one or more secondary transform coefficients and syntax elements related to the secondary transform of the coding unit to obtain a bitstream, wherein the secondary transform is a low-band non-separable transform (LFNST), the primary transform may be performed separately into a vertical transform and a horizontal transform, and the syntax elements related to the secondary transform of the coding unit are encoded if one or more predetermined conditions are satisfied, and a first condition of the one or more predetermined conditions is that an index value indicating a position of a first coefficient of the one or more secondary transform coefficients is greater than a predetermined threshold value.
[0017] Also, in this specification, the syntax element is characterized in that it includes information indicating whether the secondary transform is applied to the coding unit, and information indicating a transform kernel to be used for the secondary transform.
[0018] In this specification, the first coefficient is the last significant coefficient in a preset scan order, and the significant coefficient is a coefficient other than 0.
[0019] In addition, in this specification, the first sub-block is characterized as being the first sub-block in a preset scan order.
[0020] In addition, in this specification, a second condition among the one or more preset conditions is that the width and height of the primary transformation block are 4 pixels or more.
[0021] In addition, in this specification, the preset critical value is 0.
[0022] In addition, in this specification, the preset scan order is a top right diagonal scan order.
[0023] In addition, in this specification, the third condition among the one or more pre-set conditions is when a transform skip flag value included in the bitstream is not a specific value, and if the transform skip flag value has the specific value, the transform skip flag indicates that the first transform and the second transform are not applied to the block.
[0024] In addition, in this specification, a fourth condition among the one or more preset conditions is characterized in that at least one of the one or more secondary transform coefficients is not 0, and the one or more coefficients are present at positions other than the first position according to a preset scan order.
[0025] Also, in this specification, in a non-transitory computer-readable medium for storing a bitstream, the bitstream is encoded using a coding method including: performing a primary transform on residual samples of a block included in a coding unit to obtain a plurality of primary transform coefficients for the block; performing a secondary transform based on one or more coefficients of the plurality of primary transforms to obtain one or more secondary transform coefficients for a first sub-block that is one of sub-blocks constituting the block; and encoding information on the one or more secondary transform coefficients and syntax elements related to the secondary transform of the coding unit, wherein the secondary transform is a low-band non-separable transform (LFNST), the primary transform can be performed separately into a vertical transform and a horizontal transform, and the syntax elements related to the secondary transform are encoded if one or more predetermined conditions are satisfied, and a first condition of the one or more predetermined conditions is that an index value indicating a position of a first coefficient of the one or more secondary transform coefficients is greater than a predetermined threshold value. [Effects of the Invention]
[0026] One embodiment of the present invention provides a video signal processing method and apparatus that utilizes quadratic transformation. [Brief explanation of the drawings]
[0027] [Figure 1] 1 is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention; [Figure 2] 1 is a schematic block diagram of a video signal decoding device according to an embodiment of the present invention; [Figure 3] FIG. 1 illustrates an example of how coding tree units are divided into coding units within a picture. [Figure 4] FIG. 1 illustrates an embodiment of a method for signaling the splitting of quadtrees and multi-type trees. [Figure 5] 2 is a diagram illustrating in more detail an intra-prediction method according to an embodiment of the present invention; [Figure 6] 2 is a diagram illustrating in more detail an intra-prediction method according to an embodiment of the present invention; [Figure 7] FIG. 2 is a diagram showing in more detail how the encoder converts the residual signal. [Figure 8] 10 is a diagram illustrating in detail how the encoder and decoder inverse transform the transform coefficients to obtain the residual signal. [Figure 9] FIG. 10 illustrates basis functions for multiple transform kernels that can be used in a linear transform. [Figure 10] 1 is a block diagram showing a process of restoring a residual signal in a decoder that performs a quadratic transform according to an embodiment of the present invention; [Figure 11] 1 is a block-level diagram illustrating a process of restoring a residual signal in a decoder that performs a quadratic transform according to an embodiment of the present invention; [Figure 12] 10 illustrates a method for applying a quadratic transform using a reduced number of samples according to an embodiment of the present invention. [Figure 13] FIG. 10 illustrates a method for determining the upper right diagonal scan order according to one embodiment of the present invention. [Figure 14] FIG. 10 illustrates a diagram illustrating a top right diagonal scan order by block size according to one embodiment of the present invention. [Figure 15] FIG. 1 illustrates a method for indicating secondary transforms at the coding unit level. [Figure 16] A diagram showing a residual_coding syntax structure according to one embodiment of the present invention. [Figure 17] 1 illustrates a method for indicating a secondary transform at a coding unit level according to an embodiment of the present invention. [Figure 18] 1 illustrates a method for indicating a secondary transform at a coding unit level according to an embodiment of the present invention. [Figure 19] A diagram showing a residual_coding syntax structure according to one embodiment of the present invention. [Figure 20] FIG. 10 is a diagram illustrating a residual_coding syntax structure according to another embodiment of the present invention. [Figure 21] 10 illustrates a method for indicating a secondary transform at a coding unit level according to another embodiment of the present invention. [Figure 22] FIG. 10 is a diagram illustrating a residual_coding syntax structure according to another embodiment of the present invention. [Figure 23] FIG. 2 illustrates a method for indicating a secondary transform at the transform unit level according to an embodiment of the present invention. [Figure 24] FIG. 10 illustrates a method for indicating a secondary transform at the transform unit level according to another embodiment of the present invention. [Figure 25] FIG. 2 is a diagram illustrating a coding unit syntax according to one embodiment of the present invention. [Figure 26] FIG. 10 illustrates a method for indicating a secondary transform at the transform unit level according to another embodiment of the present invention. [Figure 27] FIG. 10 is a diagram illustrating a syntax structure regarding the position of the last significant coefficient in the scan order according to an embodiment of the present invention. [Figure 28] FIG. 10 is a diagram illustrating a residual_coding syntax structure according to another embodiment of the present invention. [Figure 29] 1 is a flowchart illustrating a video signal processing method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] The terms used in this specification are generally used as widely as possible while taking into consideration the functions of the present invention, but these may vary depending on the intentions of engineers in the field, customs, or the emergence of new technologies. In addition, in certain cases, the applicant may have arbitrarily selected terms, and in such cases, the meanings of these terms will be described in the relevant mode for carrying out the invention. Therefore, it is made clear that the terms used in this specification should be interpreted not simply as terms, but based on the substantive meanings of the terms and the overall content of this specification.
[0029] In this specification, some terms may be interpreted as follows: "Coding" may be interpreted as "Encoding" or "Decoding" in some cases. In this specification, an apparatus that encodes a video signal to generate a video signal bitstream is referred to as an encoding apparatus or encoder, and an apparatus that decodes a video signal bitstream to restore a video signal is referred to as a decoding apparatus or decoder. In this specification, "video signal processing apparatus" is used as a conceptual term that includes both an encoder and a decoder. "Information" is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, so the present invention is not limited thereto. "Unit" is used to refer to a basic unit of image processing or a specific location in a picture, and refers to an image region including at least one of a luma component and a chroma component. "Block" refers to an image region including a specific component of a luma component and a chroma component (i.e., Cb and Cr). However, depending on the embodiment, the terms "unit," "block," "partition," and "region" may be used interchangeably. In this specification, a unit is used as a concept including a coding unit, a prediction unit, and a transform unit, and a picture refers to a field or a frame, and these terms may be used interchangeably depending on the embodiment.
[0030] 1 is a schematic block diagram of a video signal encoding apparatus 100 according to an embodiment of the present invention. Referring to FIG. 1, the encoding apparatus 100 of the present specification includes a transform unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transform unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.
[0031] The transform unit 110 transforms a residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit 150, to obtain a transform coefficient value. For example, a discrete cosine transform (DCT), a discrete sine transform (DST), or a wavelet transform may be used. The discrete cosine transform and the discrete sine transform divide the input picture signal into blocks and then transform the block. During the transform, coding efficiency may vary depending on the distribution and characteristics of values within the transform domain. The quantization unit 115 quantizes the transform coefficient values output from the transform unit 110.
[0032] To improve coding efficiency, instead of directly coding the picture signal, the prediction unit 150 predicts a picture using a pre-coded region and adds the residual value between the original picture and the predicted picture to obtain a reconstructed picture. To avoid mismatch between the encoder and decoder, the encoder should use information available to the decoder when making predictions. To achieve this, the encoder performs a process of further reconstructing the coded current block. The inverse quantization unit 120 inversely quantizes the transform coefficient values, and the inverse transform unit 125 reconstructs the residual values using the inversely quantized transform coefficient values. Meanwhile, the filtering unit 130 performs filtering operations to improve the quality of the reconstructed picture and the coding efficiency. For example, the filtering unit 130 may include a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter, etc. The filtered picture is stored in the decoded picture buffer (DPB) 156 for output or use as a reference picture.
[0033] To improve coding efficiency, instead of directly coding the picture signal, the prediction unit 150 predicts a picture using a pre-coded region and adds a residual value between the original picture and the predicted picture to the predicted picture to obtain a reconstructed picture. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 predicts the current picture using a reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from a reconstructed region within the current picture and transmits intra coding information to the entropy coding unit 160. The inter prediction unit 154 includes a second motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value for the current region by referring to a specific reconstructed region. The motion estimation unit 154a transmits position information of the reference region (e.g., reference frame, motion vector) to the entropy coding unit 160 to be included in the bitstream. The motion compensation unit 154b performs inter-motion compensation using the motion vector values transmitted from the motion estimation unit 154a.
[0034] The prediction unit 150 includes an intra prediction unit 152 and an inter prediction unit 154. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 performs inter prediction to predict the current picture using a reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from reconstructed samples within the current picture and transmits intra coding information to the entropy coding unit 160. The intra coding information includes at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra coding information includes information about the reference samples. The inter prediction unit 154 includes a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value for the current region by referring to a specific region of the reconstructed reference signal picture. The motion estimation unit 154a transmits a motion information set (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector values transmitted from the motion compensation unit 154a. The inter prediction unit 154 transmits inter coding information including the motion information for the reference region to the entropy coding unit 160.
[0035] According to a further embodiment, the prediction unit 150 includes an intra block copy (BC) prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from reconstructed samples in the current picture and transmits intra BC coding information to the entropy coding unit 160. The intra BC prediction unit obtains block vector values indicating a reference region to be used for predicting the current region by referring to a specific region in the current picture. The intra BC prediction unit performs intra BC prediction using the obtained block vector values. The intra BC prediction unit transmits the intra BC coding information to the entropy coding unit 160. The intra BC prediction unit includes the block vector information.
[0036] After the picture prediction is performed, the transform unit 110 converts residual values between the original picture and the predicted picture to obtain transform coefficient values. The transform is performed in units of specific blocks within the picture, and the size of the specific blocks varies within a predetermined range. The quantization unit 115 quantizes the transform coefficient values generated by the transform unit 110 and transmits the quantized values to the entropy coding unit 160.
[0037] The entropy coding unit 160 generates a video signal bitstream by entropy coding information indicating quantized transform coefficients, intra-coding information, and inter-coding information. The entropy coding unit 160 uses a variable length coding (VLC) scheme and an arithmetic coding scheme. The VLC scheme converts input symbols into successive codewords, where the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and infrequently occurring symbols are represented by long codewords. The variable length coding scheme is a context-based adaptive variable length coding (CAVLC) scheme. The arithmetic coding scheme converts successive data symbols into a single prime number, and obtains the optimal prime number bits required to represent each symbol. The arithmetic coding scheme is a context-based adaptive binary arithmetic coding (CABAC) scheme. For example, the entropy coding unit 160 binarizes information indicating quantized transform coefficients, and then performs arithmetic coding on the binarized information to generate a bitstream.
[0038] The generated bitstream is encapsulated in Network Abstraction Layer (NAL) units as basic units. An NAL unit includes an integer number of coded coding tree units. In order for a video decoder to decode the bitstream, the bitstream must first be separated into NAL units and then each separated NAL unit must be decoded. Meanwhile, information required for decoding the video signal bitstream is transmitted via Raw Byte Sequence Payload (RBSP) of higher level sets such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS).
[0039] 1 illustrates an encoding device 100 according to one embodiment of the present invention, with separate blocks illustrating logically distinct elements of encoding device 100. Therefore, the elements of encoding device 100 described above may be implemented on a single chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of encoding device 100 described above is performed by a processor (not shown).
[0040] 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to FIG. 2, the decoding apparatus 200 of the present invention includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.
[0041] The entropy decoding unit 210 entropy codes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit 210 obtains binary codes for transform coefficient information of a specific region from the video signal bitstream. The entropy decoding unit 210 also de-binarizes the binary codes to obtain quantized transform coefficients. The inverse quantization unit 220 de-quantizes the quantized transform coefficients, and the inverse transform unit 225 restores residual values using the de-quantized transform coefficients. The video signal processing device 200 restores original pixel values by combining the residual values obtained from the inverse transform unit 225 with predicted values obtained from the prediction unit 250.
[0042] Meanwhile, the filtering unit 230 performs filtering on the picture to improve image quality. This includes a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is output or stored in the decoded picture buffer (DPB) 256 to be used as a reference picture for the next picture.
[0043] The prediction unit 250 includes an intra prediction unit 252 and an inter prediction unit 254. The prediction unit 250 generates a predicted picture using the coding type, transform coefficients for each region, intra / inter coding information, etc. decoded by the entropy decoding unit 210. To reconstruct the current block to be decoded, the current picture including the current block or a decoded region of another picture is used. A picture (or tile / slice) that uses only the current picture for reconstruction, i.e., performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that performs all of intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). Among inter-pictures (or tiles / slices), a picture (or tile / slice) that uses up to one motion vector and reference picture index to predict sample values of each block is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses up to two motion vectors and reference picture indexes is called a bi-predictive picture or B picture (or tile / slice). That is, a P picture (or tile / slice) uses up to one motion information set to predict each block, and a B picture (or tile / slice) uses up to two motion information sets to predict each block. Here, a motion information set includes one or more motion vectors and one reference picture index.
[0044] The intra prediction unit 252 generates a prediction block using intra coding information and reconstructed samples in the current picture. As described above, the intra coding information includes at least one of an intra prediction mode, a Most Probable Mode (MPM) flag, and an MPM index. The intra prediction unit 252 predicts sample values of the current block using reconstructed samples located to the left and / or above the current block as reference samples. In the present disclosure, the reconstructed samples, reference samples, and samples of the current block refer to pixels. Furthermore, sample values refer to pixel values.
[0045] In one embodiment, the reference samples are samples included in neighboring blocks of the current block. For example, the reference samples are samples adjacent to the left boundary and / or the top boundary of the current block. Furthermore, the reference samples are samples located on a line within a predetermined distance from the left boundary of the current block and / or samples located on a line within a predetermined distance from the top boundary of the current block, among samples in neighboring blocks of the current block. In this case, the neighboring blocks of the current block include at least one of the left (L) block, the top (A) block, the below left (BL) block, the above right (AR) block, and the above left (AL) block adjacent to the current block.
[0046] The inter prediction unit 254 generates a prediction block using reference pictures and inter coding information stored in the decoded picture buffer 256. The inter coding information includes a motion information set (e.g., reference picture index, motion vector, etc.) of the current block relative to the reference block. Inter prediction includes L0 prediction, L1 prediction, and bi-prediction. L0 prediction is prediction using one reference picture included in the L0 picture list, and L1 prediction is prediction using one reference picture included in the L1 picture list. This requires one set of motion information (e.g., motion vector and reference picture index). The bi-prediction method uses up to two reference regions, and these two reference regions may exist in the same reference picture or in different pictures. That is, the bi-prediction method uses up to two sets of motion information (e.g., motion vector and reference picture index), and two motion vectors may correspond to the same reference picture index or different reference picture indexes. In this case, the reference picture may be displayed (or output) either temporally before or after the current picture. In one embodiment, in a bi-predictive scheme, the two reference regions used are regions selected from the L0 picture list and the L1 picture list, respectively.
[0047] The inter prediction unit 254 obtains a current reference block using a motion vector and a reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. Furthermore, sample values of a block identified by the motion vector or their interpolated values are used as a predictor for the current block. For motion prediction with sub-pel pixel accuracy, for example, an 8-tab interpolation filter is used for the luma signal and a 4-tab interpolation filter is used for the chroma signal. However, the interpolation filters for sub-pel motion prediction are not limited thereto. In this way, the inter prediction unit 254 performs motion compensation, which predicts the texture of the current unit from a previously reconstructed picture. In this case, the inter prediction unit uses a motion information set.
[0048] According to a further embodiment, the predictor 250 includes an intra BC predictor (not shown). The intra BC predictor reconstructs the current region by referring to a specific region in the current picture that includes reconstructed samples. The intra BC predictor obtains intra BC coding information for the current region from the entropy decoding unit 210. The intra BC predictor obtains block vector values of the current region that indicate the specific region in the current picture. The intra BC predictor performs intra BC prediction using the obtained block vector values. The intra BC predictor includes block vector information.
[0049] A reconstructed video picture is generated by adding the predicted value output from the intra prediction unit 252 or the inter prediction unit 254 and the residual value output from the inverse transform unit 225. That is, the video signal decoding apparatus 200 reconstructs the current block using the predicted block generated from the prediction unit 250 and the residual value obtained from the inverse transform unit 225.
[0050] 2 shows a decoding device 200 according to one embodiment of the present invention, with separate blocks logically separating elements of the decoding device 200. Thus, the elements of the decoding device 200 described above may be implemented on a single chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the decoding device 200 described above is performed by a processor (not shown).
[0051] FIG. 3 illustrates an example in which a coding tree unit (CTU) is divided into coding units (CUs) within a picture. During video signal coding, a picture is divided into a sequence of coding tree units (CTUs). A coding tree unit consists of an NXN block of luma samples and two blocks of corresponding chroma samples. A coding tree unit is divided into multiple coding units. A coding tree unit may be a leaf node without being divided. In this case, the coding tree unit itself may be a coding unit. A coding unit refers to a basic unit for processing a picture during the above-mentioned video signal processing, i.e., intra / inter prediction, transform, quantization, and / or entropy coding. Within a picture, the size and shape of coding units are not constant. Coding units have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. In addition, in this specification, non-square blocks refer to rectangular blocks, but the present invention is not limited to this.
[0052] Referring to Figure 3, a coding tree unit is first divided into a quad tree (QT) structure. That is, in the quad tree structure, one node having a size of 2N x 2N is divided into four nodes having a size of N x N. In this specification, a quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and all nodes do not need to be divided to the same depth.
[0053] Meanwhile, the leaf node of the above-mentioned quad tree is further divided into a multi-type tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node is divided into a horizontally or vertically divided binary or ternary tree structure. That is, there are four division structures in the multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present invention, in each of the tree structures, the width and height of the node are both powers of 2. For example, in a binary tree (BT) structure, a node of size 2N×2N is divided into two N×2N nodes by vertical binary division and into two 2N×N nodes by horizontal binary division. In addition, in a ternary tree (TT) structure, a node of size 2Nx2N is divided into (N / 2)x2N, Nx2N, and (N / 2)x2N nodes by vertical ternary division, and into 2Nx(N / 2), 2NxN, and 2Nx(N / 2) nodes by horizontal ternary division. Such multi-type tree division is performed recursively.
[0054] The leaf nodes of a multi-type tree can be coding units. If a coding unit is not larger than the maximum transform length, the coding unit can be used as a unit of prediction and / or transformation without further division. In one embodiment, if the width or height of the current coding unit is larger than the maximum transform length, the current coding unit is divided into multiple transform units without explicit signaling regarding division. Meanwhile, in the above-mentioned quad tree and multi-type tree, at least one of the following parameters is predefined or transmitted via the RBSP of a higher level set such as PPS, SPS, VPS, etc. 1) CTU size: The size of the root node of the quadtree, 2) Min QT size (MinQtSize): The size of the smallest QT leaf node allowed, 3) Max BT size (MaxBtSize): The size of the largest BT root node allowed, 4) Max TT size (MaxTtSize): The size of the largest TT root node allowed, 5) Max MTT depth (MaxMttDepth): The maximum allowed depth of MTT split from the QT leaf node, 6) Min BT size (MinBtSize): The size of the smallest BT leaf node allowed, 7) Min TT size: The size of the smallest TT leaf node allowed.
[0055] 4 illustrates an embodiment of a method for signaling the splitting of a quad tree and a multi-type tree. Predefined flags are used to signal the splitting of a quad tree and a multi-type tree. Referring to FIG. 4, at least one of a flag "split_cu_flag" indicating whether a node is split, a flag "split_qt_flag" indicating whether a quad tree node is split, a flag "mtt_split_cu_vertical_flag" indicating the split direction of a multi-type tree node, and a flag "mtt_split_binarycu_flag" indicating the split pattern of a multi-type tree node is used.
[0056] According to an embodiment of the present invention, a flag 'split_cu_flag' indicating whether the current node is split is signaled first. If the value of 'split_cu_flag' is 0, it indicates that the current node is not split, and the current node becomes a coding unit. If the current node is a coding tree unit, the coding tree unit contains one unsplit coding unit. If the current node is a quad tree node 'QT node', the current node is the leaf node of the quad tree node 'QT leaf node' and becomes a coding unit. If the current node is a multi-type tree node 'MTT node', the current node is the leaf node of the multi-type tree 'MTT leaf node' and becomes a coding unit.
[0057] If the value of 'split_cu_flag' is 1, the current node is split into a quad-tree or multi-type tree node according to the value of 'split_qt_flag'. The coding tree unit is the root node of the quad-tree and is preferentially split into a quad-tree structure. In the quad-tree structure, 'split_qt_flag' is signaled for each node (QT node). If the value of 'split_qt_flag' is 1, the node is split into four square nodes. If the value of 'qt_split_flag' is 0, the node becomes a leaf node (QT leaf node) of the quad-tree and is split into a multi-type node. According to an embodiment of the present invention, quad-tree splitting may be restricted depending on the type of the current node. If the current node is a coding tree unit (root node of the quad-tree) or a quad-tree node, quad-tree splitting is allowed, but if the current node is a multi-type tree unit, quad-tree splitting is not allowed. Each quadtree leaf node, "QT leaf node," is further split into a multi-type tree structure. As mentioned above, if "split_qt_flag" is 0, the current node is split into multi-type nodes. To indicate the split direction and pattern, "mtt_split_cu_vertical_flag" and "mtt_split_cu_binary_flag" are signaled. A value of "mtt_split_cu_vertical_flag" of 1 indicates a vertical split of the node "MTT node," while a value of "mtt_split_cu_vertical_flag" of 0 indicates a horizontal split of the node "MTT node." Also, if "mtt_split_cu_binary_flag" is 1, the node "MTT node" is split into two rectangular nodes, and if "mtt_split_cu_binary_flag" is 0, the node "MTT node" is split into three rectangular nodes.
[0058] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes of a coding unit tree). Such a basic unit for prediction is hereinafter referred to as a prediction unit or a prediction block.
[0059] Hereinafter, the term "unit" used in this specification is used as an alternative term to the prediction unit, which is a basic unit for performing prediction, but the present invention is not limited thereto and can be understood as a concept including the coding unit in a broader sense.
[0060] 5 and 6 are diagrams illustrating in more detail an intra prediction method according to an embodiment of the present invention. As described above, the intra prediction unit predicts sample values of the current block using reconstructed samples located to the left and / or above the current block as reference samples.
[0061] First, Figure 5 shows an example of reference samples used to predict a current block in intra prediction mode. According to one example, the reference samples are samples adjacent to the left boundary and / or the top boundary of the current block. As shown in Figure 5, if the size of the current block is W x H and samples of a single reference line adjacent to the current block are used for intra prediction, the reference samples are set using up to 2W + 2H + 1 neighboring samples located to the left and / or above the current block.
[0062] Furthermore, if at least some samples used as reference samples have not yet been restored, the intra prediction unit performs a reference sample padding process to obtain reference samples. The intra prediction unit also performs a reference sample filtering process to reduce intra prediction errors. That is, the intra prediction unit filters the neighboring samples and / or the reference samples obtained by the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts samples of the current block using the reference samples obtained in this manner. The intra prediction unit predicts samples of the current block using unfiltered reference samples or filtered reference samples. In the present disclosure, neighboring samples include samples on at least one reference line. For example, the neighboring samples may include neighboring samples on a line adjacent to the boundary of the current block.
[0063] Next, Figure 6 illustrates an example of prediction modes used in intra prediction. For intra prediction, intra prediction mode information indicating the intra prediction direction is signaled. The intra prediction mode indicates one of a plurality of intra prediction modes constituting an intra prediction mode set. If the current block is an intra prediction block, the decoder receives the intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.
[0064] According to an embodiment of the present invention, the intra prediction mode set includes all intra prediction modes used in intra prediction (e.g., a total of 67 intra prediction modes). More specifically, the intra prediction mode set includes a planar mode, a DC mode, and a plurality of (e.g., 65) angle modes (i.e., directional modes). Each intra prediction mode is indicated by a predetermined index (i.e., intra prediction mode index). For example, as shown in FIG. 6, intra prediction mode index 0 indicates a planar mode, and intra prediction mode index 1 indicates a DC mode. In addition, intra prediction mode indexes 2 to 66 indicate different angle modes. The angle modes indicate different angles within a predetermined angle range. For example, the angle modes indicate angles within an angle range of 45 degrees to −135 degrees clockwise (i.e., a first angle range). The angle modes are defined based on 12 rotation directions. In this case, intra prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra prediction mode index 18 indicates horizontal (HOR) mode, intra prediction mode index 34 indicates diagonal (DIA) mode, intra prediction mode index 50 indicates vertical (VER) mode, and intra prediction mode index 66 indicates vertical diagonal (VDIA) mode.
[0065] Meanwhile, the preset angle ranges are set differently depending on the shape of the current block. For example, if the current block is a rectangular block, a wide-angle mode specifying an angle greater than 45 degrees or less than -135 degrees clockwise is also used. If the current block is a horizontal block, the angle mode specifies an angle within an angle range of (45 + offset1) degrees to (-135 + offset1) degrees clockwise (i.e., a second angle range). In this case, angle modes 67 to 76 outside the first angle range are also used. If the current block is a vertical block, the angle mode specifies an angle within an angle range of (45 - offset2) degrees to (-135 - offset2) degrees clockwise (i.e., a third angle range). In this case, angle modes -10 to -1 outside the first angle range are also used. According to an embodiment of the present invention, the values of offset1 and offset2 are determined differently depending on the ratio between the width and height of the rectangular block. In addition, offset1 and offset2 are positive numbers.
[0066] According to a further embodiment of the present invention, the plurality of angle modes constituting the intra prediction mode set includes a base angle mode and an extension angle mode, wherein the extension angle mode is determined based on the base angle mode.
[0067] According to one embodiment, the basic angle mode is a mode corresponding to an angle used in intra prediction of the existing High Efficiency Video Coding (HEVC) standard, and the extended angle mode is a mode corresponding to an angle newly added in intra prediction of the next-generation video codec standard. More specifically, the basic angle mode is an angle mode corresponding to one of the intra prediction modes {2, 4, 6, ..., 66}, and the extended angle mode is an angle mode corresponding to one of the intra prediction modes {3, 5, 6, ..., 65}. That is, the extended angle mode is an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode is determined based on the angle indicated by the basic angle mode.
[0068] According to another embodiment, the base angle mode is a mode corresponding to an angle within a predetermined first angle range, and the extension angle mode is a wide angle mode outside the first angle range. That is, the base angle mode is an angle mode corresponding to one of the intra prediction modes {2, 3, 4, ..., 66}, and the extension angle mode is an angle mode corresponding to one of the intra prediction modes {-10, -9, ..., -1} and {67, 68, ..., 76}. The angle indicated by the extension angle mode is determined to be the angle opposite to the angle indicated by the corresponding base angle mode. Thus, the angle indicated by the extension angle mode is determined based on the angle indicated by the base angle mode. However, the number of extension angle modes is not limited thereto, and additional extension angles may be defined depending on the size and / or pattern of the current block. For example, the extension angle mode may be defined as an angle mode corresponding to one of the intra prediction modes {-14, -13, ..., -1} and {67, 68, ..., 80}. Meanwhile, the total number of intra prediction modes included in the intra prediction mode set varies depending on the configuration of the base angle mode and the extended angle mode.
[0069] In the above embodiment, the spacing between extension angle modes is set based on the spacing between corresponding basic angle modes. For example, the spacing between extension angle modes {3, 5, 7, ..., 65} is determined based on the spacing between corresponding basic angle modes {2, 4, 6, ..., 66}. The spacing between extension angle modes {-10, -9, ..., -1} is determined based on the spacing between corresponding opposite basic angle modes {56, 57, ..., 65}, and the spacing between extension angle modes {67, 68, ..., 76} is determined based on the spacing between corresponding opposite basic angle modes {3, 4, ..., 12}. The angular spacing between extension angle modes is set to be the same as the angular spacing between corresponding basic angle modes. The number of extension angle modes in the intra prediction mode set is set to be less than or equal to the number of basic angle modes.
[0070] According to an embodiment of the present invention, an extension angle mode is signaled based on a base angle mode. For example, a wide angle mode (i.e., an extension angle mode) replaces at least one angle mode (i.e., a base angle mode) within a first angle range. The replaced base angle mode is an angle mode corresponding to the opposite side of the wide angle mode. That is, the replaced base angle mode is an angle mode corresponding to an angle opposite to the angle indicated by the wide angle mode, or an angle that is offset from the opposite angle by a preset offset index. According to an embodiment of the present invention, the preset offset index is 1. An intra-prediction mode index corresponding to the replaced base angle mode is further mapped to a wide angle mode to signal the corresponding wide angle mode. For example, wide angle modes {-10, -9, ... -1} are signaled by intra-prediction mode indexes {57, 58, ... , 66}, respectively, and wide angle modes {67, 68, ... 76} are signaled by intra-prediction mode indexes {2, 3, ... , 11}, respectively. In this way, by using the intra prediction mode index for the base angular mode to signal the extension angular mode, even if the configurations of the angular modes used for intra prediction of each block are different, the same set of intra prediction mode indexes can be used to signal the intra prediction mode, thereby minimizing signaling overhead due to changes in the intra prediction mode configuration.
[0071] Meanwhile, whether or not to use the extended angle mode is determined based on at least one of the shape and size of the current block. According to one embodiment, if the size of the current block is larger than a preset size, the extended angle mode is used for intra prediction of the current block, and if not, only the basic angle mode is used for intra prediction of the current block. According to another embodiment, if the current block is a non-square block, the extended angle mode is used for intra prediction of the current block, and if the current block is a square block, only the basic angle mode is used for intra prediction of the current block.
[0072] Meanwhile, in order to improve coding efficiency, a method is used in which, instead of directly coding the residual signal, the residual signal is transformed, the obtained transform coefficient values are quantized, and the quantized transform coefficients are coded. As described above, the transform unit transforms the residual signal to obtain transform coefficient values. In this case, the residual signal of a specific block may be distributed throughout the entire region of the current block. Therefore, the frequency domain transform of the residual signal can be performed to concentrate energy in the low-frequency region, thereby improving coding efficiency. Hereinafter, a method for transforming or inversely transforming the residual signal will be described in detail.
[0073] FIG. 7 is a diagram illustrating in detail how an encoder transforms a residual signal. As described above, a spatial-domain residual signal is transformed into a frequency domain. The encoder transforms the obtained residual signal to obtain transform coefficients. First, the encoder obtains at least one residual block containing a residual signal for a current block. The residual block is either the current block or a block divided from the current block. In this disclosure, a residual block is referred to as a residual array or residual matrix containing residual samples of the current block. In this disclosure, a residual block refers to a block having the same size as a transform unit or transform block.
[0074] Next, the encoder transforms the residual block using a transform kernel. The transform kernel used to transform the residual block is a transform kernel having separable vertical and horizontal transform characteristics. In this case, the transform of the residual block is performed separately into vertical and horizontal transforms. For example, the encoder applies a transform kernel to the vertical direction of the residual block to perform a vertical transform. Also, the encoder applies a transform kernel to the horizontal direction of the residual block to perform a horizontal transform. In this disclosure, the term "transform kernel" refers to a set of parameters used to transform the residual signal, such as a transform matrix, transform array, transform function, or transform. According to one embodiment, the transform kernel is one of a plurality of available kernels. Furthermore, transform kernels based on different transform types may be used for the vertical transform and the horizontal transform.
[0075] The encoder transmits a transform block, which is transformed from the residual block, to a quantizer for quantization. In this case, the transform block includes a plurality of transform coefficients. Specifically, the transform block is composed of a plurality of transform coefficients arranged in two dimensions. The size of the transform block is the same as that of the residual block, either the current block or one of the blocks divided from the current block. The transform coefficients transmitted to the quantizer are represented by quantized values.
[0076] The encoder also performs a further transform before quantizing the transform coefficients. As shown in FIG. 7, the above-described transform method is called a primary transform, and the further transform is called a secondary transform. The secondary transform is selective for each residual block. According to an embodiment, the encoder performs a secondary transform on a region where it is difficult to concentrate energy in the low-frequency region using only the primary transform, thereby improving coding efficiency. For example, a secondary transform may be added to a block whose residual value is significantly expressed in a direction other than the horizontal or vertical direction of the residual block. The residual value of an intra-predicted block is more likely to change in a direction other than the horizontal or vertical direction than the residual value of an inter-predicted block. Therefore, the encoder further performs a secondary transform on the residual signal of the intra-predicted block. Alternatively, the encoder may omit the secondary transform on the residual signal of the inter-predicted block.
[0077] As another example, whether to perform a secondary transform is determined depending on the size of the current block or the residual block. Also, transform kernels of different sizes are used depending on the size of the current block or the residual block. For example, an 8x8 secondary transform is applied to a block whose shorter side of either the width or height is equal to or greater than a first preset length. Also, a 4x4 secondary transform is applied to a block whose shorter side of either the width or height is equal to or greater than a second preset length but smaller than the first preset length. In this case, the first preset length may be greater than the second preset length, but the present disclosure is not limited thereto. Also, unlike the primary transform, the secondary transform does not have to be separated into a vertical transform and a horizontal transform. Such a secondary transform is called a low-band non-separable transform (LFNST).
[0078] Furthermore, in the case of a video signal of a specific region, high-frequency band energy does not decrease even after frequency transformation due to abrupt changes in brightness. This may result in a decrease in compression performance due to quantization. Furthermore, when transforming a region where residual values rarely exist, encoding and decoding times may increase unnecessarily. Therefore, the transform of the residual signal of the specific region may be omitted. Whether or not to transform the residual signal of the specific region is determined by a syntax element related to the transform of the specific region. For example, the syntax element includes transform skip information. The transform skip information is a transform skip flag. If the transform skip information for a residual block indicates a transform skip, the transform of the corresponding residual block is not performed. In this case, the encoder immediately quantizes the residual signal of the corresponding region that has not been transformed. The operation of the encoder described with reference to FIG. 7 is performed via the transform unit of FIG. 1.
[0079] The above-mentioned transform-related syntax elements are information parsed from a video signal bitstream. A decoder entropy decodes the video signal bitstream to obtain the transform-related syntax elements. An encoder entropy codes the transform-related syntax elements to generate a video signal bitstream.
[0080] FIG. 8 is a diagram illustrating in detail how an encoder and a decoder inversely transform transform coefficients to obtain a residual signal. For convenience of explanation, it will be described below that the inverse transform operation is performed by an inverse transform unit of each of the encoder and the decoder. The inverse transform unit inversely transforms the dequantized transform coefficients to obtain a residual signal. First, the inverse transform unit detects whether an inverse transform for a specific region is to be performed based on a syntax element related to the transform of the corresponding region. According to an embodiment, if a syntax element related to the transform for a specific transform block indicates a transform skip, the transform for the corresponding transform block is skipped. In this case, both the primary inverse transform and the secondary inverse transform for the transform block are skipped. Furthermore, the dequantized transform coefficients are used as a residual signal. For example, a decoder reconstructs a current block using the dequantized transform coefficients as a residual signal. The primary inverse transform described above refers to the inverse transform of a primary transform and is referred to as an inverse primary transform. The secondary inverse transform refers to the inverse transform of a secondary transform and is referred to as an inverse secondary transform or inverse LFNST. In the present invention, a linear (inverse) transform is referred to as a first (inverse) transform, and a secondary (inverse) transform is referred to as a second (inverse) transform.
[0081] In another embodiment, a syntax element related to the transform for a particular transform block may not indicate a transform skip. In this case, the inverse transform unit determines whether to perform a secondary inverse transform for the secondary transform. For example, if the transform block is an intra-predicted block, a secondary inverse transform is performed on the transform block. In addition, a secondary transform kernel to be used for the corresponding transform block is determined based on the intra-prediction mode for the transform block. As another example, whether to perform a secondary inverse transform may be determined based on the size of the transform block. The secondary inverse transform is performed after the inverse quantization process and before the primary inverse transform is performed.
[0082] The inverse transform unit performs a primary inverse transform on the dequantized transform coefficients or the secondary inverse transformed transform coefficients. The primary inverse transform is performed separately into a vertical transform and a horizontal transform, as in the primary transform. For example, the inverse transform unit performs an inverse vertical transform and an inverse horizontal transform on the transform block to obtain a residual block. The inverse transform unit inversely transforms the transform block based on the transform kernel used to transform the transform block. For example, the encoder explicitly or implicitly signals information indicating which of multiple available transform kernels is currently applied to the transform block. The decoder uses the signaled information indicating the transform kernel to select from the multiple available transform kernels to be used for the inverse transform of the transform block. The inverse transform unit reconstructs the current block using the residual signal obtained through the inverse transform of the inverse transform coefficients.
[0083] Meanwhile, the distribution of residual signals of a picture may differ from region to region. For example, the distribution of values of residual signals within a specific region may differ depending on the prediction method. When the same transform kernel is used to transform multiple different transform regions, coding efficiency may differ for each transform region depending on the distribution and characteristics of values within the transform region. Therefore, coding efficiency can be further improved by adaptively selecting a transform kernel to be used to transform a specific transform block from multiple available transform kernels. That is, the encoder and decoder are configured to use additional transform kernels other than the base transform kernel in transforming a video signal. The method of adaptively selecting a transform kernel is referred to as adaptive multiple core transform (ATM) or multiple transform selection (MTS). For convenience of description, in this disclosure, transform and inverse transform are collectively referred to as transform. Furthermore, transform kernels and inverse transform kernels are collectively referred to as transform kernels.
[0084] The residual signal, which is the difference between the original signal and the predicted signal generated through inter-frame prediction or intra-frame prediction, has energy distributed throughout the pixel domain. Therefore, if the pixel values of the residual signal are coded, the compression efficiency will decrease. Therefore, a process is required to concentrate the energy of the pixel domain residual signal in the low-frequency region of the frequency domain through transform coding.
[0085] The high-efficiency video coding (HEVC) standard mostly uses the efficient discrete cosine transform type-II (DCT-II) when the signal is uniformly distributed in the pixel domain (when adjacent pixel values are similar), and uses the discrete sine transform type-VII (DCT-VII) only for predicted 4x4 blocks within a frame to transform the pixel-domain residual signal to the frequency domain. The DCT-II transform is suitable for residual signals generated through inter-frame prediction (when energy is uniformly distributed in the pixel domain). However, for residual signals generated through intra-frame prediction, the energy of the residual signal tends to increase with distance from the reference sample due to the characteristics of intra-frame prediction, which predicts using reconstructed reference samples surrounding the current coding unit. Therefore, using only the DCT-II transform to transform the residual signal to the frequency domain does not achieve high coding efficiency.
[0086] AMT is a transformation technique that adaptively selects a transformation kernel from a number of pre-defined kernels according to a prediction method. Since the pixel domain pattern (horizontal signal characteristics, vertical signal characteristics) of the residual signal differs depending on which prediction method is used, higher coding efficiency is expected than when only DCT-II is used to transform the residual signal. In the present invention, AMT is not limited to its name and may also be referred to as MTS (multiple transform selection).
[0087] FIG. 9 is a diagram showing basis functions for a number of transformation kernels that can be used in a linear transformation.
[0088] In detail, FIG. 9 is a diagram showing the basis functions of the transform kernels used in AMT, and shows the mathematical formulas of the DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to AMT.
[0089] DCT and DST are expressed by cosine and sine functions, respectively. When the basis function of the transform kernel for the number of samples N is represented by Ti(j), the index i indicates an index in the frequency domain, and the index j indicates an index within the basis function. That is, a smaller i indicates a lower-frequency basis function, and a larger i indicates a higher-frequency basis function. When expressed as a two-dimensional matrix, the basis function Ti(j) indicates the j-th element of the i-th row. Since all of the transform kernels shown in Figure 9 have separable properties, they can perform horizontal and vertical transforms on the residual signal X. In other words, if the residual signal block is X and the transform kernel matrix is T, the transform for the residual signal X is represented by TxT'. In this case, T' refers to the transpose matrix of the transform kernel matrix T.
[0090] The transform matrix values defined by the basis functions shown in FIG. 9 are in prime number form rather than integer form. Prime number form values may be difficult to implement in hardware in a video encoding device and a video decoding device. Therefore, integer-approximated transform kernels are used in encoding and decoding video signals from original transform kernels containing prime number form values. The approximated transform kernels containing integer form values are generated by scaling and rounding the original transform kernel. The integer values contained in the approximated transform kernels are within a range that can be represented by a predetermined number of bits. The predetermined number of bits is 8-bit or 10-bit. The orthogonal properties of the DCT and DST may not be maintained due to the approximation. However, since the resulting loss in coding efficiency is not significant, approximating the transform kernels to integer form is advantageous in terms of hardware implementation.
[0091] In the linear transform domain and inverse linear transform described with reference to FIGS. 7 and 8, the separable transform kernels are represented as two-dimensional matrices and transformed in the vertical and horizontal directions, respectively, so two-dimensional matrix multiplication operations are considered to be performed. This can be problematic from an implementation perspective due to the large amount of computation required. Therefore, from an implementation perspective, important issues include whether the amount of computation can be reduced by using a combined structure of a butterfly structure or half butterfly structure and a half matrix multiplier, as in DCT-II, or whether the corresponding transform kernel can be decomposed into a transform kernel with low implementation complexity (whether the corresponding channel can be represented by a matrix product with low complexity). Furthermore, because the elements of the transform kernel (the matrix elements of the transform kernel) must be stored in memory for computation, the memory capacity for storing the kernel matrix must also be considered when implementing the transform kernel. From this perspective, since the implementation complexity of DST-VII and DCT-VIII is relatively high, a transform that exhibits similar characteristics to DST-VII and DCT-VIII but has low implementation complexity can replace DST-VII and DCT-VIII.
[0092] DST-IV (discrete sine transform type-IV) and DCT-IV (discrete cosine transform type-IV) are considered candidates to replace DST-VII and DCT-VIII, respectively. The DCT-II kernel for 2N samples contains the DCT-IV kernel for N samples, and the DST-IV kernel for N samples can be implemented by simply negating the sign of the DCT-IV kernel for N samples and rearranging the corresponding basis functions in reverse order. Therefore, DST-IV and DCT-IV for N samples can be easily derived from the DCT-II for 2N samples.
[0093] Since the residual signal, which is the difference between an original signal and a predicted signal, exhibits a characteristic in which the signal energy distribution changes depending on the prediction method, adaptively selecting a transform kernel according to the prediction method, such as AMT or MTS, can improve coding efficiency. Furthermore, as described with reference to FIGS. 7 and 8, coding efficiency can be improved by performing an additional transform, a secondary transform and an inverse secondary transform (an inverse transform corresponding to the secondary transform), in addition to the primary transform and the inverse primary transform (an inverse transform corresponding to the primary transform). Such a secondary transform improves energy compaction, particularly for intra-predicted residual signal blocks in which strong energy is likely to exist in directions other than the horizontal and vertical directions of the residual signal. As described above, such a secondary transform is called a low-band non-separable transform (LFNST). The primary transform is also called a core transform.
[0094] FIG. 10 is a block diagram showing a process of reconstructing a residual signal in a decoder that performs a secondary transform according to an embodiment of the present invention. First, an entropy coder parses syntax elements related to the residual signal from a bitstream and obtains quantized coefficients through de-binarization. The decoder performs inverse quantization on the reconstructed quantized coefficients to obtain transform coefficients and then performs an inverse transform on the transform coefficients to reconstruct a residual signal block. The inverse transform is applied to blocks to which transform skip (TS) is not applied. The decoder performs the inverse transform in the order of a secondary inverse transform and a primary inverse transform. In this case, the secondary inverse transform may be omitted. The secondary inverse transform may not be performed on inter-predicted blocks and may be omitted. Alternatively, the secondary inverse transform may be omitted depending on the block size. The reconstructed residual signal contains quantization error, and the secondary transform changes the energy distribution of the residual signal, thereby reducing the quantization error compared to when only a primary transform is performed.
[0095] FIG. 11 illustrates a block-level process of restoring a residual signal in a decoder that performs a secondary transform according to an embodiment of the present invention. Restoration of the residual signal is performed in units of transform units (TUs) or sub-blocks within a TU. FIG. 11 illustrates a process of restoring a residual signal block to which a secondary transform is applied, in which a secondary inverse transform is first performed on a dequantized transform coefficient block. The decoder may perform a secondary inverse transform on all W×H (W: width, number of horizontal samples; H: height, number of vertical samples) samples within the TU, or, in consideration of complexity, may perform the secondary inverse transform only on the top-left sub-block of size W'×H', which is the most influential low-frequency region. In this case, W' may be equal to or smaller than W, and H' may be equal to or smaller than H. The size of the top-left sub-block, W'×H', is set differently depending on the TU size. For example, if min(W,H)=4, both W' and H' are set to 4. If min(W,H)>=8, both W' and H' are set to 8. min(x,y) indicates an operation that returns x if x is equal to or smaller than y, and returns y if x is equal to or smaller than y. After performing a second-order inverse transform, the decoder obtains the left-to-top W'×H' size sub-block transform coefficients within the TU, and performs a first-order inverse transform on the entire W×H size transform coefficient block to restore the residual signal block.
[0096] Whether secondary transformation is active or applicable is indicated by a 1-bit flag in at least one of the High Level Syntax (HLS) RBSPs, such as the sequence parameter set (SPS), picture parameter set (PPS), picture header, slice header, and tile group header. Furthermore, if secondary transformation is applicable, the size of the upper left sub-block to be considered in the secondary transformation may be indicated by a 1-bit flag in at least one of the HLS RBSPs. For example, whether an 8x8 sub-block can be used in a secondary transformation that considers 4x4 and 8x8 sub-blocks is indicated by a 1-bit flag in at least one of the HLS RBSPs.
[0097] If activation or applicability of secondary transform is indicated at a higher level (e.g., HLS), whether secondary transform is provided is indicated by a 1-bit flag at the coding unit (CU) level. Also, if secondary transform is applied to the current block, an index indicating a transform kernel to be used for secondary transform is indicated at the coding unit (CU) level. The decoder performs a secondary inverse transform on the block to which secondary transform is applied, using a transform kernel indicated by the corresponding index within a transform kernel set pre-defined according to the prediction mode. The index indicating the transform kernel is binarized using a truncated unary or fixed-length binarization method. The 1-bit flag indicating whether secondary transform is applied at the CU level and the index indicating the transform kernel to be used for secondary transform may be indicated using a single syntax element, which is referred to as lfnst_idx[x0][y0] or lfnst_idx in the present invention, but is not limited thereto. As an example, the first bit of lfnst_idx[x0][y0] indicates whether a secondary transform is applied at the CU level. The remaining bits indicate an index indicating the transform kernel used for the secondary transform. That is, lfnst_idx[x0][y0] indicates whether a secondary transform (LFNST) is applied and, if a secondary transform is applied, an index indicating the transform kernel used. lfnst_idx[x0][y0] is coded using an entropy coder, such as CABAC (context-based adaptive binary arithmetic coding) or CAVLC (context-based adaptive variable length coding), which adaptively codes according to the context. If the current CU is divided into multiple TUs smaller than the CU size, the secondary transform is not applied, and the syntax element lfnst_idx[x0][y0] related to the secondary transform is set to 0 without signaling. For example, lfnst_idx[x0][y0] = 0 indicates that the secondary transform is not applied.On the other hand, if lfnst_idx[x0][y0] is greater than 0, it indicates that a secondary transformation is applied, and the transformation kernel used for the secondary transformation is selected based on lfnst_idx[x0][y0].
[0098] As described above, a coding tree unit, a leaf node of a quad tree, or a leaf node of a multi-type tree can be a coding unit. If a coding unit is not larger than the maximum transform length, the coding unit is used as a unit of prediction and / or transformation without further division. In one embodiment, if the width or height of a current coding unit is larger than the maximum transform length, the current coding unit is divided into multiple transform units without explicit signaling regarding division. If the size of a coding unit is larger than the maximum transform size, it is divided into multiple transform blocks without signaling. In this case, the maximum coding block (or the maximum size of a coding block) to which a secondary transform is applied is limited because applying a secondary transform reduces performance and increases complexity. The size of the maximum coding block is the same as the maximum transform size. Alternatively, the size of the maximum coding block is defined as a preset coding block size. In one embodiment, the preset value may be 64, 32, or 16, but the present invention is not limited thereto. In this case, the value compared with the preset value (or the maximum transform size) is defined as the length of the long side or the number of samples.
[0099] Meanwhile, transform kernels based on DCT-II, DCT-VII, and DCT-VIII basis functions used in primary transforms have separable characteristics. Therefore, two vertical and horizontal transforms are performed on samples in an NxN residual block, resulting in a transform kernel size of NxN. In contrast, secondary transforms have non-separable characteristics. Therefore, if the number of samples considered in the secondary transform is nxn, one transform is performed. In this case, the size of the transform kernel is (n^2) x (n^2). For example, when performing a secondary transform on a 4x4 coefficient block at the top left, a 16x16 transform kernel is applied. When performing a secondary transform on an 8x8 coefficient block at the top left, a 64x64 transform kernel is applied. A 64x64 transform kernel requires a large number of multiplication operations, which can place a heavy burden on the encoder and decoder. Therefore, reducing the number of samples considered in the secondary transform can reduce the amount of calculations and the memory required to store the transform kernels.
[0100] FIG. 12 illustrates a method for applying a quadratic transform using a reduced number of samples according to an embodiment of the present invention. According to an embodiment of the present invention, the quadratic transform is expressed as the product of a quadratic transform kernel matrix and a linearly transformed coefficient vector, and is interpreted as mapping the linearly transformed coefficients to another space. In this case, by reducing the number of coefficients to be doubly transformed, i.e., the number of basis vectors constituting the quadratic transform kernel, the amount of calculation required for the quadratic transform and the memory capacity required to store the transform kernel can be reduced. For example, when performing a quadratic transform on a left-top 8x8 coefficient block, if the number of coefficients to be doubly transformed is reduced to 16, a quadratic transform kernel of 16 (rows) x 64 (columns) size (or 16 (rows) x 48 (columns) size) is applied. The transform unit of the encoder obtains a quadratically transformed coefficient vector by performing an inner product of each row vector constituting the transform kernel matrix and the linearly transformed coefficient vector. The inverse transform units of the encoder and decoder obtain linearly transformed coefficient vectors by performing an inner product of each column vector constituting the transform kernel matrix and the quadratically transformed coefficient vector.
[0101] Referring to FIG. 12, the encoder first performs a forward primary transform on the residual signal block to obtain a primary-transformed coefficient block. Assuming that the size of the primary-transformed coefficient block is M×N, for an intra-predicted block with min(M,N) equal to 4, a 4×4 forward secondary transform is performed on the left-to-top 4×4 samples of the primary-transformed coefficient block. For an intra-predicted block with min(M,N) equal to or greater than 8, an 8×8 secondary transform is performed on the left-to-top 8×8 samples of the primary-transformed coefficient block. Because the 8×8 secondary transform requires a large amount of computation and memory, only a portion of the 8×8 samples may be used. In one embodiment, to improve coding efficiency, for a rectangular block with min(M,N) equal to 4 and M or N greater than 8 (e.g., a rectangular block of 4×16 or 16×4 size), a 4×4 secondary transform may be performed on each of the two left-to-top 4×4 sub-blocks of the primary-transformed coefficient block.
[0102] Since a quadratic transform is calculated by multiplying a quadratic transform kernel matrix by an input vector, the encoder first organizes the coefficients in the left-top sub-block of the linearly transformed coefficient block into vectors. The vector organization method depends on the intra prediction mode. For example, if the intra prediction mode is angle mode 34 or lower among the intra prediction modes shown in FIG. 6, the encoder horizontally scans the left-top sub-block of the linearly transformed coefficient block to organize the coefficients into vectors. If the element in the ith row and jth column of the left-top nxn block of the linearly transformed coefficient block is represented as x(i,j), the vectorized coefficients can be expressed as [X(0,0), X(0,1), ..., X(0,n-1), X(1,0), X(1,1), ..., X(1,n-1), ..., X(n-1,0), X(n-1,1), ..., X(n-1,n-1)]. On the other hand, if the intra prediction mode is greater than the 34th angle mode, the left-to-top sub-block of the linearly transformed coefficient block is scanned vertically to construct a vector of coefficients. The vectorized coefficients are represented as [X(0,0), X(1,0), ..., X(n-1,0), X(0,1), X(1,1), ..., X(n-1,1), ..., X(0,n-1), X(1,n-1), ..., X(n-1,n-1)]. To reduce the amount of computation, if only a portion of the 8x8 samples are used in the 8x8 quadratic transform, coefficients x_ij where i>3 and j>3 may not be included in the above vector construction method. In this case, 16 linearly transformed coefficients may be input to the secondary transform in a 4x4 quadratic transform. 48 linearly transformed coefficients may be input to the secondary transform in an 8x8 quadratic transform.
[0103] The encoder obtains quadratically transformed coefficients by multiplying the left-to-top sub-block samples of the vectorized primary transform coefficient block by a quadratically transformed kernel matrix. The quadratically transformed kernel applied to the quadratically transformed coefficients is determined according to the size of the transform unit or transform block, the intra mode, and a syntax element indicating the transform kernel. As described above, reducing the number of coefficients to be quadratically transformed reduces the amount of computation and the memory required to store the transform kernels. Therefore, the number of coefficients to be quadratically transformed is determined according to the size of the current transform block. For example, for a 4x4 block, the encoder obtains a length 8 coefficient vector by multiplying a length 16 vector by an 8(row) x 16(column) transform kernel matrix. The 8(row) x 16(column) transform kernel matrix is obtained based on the first to eighth basis vectors constituting the 16(row) x 16(column) transform kernel matrix. For 4xN or Mx4 blocks (where N and M are 8 or greater), the encoder obtains a length-16 coefficient vector by multiplying a length-16 vector with a 16 (row) x 16 (column) transformation kernel matrix. For 8x8 blocks, the encoder obtains a length-8 coefficient vector by multiplying a length-48 vector with an 8 (row) x 48 (column) transformation kernel matrix. The 8 (row) x 48 (column) transformation kernel matrix is obtained based on the first to eighth basis vectors that make up the 16 (row) x 48 (column) transformation kernel matrix. For MxN blocks (where M and N are 8 or greater) other than 8x8, the encoder obtains a length-16 coefficient vector by multiplying a length-48 vector with a 16 (row) x 48 (column) transformation kernel matrix.
[0104] According to one embodiment of the present invention, the quadratically transformed coefficients are in the form of vectors and are therefore represented as two-dimensional data. The quadratically transformed coefficients are organized into upper-left coefficient sub-blocks according to a preset scan order. In one embodiment, the preset scan order is a diagonal scan order from the top right corner. However, the present invention is not limited thereto, and the diagonal scan order from the top right corner is determined based on the method described below with reference to FIGS. 13 and 14.
[0105] According to an embodiment of the present invention, the transform coefficients of the entire transform unit including the secondarily transformed coefficients are quantized and then transmitted in a bitstream. The bitstream includes syntax elements related to the secondarily transformed block. Specifically, the bitstream includes information indicating whether the secondarily transformed block is to be subjected to the secondarily transformed block and information indicating the transform kernel to be applied to the current block.
[0106] A decoder first parses quantized transform coefficients from the bitstream and obtains transform coefficients through dequantization. Dequantization is also called scaling. The decoder determines whether a secondary inverse transform is to be performed on the current block based on the syntax element for the secondary transform. If a secondary inverse transform is applied to the current transform unit or transform block, 8 or 16 transform coefficients can be input to the secondary inverse transform depending on the size of the transform unit or transform block. The number of coefficients input to the secondary inverse transform corresponds to the number of coefficients output by the secondary transform of the encoder. For example, if the size of the transform unit or transform block is 4x4 or 8x8, 8 transform coefficients are input to the secondary inverse transform; otherwise, 16 transform coefficients are input to the secondary inverse transform. If the size of the transform unit is MxN, for an intra-predicted block with min(M,N) being 4, a 4x4 secondary inverse transform is performed on 16 or 8 coefficients in the left-top 4x4 sub-block of the transform coefficient block. For intra-predicted blocks where min(M,N) is 8 or greater, an 8x8 quadratic transform is performed on the 16 or 8 coefficients of the top-left 4x4 sub-blocks of the transform coefficient block. In one embodiment, to improve coding efficiency, if min(M,N) is 4 and M or N is greater than 8 (e.g., a rectangular block of 4x16 or 16x4 size), an inverse 4x4 quadratic transform may be performed on each of the two top-left 4x4 sub-blocks in the coefficient block.
[0107] According to one embodiment of the present invention, since the second-order inverse transform is calculated by multiplying the second-order inverse transform kernel matrix and the input vector, the decoder configures the previously input dequantized transform coefficient block into the form of a vector according to a preset scan order. According to one embodiment, the preset scan order is the upper right diagonal scan order, but the present invention is not limited thereto, and the upper right diagonal scan order is determined based on the method described below with reference to FIGS. 13 and 14.
[0108] According to another embodiment of the present invention, a decoder obtains linearly transformed coefficients by multiplying vectorized transform coefficients with a secondary inverse transform kernel matrix. The secondary inverse transform kernel is determined based on the size of the transform unit or transform block, the intra mode, and a syntax element indicating the transform kernel. The secondary inverse transform kernel matrix is the transpose of the secondary transform kernel matrix. Considering implementation complexity, the elements of the kernel matrix are integers expressed with 10-bit or 8-bit precision. The length of the vector output from the secondary inverse transform is determined based on the size of the current transform block. For example, for a 4x4 block, a length 16 coefficient vector is obtained by multiplying a length 8 vector with an 8 (row) x 16 (column) transform kernel matrix. The 8 (row) x 16 (column) transform kernel matrix is obtained based on the first to eighth basis vectors constituting the 16 (row) x 16 (column) transform kernel matrix. For 4xN or MxN blocks (N and M are 8 or greater), a 16-length coefficient vector is obtained by multiplying a 16-length vector with a 16 (row) x 16 (column) transformation kernel matrix. For 8x8 blocks, a 48-length coefficient vector is obtained by multiplying a 8-length vector with an 8 (row) x 48 (column) transformation kernel matrix. The 8 (row) x 48 (column) transformation kernel matrix is obtained based on the first to eighth basis vectors that make up the 16 (row) x 48 (column) transformation kernel matrix. For MxN blocks (M and N are 8 or greater) other than 8x8, a 48-length coefficient vector is obtained by multiplying a 16-length vector with a 16 (row) x 48 (column) transformation kernel matrix.
[0109] In one embodiment, since the primary conversion coefficients obtained through the secondary inverse conversion are in vector form, the decoder can further represent them as two-dimensional data, which is intra-mode dependent. At this time, the mapping relationship based on the intra-mode applied by the encoder is also applied. As described above, if the intra prediction mode is 34th angle mode or less, the decoder scans the coefficient vector obtained by the secondary inverse conversion horizontally to obtain a two-dimensional conversion coefficient array. If the intra prediction mode is greater than the 34th angle mode, the decoder scans the coefficient vector obtained by the secondary inverse conversion vertically to obtain a two-dimensional conversion coefficient array. The decoder performs a primary inverse conversion on all conversion units or conversion coefficient blocks of the conversion block size including the conversion coefficients obtained by performing the secondary inverse conversion to obtain a residual signal.
[0110] Although not shown in FIG. 12, a scaling process using a bit shift operation may be included when applying the conversion or inverse conversion to correct the scale increased by the conversion kernel after the change or inverse conversion.
[0111] FIG. 13 is a diagram showing a method for determining the upper right diagonal scan order according to an embodiment of the present invention. According to an embodiment of the present invention, during encoding or decoding, a process of initializing the scan order is performed. Initialization of an array including scan order information is performed according to the block size. Specifically, for the combination of log2BlockWidth and log2BlockHeight, the array initialization process of the upper right diagonal scan order shown in FIG. 13 with 1<<log2BlockWidth and 1<<log2BlockHeight as inputs is called (or performed). The output of the array initialization process of the upper right diagonal scan order is assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight]. Here, log2BlockWidth and log2BlockHeight are variables representing the base-2 logarithm values of the width and height of the block, respectively, and are values in the range [0, 4].
[0112] Through the array initialization process of the upper right diagonal scan order shown in FIG. 13, the encoder / decoder outputs the array diagScan[sPos][sComp] for blkWidth, the width of the input block, and blkHeight, the height of the block. The sPos, which is the index of the array, indicates the scan position (scan index) and is a value in the range [0, blkWidth*blkHeight-1]. If the sComp, which is the index of the array, is 0, sPos indicates the horizontal component (x), and if sComp is 1, sPos indicates the vertical component (y). The algorithm shown in FIG. 13 is interpreted such that the x-coordinate value and y-coordinate value on the two-dimensional coordinates at the scan position sPos in the upper right diagonal scan order are assigned to diagScan[sPos][0] and diagScan[sPos][1], respectively. That is, the value stored in the DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][sComp] array (or array) means the coordinate value corresponding to sComp at the sPos scan position (scan index) in the upper right diagonal scan order of the block whose block width and height are 1<<log2BlockWidth and 1<<log2BlockHeight, respectively.
[0113] FIG. 14 illustrates the upper right diagonal scan order according to one embodiment of the present invention, depending on the block size. Referring to FIG. 14(a), if log2BlockWidth and log2BlockHeight are both 2, this indicates a 4x4 block. Referring to FIG. 14(b), if log2BlockWidth and log2BlockHeight are both 3, this indicates an 8x8 block. In FIG. 14, the numbers in the gray shaded area indicate the scan position (scan index) sPos. The x and y coordinate values at the sPos position are assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][0] and DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][1], respectively.
[0114] The encoder / decoder codes the transform coefficient information based on the above-mentioned scan order. Although the present invention will be described mainly based on an embodiment in which a top-right scan method is used, the present invention is not limited thereto and can be applied to other known scan methods.
[0115] The decoding process for the quadratic transform will now be described in detail. For convenience of explanation, the decoder will be mainly described for the process for the quadratic transform, but the embodiments described below apply to the encoder in substantially the same manner.
[0116] FIG. 15 illustrates a method for indicating a second-order transform at the coding unit level. The second-order transform is indicated at the coding unit level, and syntax elements related to the second-order transform are included in the coding_unit syntax structure. The coding_unit syntax structure includes syntax elements related to the coding unit. The coding_unit syntax structure includes inputs such as (x0, y0), which are the coordinates of the left-most and top-most luma samples of the current block based on the left-most and top-most luma samples of the picture, cbWidth, which is the block width, cbHeight, which is the block height, and treeType, which is a variable indicating the type of coding tree. Because there is a correlation between luma and chroma, encoding luma and chroma using the same coding structure enables efficient video compression. To improve coding efficiency, luma and chroma may be encoded using different coding structures. If the variable treeType is SINGLE_TREE, it means that luma and chroma are encoded using the same coding tree structure, and the coding unit includes a luma coding block and a chroma coding block depending on the color format. If treeType is DUAL_TREE_LUMA, it means that luma and chroma are coded using different coding trees, and the currently processed tree is the tree for luma. In this case, the coding unit includes only luma coding blocks. If treeType is DUAL_TREE_CHROMA, it means that luma and chroma are coded using different coding trees, and the currently processed tree is the tree for chroma. In this case, the coding unit includes chroma coding blocks according to the color format.
[0117] The coding_unit syntax structure indicates a prediction method for the current coding unit, and the variable CuPredMode[x0][y0] indicates the prediction method for the current block. If CuPredMode[x0][y0] is MODE_INTRA, it indicates that the intra prediction method is applied to the current block, and if it is MODE_INTER, it indicates that the inter prediction method is applied to the current block. Also, if CuPredMode[x0][y0] is MODE_IBC, it indicates that the current block is subjected to Intra Block Copy (IBC) prediction, which generates a reference block from a region where reconstruction of the current picture is completed and performs prediction. Syntax elements related to the prediction method are processed according to the value of the variable CuPredMode[x0][y0]. For example, if the variable CuPredMode[x0][y0] indicates intra prediction, the decoder parses syntax elements including information about the intra prediction mode, reference line index, and ISP (Intra Sub-Partitions) prediction, or sets variables related to the intra prediction mode according to a preset method.
[0118] After processing the syntax elements related to the prediction method, the syntax elements related to the residual signal are processed. The transform_tree() syntax structure is a syntax structure for a transform tree, where the transform tree is divided into nodes of smaller size than the root node, with the root node being the same size as the coding unit, and the leaf nodes of the transform tree are transform units. The transform_tree syntax structure contains information about the division of the transform tree.
[0119] One of the intra prediction methods is PCM (Pulse Code Modulation) prediction. When PCM prediction is used for predicting the current coding unit, no transform or quantization is performed, and therefore no transform_tree syntax structure exists. In other words, since no transform_tree syntax structure exists, the decoder does not perform operations related to the transform_tree syntax structure. When intra prediction is indicated for the current coding unit, PCM prediction is indicated by pcm_flag[x0][y0]. In other words, if pcm_flag[x0][y0] is 1, the decoder does not perform operations related to the transform_tree syntax structure. Meanwhile, the presence or absence of a transform_tree syntax structure for the current coding unit is indicated by a 1-bit flag, which is referred to as, but not limited to, cu_cbf in the present invention. The decoder parses cu_cbf, or, if cu_cbf is not parsed, sets cu_cbf according to a preset method. If cu_cbf is 1, the decoder operates according to the transform_tree syntax structure. If inter prediction or IBC prediction is used to predict the current coding unit, merge prediction can also be used to predict the current coding unit. Whether merge prediction is used is indicated by merge_flag[x0][y0]. If merge prediction is indicated to be used for the current block (merge_flag[x0][y0] == 1), cu_cbf is not parsed, and the value of cu_cbf is determined according to a preset method. The preset method is based on cu_skip_flag[x0][y0], which indicates skip mode. For example, if cu_skip_flag[x0][y0] is 1, cu_cbf is inferred to 0; otherwise, cu_cbf is inferred to 1. If cu_cbf is 1, the transform_tree syntax structure is processed and the counter value for measuring the number of non-zero significant coefficients is initialized to 0.
[0120] The numSigCoeff variable indicates the number of non-zero quantization coefficients present in the transform unit of the current coding unit, and the processing of syntax elements related to secondary transform may differ depending on the value of numSigCoeff.
[0121] The numZeroOutSigCoeff variable indicates the number of non-zero quantization coefficients present at a specific position within a transform unit included in the current coding unit, and the processing of syntax elements related to secondary transforms may differ depending on the value of numZeroOutSigCoeff.
[0122] In transform_tree, the transform tree is divided, and the leaf nodes of the transform tree are transform units. Transform_tree includes a transform_unit syntax structure, which is a syntax structure related to the transform unit, which is the leaf node. Transform_unit processes syntax elements related to the transform unit, and if the corresponding transform unit contains one or more non-zero coefficients, it includes a residual_coding syntax structure. The residual_coding syntax structure includes a syntax structure related to quantized transform coefficients and related processing. The transform blocks that make up the transform unit may vary depending on the type of tree currently being processed. If treeType is SINGLE_TREE, the current transform unit includes a luma transform block and, depending on the color format, a chroma transform block. If treeType is DUAL_TREE_LUMA, the current transform unit includes a luma transform block. If treeType is DUAL_TREE_CHROMA, the current transform unit includes a chroma transform block. The transform_unit syntax structure includes coded block flag (CBF) information, which indicates whether the transform block of the current transform unit includes one or more non-zero coefficients, depending on the treeType. The CBF information is information indicated for each color component. For example, if the CBF value for the luma transform block of the current transform unit indicates that the luma transform block does not include one or more non-zero coefficients, the residual_coding syntax structure for the luma transform block is not processed because all coefficients of the luma transform block are zero. As another example, if the CBF value for the chroma Cb transform block of the current transform unit indicates that the chroma Cb transform block includes one or more non-zero coefficients, the residual_coding syntax structure for the Cb transform block of the current transform unit exists.
[0123] Whether a secondary transform is applied to the current block is indicated at the CU level. If a secondary transform is applied, an index indicating a transform kernel used for the secondary transform may also be indicated. As described with reference to FIG. 11, whether a secondary transform is applied to the current block is indicated using the lfnst_idx[x0][y0] syntax element. The first bit of lfnst_idx[x0][y0] indicates whether a secondary transform is applied to the current coding unit. If the first bit of lfnst_idx[x0][y0] is 0, that is, if lfnst_idx[x0][y0] is 0, it indicates that a secondary transform is not applied to the current block. On the other hand, if the first bit of lfnst_idx[x0][y0] is 1, that is, if lfnst_idx[x0][y0] is greater than 0 (lfnst_idx[x0][y0]>0), it indicates that a secondary transform is applied to the current block. In this case, an additional bit is used to indicate the transform kernel used for the secondary transform, and an index indicating the secondary transform kernel is signaled via the additional bit.
[0124] The lfnst_idx[x0][y0] syntax element is parsed if the conditions described below are met, whereas if the conditions described below are not met, lfnst_idx[x0][y0] is not present in the current coding unit and lfnst_idx[x0][y0] is set to 0.
[0125] In other words, if the conditions described in the first to fourth embodiments, including the lfnst_idx[x0][y0] syntax element parsing conditions described below, are satisfied, the encoder generates a bitstream including the lfnst_idx[x0][y0] syntax element for the current coding unit. On the other hand, if the conditions described below are not satisfied, the bitstream generated by the encoder does not include the lfnst_idx[x0][y0] syntax element for the current coding unit, and lfnst_idx[x0][y0] is set to 0. A decoder receiving such a bitstream parses the lfnst_idx[x0][y0] syntax element based on the conditions described below.
[0126] Parsing conditions for lfnst_idx[x0][y0] syntax element
[0127] i) Min(lfnWidth,lfnHeight)>=4
[0128] The first condition relates to the size of the block. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0129] Specifically, the decoder checks the block size conditions for which secondary transforms can be applied. The variables SubWidthC and SubHeightC are set according to the color format and indicate the ratio of the width and height of the chroma component to the width and height of the luma component of the picture, respectively. For example, a 4:2:0 color format image has a structure including one chroma sample per four luma samples, so SubWidthC and SubHeightC are both set to 2. As another example, a 4:4:4 color format image has a structure including one chroma sample per luma sample, so SubWidthC and SubHeightC are both set to 1. The number of horizontal samples of the current block, lfnWidth, and the number of vertical samples, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the coding unit includes only chroma components, so the number of horizontal samples of the chroma coding block is equal to the value obtained by dividing cbwidth, the width of the luma coding block, by SubWidthC. Similarly, the number of vertical samples in a chroma coding block is equal to the height of the luma coding block, cbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the coding unit includes the luma component, so lfnWidth and lfnHeight are set to cbWidth and cbHeight, respectively. Since the minimum condition for a block to which a degree 22 transform can be applied is 4x4, lfnst_idx[x0][y0] is parsed if Min(lfnWidth,lfnHeight)>=4 is satisfied.
[0130] ii)sps_lfnst_enabled_flag==1
[0131] The second condition relates to the flag value indicating whether a secondary transform is activated or applicable. If the value of the flag (sps_lfnst_enabled_flag) indicating whether a secondary transform is activated or applicable is set to 1, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0132] In detail, the secondary transform is indicated by the higher level syntax RBSP. At least one of the SPS, PPS, VPS, tile group header, and slice header includes a 1-bit flag indicating whether the secondary transform is activated and applicable. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element is present in the coding unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element is not present in the coding unit syntax.
[0133] iii)CuPredMode[x0][y0]==MODE_INTRA
[0134] The third condition concerns the prediction mode: the secondary transform is only applied to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0135] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0136] The fourth condition relates to whether the ISP prediction method is applied. If the ISP prediction method is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0137] More specifically, as described with reference to FIG. 11, when the current CU is divided into a number of transform units smaller than the CU size, secondary transform is not applied to the divided transform units. In this case, lfnst_idx[x0][y0], a syntax element related to secondary transform, is set to 0 without being parsed. When the current CU is divided into a number of transform units whose CU size is smaller than that of the transform tree, this includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that divides the transform tree into a number of transform units smaller than the CU size according to a preset division method when intra prediction is applied to the current coding unit. An ISP prediction mode is indicated at the coding unit level, and the variable IntraSubPartitionsSplitType is set based on the ISP prediction mode. In this case, if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Although secondary transform is indicated at the coding unit level, the actual secondary transform is applied at the transform unit level. Therefore, when the transform tree is divided into a number of transform units, it is inefficient to apply the same secondary transform kernel to all divided transform units. Furthermore, due to the characteristics of intra prediction, which generates predicted samples at the transform unit level, the prediction accuracy is higher when the transform tree is divided into multiple transform units than when it is not divided. Therefore, if the transform tree is divided into multiple transform units, the energy of the residual signal is likely to be efficiently compressed even if secondary transforms are not applied to the divided multiple transform units. Furthermore, if the size of the current CU is larger than the size of the maximum luma transform block (MaxTbSizeY) (i.e., cbWidth > MaxTbSizeY || cbHeight > MaxTbSizeY), the transform tree is divided into multiple transform units smaller than the CU size. Although not shown in FIG. 15, secondary transforms are not applied when the size of the current CU is larger than the maximum luma transform block size (MaxTbSizeY).Therefore, the fourth condition may be expressed as IntraSubPartitionsSplitType==ISP_NO_SPLIT&&cbWidth<=MaxTbSizeY&&cbHeight<=MaxTbSize, where MaxTbSizeY is a natural number expressed in the form of a power of 2. MaxTbSizeY may be included and indicated in higher-level syntax RBSP such as SPS, PPS, slice header, or tile group header, or the encoder and decoder may use the same preset value. For example, the preset value may be 64 (2^6).
[0138] v)!intra_mip_flag[x0][y0]
[0139] The fifth condition relates to the intra prediction method. If MIP (Matrix-based Intra Prediction) is not applied to the prediction of the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0140] Specifically, MIP is used as one method of intra prediction, and whether MIP is applicable is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and prediction is performed by multiplying reconstructed samples around the current block by a preset matrix. When MIP is applied, the residual signal exhibits properties that differ from those of general intra prediction, which performs directional or non-directional prediction, and therefore, when MIP is applied, secondary transformation may not be applied to the transform block.
[0141] vi)numSigCoeff>((treeType==SIGNLE_TREE)?2:1)
[0142] The sixth condition relates to the treeType and coefficients.
[0143] Specifically, if treeType is SINGLE_TREE, then if the value of the variable numSigCoeff is greater than 2, a quadratic transform is applied to the current block and the decoder parses the lfnst_idx[x0][y0] syntax element.
[0144] When treeType is DUAL_TREE_LUMA or DUAL_TREE_CHROMA, if the value of the variable numSigCoeff is greater than 1, a secondary transform is applied to the current block, and lfnst_idx[x0][y0] is parsed. In this case, numSigCoeff refers to a variable indicating the number of significant coefficients present in the current coding unit. If numSigCoeff is smaller than a critical value, efficient coding may not be achieved even if a secondary transform is applied to the current block. This is because if the number of significant coefficients is small, the overhead of signaling lfnst_idx[x0][y0] compared to the bits required for coefficient coding is relatively large. In this case, a significant coefficient refers to a coefficient that is not zero. Hereinafter, a significant coefficient referred to in the present invention refers to a coefficient that is not zero, as described above.
[0145] vii)numZeroOutSigCoeff==0
[0146] The seventh condition relates to the effectiveness coefficients present at specific locations.
[0147] In more detail, if a secondary transform is applied to the current block, the transform coefficients quantized by the decoder are always 0 at a specific position. Therefore, if a non-zero (quantized) coefficient exists at a specific position, it means that a secondary transform has not been applied to the current block, and whether to parse lfnst_idx[x0][y0] is determined according to the number of significant coefficients at the specific position. For example, if numZeroOutSigCoeff is not 0, it means that a significant coefficient exists at the specific position, so lfnst_idx[x0][y0] is not parsed and is set to 0. On the other hand, if numZeroOutSigCoeff is 0, it means that no significant coefficient exists at the specific position, so lfnst_idx[x0][y0] is parsed.
[0148] FIG. 16 is a diagram illustrating a residual_coding syntax structure according to one embodiment of the present invention.
[0149] The residual_coding syntax structure is a syntax structure related to quantization coefficients and receives x0, y0, log2TbWidth, and log2TbHeight as inputs. At this time, x0 and y0 represent the left-upper coordinates (x0, y0) of the conversion block, log2TbWidth is the value obtained by taking the base-2 logarithm of the width of the conversion block, and log2TbHeight is the value obtained by taking the base-2 logarithm of the height of the conversion block. The number within the conversion block is coded in sub-block units, and the values of the coefficients within each sub-block are determined based on various syntax elements including sig_coeff_flag. At this time, the coefficients in sub-block units may be expressed as a coefficient group (Coefficient Group, CG). sig_coeff_flag[xC][yC] indicates whether the coefficient value at the (xC, yC) position within the current block is 0. If sig_coeff_flag[xC][yC] is 1, it indicates that the coefficient value at the corresponding position is not 0, and if sig_coeff_flag[xC][yC] is 0, it indicates that the coefficient value at the corresponding position is 0. In residual_coding, the x coordinate value and y coordinate value of the last significant coefficient in the scan order are indicated. Based on the x coordinate value and y coordinate value of the last significant coefficient in the scan order, the index (lastSubBlock) of the sub-block including the last significant coefficient in the scan order is determined. The index of the sub-block is also indexed based on the scan order. The scan order is the upper-right diagonal scan order described in FIG. 13. In the coefficient coding in sub-block units, the indexes xC and yC indicating the coefficient position (coordinate value) are determined based on the left-upper coordinates (xS<<log2SbW, yS<<log2SbH) of the sub-block and the upper-right diagonal scan order (DiagScanOrder). At this time, xS and yS indicate the horizontal index and vertical index, respectively. log2SbW and log2SbH are the values obtained by taking the base-2 logarithm of the width and height of the sub-block, respectively.
[0150] If the value of sig_coeff_flag[xC][yC] is 1 (i.e., the number of positions (xC, yC) is not 0), and if transform skip is not applied to the current block (i.e., !transform_skip_flag[x0][y0]), numSigCoeff is counted. Since quadratic transforms may not be applied when transform skip is applied, numSigCoeff, which is used to parse lfnst_idx[x0][y0], counts the number of valid coefficients in blocks to which transform skip is not applied.
[0151] Also, as described in Figure 15, when a quadratic transform is applied to a transform block, no significant coefficients exist in a specific region within the transform block. Therefore, a numZeroOutSigCoeff counter counts the number of significant coefficients (numZeroOutSigCoeff) existing in the specific region, and lfnst_idx[x0][y0] is not parsed unless numZeroOutSigCoeff is 0. In particular, when a quadratic transform is applied to a transform block, the region where significant coefficients cannot exist is determined by the size of the transform block.
[0152] For example, if the size of the transform block is 4x4 (i.e., log2TbWidth==2&& log2TbHeight==2) to apply a quadratic transform, the transform block is divided into an index region [0,7] and an index region [8,15] in the scan order, and significant coefficients exist in the [0,7] region, but not in the [8,15] region. The 4x4 transform block includes one sub-block. Therefore, when the size of the transform block is 4x4, if the scan position is 8 or greater and the sub-block index is 0 (i.e., n>=8&&i==0), the number of significant coefficients is counted. In this case, the scan order is the upper right diagonal scan order.
[0153] As another example, if the size of the transform block is 8x8 (i.e., log2TbWidth==3&& log2TbHeight==3), significant coefficients exist only in the first sub-block of the transform block, and significant coefficients cannot exist in the remaining sub-blocks (e.g., the second and third sub-blocks). Even in the first sub-block, significant coefficients exist in the index [0,7] range in scan order, but not in the index [8,15] range. Therefore, if the size of the transform block is 8x8, the number of significant coefficients is counted if the scan position in the first sub-block is 8 or greater (i.e., n>=8&&i==0) or if the scan position exists in the remaining sub-blocks excluding the first sub-block (e.g., exists in the second and third sub-blocks, i==1||i==2).
[0154] Finally, if the size of the transform block is greater than 8x8, only the first sub-block in the transform block contains significant coefficients, and the remaining sub-blocks (e.g., the second and third sub-blocks) cannot contain significant coefficients. Therefore, if the sub-block is the second or third sub-block (i.e., i==1||i==2), the number of significant coefficients is counted. Like the numSigCoeff counter, the numZeroOutSigCoeff counter counts the number of significant coefficients only when sig_coeff_flag[xC][yC] is 1 and transform_skip_flag[x0][y0] is 0. In this case, the sub-blocks are indexed in the upper right diagonal scan order described in FIG. 13.
[0155] In other words, if a non-zero coefficient exists in an area (specific area) where no significant coefficients can exist, it means that a secondary transformation has not been performed, so significant coefficients are counted to check whether a non-zero coefficient exists in the specific area.
[0156] FIG. 17 illustrates a method for indicating secondary transformation at a coding unit level according to an embodiment of the present invention.
[0157] As described in Figures 15 and 16, whether or not a secondary transform is applied is indicated by the lfnst_idx[x0][y0] syntax element at the coding unit level, and two significant coefficient counters (i.e., numSigCoeff counter and numZeroOutSigCoeff counter) are required to parse lfnst_idx[x0][y0]. In particular, in the case of numSigCoeff, since the numSigCoeff counter must count the number of significant coefficients present within the entire area of the coding unit, the throughput of coefficient coding may be reduced. Therefore, a method is needed to reduce the number of counters or not use counters.
[0158] The secondary conversion instruction method shown in Fig. 17 is a method of parsing lfnst_idx[x0][y0] regardless of numSigCoeff. In other words, if all of i), ii), iii), iv), and v) of the conditions described in Fig. 15 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. Also, since the value of numSigCoeff is not referenced, the operation of the numSigCoeff counter described in Fig. 16 is not performed.
[0159] Hereinafter, this specification will describe a method for instructing secondary transform based on position information of the last significant coefficient in scan order. As with when the number of significant coefficients is small, if the position (scan index) of the last significant coefficient in scan order is small, the coding efficiency of secondary transform is low. Therefore, it is necessary to efficiently instruct secondary transform based on position information of the last significant coefficient in scan order without using a counter.
[0160] (First Example)
[0161] FIG. 18 illustrates a method for indicating secondary transformation at a coding unit level according to an embodiment of the present invention.
[0162] FIG. 18 is a diagram illustrating a method of parsing lfnst_idx[x0][y0] using the position information of the last significant coefficient in the scan order obtained by residual_coding instead of the numSigCoeff counter.
[0163] As shown in FIG. 18, the numSigCoeff counter is not used, so the numSigCoeff value does not need to be initialized. Instead, lfnLastScanPos, a variable related to the position of the position information of the last significant coefficient in scan order, is initialized to 1. If the value of lfnLastScanPos is 1, this indicates that the position (scan index) of the last significant coefficient in scan order is less than a critical value, or all transform coefficients in the block are 0. Conversely, if the value of lfnLastScanPos is 0, this indicates that there is at least one significant coefficient in the block, and the position (scan index) of the last significant coefficient in scan order is greater than or equal to the critical value. Therefore, if the value of lfnLastScanPos is 1, lfnst_idx[x0][y0] is not parsed, and if the value of lfnLastScanPos is 0, lfnst_idx[x0][y0] is parsed. Additionally, lfnst_idx[x0][y0] may be parsed if the value of lfnLastScanPos is 0 and all of the conditions i), ii), iii), iv), v), and vii) described in Figure 15 are satisfied (all are true).
[0164] In other words, if there is at least one significant coefficient in the current block and the position (scan index) of the last significant coefficient in the scan order is equal to or greater than a threshold, lfnst_idx[x0][y0] is parsed. Here, the threshold is an integer equal to or greater than 0, as will be described later. For example, assuming the threshold is 1, the fact that the position (scan index) of the last significant coefficient in the scan order is equal to or greater than the threshold means that a significant coefficient exists in a position other than the upper left corner of the block. In other words, lfnst_idx[x0][y0] is parsed only in the remaining cases, excluding the cases where a significant coefficient does not exist in the current block or exists only in the upper left corner of the current block, i.e., only when a significant coefficient exists in a position other than the upper left corner of the current block. The presence of a significant coefficient in a position other than the upper left corner of the current block may be expressed as "LfnstDConly==0." The upper left corner of a block described in the present invention may mean that the vertical coordinate value is (0,0), or may mean the first position according to a preset scan order (e.g., upper right diagonal order), or may be referred to as DC.
[0165] FIG. 19 is a diagram illustrating a residual_coding syntax structure according to an embodiment of the present invention.
[0166] FIG. 19 shows the residual_coding syntax structure according to FIG. 18. In residual_coding, syntax elements related to the x- and y-coordinates of the last significant coefficient in scan order are parsed, and LastSignificantCoeffX and LastSignificantCoeffY variables are set. LastSignificantCoeffX indicates the x-coordinate of the last significant coefficient in scan order, and LastSignificantCoeffY indicates the y-coordinate of the last significant coefficient in scan order. Based on LastSignificantCoeffX and LastSignificantCoeffY, the LastScanPos variable, which is the scan index of the last significant coefficient in scan order, and the index of the sub-block containing the last significant coefficient (lastSubBlock) are determined. In this case, as described in FIG. 16, when a secondary transform is applied to the current block, significant coefficients exist only in the first sub-block. In other words, if significant coefficients exist only in the first sub-block, a secondary transform is applied.
[0167] For example, in the 4x4 block of FIG. 14(a), if LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, LastScanPos is determined to be 13. Since the 4x4 block is composed of one sub-block, the index of the sub-block containing the last significant coefficient (lastSubBlock) is determined to be 0. As another example, the 8x8 block of FIG. 14(b) is divided into 4x4 sub-blocks. In detail, in FIG. 14(b), the 4x4 block corresponding to x coordinates 0 to 3 and y coordinates 0 to 3 is set as the first sub-block, the 4x4 block corresponding to x coordinates 0 to 3 and y coordinates 4 to 37 is set as the second sub-block, the 4x4 block corresponding to x coordinates 4 to 7 and y coordinates 0 to 34 is set as the third sub-block, and the 4x4 block corresponding to x coordinates 4 to 7 and y coordinates 4 to 37 is set as the fourth sub-block. In this case, the first sub-block is indexed as 0, the second sub-block as index 1, the third sub-block as index 2, and the fourth sub-block as index 3. The sub-blocks are indexed in the upper right diagonal scan order described in FIG. 13. In this case, if LastScanPos is 2 and LastScanPos is 3, lastScanPos is determined to be 13. Since lastScanPos is 13, the sub-block containing lastScanPos13 is the first sub-block (i.e., sub-block index 0), and the index of the sub-block containing the last significant coefficient (lastSubBlock) is determined to be 0.
[0168] lfnstLastScanPos is determined based on the above-mentioned lastScanPos. More specifically, if the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in the following equation 1. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in the following equation 1. In this case, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block. More specifically, the flag transform_skip_flag[x0][y0] described in this specification indicates whether a linear transform and a secondary transform are applied to the transform block. For example, if the value of the transform_skip_flag[x0][y0] is 1, it indicates that no linear or secondary transform is applied to the transform block (i.e., transform skip is applied), and if the value of the transform_skip_flag[x0][y0] is 0, it indicates that no linear or secondary transform is applied to the transform block (i.e., transform skip is not applied).
[0169]
number
[0170] As mentioned above, the initialization value of lfnstLastScanPos is set to 1.
[0171] In Equation 1, cIdx represents a variable indicating the color component of the current transform block. For example, if cIdx is 0, it indicates that the transform block processed by residual_coding is the luma Y component. If cIdx is 1, it indicates that the transform block processed by residual_coding is the chroma Cb component, and if cIdx is 2, it indicates that the transform block processed is the chroma Cr component. lfnstLastScanPosTh[cIdx], a threshold value for lastScanPos, is set to different values depending on the color component.
[0172] According to Equation 1, if lfnstLastScanPos of a line is 1 and lastScanPos is smaller than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 1. On the other hand, if lfnstLastScanPos of a line is 0 or lastScanPos is equal to or larger than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 0. In other words, if lastScanPos of all transform blocks included in a coding unit are smaller than a threshold value or the number of all transform blocks is 0, lfnstLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being parsed according to the lfnst_idx[x0][y0] parsing condition of FIG. Setting lfnst_idx[x0][y0] to 0 without parsing indicates that a secondary transform is not applied to the current block. On the other hand, if the LastScanPos of any one of the transform blocks included in the coding unit is equal to or greater than the threshold, lfnstLastScanPos is set to 0. If all of conditions i), ii), iii), iv), v), and vii) described in FIG. 15 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if a secondary transform is applied, it checks / determines the transform kernel to be used for the secondary transform.
[0173] In Equation 1, lfnstLastScanPosTh[cIdx] is a preset integer value equal to or greater than 0, and both the encoder and decoder use the same value. Alternatively, the same threshold value may be used for all color components. In this case, lfnstLastScanPos is set as shown in Equation 2 below. A coding unit described herein is composed of multiple coding blocks, each of which has a corresponding transform block. The transform blocks are transform blocks having luma and chroma components. More specifically, they are Y transform blocks, Cb transform blocks, and Cr transform blocks. In this regard, whether lfnst_idx[x0][y0] described herein should be parsed is determined for each transform block corresponding to each coding block. In other words, if any one of the Y transform block, Cb transform block, and Cr transform block satisfies the conditions described herein, lfnst_idx[x0][y0] is parsed.
[0174]
number
[0175] lfnstLastScanPosTh is a preset integer value greater than or equal to 0, and both the encoder and decoder use the same value. For example, lfnstLastScanPosTh may be 1. That is, if lastScanPos is greater than or equal to 1, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is parsed. In this case, since the threshold value (lfnstLastScanPosTh) is an integer value, if lastScanPos is greater than or equal to 1, it has the same meaning as if lastScanPos is greater than 0. Although the case where the threshold value is 1 has been described as an example of the present invention, the present invention is not limited thereto.
[0176] In other words, whether to parse lfnst_idx[x0][y0] is determined based on lastScanPos. More specifically, when a quadratic transform is applied as described above, the last valid coefficient in scan order is only present in the first sub-block of the transform block. Therefore, if the index (lastSubBlock) of the sub-block containing the last valid coefficient in scan order (where the index indicated by lastScanPos is located) is 0, the width of the transform block is 4 or more (log2TbWidth>=2), the height of the transform block is 4 or more (log2TbHeight>=2), transform_skip_flag[x0][y0] is 0 (transform skip is not applied), and LastScanPos is greater than 0 (LastScanPos is 1 or greater), lfnst_idx[x0][y0] is parse. This can be expressed as Equation 3 below.
[0177]
number
[0178] On the other hand, in the first embodiment described above, the numSigCoeff counter is not used for parsing lfnst_idx[x0][y0], so the number of valid coefficients (numSigCoeff) is not counted.
[0179] (Second Example)
[0180] FIG. 20 is a diagram illustrating a residual_coding syntax structure according to another embodiment of the present invention.
[0181] FIG. 20 shows a method for setting a threshold value for LastScanPos according to treeType when residual_coding is further input to FIG. 19 and a treeType variable is added.
[0182] If the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in the following formula 4. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in the following formula 4. In this case, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.
[0183]
number
[0184] In Equation 4, lfnstLastScanPosTh means a threshold value for lastScanPos, and its value is set according to treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, lfnstLastScanPosTh is set to val1, val2, or val3, respectively. If lfnstLastScanPos of a line is 1 and lastScanPos is smaller than lfnstLastScanPosTh, lfnstLastScanPos is updated to 1. On the other hand, if lfnstLastScanPos of a line is 0 or lastScanPos is equal to or greater than lfnstLastScanPosTh, lfnstLastScanPos is updated to 0.
[0185] As a result, in Equation 4, if the lastScanPos of all transform blocks included in the coding unit are smaller than the threshold value or the number of all transform blocks is 0, lfnstLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is not parsed and is set to 0 according to the lfnst_idx[x0][y0] parsing condition of Figure 18. This indicates that secondary transform is not applied to the current block. On the other hand, if the LastScanPos of any one of the transform blocks included in the coding unit is equal to or greater than the threshold value, lfnstLastScanPos is determined to be 0, and if all of i), ii), iii), iv), v), and vii) described in Figure 15 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if so, determines the transform kernel to be used for the secondary transform.
[0186] val1, val2, and val3 are preset integer values greater than or equal to 0, and both the encoder and decoder use the same value. If treeType is SINGLE_TREE, both luma and chroma components are included, so val1, which is the value of lfnstLastScanPosTh, may be expressed as the sum of val2 and val3.
[0187] In the second embodiment, the numSigCoeff counter is not used for parsing lfnst_idx[x0][y0], so the number of valid coefficients (numSigCoeff) is not counted.
[0188] (Third Example)
[0189] FIG. 21 is a diagram illustrating a method for indicating a secondary transform at a coding unit level according to another embodiment of the present invention.
[0190] According to FIG. 21, lfnst_idx[x0][y0] is parsed using the position information of the last significant coefficient in the scan order obtained by residual_coding instead of the numSigCoeff counter.
[0191] Since the numSigCoeff counter is not used, numSigCoeff does not need to be initialized, and lfnLastScanPos, a variable related to the position information of the last significant coefficient in scan order, is initialized to 0. The lfnstLastScanPos variable in FIG. 21 is a value obtained by adding the lastScanPos of the transform blocks included in the coding unit. In this case, if lfnLastScanPos is greater than a critical value and all of conditions i), ii), iii), iv), v), and vii) described in FIG. 15 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if a secondary transform is applied, it checks / determines the transform kernel to be used for the secondary transform. On the other hand, if lfnLastScanPos is less than or equal to the critical value, lfnst_idx[x0][y0] is not parsed and is set to 0. This indicates that no quadratic transformation is applied.
[0192] The threshold is set according to the treeType. If the treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the thresholds are set to Th1, Th2, and Th3, respectively. Th1, Th2, and Th3 are preset integers greater than or equal to 0, and both the encoder and decoder use the same value. If the treeType is SINGLE_TREE, both luma and chroma components are included, so the threshold Th1 may be expressed as the sum of Th2 and Th3.
[0193] FIG. 22 is a diagram illustrating a residual_coding syntax structure according to another embodiment of the present invention.
[0194] Figure 22 shows the residual_coding syntax structure according to Figure 21 above, and if the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in the following Equation 5. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in the following Equation 5. In this case, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.
[0195]
number
[0196] In Equation 5, lfnLastScanPos is the sum of all lastScanPos of the transform blocks included in the coding unit, and as described in FIG. 21, whether or not to parse lfnst_idx[x0][y0] is determined by comparing lfnLastScanPos with a threshold value.
[0197] In the third embodiment, the numSigCoeff counter is not used for parsing lfnst_idx[x0][y0], so the number of valid coefficients (numSigCoeff) is not counted.
[0198] Meanwhile, a coding unit includes transform units divided by a transform tree with the same size as the coding unit as the root node. In this case, the transform unit includes a transform block for each color component. If a secondary transform is specified at the coding unit level, residual coding is performed on all transform blocks included in the coding unit, and then lfnst_idx[x0][y0] is parsed based on coefficient information. In another embodiment, the secondary transform may be specified at the transform unit level. If a secondary transform is specified at the transform unit level, each transform unit included in the coding unit uses a different lfnst_idx[x0][y0]. Therefore, the encoder can find the optimal lfnst_idx[x0][y0] for each transform unit, thereby further improving coding efficiency. Also, if a secondary transform is specified at the coding unit level and the coding unit includes four transform units, residual coding must be performed on all transform blocks included in the four transform units in order to parse lfnst_idx[x0][y0]. That is, even if the decoder obtains the transform coefficients for the first transform unit through residual coding, the decoder cannot perform the inverse transform for the first transform unit because it cannot obtain the lfnst_idx[x0][y0] value, which not only increases the decoder buffer size but also may cause excessive delay in the decoder.
[0199] The first to third embodiments described with reference to Figures 18 to 22 can also be applied to cases where secondary transform is specified at the transform unit level. If secondary transform is specified at the coding unit level, the first to third embodiments determine whether to parse lfnst_idx[x0][y0] based on the position of the last significant coefficient in the scan order of the transform block included in the coding unit. If secondary transform is specified at the transform unit level, the first to third embodiments determine whether to parse lfnst_idx[x0][y0] based on the position of the last significant coefficient in the scan order of the transform block included in the transform unit.
[0200] Hereinafter, the specific manner in which secondary transforms are indicated at the transform unit level will be described in this specification.
[0201] FIG. 23 is a diagram illustrating a method for directing secondary transforms at the transform unit level according to an embodiment of the present invention.
[0202] According to FIG. 12, lfnst_idx[x0][y0] is parsed using the position information of the last significant coefficient in the scan order obtained by residual_coding instead of the numSigCoeff counter.
[0203] First, before performing residual_coding, lfnLastScanPos, a variable related to the position of the last significant coefficient in scan order, is initialized to 1. If the lfnLastScanPos variable is 1, it indicates that the position (scan index) of the last significant coefficient in scan order for all transform blocks included in the transform unit is less than a critical value, or all transform coefficients in the block are zero. If the lfnLastScanPos variable is 0, it indicates that one or more transform blocks included in the transform unit have at least one significant coefficient in the block, and the position (scan index) of the last significant coefficient in scan order is greater than or equal to a critical value. According to the first embodiment described above, if lfnLastScanPos, which is set based on the position of the last significant coefficient in scan order for the transform block, is 0, and all of the following conditions i), ii), iii), iv), v), and vi) are satisfied (all are true), the decoder parses lfnst_idx[x0][y0].
[0204] Parsing conditions for lfnst_idx[x0][y0] syntax element
[0205] i) Min(lfnWidth,lfnHeight)>=4
[0206] The first condition relates to the size of the block. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0207] Specifically, the decoder checks the block size conditions for which secondary transforms can be applied. The variables SubWidthC and SubHeightC are set according to the color format and indicate the ratio of the width and height of the chroma components to the width and height of the luma components of the picture, respectively. For example, a 4:2:0 color format video has a structure including one chroma sample per four luma samples, so SubWidthC and SubHeightC are both set to 2. As another example, a 4:4:4 color format video has a structure including one chroma sample per luma sample, so SubWidthC and SubHeightC are both set to 1. The number of horizontal samples of the current block, lfnWidth, and the number of vertical samples, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the transform unit includes only chroma components, so the number of horizontal samples of the chroma transform block is equal to the value obtained by dividing tbwidth, the width of the luma transform block, by SubWidthC. Similarly, the number of vertical samples in a chroma transform block is equal to the height of the luma transform block, tbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the transform unit includes the luma component, so lfnWidth and lfnHeight are set to tbwidth and tbHeight, respectively. The minimum condition for a block to which a quadratic transform can be applied is 4x4, so if Min(lfnWidth,lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is parsed.
[0208] ii)sps_lfn_enabled_flag==1
[0209] The second condition relates to the flag value indicating whether the secondary transform is activated or applicable. If the value of the flag (sps_lfnst_enabled_flag) indicating whether the secondary transform is activated or applicable is set to 1, the decoder parses lfnst_idx[x0][y0].
[0210] In detail, secondary transforms are indicated by the high-level syntax RBSP. At least one of the SPS, PPS, VPS, tile group header, and slice header includes a 1-bit flag indicating whether the secondary transform is activated and applicable. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element is present in the transform unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element is not present in the transform unit syntax.
[0211] iii)CuPredMode[x0][y0]==MODE_INTRA
[0212] The third condition is related to the prediction mode: the secondary transform is only applied to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].
[0213] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0214] The fourth condition relates to whether the ISP prediction method is applied. If the ISP prediction method is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0215] More specifically, as described with reference to FIG. 11, when a current CU is divided into a number of transform units smaller than the CU size, secondary transform is not applied to the divided transform units. In this case, lfnst_idx[x0][y0], a syntax element related to secondary transform, is not parsed and is set to 0. When a current CU is divided into a number of transform units whose CU size is smaller than that of the transform tree, ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that divides the transform tree into a number of transform units smaller than the CU size according to a preset division method when intra prediction is applied to the current coding unit. An ISP prediction mode is indicated at the coding unit level, and the IntraSubPartitionsSplitType variable is set accordingly. In this case, IntraSubPartitionsSplitType being ISP_NO_SPLIT indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction, which generates prediction samples at the transform unit level, when the transform tree is divided into a number of transform units, prediction accuracy is higher than when the transform tree is not divided. Therefore, even if a secondary transform is not applied to the divided multiple transform units, the energy of the residual signal is likely to be efficiently compressed.
[0216] v)!intra_mip_flag[x0][y0]
[0217] The fifth condition relates to the intra prediction method. If MIP (Matrix-based Intra Prediction) is not applied to the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0218] Specifically, MIP is used as one method of intra prediction, and whether MIP is applicable is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and prediction is performed by multiplying reconstructed samples around the current block by a preset matrix. When MIP is applied, the residual signal exhibits properties that differ from those of general intra prediction, which performs directional or non-directional prediction, and therefore, when MIP is applied, secondary transformation may not be applied to the transform block.
[0219] vi)numZeroOutSigCoeff==0
[0220] The sixth condition relates to the effective coefficients present at a particular location.
[0221] In more detail, if a secondary transform is applied to the current block, the transform coefficients quantized by the decoder are always 0 at a specific position. Therefore, if a non-zero quantized coefficient exists at a specific position, it means that a secondary transform has not been applied, and lfnst_idx[x0][y0] is parsed according to the number of significant coefficients at the specific position. For example, if numZeroOutSigCoeff is not 0, it means that a significant coefficient exists at the specific position, so lfnst_idx[x0][y0] is not parsed and is set to 0. On the other hand, if numZeroOutSigCoeff is 0, it means that no significant coefficient exists at the specific position, so lfnst_idx[x0][y0] is parsed.
[0222] If whether or not a secondary transform is applied to the current block is indicated at the transform unit level according to the first embodiment described above, the residual_coding method described in FIG. 19 is followed. According to Equation 1 for determining lfnLastScanPos described in FIG. 19, if the lastScanPos of all transform blocks included in the transform unit are less than a threshold value or the number of all transform blocks is 0, lfnLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that a secondary transform is not applied to the current block. On the other hand, if the LastScanPos of any one of the transform blocks included in the transform unit is equal to or greater than the threshold value, lfnstLastScanPos is determined to be 0, and if all of conditions i), ii), iii), iv), v), and vi) described in FIG. 23 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if so, determines the transform kernel to be used for the secondary transform.
[0223] If whether or not a secondary transform is applied is indicated at the transform unit level according to the second embodiment described above, the transform unit syntax structure described in FIG. 23 is applied, and the residual_coding method described in FIG. 20 is used. According to Equation 4 for determining lfnLastScanPos described in FIG. 20, if the lastScanPos of all transform blocks included in the transform unit are less than a threshold value or the number of all transform blocks is 0, lfnLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that a secondary transform is not applied to the current block. On the other hand, if the LastScanPos of any one transform block included in the transform unit is equal to or greater than the threshold value, lfnstLastScanPos is determined to be 0, and if all of conditions i), ii), iii), iv), v), and vi) described in FIG. 23 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if so, determines the transform kernel to be used for the secondary transform.
[0224] FIG. 24 is a diagram illustrating a method for directing a secondary transform at the transform unit level according to another embodiment of the present invention.
[0225] According to the third embodiment described above, lfnst_idx[x0][y0] is parsed using the position information of the last significant coefficient in the scan order obtained by residual_coding instead of the numSigCoeff counter.
[0226] Before performing residual_coding, the variable lfnLastScanPos, which is related to the position of the last significant coefficient in scan order, is initialized to 0. The variable lfnLastScanPos is the sum of the lastScanPos of the transform blocks included in the transform unit. If lfnLastScanPos is greater than a critical value and all of conditions i), ii), iii), iv), v), and vii) described in Figure 23 are satisfied (all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block. If a secondary transform is applied, the decoder checks / determines the transform kernel to be used for the secondary transform. On the other hand, if lfnLastScanPos is less than or equal to the critical value, lfnst_idx[x0][y0] is not parsed and is set to 0, indicating that a secondary transform is not applied.
[0227] The threshold is set according to the treeType. If the treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the thresholds are set to Th1, Th2, and Th3, respectively. Th1, Th2, and Th3 are preset integers greater than or equal to 0, and both the encoder and decoder use the same value. If the treeType is SINGLE_TREE, both luma and chroma components are included, so the threshold Th1 may be expressed as the sum of Th2 and Th3.
[0228] If whether or not a secondary transform is applied is indicated at the transform unit level based on the third embodiment described above, the residual_coding method described in Figure 22 is used. According to Equation 5 for determining lfnLastScanPos described in Figure 22, the variable lfnLastScanPos is set to the sum of all lastScanPos of the transform blocks included in the transform unit. Then, lfnLastScanPos is compared with a threshold value to determine whether or not to parse lfnst_idx[x0][y0].
[0229] On the other hand, if a secondary transform is indicated at the transform unit level, there is a possibility that there is a high correlation between the transform units included in a coding unit. This is because the prediction method is determined at the coding unit level. Therefore, lfnst_idx[x0][y0] is signaled only in the first transform unit included in a coding unit, and the signaled lfnst_idx[x0][y0] is shared with the remaining transform units. That is, only when subTuIndex, which indicates the index of a transform unit, is 0, lfnst_idx[x0][y0] may be parsed using the first to third embodiments described above. If subTuIndex is greater than 0, the corresponding transform unit does not parse lfnst_idx[x0][y0], and uses the value of lfnst_idx[x0][y0] of the first shared transform unit.
[0230] On the other hand, although a counter is used to count the significant coefficients, the decoder decides whether to parse lfnst_idx[x0][y0] by considering only the significant coefficients in the upper left sub-block of the transform block, in order to reduce the amount of calculation.
[0231] On the other hand, if a secondary transform is specified at the transform unit level, the decoder delay time is reduced compared to when it is specified at the coding unit level, but other delay times may occur. For example, even if a secondary transform is specified at the transform unit level, the secondary transform is specified after all of the coding of the luma transform coefficients, Cb transform coefficients, and Cr transform coefficients is completed. Therefore, even if all of the coding (processing) of the luma transform coefficients is completed, the inverse transform processing of the luma transform coefficients is performed after all of the coding (processing) of the Cb transform coefficients and Cr transform coefficients is completed. This results in other delay times in the decoder.
[0232] Hereinafter, in this specification, a method for specifying a secondary transformation that can minimize the delay time of a decoder will be described.
[0233] (Fourth Example)
[0234] One example of a method for indicating a secondary transform that can minimize decoder latency is a method in which the secondary transform is indicated at the transform unit level, but lfnst_idx[x0][y0], a syntax element related to the secondary transform, is parsed before coding the luma transform coefficients. Thus, the decoder can perform an inverse transform process on the luma transform coefficients immediately after the coding of the luma transform coefficients is completed, without waiting for the Cb and Cr transform coefficients. Similarly, the decoder can perform an inverse transform process on the Cb transform coefficients immediately after the coding of the Cb transform coefficients is completed, without waiting for the coding of the Cr transform coefficients. This method of indicating a secondary transform can minimize decoder latency and solve pipeline problems.
[0235] FIG. 25 is a diagram illustrating a coding unit syntax according to one embodiment of the present invention.
[0236] As can be seen from Figure 25, secondary transforms are indicated at the transform unit level, so the syntax for secondary transforms, lfnst_idx[x0][y0], is not parsed at the coding unit level, but at the transform unit level divided by the transform_tree.
[0237] FIG. 26 is a diagram illustrating a method for directing a secondary transform at the transform unit level according to another embodiment of the present invention.
[0238] Referring to Figure 26, the secondary transform is indicated at the transform unit level, and the syntax element lfnst_idx[x0][y0] related to the secondary transform is parsed first before the luma and chroma transform coefficient coding (residual_coding). For example, if lfnst_idx[x0][y0] is parsed before acquiring the transform coefficients, the inverse transform of the Y, Cb, and Cr transform coefficients is performed immediately after the coefficient coding of each color component, Y, Cb, and Cr, is completed. For example, the inverse transform of the luma (Y) transform coefficient is performed immediately after the transform coefficient coding of the Y component is completed. Similarly, the inverse transform of the Cb transform coefficient is performed immediately after the transform coefficient coding of the Cb component (residual_coding) is completed, and the inverse transform of the Cr transform coefficient is performed immediately after the transform coefficient coding of the Cr component (residual_coding) is completed.
[0239] If lfnst_idx[x0][y0] is parsed after transform coefficient coding (residual_coding) for Y, Cb, and Cr, even if transform coefficient coding (residual_coding) for Y is completed, the inverse transform for the Y transform coefficient is not performed / processed until transform coefficient coding (residual_coding) for Cb and Cr is completed / processed. Therefore, even if transform coefficient coding (residual_coding) for Y is completed, the decoder cannot perform the inverse transform for the Y transform coefficient until transform coefficient coding (residual_coding) for other components (Cb, Cr) is completed, resulting in unnecessary delay. However, as described above, if lfnst_idx[x0][y0] is parsed before transform coefficient coding (residual_coding), inverse transform for each color component is performed immediately after transform coefficient coding (residual_coding) for each color component (Y, Cb, Cr) is completed, thereby minimizing decoder delay.
[0240] In the transform_unit() syntax structure, tu_cbf_luma[x0][y0], tu_cbf_cb[x0][y0], tu_cbf_cr[x0][y0], transform_skip_flag[x0][y0], etc. are parsed.
[0241] In detail, tu_cbf_luma[x0][y0] is an element indicating whether the current luma transform block includes one or more non-zero transform coefficients. When tu_cbf_luma[x0][y0] is 1, it indicates that the current luma transform block includes one or more non-zero transform coefficients. When tu_cbf_luma[x0][y0] is 0, it indicates that all the transform coefficients of the current luma transform block are zero. When tu_cbf_cb[x0][y0] is 1, it indicates that the current chroma Cb transform block includes one or more non-zero transform coefficients. When tu_cbf_cb[x0][y0] is 0, it indicates that all the transform coefficients of the current chroma Cb transform block are zero. tu_cbf_cr[x0][y0] is an element that indicates whether the current chroma Cr transform block contains one or more non-zero transform coefficients. If tu_cbf_cr[x0][y0] is 1, it indicates that the current chroma Cr transform block contains one or more non-zero transform coefficients. If tu_cbf_cr[x0][y0] is 0, it indicates that all transform coefficients of the current chroma Cr transform block are zero. transform_skip_flag[x0][y0] is a syntax element related to transform skip. If transform_skip_flag[x0][y0] is 1, it indicates that no inverse transform is applied to the current luma transform block. If transform_skip_flag[x0][y0] is 0, it indicates that whether an inverse transform is applied to the current luma transform block is determined by other syntax elements.
[0242] As an example of a method for indicating a secondary transform according to Figure 26, the syntax element lfnst_idx[x0][y0] for the secondary transform is parsed based on the position of the last significant coefficient in the scan order, rather than based on the number of non-zero transform coefficients.
[0243] First, the lfnLastScanPos variable is initialized to 1. As described in Figure 23, the variable lfnLastScanPos indicates the position information of the last significant coefficient in scan order of the transform block included in the current transform unit. In more detail, if lfnLastScanPos is 1, it indicates that the position (scan index) of the last significant coefficient in scan order for all transform blocks included in the transform unit is less than a critical value, or all transform coefficients in the block are 0. If lfnLastScanPos is 0, it indicates that for one or more transform blocks included in the transform unit, there is one or more significant coefficients in the block, and the position (scan index) of the last significant coefficient in scan order is greater than or equal to the critical value.
[0244] Next, the variable numZeroOutSigCoeff is initialized to 0. If a secondary transform is applied to a transform block, the last significant coefficient in the scan order cannot exist. Therefore, the variable numZeroOutSigCoeff indicates whether a significant coefficient exists at a specific position, and whether a secondary transform is applied is determined based on that. For example, assume that only up to 16 significant coefficients are allowed when a secondary transform is applied to a transform block. For a transform block of 4x4 or 8x8 size, significant coefficients can exist in the index range [0,7] in the scan order (allowing up to 8 non-zero transform coefficients). On the other hand, for a transform block of a size other than 4x4 or 8x8, significant coefficients can exist in the index range [0,15] in the scan order (allowing up to 16 non-zero transform coefficients). Therefore, if the position of the last significant coefficient in the scan order (scan index) is outside the range where significant coefficients can exist, the decoder can automatically recognize that a secondary transform is not applied to the current transform block.
[0245] Whether to parse the syntax element lfnst_idx[x0][y0] related to the secondary transform before coefficient coding (residual_coding) is determined based on the position of the last significant coefficient in the scan order (scan index). Therefore, the decoder processes information about the position of the last significant coefficient in the scan order before coefficient coding (residual_coding).
[0246] In detail, if the current luma transform block contains one or more significant coefficients that are not zero (tu_cbf_luma[x0][y0] == 1), and if transform skip is not applied to the current luma transform block (transform_skip_flag[x0][y0] == 0), last_significant_pos, a syntax structure relating to the position of the last significant coefficient in the luma scan order, is processed.
[0247] If the value of tu_cbf_luma[x0][y0] is 0 (tu_cbf_luma[x0][y0]==0), it indicates that all coefficients of the corresponding transform block are 0, which means that coefficient coding (residual_coding) is not performed. Therefore, there is no need to process the position information of the last significant coefficient in the scan order.
[0248] If transform_skip_flag[x0][y0] is 1, it indicates that the inverse transform is not applied to the current luma transform block, and therefore, coefficient coding (residual_coding) is performed without relying on the position information of the last significant coefficient in the scan order.
[0249] If the current chroma-Cb transform block contains at least one significant coefficient (tu_cbf_cb[x0][y0] == 1), the syntax structure last_significant_pos, which indicates the position of the last non-zero coefficient in the scan order of the current chroma-Cb transform block, is processed. The last_significant_pos syntax structure receives as input the top-left coordinates of the transform block (x0, y0), the base 2 logarithm of the width of the transform block, the base 2 logarithm of the height of the transform block, and cIdx, a variable indicating which color component the transform block represents. For example, cIdx = 0 indicates a luma-Y transform block, cIdx = 1 indicates a chroma-Cb transform block, and cIdx = 2 indicates a chroma-Cr transform block. If the value of tu_cbf_cb[x0][y0] is 0 (tu_cbf_cb[x0][y0]==0), it indicates that all coefficients of the corresponding transform block are 0. This means that coefficient coding (residual_coding) is not performed, so there is no need to process the position information of the last coefficient that is not 0 in the scan order.
[0250] On the other hand, if the current luma Cr transform block contains at least one significant coefficient (tu_cbf_cr[x0][y0] == 1), tu_joint_cbcr_residual[x0][y0], a syntax element that indicates whether chroma Cb and Cr are represented as a single residual signal, is parsed before processing last_significant_pos. For example, if tu_joint_cbcr_residual[x0][y0] is 1, coefficient coding (residual_coding) for Cr is not processed, and the residual signal for Cr is derived from the restored residual signal of Cb. On the other hand, if tu_joint_cbcr_residual[x0][y0] is 0, coefficient coding (residual_coding) for Cr is performed according to the value of tu_cbf_cr[x0][y0]. If the current chroma Cr transform block contains at least one significant coefficient (tu_cbf_cr[x0][y0] == 1), last_significant_pos, a syntax structure related to the position of the last significant coefficient in the scan order of chroma Cr, is processed. If the value of tu_cbf_cbr[x0][y0] is 0 (tu_cbf_cr[x0][y0] == 0), it indicates that all coefficients of the chroma Cr transform block are 0. This means that coefficient coding (residual_coding) is not performed, so there is no need to process the position information of the last non-zero coefficient in the scan order.
[0251] By processing last_significant_pos for each color component, the position (scan index) of the last significant coefficient in the scan order for each color component is obtained, and the lfnLastScanPos and numZeroOutSigCoeff values are updated based on that.
[0252] Then, if all of the following conditions i), ii), iii), iv), v), vi), and vii) are satisfied (all are true), the decoder parses lfnst_idx[x0][y0] before coefficient coding (residual_coding).
[0253] Parsing conditions for the lfnst_idx[x0][y0] syntax element before coefficient coding (residual_coding)
[0254] i) Min(lfnWidth,lfnHeight)>=4
[0255] The first condition relates to the size of the block. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0256] Specifically, the decoder checks the block size conditions for which secondary transforms can be applied. The variables SubWidthC and SubHeightC are set according to the color format and indicate the ratio of the width and height of the chroma components to the width and height of the luma components of the picture, respectively. For example, a 4:2:0 color format video has a structure including one chroma sample per four luma samples, so SubWidthC and SubHeightC are both set to 2. As another example, a 4:4:4 color format video has a structure including one chroma sample per luma sample, so SubWidthC and SubHeightC are both set to 1. The number of horizontal samples of the current block, lfnWidth, and the number of vertical samples, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the transform unit includes only chroma components, so the number of horizontal samples of the chroma transform block is equal to the value obtained by dividing tbwidth, the width of the luma transform block, by SubWidthC. Similarly, the number of vertical samples in a chroma transform block is equal to the height of the luma transform block, tbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the transform unit includes the luma component, so lfnWidth and lfnHeight are set to tbwidth and tbHeight, respectively. The minimum condition for a block to which a quadratic transform can be applied is 4x4, so if Min(lfnWidth,lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is parsed.
[0257] ii)sps_lfnst_enabled_flag==1
[0258] The second condition is related to the flag value indicating whether the secondary transformation is activated or applicable. sps_lfnst_enabled_flag ) is set to 1, the decoder parses lfnst_idx[x0][y0].
[0259] In detail, the secondary transform is indicated by the high-level syntax RBSP. At least one of the SPS, PPS, VPS, tile group header, and slice header includes a 1-bit flag indicating whether the secondary transform is activated and applicable. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element is present in the transform unit syntax, and if sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element is not present in the transform unit syntax.
[0260] iii)CuPredMode[x0][y0]==MODE_INTRA
[0261] The third condition is related to the prediction mode: the secondary transform is only applied to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].
[0262] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0263] The fourth condition relates to whether the ISP prediction method is applied. If the ISP prediction method is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0264] More specifically, as described with reference to FIG. 11, when a current CU is divided into a number of transform units smaller than the CU size, secondary transform is not applied to the divided transform units. In this case, lfnst_idx[x0][y0], a syntax element related to secondary transform, is set to 0 without being parsed. When a current CU is divided into a number of transform units whose CU size is smaller than that of the transform tree, ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that divides the transform tree into a number of transform units smaller than the CU size according to a preset division method when intra prediction is applied to the current coding unit. An ISP prediction mode is indicated at the coding unit level, and the IntraSubPartitionsSplitType variable is set accordingly. IntraSubPartitionsSplitType being ISP_NO_SPLIT indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction, which generates prediction samples at the transform unit level, when the transform tree is divided into a number of transform units, prediction accuracy is higher than when the transform tree is not divided. Therefore, even if a secondary transform is not applied to the divided multiple transform units, the energy of the residual signal is likely to be efficiently compressed.
[0265] v)!intra_mip_flag[x0][y0]
[0266] The fifth condition relates to the intra prediction method. If MIP (Matrix-based Intra Prediction) is not applied to the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0267] Specifically, MIP is used as one method of intra prediction, and whether MIP is applicable is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and prediction is performed by multiplying reconstructed samples around the current block by a preset matrix. When MIP is applied, the residual signal exhibits properties that differ from those of general intra prediction, which performs directional or non-directional prediction, and therefore, when MIP is applied, secondary transformation may not be applied to the transform block.
[0268] vi)lfnLastScanPos==0 The sixth condition relates to the last significant coefficient in scan order of the transform block.
[0269] Specifically, if the position information (scan index) of the last significant coefficient in scan order of a transform block included in the current transform unit is smaller than a preset critical value, the gain in coding efficiency obtained by secondary transform may be small. Therefore, in such a case, the encoder is likely not to apply secondary transform to the transform block (lfnst_idx[x0][y0] is 0), and therefore, it is considered that signaling lfnst_idx[x0][y0] by the encoder would result in a large overhead. Therefore, lfnst_idx[x0][y0] is parsed only if the position (scan index) of the last significant coefficient in scan order for at least one transform block included in the transform unit is equal to or greater than a preset critical value.
[0270] In other words, as described above, the threshold is an integer equal to or greater than 0. For example, assuming that the threshold is 1, the fact that the position (scan index) of the last significant coefficient in the scan order is equal to or greater than the threshold means that the significant coefficient is located at a position other than the upper left corner of the block (scan index 0, DC). In this case, the fact that the position of the last significant coefficient in the scan order of the transform block is equal to or greater than the threshold may be expressed as "lfnLastScanPos==".
[0271] vii)numZeroOutSigCoeff==0
[0272] The seventh condition relates to the effectiveness coefficients present at a particular location.
[0273] Specifically, if a secondary transform is applied to the current block, a significant coefficient cannot exist at a specific position in the scan order. That is, the numZeroOutSigCoeff variable indicates whether a non-zero transform coefficient exists at a specific position. For example, assume that a maximum of 16 significant coefficients are allowed when a secondary transform is applied to the current block. For a 4x4 or 8x8 transform block, significant coefficients can exist in the index range [0,7] in the scan order (allowing a maximum of 8 non-zero transform coefficients). On the other hand, for a transform block having a size other than 4x4 or 8x8, significant coefficients can exist in the index range [0,15] in the scan order (allowing a maximum of 16 non-zero transform coefficients). Therefore, if the position of the last significant coefficient in the scan order (scan index) is outside the range where significant coefficients can exist, the decoder can automatically determine that a secondary transform is not applied to the current block. Therefore, if numZeroOutSigCoeff>0, this means that a secondary transform is not applied to the current block, and lfnst_idx[x0][y0] is set to 0 without being parsed.
[0274] In other words, if numZeroOutSigCoeff is not 0, it means that a valid coefficient exists at the specific position, so lfnst_idx[x0][y0] is not parsed and is set to 0. On the other hand, if numZeroOutSigCoeff is 0, it means that no valid coefficient exists at the specific position, so lfnst_idx[x0][y0] is parsed.
[0275] If all of the above conditions i) to vii) are true, lfnst_idx[x0][y0] is parsed, otherwise lfnst_idx[x0][y0] is not parsed and is set to 0.
[0276] FIG. 27 is a diagram showing a syntax structure relating to the position of the last significant coefficient in the scan order according to an embodiment of the present invention.
[0277] Referring to Figure 27, the last_significant_pos syntax structure refers to a syntax structure including position information of the last significant coefficient in scan order for each color component Y, Cb, and Cr transform block. The last_significant_pos syntax structure receives as input (x0, y0), which are the left-to-top coordinates of the transform block, log2TbWidth, which is the logarithm of the width of the transform block to base 2, log2TbHeight, which is the logarithm of the height of the transform block to base 2, and cIdx, which indicates which color component the transform block represents. If cIdx is 0, it indicates a luma transform block; if cIdx is 1, it indicates a chroma-Cb transform block; and if cIdx is 2, it indicates a chroma-Cr transform block.
[0278] The last_significant_pos syntax structure parses a syntax element related to the position information of the last significant coefficient in scan order. Specifically, it parses syntax elements related to the x-coordinate and y-coordinate of the last significant coefficient in scan order. At this time, each coordinate value is divided into prefix information and suffix information and specified. The decoder sets a LastSignificantCoeffX variable, which is the x-coordinate of the last significant coefficient in scan order, based on the prefix information and suffix information for the x-coordinate. Similarly, the decoder sets a LastSignificantCoeffY variable, which is the y-coordinate of the last significant coefficient in scan order, based on the prefix information and suffix information for the y-coordinate. As shown in FIG. 27, the decoder sets lastScanPos, which is the scan index of the last significant coefficient in scan order, in the do{}while() structure based on LastSignificantCoeffX, LastSignificantCoeffY, and DiagScanOrder. In addition, the decoder updates numZeroOutSigCoeff and lfnstLastScanPos, which are variables used in the parsing conditions of lfnst_idx[x0][y0], which is a syntax element related to secondary transformation, based on lastScanPos.
[0279] When a secondary transform is applied to the current block, significant coefficients cannot exist at certain scan positions. The numZeroOutSigCoeff variable indicates whether non-zero transform coefficients exist at such positions. For example, assume that a maximum of 16 significant coefficients are allowed when a secondary transform is applied to the current block. For a 4x4 or 8x8 transform block, significant coefficients can exist in the [0,7] range in the scan order (allowing a maximum of 8 non-zero transform coefficients). On the other hand, for a transform block with a size other than 4x4 or 8x8, significant coefficients can exist in the [0,15] range in the scan order (allowing a maximum of 16 non-zero transform coefficients). Therefore, if the position of the last significant coefficient in the scan order (scan index) is outside the range where significant coefficients can exist, the decoder automatically recognizes that a secondary transform is not applied to the current block. The minimum size of a block to which a secondary transform can be applied is 4x4. If a transform skip is applied (transform_skip_flag[x0][y0]==1), the secondary transform is not applied. Therefore, numZeroOutSigCoeff is updated for transform blocks whose width is 4 or greater (log2TbWidth>=2), whose height is 4 or greater (log2TbHeight>=2), and whose transform skip is not applied (transform_skip_flag[x0][y0]==0). If a quadratic transform is applied, for 4x4 and 8x8 size transform blocks, non-zero transform coefficients can exist only in the scan order index range [0,7]. Therefore, if the transform block is 4x4 or 8x8, ((log2TbWidth==2||log2TbHeight==3)&&(log2TbWidth==log2TbHeight)), and lastScanPos is greater than 7 (lastScanPos>7), numZeroOutSigCoeff is incremented by 1. For transform blocks other than those of 4x4 and 8x8 size, which can be applied with quadratic transforms, non-zero transform coefficients can exist only in the scan order index range [0,15]. Therefore, if lastScanPos is greater than 15 (lastScanPos>15), numZeroOutSigCoeff is incremented by 1.
[0280] The decoder determines lfnstLastScanPos based on lastScanPos. More specifically, if the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in Equation 6 below. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in Equation 1 below. In this case, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.
[0281]
number
[0282] As mentioned above, the initialization value of lfnstLastScanPos is set to 1.
[0283] In Equation 6, cIdx represents a variable that indicates the color component of the current transform block, as described above.
[0284] According to Equation 6, if lfnstLastScanPos of a line is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 1. On the other hand, if lfnstLastScanPos of a line is 0 or lastScanPos is greater than or equal to lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 0.
[0285] In other words, if the lastScanPos of all transform blocks included in the transform unit is less than the critical value or the number of all transform blocks is 0, lfnstLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is not parsed and is set to 0 according to the lfnst_idx[x0][y0] parsing condition in Figure 26. This indicates that secondary transformation is not applied to the current block. On the other hand, if the LastScanPos of any one of the transform blocks included in the transform unit is greater than or equal to the critical value, lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vii) described in Figure 26 are satisfied (true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to determine whether a secondary transform is applied to the current block, and if a secondary transform is applied to the current block, it determines / determines the transform kernel to be used for the secondary transform.
[0286] In Equation 6, lfnstLastScanPosTh[cIdx] is a preset integer value equal to or greater than 0, and both the encoder and decoder use the same value. Alternatively, all color components may use the same threshold value. In this case, lfnstLastScanPos is set as shown in Equation 7.
[0287]
number
[0288] lfnstLastScanPosTh is a preset integer value greater than or equal to 0, and both the encoder and decoder use the same value. For example, lfnstLastScanPosTh may be 1. That is, if lastScanPos is greater than or equal to 1, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is parsed. In this case, since the threshold value (lfnstLastScanPosTh) is an integer value, if lastScanPos is greater than or equal to 1, it has the same meaning as if lastScanPos is greater than 0. Although Figure 27 describes the case where all color components have the same threshold value of 1, the present invention is not limited to this.
[0289] FIG. 28 is a diagram illustrating a residual_coding syntax structure according to an embodiment of the present invention.
[0290] Referring to Figure 28, position information of the last significant coefficient in scan order is indicated in the eye of coefficient coding (residual_coding). Therefore, the syntax structure of coefficient coding (residual_coding) does not need to include a syntax structure related to position information of the last significant coefficient in scan order. For example, the position information of the last significant coefficient in scan order is a prefix, a suffix for the x-coordinate, and a prefix, a suffix for the y-coordinate of the last significant coefficient in scan order. Looking at the syntax structure of coefficient coding (residual_coding) in Figure 28, coefficient coding (residual_coding) is performed based on LastSignificantCoeffX and LastSignificantCoeffY, which are the x-coordinate and y-coordinate of the last significant coefficient in scan order determined before coefficient coding (residual_coding).
[0291] The method for specifying a secondary transform according to the fourth embodiment does not use the numSigCoeff counter. Therefore, even if the coefficient at the (xC, yC) position is a valid coefficient (sig_coeff_flag[xC][yC]==1), numSigCoeff is not updated. In other words, the method for specifying a secondary transform according to the fourth embodiment does not use a counter for valid coefficients. Also, according to the method for specifying a secondary transform according to the fourth embodiment, since the numZeroOutSigCoeff variable is set based on lastScanPos, a counter based on sig_coeff_flag does not need to be used in coefficient coding (residual_coding).
[0292] FIG. 29 is a flowchart illustrating a video signal processing method according to an embodiment of the present invention.
[0293] The following describes a video signal processing method and apparatus based on the embodiments described with reference to FIGS.
[0294] The video signal decoding device includes a processor that performs the video signal processing method described in FIG.
[0295] First, the processor receives a bitstream containing syntax elements related to a secondary transformation of a coding unit.
[0296] The processor checks whether one or more predetermined conditions are satisfied, and if the one or more predetermined conditions are satisfied, parses syntax elements related to secondary transformation of the coding unit (S2910, S2920). On the other hand, if the one or more predetermined conditions are not satisfied, the processor does not parse syntax elements related to secondary transformation of the coding unit (S2930). In this case, the value of the syntax element related to secondary transformation is set to 0.
[0297] The syntax element related to the secondary transformation of the coding unit described in Figure 29 is lfnst_idx[x0][y0], which is a syntax element indicating whether the secondary transformation of the transform block included in the current coding unit described in Figures 5 to 28 is applied.
[0298] The processor parses syntax elements related to secondary transformation of the coding unit via step S2920, and determines whether the secondary transformation is applied to the transformation block included in the coding unit based on the parsed syntax elements (S2940).
[0299] In this case, if the secondary transform has been applied to the transform block, the processor performs a secondary inverse transform based on one or more coefficients of a first sub-block, which is one of one or more sub-blocks constituting the transform block, and identifies one or more inverse transform coefficients for the first sub-block (S2950).
[0300] The processor then performs a first inverse transform based on the one or more inverse transform coefficients obtained in step S2950, and identifies residual samples for the transform block in step S2960.
[0301] The secondary transform is a low-band non-separable transform (LFNST). The transform block is a block to which a primary transform, which is separately performed into a vertical transform and a horizontal transform, is applied. In this case, the primary inverse transform is an inverse transform of the primary transform, and the secondary inverse transform is an inverse transform of the secondary transform.
[0302] The syntax element related to the secondary transform of the coding unit includes information indicating whether the secondary transform is applied to the coding unit and information indicating the transform kernel used for the secondary transform.
[0303] The first sub-block is the first sub-block in a preset scan order, and the index of the first sub-block is 0.
[0304] A first condition among the one or more preset conditions is that an index value indicating a position of a first coefficient among the one or more coefficients of the first sub-block is greater than a preset threshold value. In this case, the first coefficient is the last significant coefficient according to a preset scan order, and the significant coefficient means a coefficient that is not 0. The preset threshold value is 0. The preset scan order is the upper right diagonal scan order described with reference to FIGS. 13 and 14.
[0305] A second condition among the one or more preset conditions is that the width and height of the transformation block are equal to or greater than 4 pixels.
[0306] A third condition among the one or more preset conditions is when a transform skip flag value included in the bitstream is not a specific value, in which case, if the transform skip flag value has a specific value, the transform skip flag indicates that the primary transform and the secondary transform are not applied to the transform block.
[0307] A fourth condition among the one or more preset conditions is that at least one of the one or more coefficients of the first sub-block is not 0 and the at least one coefficient is located at a position other than the first position in a preset scan order, where the first position in the preset scan order refers to a position where the horizontal and vertical coordinate values are (0,0) as described above, or refers to the first position in a preset scan order (e.g., upper right diagonal order).
[0308] The coding unit is composed of a plurality of coding blocks, and if at least one of the transform blocks corresponding to each of the plurality of coding blocks satisfies one or more predetermined conditions, the syntax elements related to the secondary transform are parsed.
[0309] On the other hand, if a syntax element related to a secondary transform is not parsed or is set to 0 (S2930), or if it is determined in step S2940 that the secondary transform is not applied to the transform block included in the coding unit, the processor performs a primary inverse transform based on one or more coefficients of the transform block to obtain a residual sample for the transform block (S2970).
[0310] In this case, the above-mentioned primary inverse transform and secondary inverse transform are inverse transforms of the primary transform and secondary transform, respectively.
[0311] The video signal processing method performed in the video signal decoding apparatus described with reference to FIG. 29, or a method similar thereto, is performed in the video signal encoding apparatus.
[0312] The video signal encoding device includes a processor for encoding the video signal.
[0313] In this case, the processor performs a primary transform on residual samples of a block included in a coding unit to obtain a plurality of primary transform coefficients for the block, performs a secondary transform based on one or more of the plurality of primary transform coefficients to obtain one or more secondary transform coefficients for a first sub-block that is one of the sub-blocks constituting the block, and encodes information on the one or more secondary transform coefficients and syntax elements related to the secondary transform of the coding unit to obtain a bitstream.
[0314] The secondary transform may be a low-band non-separable transform (LFNST), and the primary transform may be separated into a horizontal transform and a vertical transform.
[0315] Furthermore, the syntax element related to the secondary transform is coded if it satisfies one or more predetermined conditions. The syntax element related to the secondary transform includes information indicating whether the secondary transform is applied to the coding unit and information indicating a transform kernel used for the secondary transform. In this case, the syntax element related to the secondary transform is lfnst_idx[x0][y0], which is the syntax element described with reference to Figures 15 to 28.
[0316] The first sub-block is the first sub-block in a preset scan order, and has an index of 0.
[0317] A first condition among the one or more preset conditions is that an index value indicating a position of a first coefficient among the one or more secondary transform coefficients is greater than a preset critical value. In this case, the first coefficient is the last significant coefficient according to a preset scan order, and the significant coefficient means a coefficient that is not 0. The preset critical value is 0. The preset scan order is the upper right diagonal scan order described with reference to FIGS. 13 and 14.
[0318] A second condition among the one or more preset conditions is that the width and height of the primary transformation block are 4 pixels or more.
[0319] A third condition among the one or more preset conditions is when a transform skip flag value included in the bitstream is not a specific value, in which case, if the transform skip flag value has a specific value, the transform skip flag indicates that the primary transform and the secondary transform are not applied to the block.
[0320] A fourth condition among the one or more preset conditions is when at least one of the one or more secondary transform coefficients is not 0 and the one or more coefficients are located at positions other than the first position in a preset scan order, where the first position in the preset scan order refers to a position where the horizontal and vertical coordinate values are (0,0) as described above, or refers to the first position in a preset scan order (e.g., upper right diagonal order).
[0321] The coding unit is composed of a plurality of coding blocks, and if at least one of the (transform) blocks included in the coding unit corresponding to each of the plurality of coding blocks satisfies one or more of the predetermined conditions, the syntax element related to the secondary transform is coded.
[0322] The video signal encoding device may also include a video signal decoding processor that performs the video signal processing method described in FIG.
[0323] As described above, the bitstream includes syntax elements related to the secondary transform of the coding unit described in Figures 15 to 29. In this case, the bitstream is stored in a non-transitory computer-readable medium. Meanwhile, if one or more of the above-mentioned preset conditions are not satisfied, the video signal encoding device does not include the syntax elements related to the secondary transform in the bitstream or sets the syntax elements related to the secondary transform to 0. The bitstream is decoded by the video signal decoding device described with reference to Figure 29 or encoded by the above-mentioned video signal encoding device.
[0324] A method for encoding such a bitstream includes, for example, a process of performing a primary transform on residual samples of a block included in a coding unit to obtain a plurality of primary transform coefficients for the block, performing a secondary transform based on one or more of the plurality of primary transform coefficients to obtain one or more secondary transform coefficients for a first sub-block, which is one of the sub-blocks constituting the block, and encoding information on the one or more secondary transform coefficients and syntax elements related to the secondary transform of the coding unit.
[0325] Obtaining a coefficient as described herein means obtaining a pixel / block for the coefficient, and obtaining a residual sample means obtaining a residual signal / pixel / block for the residual sample.
[0326] The above-described embodiments of the present invention may be implemented in various ways, for example, in hardware, firmware, software, or a combination thereof.
[0327] In the case of a hardware implementation, the method according to an embodiment of the present invention may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSDPs (Digital Signal Processing Devices), PDLs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.
[0328] In the case of implementation by firmware or software, the methods according to the embodiments of the present invention may be implemented in the form of modules, procedures, or functions that perform the functions or operations described above. The software code is stored in a memory and executed by a processor. The memory may be located inside or outside the processor and exchange data with the processor through various means known in the art.
[0329] Some embodiments may also be embodied in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media are any available media that can be accessed by a computer, including both volatile and non-volatile media, and both detachable and non-detachable media. Computer-readable media also include both storage media and communication media. Computer storage media include both volatile and non-volatile media, and both detachable and non-detachable media embodied in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as a program module, or other transmission mechanism, and include any information delivery media.
[0330] The above description of the present invention is for illustrative purposes only, and those skilled in the art will understand that the present invention can be easily modified into other specific forms without changing the technical spirit or essential features of the present invention. Therefore, it should be understood that the above-described embodiments are illustrative in all respects and are not limiting. For example, each component described as a single component may be implemented in a distributed form, and components described as distributed may also be implemented in a combined form.
[0331] The scope of the present invention is indicated by the claims that follow rather than by the above detailed description, and all modifications and variations that fall within the meaning and scope of the claims and their equivalents should be interpreted as being included within the scope of the present invention. [Explanation of symbols]
[0332] 110 Conversion unit 115 Quantization section 120 Inverse quantization section 125 Inverse conversion unit 130 Filtering section 150 Prediction Department 152 Intra prediction unit 154 Inter Prediction Unit 154a Motion estimation unit 154b Motion Compensation Unit 160 Entropy Coding Unit 210 Entropy Decoding Unit 220 Inverse quantization section 225 Inverse conversion unit 230 Filtering Section 250 Prediction Department 252 Intra prediction unit 254 Inter Prediction Unit
Claims
1. 1. A video decoding method performed by an apparatus, comprising: obtaining syntax information for a secondary transformation of the coding unit; determining whether the secondary transform is applied to a transform block included in the coding unit based on the acquired syntax information; if the secondary transform is applied to the transform block, obtaining one or more inverse transform coefficients based on the secondary transform; and obtaining a residual sample for the transform block based on the one or more inverse transform coefficients, The second-order transform is a low frequency non-separable transform (LFNST), The syntax information is obtained if one or more conditions are met; A first condition among the one or more conditions is that an index of a last significant coefficient in a sub-block according to a preset scan order is greater than a predetermined value; a second condition among the one or more conditions is that intra prediction is applied to the coding unit; The video decoding method, wherein the sub-block is a sub-block having sub-block index 0 in the transform block according to a preset scan order.
2. The video decoding method of claim 1 , wherein the default value is zero.
3. The video decoding method of claim 1, wherein the preset scan order is an up-right diagonal scan order.
4. indices of the one or more inverse transform coefficients are determined based on the preset scan order; an index of a first coefficient among the one or more inverse transform coefficients is 0; 4. The video decoding method of claim 3, wherein the last significant coefficient is a non-zero coefficient.
5. The video decoding method of claim 1 , wherein the residual samples are obtained by performing an inverse transform of a linear transform based on the one or more inverse transform coefficients.
6. 1. A video encoding device, comprising: at least one processor; The at least one processor configured to generate a bitstream using an encoding method; The encoding method comprises: encoding syntax information related to a secondary transformation of the coding unit; obtaining one or more first-order transform coefficients based on residual samples for a transform block included in the coding unit; if the secondary transform is applied to the transform block according to the syntax information, obtaining one or more secondary transform coefficients using the one or more primary transform coefficients based on the secondary transform; and encoding the one or more secondary transform coefficients, the second-order transform is a low frequency non-separable transform (LFNST); The syntax information is encoded if one or more conditions are met; A first condition among the one or more conditions is that an index of a last significant coefficient in a sub-block according to a preset scan order is greater than a predetermined value; a second condition among the one or more conditions is that intra prediction is applied to the coding unit; The sub-block is a sub-block having sub-block index 0 in the transform block according to a preset scan order.
7. The video encoding device of claim 6 , wherein the default value is 0.
8. 7. The video encoding apparatus of claim 6, wherein the preset scan order is an up-right diagonal scan order.
9. indices of the one or more primary transform coefficients are determined based on the preset scan order; an index of a first coefficient among the one or more primary transform coefficients is 0; 9. The video encoding apparatus of claim 8, wherein the last significant coefficient is a non-zero coefficient.
10. The video encoding apparatus of claim 6 , wherein the one or more first-order transform coefficients are obtained by performing a first-order transform based on the residual samples.
11. A method for storing a bitstream, comprising: the bitstream is generated by an encoding method; The encoding method comprises: encoding syntax information related to a secondary transformation of the coding unit; obtaining one or more first-order transform coefficients based on residual samples for a transform block included in the coding unit; if the secondary transform is applied to the transform block according to the syntax information, obtaining one or more secondary transform coefficients using the one or more primary transform coefficients based on the secondary transform; and encoding the one or more secondary transform coefficients, the second-order transform is a low frequency non-separable transform (LFNST); The syntax information is encoded if one or more conditions are met; A first condition among the one or more conditions is that an index of a last significant coefficient in a sub-block according to a preset scan order is greater than a predetermined value; a second condition among the one or more conditions is that intra prediction is applied to the coding unit; The sub-block is a sub-block having a sub-block index of 0 in the transform block according to a preset scan order.
Citation Information
Patent Citations
Binarizing secondary transform index
US20170324643A1
Method and device for coding residual signal in video coding system
US20180288409A1
Transform method in image coding system and apparatus for same
US20200177889A1