Video signal processing method and apparatus using a second conversion

The implementation of Low Frequency Non-Separable Transform (LFNST) in video signal processing addresses inefficiencies in existing methods, enhancing coding efficiency through optimized transformation techniques.

JP7715906B2Active Publication Date: 2025-07-30SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024185656
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-07
Filing Date
2024-10-22
Publication Date
2025-07-30
Estimated Expiration
2040-06-25

AI Technical Summary

Technical Problem

Existing video signal processing methods lack efficiency in coding, particularly in handling spatial and temporal correlations, necessitating improved techniques for enhanced compression encoding.

Method used

Implementing a secondary conversion method using Low Frequency Non-Separable Transform (LFNST) on video signal blocks, with specific conditions for applying transformations based on syntax elements and coefficient positions, to enhance coding efficiency.

Benefits of technology

The method significantly improves coding efficiency by optimizing transformation processes, leading to more effective video signal compression and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715906000008
    Figure 0007715906000008
  • Figure 0007715906000009
    Figure 0007715906000009
  • Figure 0007715906000010
    Figure 0007715906000010
Patent Text Reader

Abstract

To provide a video signal processing method using a secondary transform.SOLUTION: A video signal decoding apparatus comprises a processor configured to: parse a syntax element related to a secondary transform of a coding unit from a bitstream of a video signal when one or more preset conditions are satisfied; check whether the secondary transform is applied to a transform block included in the coding unit based on the parsed syntax element; obtain one or more inverse transform coefficients for a first sub-block by performing an inverse secondary transform based on one or more coefficients of the first sub-block, which is one of one or more sub-blocks constituting the transform block, when the secondary transform is applied to the transform block; and obtain a residual sample for the transform block by performing an inverse primary transform based on the one or more inverse transform coefficients.SELECTED DRAWING: Figure 29
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video signal processing method and apparatus, and more particularly, to a video signal processing method and apparatus for encoding or decoding a video signal.

Background Art

[0002] Compression encoding means a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc., and in particular, the technique of performing compression encoding on video is called video compression. Compression encoding of a video signal is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, due to the recent development of various media and data transmission media, a more efficient video signal processing method and apparatus are required.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present invention is to increase the coding efficiency of a video signal.

[0004] The present invention has an object to increase the coding efficiency through secondary conversion.

Means for Solving the Problems

[0005] This specification provides a video signal processing method using secondary conversion.

[0006] Specifically, in a video signal decoding apparatus, it includes a processor. If the processor satisfies one or more preset conditions, it parses syntax elements related to the second transformation of a coding unit from a bitstream of a video signal, and based on the parsed syntax elements, it checks whether the second transformation is to be applied to a transformation block included in the coding unit. If the second transformation is applied to the transformation block, it performs an inverse second transformation based on one or more coefficients of a first sub-block, which is one of one or more sub-blocks constituting the transformation block, to obtain one or more inverse transformation coefficients for the first sub-block, and performs a first inverse transformation based on the one or more inverse transformation coefficients to obtain residual samples for the transformation block. The second transformation is a Low Frequency Non-Separable Transform (LFNST), and the transformation block is a block to which a separable first transformation, which can be performed separately for vertical and horizontal transformations, is applied. Among the one or more preset conditions, the first condition is that an index value indicating the position of a first coefficient among the one or more coefficients of the first sub-block is greater than a preset threshold value.

[0007] Also, in this specification, the syntax elements include information indicating whether the second transformation is to be applied to the coding unit and information indicating a transformation kernel used for the second transformation.

[0008] Also, in this specification, the first coefficient is the last significant coefficient according to a preset scan order, and the significant coefficient is a non-zero coefficient.

[0009] Also, in this specification, the first sub-block is the first sub-block according to a preset scan order.

[0010] Also, in this specification, the second condition among the one or more preset conditions is that the width and height of the conversion block are 4 pixels or more.

[0011] Also, in this specification, the preset threshold value is 0.

[0012] Also, in this specification, the preset scan order is the up - right diagonal scan order.

[0013] Also, in this specification, the third condition among the one or more preset conditions is when the conversion skip flag value included in the bitstream is not a specific value. If the conversion skip flag value has the specific value, the conversion skip flag supports that the first - order conversion and the second - order conversion are not applied to the conversion block.

[0014] Also, in this specification, the fourth condition among the one or more preset conditions is that at least one of the one or more coefficients of the first sub - block is not 0, and the at least one or more coefficients exist at a position except for the first position according to the preset scan order.

[0015] Also, in this specification, the coding unit is composed of a plurality of coding blocks. If at least any one of the conversion blocks corresponding to the plurality of coding blocks satisfies the one or more preset conditions, the syntax elements related to the second - order conversion are parsed.

[0016] Also, in this specification, in a video signal decoding apparatus, it includes a processor, and the processor performs a first-order transformation on the residual samples of the blocks included in the coding unit to obtain a plurality of first-order transformation coefficients for the blocks, and performs a second-order transformation based on any one or more of the coefficients of the plurality of first-order transformations to obtain one or more second-order transformation coefficients for a first sub-block that is one of the sub-blocks constituting the block, and encodes information regarding the one or more second-order transformation coefficients and syntax elements regarding the second-order transformation of the coding unit to obtain a bitstream, wherein the second-order transformation is a low-frequency non-separable transform (LFNST), the first-order transformation can be separately performed for vertical transformation and horizontal transformation, the syntax elements regarding the second-order transformation of the coding unit are encoded if one or more preset conditions are satisfied, and a first condition among the one or more preset conditions is that an index value indicating the position of a first coefficient among the one or more second-order transformation coefficients is greater than a preset threshold value.

[0017] Also, in this specification, the syntax elements include information indicating whether the second-order transformation is applied to the coding unit and information indicating a transform kernel used for the second-order transformation.

[0018] Also, in this specification, the first coefficient is the last valid coefficient in a preset scan order, and the valid coefficient is a coefficient that is not 0.

[0019] Also, in this specification, the first sub-block is the first sub-block in a preset scan order.

[0020] Also, in this specification, a second condition among the one or more preset conditions is that the width and height of the first-order transformation block are 4 pixels or more.

[0021] In addition, in this specification, it is characterized in that the preset critical value is 0.

[0022] In addition, in this specification, it is characterized in that the preset scan order is the upper right diagonal scan order.

[0023] In addition, in this specification, the third condition among the one or more preset conditions is a case where the conversion skip flag value included in the bit stream is not a specific value. If the conversion skip flag value has the specific value, the conversion skip flag supports that the first conversion and the second conversion are not applied to the block.

[0024] In addition, in this specification, the fourth condition among the one or more preset conditions is a case where at least one of the one or more second conversion coefficients is not 0 and the one or more coefficients exist except at the first position according to the preset scan order.

[0025] Also, in this specification, in a non-transitory computer-readable medium storing a bitstream, the bitstream includes: a step of performing a primary transformation on residual samples of blocks included in a coding unit to obtain a plurality of primary transformation coefficients for the blocks; a step of performing a secondary transformation based on any one or more of the plurality of primary transformation coefficients to obtain one or more secondary transformation coefficients for a first sub-block that is one of the sub-blocks constituting the block; and a step of encoding information regarding the one or more secondary transformation coefficients and syntax elements regarding the secondary transformation of the coding unit. However, the secondary transformation is a low-frequency non-separable transform (LFNST), the primary transformation can be separately performed as a vertical transformation and a horizontal transformation, the syntax elements regarding the secondary transformation are encoded if they satisfy one or more preset conditions, and a first condition among the one or more preset conditions is that an index value indicating the position of a first coefficient among the one or more secondary transformation coefficients is greater than a preset threshold value.

Advantages of the Invention

[0026] One embodiment of the present invention provides a video signal processing method using a secondary transformation and an apparatus therefor.

Brief Description of the Drawings

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Best Mode for Carrying Out the Invention

[0028] The terms used in this specification are selected as generally widely used terms as much as possible while considering the functions in the present invention, but this may vary depending on the intentions, conventions of those skilled in the art, or the emergence of new technologies. Also, in certain cases, there are terms arbitrarily selected by the applicant, and in this case, the meaning is described in the part of the embodiment for carrying out the corresponding invention. Therefore, it is clarified that the terms used in this specification should be interpreted based not only on the name of the terms but also on the substantial meaning of the terms and the content throughout this specification.

[0029] In this specification, some terms are interpreted as follows. Coding may be interpreted as encoding or decoding in some cases. In this specification, a device that performs encoding (encoding) of a video signal to generate a bitstream of the video signal is referred to as an encoding device or an encoder, and a device that performs decoding (decoding) of the video signal bitstream to restore the video signal is referred to as a decoding device or a decoder. Also, in this specification, a video signal processing device is used as a term of a concept that includes both an encoder and a decoder. Information is a term that includes all of values, parameters, coefficients, elements, etc., and may be interpreted to have different meanings in some cases, so the present invention is not limited thereto. "Unit" is used to mean a basic unit of video processing or a specific position of a picture, and refers to an image region including at least one of a luminance component and a chrominance component. Also, "block" refers to an image region including a specific component among a luminance component and chrominance components (that is, Cb and Cr). However, depending on the embodiment, terms such as "unit", "block", "partition", and "region" may be used in combination with each other. Also, in this specification, a unit is used as a concept that includes all of a coding unit, a prediction unit, and a transform unit. A picture refers to a field or a frame, and depending on the embodiment, the terms are used interchangeably with each other.

[0030] FIG. 1 is a schematic block diagram of a video signal encoding device 100 according to an embodiment of the present invention. Referring to FIG. 1, the encoding device 100 of this specification includes a transform unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transform unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.

[0031] The conversion unit 110 converts the residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit 150, to obtain conversion coefficient values. For example, a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Wavelet Transform may be used. The Discrete Cosine Transform and the Discrete Sine Transform perform the conversion by dividing the input picture signal into blocks. In the conversion, the coding efficiency may vary depending on the distribution and characteristics of the values within the conversion region. The quantization unit 115 quantizes the values of the conversion coefficients output within the conversion unit 110.

[0032] To improve the coding efficiency, instead of coding the picture signal as it is, a method is used in which the picture is predicted using a region that has been pre-coded via the prediction unit 150, and the residual value between the original picture and the predicted picture is added to the predicted picture to obtain a restored picture. To prevent a mismatch between the encoder and the decoder, information that can also be used by the decoder should be used when performing prediction in the encoder. For this purpose, the encoder performs a process of further restoring the encoded current block. The inverse quantization unit 120 inverse quantizes the conversion coefficient values, and the inverse conversion unit 125 restores the residual values using the inverse quantized conversion coefficient values. On the other hand, the filtering unit 130 performs filtering operations for improving the quality of the restored picture and enhancing the coding efficiency. For example, it may include a deblocking filter, a Sample Adpative Offset (SAO), and an adaptive loop filter. The picture that has undergone filtering is stored in a Decoded Picture Buffer (DPB) 156 to be output or used as a reference picture.

[0033] In order to improve coding efficiency, instead of directly coding the picture signal, a method is used where the picture is predicted by utilizing a region that has been previously coded via the prediction unit 150, and a residual value between the original picture and the predicted picture is added to the predicted picture to obtain a restored picture. In the intra prediction unit 152, intra prediction is performed within the current picture, and in the inter prediction unit 154, the current picture is predicted by using the reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from the restored region within the current picture and transmits the intra coding information to the entropy coding unit 160. The inter prediction unit 154 is configured to include a motion estimation unit 154a and a motion compensation unit 154b again. In the motion estimation unit 154a, a motion vector value of the current region is obtained by referring to the restored specific region. In the motion estimation unit 154a, position information of the reference region (such as the reference frame, motion vector, etc.) is transmitted to the entropy coding unit 160 so as to be included in the bitstream. Using the motion vector value transmitted from the motion estimation unit 154a, the motion compensation unit 154b performs inter motion compensation.

[0034] The prediction unit 150 includes an intra prediction unit 152 and an inter prediction unit 154. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 performs inter prediction for predicting the current picture by using the reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from the restored samples within the current picture and transmits the intra-coded information to the entropy coding unit 160. The intra-coded information includes at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra-coded information includes information regarding the reference samples. The inter prediction unit 154 includes a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value of the current region by referring to a specific region of the restored reference signal picture. The motion estimation unit 154a transmits a set of motion information (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation by using the motion vector value transmitted from the motion compensation unit 154a. The inter prediction unit 154 transmits inter-coded information including motion information for the reference region to the entropy coding unit 160.

[0035] According to a further embodiment, the prediction unit 150 includes an intra block copy (BC) prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from the restored samples within the current picture and transmits the intra BC-coded information to the entropy coding unit 160. The intra BC prediction unit obtains a block vector value indicating a reference region used for predicting the current region by referring to a specific region within the current picture. The intra BC prediction unit performs intra BC prediction by using the obtained block vector value. The intra BC prediction unit transmits the intra BC-coded information to the entropy coding unit 160. The intra BC prediction unit includes block vector information.

[0036] If the above-described picture prediction is performed, the conversion unit 110 converts the residual value between the original picture and the predicted picture to obtain a conversion coefficient value. At this time, the conversion is performed in units of specific blocks within the picture, and the size of the specific block varies within a preset range. The quantization unit 115 quantizes the value of the conversion coefficient generated by the conversion unit 110 and transmits it to the entropy coding unit 160.

[0037] The entropy coding unit 160 performs entropy coding on information indicating the quantized conversion coefficient, intra-coding information, inter-coding information, etc. to generate a video signal bitstream. In the entropy coding unit 160, a variable length coding (VLC) method, an arithmetic coding method, etc. are used. The variable length coding (VLC) method converts the input symbol into a continuous codeword, and the length of the codeword is variable. For example, frequently occurring symbols are represented by short codewords, and infrequently occurring symbols are represented by long codewords. As the variable length coding method, a context-based adaptive variable length coding (CAVLC) method is used. Arithmetic coding converts a continuous data symbol into a single prime number, and arithmetic coding obtains the optimal prime number bits required to represent each symbol. As the arithmetic coding, a context-based adaptive binary arithmetic coding (CABAC) method is used. For example, the entropy coding unit 160 binaryizes the information indicating the quantized conversion coefficient. Further, the entropy coding unit 160 performs arithmetic coding on the binaryized information to generate a bitstream.

[0038] The generated bitstream is encapsulated in units of NAL (Network Abstraction Layer) units. An NAL unit contains an integral number of encoded coding tree units. In order to decode the bitstream in a video decoder, first the bitstream should be separated into NAL units, and then each of the separated NAL units should be decoded. On the other hand, information necessary for decoding the video signal bitstream is transmitted via upper-level sets of RBSP (Raw Byte Sequence Payload) such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS).

[0039] On the other hand, the block diagram of FIG. 1 shows an encoding apparatus 100 according to an embodiment of the present invention, and the separately shown blocks logically distinguish the elements of the encoding apparatus 100. Therefore, the elements of the encoding apparatus 100 described above are attached to one chip or a plurality of chips according to the design of the device. According to one embodiment, the operations of each element of the encoding apparatus 100 described above are performed by a processor (not shown).

[0040] FIG. 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to FIG. 2, the decoding apparatus 200 in this specification includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.

[0041] The entropy decoding unit 210 entropy-decodes the video signal bitstream and extracts conversion coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit 210 obtains a binary code for the conversion coefficient information of a specific region from the video signal bitstream. Also, the entropy decoding unit 210 inverse-binary-decodes the binary code to obtain the quantized conversion coefficients. The inverse quantization unit 220 inverse-quantizes the quantized conversion coefficients, and the inverse conversion unit 225 restores the residual value using the inverse-quantized conversion coefficients. The video signal processing apparatus 200 restores the original pixel value by adding the residual value obtained from the inverse conversion unit 225 to the prediction value obtained from the prediction unit 250.

[0042] On the other hand, the filtering unit 230 performs filtering on the picture to improve the picture quality. This includes a deblocking filter for reducing the block distortion phenomenon and / or an adaptive loop filter for removing the distortion of the entire picture. The picture that has undergone filtering is output or stored in the decoded picture buffer (DPB) 256 for use as a reference picture for the next picture.

[0043] The prediction unit 250 includes an intra prediction unit 252 and an inter prediction unit 254. The prediction unit 250 generates a prediction picture by utilizing the encoded type decoded via the entropy decoding unit 210 described above, the conversion coefficients for each region, the intra / inter encoding information, etc. To restore the current block for which decoding is performed, the current picture including the current block or the decoded regions of other pictures is used. A picture (or tile / slice) that uses only the current picture for restoration, that is, a picture (or tile / slice) that performs intra prediction or intra BC prediction is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that performs both intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). A picture (or tile / slice) that uses at most one motion vector and a reference picture index to predict the sample values of each block in an inter picture (or tile / slice) is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and a reference picture index is called a Bi-predictive picture or B picture (or tile / slice). That is, a P picture (or tile / slice) uses at most one set of motion information to predict each block, and a B picture (or tile / slice) uses at most two sets of motion information to predict each block. Here, a set of motion information includes one or more motion vectors and one reference picture index.

[0044] The intra prediction unit 252 generates a prediction block by using the intra-encoded information and the restored samples within the current picture. As described above, the intra-encoded information includes at least one of an intra prediction mode, an MPM (MOST Probable Mode) flag, and an MPM index. The intra prediction unit 252 predicts the sample values of the current block by using the restored samples located on the left side and / or the upper side of the current block as reference samples. In the present disclosure, the restored samples, the reference samples, and the samples of the current block represent pixels. Also, the sample value represents a pixel value.

[0045] In one embodiment, the reference sample is a sample included in a peripheral block of the current block. For example, the reference sample is a sample adjacent to the left boundary of the current block and / or a sample adjacent to the upper boundary of the current block. Also, the reference sample is a sample located on a line within a preset distance from the left boundary of the current block and / or a sample located on a line within a preset distance from the upper boundary of the current block among the samples of the peripheral block of the current block. At this time, the peripheral block of the current block includes at least one of a left (L) block adjacent to the current block, an upper (A) block, a below left (BL) block, an above right (AR) block, or an above left (AL) block.

[0046] The inter prediction unit 254 generates a prediction block by using the reference pictures and inter-coding information stored in the decoded picture buffer 256. The inter-coding information includes a set of motion information (such as a reference picture index, a motion vector, etc.) of the current block with respect to the reference block. There are L0 prediction, L1 prediction, and bi-prediction in inter prediction. The L0 prediction is a prediction using one reference picture included in the L0 picture list, and the L1 prediction means a prediction using one reference picture included in the L1 picture list. For this purpose, a set of motion information (for example, a motion vector and a reference picture index) is required. In the bi-prediction method, a maximum of two reference areas are used, and these two reference areas may exist in the same reference picture or in different pictures respectively. That is, in the bi-prediction method, a maximum of two sets of motion information (for example, a motion vector and a reference picture index) are used, and the two motion vectors may correspond to the same reference picture index or to different reference picture indexes. At this time, the reference picture is displayed (or output) either before or after the current picture in time. According to one embodiment, in the bi-prediction method, the two reference areas used are areas selected from the L0 picture list and the L1 picture list respectively.

[0047] The inter prediction unit 254 acquires the current reference block by using the motion vector and the reference picture index. The reference block exists in the reference picture corresponding to the reference picture index. Also, the sample value of the block specified by the motion vector or its interpolated value is used as the predictor of the current block. For motion prediction with pixel accuracy in sub-pel units, for example, an 8-tap interpolation filter is used for the luminance signal and a 4-tap interpolation filter is used for the chrominance signal. However, the interpolation filter for motion prediction in sub-pel units is not limited to this. Thus, the inter prediction unit 254 performs motion compensation for predicting the texture of the current unit from the previously restored picture. At this time, the inter prediction unit uses a set of motion information.

[0048] According to a further embodiment, the prediction unit 250 includes an intra BC prediction unit (not shown). The intra BC prediction unit restores the current region by referring to a specific region including the restored samples in the current picture. The intra BC prediction unit acquires the intra BC coding information for the current region from the entropy decoding unit 210. The intra BC prediction unit acquires the block vector value of the current region that indicates a specific region in the current picture. The intra BC prediction unit performs intra BC prediction by using the acquired block vector value. The intra BC prediction unit includes block vector information.

[0049] A restored video picture is generated by adding the prediction value output from the intra prediction unit 252 or the inter prediction unit 254 and the residual value output from the inverse conversion unit 225. That is, the video signal decoding device 200 restores the current block by using the prediction block generated from the prediction unit 250 and the residual value acquired from the inverse conversion unit 225.

[0050] On the one hand, the block diagram of FIG. 2 shows a decoding device 200 according to an embodiment of the present invention, and the blocks shown separately logically distinguish the elements of the decoding device 200. Therefore, the elements of the decoding device 200 described above are attached to one chip or multiple chips according to the design of the device. According to an embodiment, the operations of each element of the decoding device 200 described above are performed by a processor (not shown).

[0051] FIG. 3 shows an embodiment in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. In the coding process of a video signal, a picture is divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit consists of an N×N block of luminance samples and two blocks of corresponding chrominance samples. A Coding Tree Unit is divided into a plurality of Coding Units. The Coding Tree Unit may become a leaf node without being divided. In this case, the Coding Tree Unit itself can become a Coding Unit. A Coding Unit refers to a basic unit for processing a picture in the above-described video signal processing process, that is, processes such as intra / inter prediction, transformation, quantization, and / or entropy coding. Within one picture, the size and pattern of the Coding Unit are not constant. The Coding Unit has a square or rectangular pattern. A rectangular Coding Unit (or rectangular block) includes a vertical Coding Unit (or vertical block) and a horizontal Coding Unit (or horizontal block). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Also, in this specification, a non-square block refers to a rectangular block, but the present invention is not limited thereto.

[0052] Referring to FIG. 3, the coding tree unit is first divided into a Quad Tree (QT) structure. That is, in the quad tree structure, one node having a size of 2N×2N is divided into four nodes having a size of N×N. In this specification, the quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and it is not necessary for all nodes to be divided to the same depth.

[0053] On the other hand, the leaf node of the above-mentioned quad tree is further divided into a Multi-Type Tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node is divided into a binary (binary) or ternary (ternary) tree structure of horizontal or vertical division. That is, there are four division structures in the multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present invention, in each of the above tree structures, both the width and height of the node have a value that is a power of 2. For example, in a binary Tree (BT) structure, a node with a size of 2N×2N is divided into two N×2N nodes by vertical binary division and into two 2N×N nodes by horizontal binary division. Also, in a ternary Tree (TT) structure, a node with a size of 2N×2N is divided into (N / 2)×2N, N×2N, and (N / 2)×2N nodes by vertical ternary division and into 2N×(N / 2), 2N×N, and 2N×(N / 2) nodes by horizontal ternary division. Such multi-type tree division is performed recursively.

[0054] The leaf nodes of a multi-type tree can be coding units. If a coding unit is not larger than the maximum transform length, the corresponding coding unit can be used as a prediction and / or transformation unit without further splitting. As an example, if the width or height of the current coding unit is larger than the maximum transform length, the current coding unit is split into multiple transform units without explicit signaling regarding the split. On the other hand, in the quad tree and multi-type tree described above, at least one of the following parameters is predefined or transmitted via a higher-level set of RBSPs such as PPS, SPS, VPS, etc. 1) CTU size: the size of the root node of the quad tree, 2) minimum QT size (MinQtSize): the size of the smallest allowable QT leaf node, 3) maximum BT size (MaxBtSize): the size of the largest allowable BT root node, 4) maximum TT size (MaxTtSize): the size of the largest allowable TT root node, 5) maximum MTT depth (MaxMttDepth): the maximum allowable depth of MTT splitting from the leaf node of the QT, 6) minimum BT size (MinBtSize): the size of the smallest allowable BT leaf node, 7) minimum TT size: the size of the smallest allowable TT leaf node.

[0055] Figure 4 is a diagram showing an example of a method for signaling the splitting of a quad tree and a multi-type tree. To signal the splitting of the described quad tree and multi-type tree, a preset flag is used. Referring to Figure 4, at least one of the flag "split_cu_flag" indicating whether a node can be split, the flag "split_qt_flag" indicating whether a quad tree node can be split, the flag "mtt_split_cu_vertical_flag" indicating the splitting direction of a multi-type tree node, or the flag "mtt_split_binarycu_flag" indicating the splitting pattern of a multi-type tree node is used.

[0056] According to an embodiment of the present invention, a flag "split_cu_flag" that indicates whether the current node can be split is signaled first. If the value of "split_cu_flag" is 0, it indicates that the current node is not split, and the current node becomes a coding unit. If the current node is a coding tree unit, the coding tree unit includes one coding unit that is not split. If the current node is a quad tree node "QT node", the current node is a leaf node of the quad tree node "QT leaf node" and becomes a coding unit. If the current node is a multi-type tree node "MTT node", the current node is a leaf node of the multi-type tree "MTT leaf node" and becomes a coding unit.

[0057] If the value of 「split_cu_flag」 is 1, the current node is split into a node of a quad tree or a multi-type tree according to the value of 「split_qt_flag」. The coding tree unit is the root node of the quad tree and is preferentially split into a quad tree structure. In the quad tree structure, 「split_qt_flag」 is signaled for each node 「QT node」. If the value of 「split_qt_flag」 is 1, the corresponding node is split into four square nodes. If the value of 「qt_split_flag」 is 0, the corresponding node becomes a leaf node 「QT leaf node」 of the quad tree, and the corresponding node is split into a multi-type node. According to the embodiments of the present invention, the quad tree split may be restricted according to the type of the current node. If the current node is a coding tree unit (the root node of the quad tree) or a quad tree node, the quad tree split is allowed. If the current node is a multi-type tree unit, the quad tree split is not allowed. Each quad tree leaf node 「QT leaf node」 is further split into a multi-type tree structure. As described above, if 「split_qt_flag」 is 0, the current node is split into a multi-type node. To indicate the split direction and split pattern, 「mtt_split_cu_vertical_flag」 and 「mtt_split_cu_binary_flag」 are signaled. If the value of 「mtt_split_cu_vertical_flag」 is 1, the vertical split of the node 「MTT node」 is indicated. If the value of 「mtt_split_cu_vertical_flag」 is 0, the horizontal split of the node 「MTT node」 is indicated. Also, if the value of 「mtt_split_cu_binary_flag」 is 1, the node 「MTT node」 is split into two rectangular nodes. If the value of 「mtt_split_cu_binary_flag」 is 0, the node 「MTT node」 is split into three rectangular nodes.

[0058] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes of the coding unit tree). The basic unit for performing such prediction is hereinafter referred to as a prediction unit or a prediction block.

[0059] Hereinafter, the term "unit" used in this specification is used as a term to replace the above-mentioned prediction unit, which is the basic unit for performing prediction. However, the present invention is not limited thereto, and in a broader sense, it is understood as a concept including the above-mentioned coding unit.

[0060] FIGS. 5 and 6 are diagrams showing in more detail an intra prediction method according to an embodiment of the present invention. As described above, the intra prediction unit uses the restored samples located on the left side and / or the upper side of the current block as reference samples to predict the sample values of the current block.

[0061] First, FIG. 5 shows an example of reference samples used for predicting the current block in the intra prediction mode. According to one embodiment, the reference samples are samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary of the current block. As shown in FIG. 5, if the size of the current block is W×H and the samples of a single reference line adjacent to the current block are used for intra prediction, the reference samples are set using a maximum of 2W+2H+1 peripheral samples located on the left side and / or the upper side of the current block.

[0062] Also, if at least some of the samples used as reference samples have not yet been restored, the intra prediction unit performs a reference sample padding process to obtain reference samples. Further, the intra prediction unit performs a reference sample filtering process to reduce the error of intra prediction. That is, filtering is performed on the reference samples obtained by the peripheral samples and / or the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts the samples of the current block using the reference samples thus obtained. The intra prediction unit predicts the samples of the current block using the unfiltered reference samples or the filtered reference samples. In the present disclosure, the peripheral samples include samples on at least one reference line. For example, the peripheral samples may include adjacent samples on a line adjacent to the boundary of the current block.

[0063] Next, FIG. 6 is a diagram showing an example of a prediction mode used for intra prediction. For intra prediction, intra prediction mode information indicating the intra prediction direction is signaled. The intra prediction mode indicates any one of a plurality of intra prediction modes constituting the intra prediction mode set. If the current block is an intra prediction block, the decoder receives the intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.

[0064] According to an embodiment of the present invention, the intra prediction mode set includes all intra prediction modes used for intra prediction (for example, a total of 67 intra prediction modes). More specifically, the intra prediction mode set includes a planar mode, a DC mode, and a plurality of (for example, 65) angular modes (i.e., direction modes). Each intra prediction mode is indicated via a preset index (i.e., an intra prediction mode index). For example, as shown in FIG. 6, the intra prediction mode index 0 indicates the planar mode, and the intra prediction mode index 1 indicates the DC mode. Also, the intra prediction mode indices 2 to 66 each indicate different angular modes. The angular modes each indicate different angles within a preset angle range. For example, the angular modes indicate angles within an angle range (i.e., a first angle range) between 45 degrees and -135 degrees in the clockwise direction. The angular modes are defined based on 12 holding directions. At this time, the intra prediction mode index 2 indicates the Horizontal Diagonal (HDIA) mode, the intra prediction mode index 18 indicates the Horizontal (HOR) mode, the intra prediction mode index 34 indicates the Diagonal (DIA) mode, the intra prediction mode index 50 indicates the Vertical (VER) mode, and the intra prediction mode index 66 indicates the Vertical Diagonal (VDIA) mode.

[0065] On the one hand, the preset angle ranges are set to be different from each other according to the pattern of the current block. For example, if the current block is a rectangular block, a wide-angle mode that indicates an angle exceeding 45 degrees clockwise or less than -135 degrees is further used. If the current block is a horizontal block, the angle mode indicates an angle within an angle range (i.e., the second angle range) between (45 + offset1) degrees and (-135 + offset1) degrees clockwise. At this time, angle modes 67 to 76 that deviate from the first angle range are further used. Also, if the current block is a vertical block, the angle mode indicates an angle within an angle range (i.e., the third angle range) between (45 - offset2) degrees and (-135 - offset2) degrees clockwise. At this time, angle modes -10 to -1 that deviate from the first angle range are further used. According to an embodiment of the present invention, the values of offset1 and offset2 are determined to be different from each other according to the ratio between the width and height of the rectangular block. Also, offset1 and offset2 are positive numbers.

[0066] According to a further embodiment of the present invention, the plurality of angle modes constituting the intra prediction mode set include a basic angle mode and an extended angle mode. At this time, the extended angle mode is determined based on the basic angle mode.

[0067] According to one embodiment, the basic angle mode is a mode corresponding to the angle used in the intra prediction of the conventional HEVC (High Efficiency Video Coding) standard, and the extended angle mode is a mode corresponding to the angle newly added in the intra prediction of the next-generation video codec standard. More specifically, the basic angle mode is an angle mode corresponding to any one of the intra prediction modes {2, 4, 6,..., 66}, and the extended angle mode is an angle mode corresponding to any one of the intra prediction modes {3, 5, 6,..., 65}. That is, the extended angle mode is an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode is determined based on the angle indicated by the basic angle mode.

[0068] According to another embodiment, the basic angle mode is a mode corresponding to an angle within a preset first angle range, and the extended angle mode is a wide-angle mode deviating from the first angle range. That is, the basic angle mode is an angle mode corresponding to any one of the intra prediction modes {2, 3, 4, …, 66}, and the extended angle mode is an angle mode corresponding to any one of the intra prediction modes {-10, -9, …, -1} and {67, 68, …, 76}. The angle indicated by the extended angle mode is determined to be the angle on the opposite side of the angle indicated by the corresponding basic angle mode. Therefore, the angle indicated by the extended angle mode is determined based on the angle indicated by the basic angle mode. On the other hand, the number of extended angle modes is not limited to this, and further extended angles are defined according to the size and / or pattern of the current block. For example, the extended angle mode may be defined as an angle mode corresponding to any one of the intra prediction modes {-14, -13, …, -1} and {67, 68, …, 80}. On the other hand, the total number of intra prediction modes included in the intra prediction mode set varies according to the configurations of the basic angle mode and the extended angle mode described above.

[0069] In the above embodiment, the interval between the extended angle modes is set based on the interval between the corresponding basic angle modes. For example, the interval between the extended angle modes {3, 5, 7, …, 65} is determined based on the interval between the corresponding basic angle modes {2, 4, 6, …, 66}. Also, the interval between the extended angle modes {-10, -9, …, -1} is determined based on the interval between the corresponding opposite basic angle modes {56, 57, …, 65}, and the interval between the extended angle modes {67, 68, …, 76} is determined based on the interval between the corresponding opposite basic angle modes {3, 4, …, 12}. The angular interval between the extended angle modes is set in the same way as the angular interval between the corresponding basic angle modes. Also, in the intra prediction mode set, the number of extended angle modes is set to be less than or equal to the number of basic angle modes.

[0070] According to an embodiment of the present invention, the extended angle mode is signaled based on the basic angle mode. For example, the wide angle mode (i.e., the extended angle mode) replaces at least one angle mode (i.e., the basic angle mode) within the first angle range. The basic angle mode to be replaced is the angle mode corresponding to the opposite side of the wide angle mode. That is, the basic angle mode to be replaced corresponds to the angle in the opposite direction of the angle indicated by the wide angle mode, or the angle mode corresponding to the angle with a difference of a preset offset index from the angle in the opposite direction. According to an embodiment of the present invention, the preset offset index is 1. The intra prediction mode index corresponding to the basic angle mode to be replaced is further mapped to the wide angle mode to signal the corresponding wide angle mode. For example, the wide angle modes {-10, -9, … -1} are respectively signaled by the intra prediction mode indices {57, 58, …, 66}, and the wide angle modes {67, 68, … 76} are respectively signaled by the intra prediction mode indices {2, 3, …, 11}. By making the intra prediction mode index for the basic angle mode signal the extended angle mode in this way, even if the configurations of the angle modes used for intra prediction of each block are different from each other, the same set of intra prediction mode indices can be used for signaling the intra prediction mode. Therefore, the signaling overhead due to the change in the configuration of the intra prediction mode is minimized.

[0071] On the other hand, whether the extended angle mode can be used is determined based on at least one of the pattern and size of the current block. According to one embodiment, if the size of the current block is larger than a preset size, the extended angle mode is used for intra prediction of the current block, otherwise only the basic angle mode is used for intra prediction of the current block. According to other embodiments, if the current block is a non-square block, the extended angle mode is used for intra prediction of the current block, and if the current block is a square block, only the basic angle mode is used for intra prediction of the current block.

[0072] On the one hand, in order to improve the coding efficiency, instead of directly coding the above-described residual signal, a method is used in which the conversion coefficient values obtained by converting the residual signal are quantized and the quantized conversion coefficients are coded. As described above, the conversion unit converts the residual signal to obtain the conversion coefficient values. At this time, the residual signal of a specific block may be dispersed over the entire area of the current block. Thereby, the energy can be concentrated in the low-frequency region through the frequency-domain conversion of the residual signal, and the coding efficiency can be improved. Hereinafter, the method by which the residual signal is converted or inversely converted will be described in detail.

[0073] FIG. 7 is a diagram showing in detail how the encoder converts the residual signal. As described above, the residual signal in the spatial domain is converted into the frequency domain. The encoder converts the obtained residual signal to obtain the conversion coefficients. First, the encoder obtains at least one residual block including the residual signal for the current block. The residual block is either the current block or one of the blocks divided from the current block. In the present disclosure, the residual block is referred to as a residual array or a residual matrix including the residual samples of the current block. Also, in the present disclosure, the residual block indicates a block having the same size as the size of the conversion unit or the conversion block.

[0074] Next, the encoder uses a conversion kernel to convert the residual block. The conversion kernel used for the conversion of the residual block is a conversion kernel having separable characteristics of vertical conversion and horizontal conversion. In this case, the conversion of the residual block is performed separately for vertical conversion and horizontal conversion. For example, the encoder applies a conversion kernel in the vertical direction of the residual block to perform vertical conversion. Also, the encoder applies a conversion kernel in the horizontal direction of the residual block to perform horizontal conversion. In the present disclosure, the conversion kernel is used as a term referring to a parameter set used for the conversion of a residual signal such as a conversion matrix, a conversion array, a conversion function, and a conversion. According to one embodiment, the conversion kernel is any one of a plurality of available kernels. Also, conversion kernels based on different conversion types may be used for each of vertical conversion and horizontal conversion.

[0075] The encoder transmits the converted conversion block from the residual block to the quantization unit for quantization. In this case, the conversion block includes a plurality of conversion coefficients. Specifically, the conversion block consists of a plurality of conversion coefficients arranged in a two-dimensional array. The size of the conversion block is the same as any one of the current block or a block divided from the current block, which is the same as the residual block. The conversion coefficients transmitted to the quantization unit are represented by quantized values.

[0076] Also, the encoder performs a further transformation before the transformation coefficient is quantized. As shown in FIG. 7, the above-described transformation method is referred to as a primary transform, and the further transformation is referred to as a secondary transform. The secondary transform is selective for each residual block. According to one embodiment, the encoder can perform a secondary transform on a region where it is difficult to concentrate energy in the low-frequency region only by the primary transform, thereby improving the coding efficiency. For example, a secondary transform may be added to a block in which the residual value is large in a direction other than the horizontal or vertical direction of the residual block. The residual value of an intra-predicted block has a higher probability of changing in a direction other than the horizontal or vertical direction compared to the residual value of an inter-predicted block. Accordingly, the encoder further performs a secondary transform on the residual signal of the intra-predicted block. Also, the encoder may omit the secondary transform on the residual signal of the inter-predicted block.

[0077] As another example, whether to perform a secondary transform is determined according to the size of the current block or the residual block. Also, transform kernels having different sizes are used according to the size of the current block or the residual block. For example, an 8×8 secondary transform is applied to a block whose short side length of the width or height is the same as or greater than a first preset length. Also, a 4×4 secondary transform is applied to a block whose short side length of the width or height is the same as or greater than a second preset length and smaller than the first preset length. At this time, the first preset length may be a value greater than the second preset length, but the present disclosure is not limited thereto. Also, unlike the primary transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform is referred to as a low-band non-separable transform (LFNST).

[0078] Also, in the case of a video signal in a specific region, even if frequency conversion is performed due to a rapid change in brightness, the high-frequency band energy does not decrease. As a result, there is a risk that the compression performance due to quantization will deteriorate. Also, when conversion is performed on a region where residual values rarely exist, the encoding and decoding times may increase unnecessarily. Therefore, the conversion for the residual signal in the specific region may be omitted. Whether to perform the conversion for the residual signal in the specific region is determined by a syntax element related to the conversion in the specific region. For example, the syntax element includes transform skip information. The transform skip information is a transform skip flag. If the transform skip information for a residual block indicates a transform skip, the conversion for the corresponding residual block is not performed. In this case, the encoder immediately quantizes the residual signal for which the conversion in the corresponding region has not been performed. The operation of the encoder described with reference to FIG. 7 is performed via the conversion unit in FIG. 1.

[0079] The syntax element related to the conversion described above is information parsed from the video signal bitstream. The decoder entropy decodes the video signal bitstream to obtain the syntax element related to the conversion. Also, the encoder entropy encodes the syntax element related to the conversion to generate the video signal bitstream.

[0080] FIG. 8 is a diagram showing in detail a method in which an encoder and a decoder inverse-transform conversion coefficients to obtain a residual signal. Hereinafter, for convenience of explanation, it will be described that the inverse-transform operation is performed through the inverse-transform units of the encoder and the decoder, respectively. The inverse-transform unit inverse-transforms the inverse-quantized conversion coefficients to obtain a residual signal. First, the inverse-transform unit detects whether an inverse transform for a corresponding region is to be performed from syntax elements related to the transform of a specific region. According to one embodiment, if a syntax element related to the transform for a specific transform block indicates a transform skip, the transform for the corresponding transform block is omitted. In this case, both the primary inverse transform and the secondary inverse transform for the transform block are omitted. Also, the inverse-quantized conversion coefficients are used as the residual signal. For example, the decoder uses the inverse-quantized conversion coefficients as the residual signal to restore the current block. The primary inverse transform described above indicates the inverse transform for the primary transform and is referred to as an inverse primary transform. The secondary inverse transform indicates the inverse transform for the secondary transform and is referred to as an inverse secondary transform or an inverse LFNST. In the present invention, the primary (inverse) transform is referred to as the first (inverse) transform, and the secondary (inverse) transform is referred to as the second (inverse) transform.

[0081] According to another embodiment, a syntax element related to the transform for a specific transform block may not indicate a transform skip. In this case, the inverse-transform unit determines whether to perform a secondary inverse transform for the secondary transform. For example, if the transform block is a transform block of an intra-predicted block, a secondary inverse transform for the transform block is performed. Also, based on the intra-prediction mode for the transform block, the secondary transform kernel used for the corresponding transform block is determined. As another example, it may be determined whether to perform a secondary inverse transform according to the size of the transform block. The secondary inverse transform is performed after the inverse-quantization process and before the primary inverse transform is performed.

[0082] The inverse transformation unit performs a first inverse transformation on the inverse quantized transformation coefficients or the secondarily inverse transformed transformation coefficients. In the case of the first inverse transformation, similar to the first transformation, it is performed separately for the vertical transformation and the horizontal transformation. For example, the inverse transformation unit performs a vertical inverse transformation and a horizontal inverse transformation on the transformation block to obtain a residual block. The inverse transformation unit inverse-transforms the transformation block based on the transformation kernel used for the transformation of the transformation block. For example, the encoder explicitly or implicitly signals information indicating the transformation kernel currently applied to the current transformation block among the plurality of available transformation kernels. The decoder uses the information indicating the signaled transformation kernel to select the transformation kernel to be used for the inverse transformation of the transformation block among the plurality of available transformation kernels. The inverse transformation unit restores the current block using the residual signal obtained through the inverse transformation of the inverse transformation coefficients.

[0083] On the other hand, the distribution of the residual signal of a picture may vary by region. For example, the residual signal within a specific region may have a different value distribution depending on the prediction method. When performing transformation using the same transformation kernel for a plurality of mutually different transformation regions, the coding efficiency may vary by transformation region according to the value distribution and characteristics within the transformation region. Thereby, if the transformation kernel to be used for the transformation of a specific transformation block is adaptively selected among the plurality of available transformation kernels, the coding efficiency can be further improved. That is, the encoder and the decoder are set so that transformation kernels other than the basic transformation kernel can be further used in the transformation of the video signal. The method of adaptively selecting the transformation kernel is referred to as adaptive multiple core transform (ATM) or multiple transform selection (MTS). In the present disclosure, for convenience of explanation, the transformation and the inverse transformation are collectively referred to as transformation. Also, the transformation kernel and the inverse transformation kernel are collectively referred to as the transformation kernel.

[0084] The residual signal, which is the signal of the difference between the original signal and the predicted signal generated through inter-picture prediction or intra-picture prediction, has energy dispersed across the entire pixel domain. If the pixel values of the residual signal itself are encoded, there will be a problem of reduced compression efficiency. Therefore, a process is required to concentrate the energy in the low-frequency region of the frequency domain through transform coding of the residual signal in the pixel domain.

[0085] In the HEVC (high efficiency video coding) standard, when the signal is uniformly distributed in the pixel domain (when adjacent pixel values are similar), DCT-II (discrete cosine transform type-II) is mostly used, and DST-VII (discrete sine transform type-VII) is only used limitedly for the predicted 4×4 blocks within the picture to transform the residual signal in the pixel domain into the frequency domain. The DCT-II transform is suitable for the residual signal generated through inter-picture prediction (when energy is uniformly distributed in the pixel domain). However, in the case of the residual signal generated through intra-picture prediction, due to the characteristics of intra-picture prediction that predicts using the restored reference samples around the current coding unit, the energy of the residual signal tends to increase as the distance from the reference samples increases. Therefore, when only DCT-II transform is used to transform the residual signal into the frequency domain, high coding efficiency cannot be achieved.

[0086] AMT is a conversion technique that adaptively selects a conversion kernel from a number of pre-set kernels according to the prediction method. Since the pattern (signal characteristics in the horizontal direction, signal characteristics in the vertical direction) in the pixel domain of the residual signal varies depending on which prediction method is used, higher coding efficiency is expected than when only DCT-II is used for the conversion of the residual signal. In the present invention, AMT may be referred to as MTS (multiple transform selection) in addition to its name.

[0087] FIG. 9 is a diagram showing basis functions for a plurality of conversion kernels that can be used in the first conversion.

[0088] Specifically, FIG. 9 is a diagram showing basis functions of conversion kernels used in AMT, and shows mathematical expressions of DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to AMT.

[0089] Although DCT and DST are represented by cosine and sine functions respectively, when the basis function of the conversion kernel with respect to the number of samples N is represented by Ti(j), the index i indicates the index in the frequency domain, and the index j indicates the index within the basis function. That is, the smaller i is, the lower frequency basis function is indicated, and the larger i is, the higher frequency basis function is indicated. The basis function Ti(j) indicates the element at the j-th position in the i-th row when represented as a two-dimensional matrix. However, since all the conversion kernels shown in FIG. 9 have separable characteristics, the conversion can be performed on the residual signal X in the horizontal and vertical directions respectively. That is, if the residual signal block is X and the conversion kernel matrix is T, the conversion on the residual signal X is represented by TXT’. Here, T’ means the transpose matrix of the conversion kernel matrix T.

[0090] The conversion matrix values defined by the basis functions shown in FIG. 9 are in complex form rather than integer form. The complex form values may be difficult to be implemented hardware-wise in the video encoding device and the decoding device. Therefore, a conversion kernel approximated to an integer from the original conversion kernel including complex form values is used in the encoding and decoding of video signals. The approximated conversion kernel including integer form values is generated through scaling and rounding with respect to the original conversion kernel. The integer values included in the approximated conversion kernel are values within a range that can be represented by a preset number of bits. The preset number of bits is 8-bit or 10-bit. The orthogonality properties of DCT and DST may not be maintained due to the approximation. However, since the loss of coding efficiency due to this is not significant, approximating the conversion kernel to an integer form is advantageous in terms of hardware implementation.

[0091] In the case of the primary conversion region and the inverse primary conversion described with reference to FIGS. 7 to 8, since separable conversion kernels are represented by two-dimensional matrices and the conversions are performed in the vertical and horizontal directions respectively, it is considered that the two-dimensional matrix product operation is performed twice. Since this involves a large amount of computation, it can be a problem from the perspective of implementation. Therefore, from the perspective of implementation, can the amount of computation be reduced by using a combination structure of a butterfly structure or a half butterfly structure like DCT-II and a half matrix multiplier, or can the corresponding conversion kernel be decomposed into a conversion kernel with a lower implementation complexity (can the corresponding channel be represented by the product of matrices with a lower complexity)? And since the elements of the conversion kernel (the matrix elements of the conversion kernel) should be stored in memory for computation, the memory capacity for storing the kernel matrix should also be considered during implementation. From such a perspective, since the implementation complexity of DST-VII and DCT-VIII is high, a conversion with a lower implementation complexity while showing characteristics similar to DST-VII and DCT-VIII can replace DST-VII and DCT-VIII.

[0092] DST-IV (discrete sine transform type-IV) and DCT-IV (discrete cosine transform type-IV) are considered to be candidates that can replace DST-VII and DCT-VIII respectively. The DCT-II kernel for 2N samples includes the DCT-IV kernel for N samples, and since the DST-IV kernel for N samples can be implemented by simply inverting the sign and rearranging the corresponding basis functions in reverse order, which is a simple operation from the DCT-IV kernel for N samples, DST-IV and DCT-IV for N samples can be easily derived from DCT-II for 2N samples.

[0093] Since the residual signal, which is the difference between the original signal and the predicted signal, exhibits the characteristic that the energy distribution of the signal changes depending on the prediction method, the coding efficiency can be improved by adaptively selecting the conversion kernel according to the prediction method, such as AMT or MTS. Also, as described with reference to FIGS. 7 to 8, in addition to the primary conversion and the inverse primary conversion (the inverse conversion corresponding to the primary conversion), the secondary conversion and the inverse secondary conversion (the inverse conversion corresponding to the secondary conversion), which are additional conversions, can be performed to improve the coding efficiency. Such a secondary conversion improves energy compaction for an in-predicted residual signal block in which there is a high possibility that strong energy exists in a direction other than the horizontal and vertical directions of the residual signal. As described above, such a secondary conversion is referred to as a low-frequency non-separable transform (LFNST). And the primary conversion is referred to as a core transform.

[0094] FIG. 10 is a block diagram showing a process of restoring a residual signal by a decoder that performs a secondary conversion according to an embodiment of the present invention. First, the entropy coder parses the syntax elements regarding the residual signal from the bitstream, and the quantized coefficients are obtained through de-binarization. The decoder performs inverse quantization on the restored quantized coefficients to obtain the transform coefficients, and performs inverse transform on the transform coefficients to restore the residual signal block. The inverse transform is applied to blocks to which transform skip (TS) is not applied. The inverse transform is performed in the decoder in the order of the secondary inverse transform and the primary inverse transform. At this time, the secondary inverse transform may be omitted. The secondary inverse transform may be omitted without being performed on the inter-predicted block. Or, the secondary inverse transform may be omitted according to the block size condition. The restored residual signal contains quantization error, and the secondary conversion can reduce the quantization error more than when only the primary conversion is performed by changing the energy distribution of the residual signal.

[0095] FIG. 11 is a diagram showing at a block level a process of restoring a residual signal by a decoder performing secondary conversion according to an embodiment of the present invention. Restoration of the residual signal is performed in units of a transform unit (TU) or sub-blocks within the TU. FIG. 11 shows a process of restoring a residual signal block to which secondary conversion is applied, and secondary inverse conversion is first performed on the inverse quantized transform coefficient block. The decoder may perform secondary inverse conversion on all W×H (W: width, number of horizontal samples, H: height, number of vertical samples) samples within the TU, but considering complexity, may perform secondary inverse conversion only on a left-upper W’×H’ size sub-block which is the most influential low-frequency region. At this time, W’ is the same as or smaller than W. H’ is the same as or smaller than H. The left-upper sub-block size W’×H’ is set to be different according to the TU size. For example, if min(W,H)=4, both W’ and H’ are set to 4. If min(W,H)>=8, both W’ and H’ are set to 8. min(x,y) represents an operation that returns x if x is the same as or smaller than y, and returns y if x is larger than y. After the decoder performs secondary inverse conversion, it obtains sub-block transform coefficients of the left-upper W’×H’ size within the TU, and performs primary inverse conversion on the entire W×H size transform coefficient block to restore the residual signal block.

[0096] Activation or applicability of secondary conversion is indicated in the form of a 1-bit flag in at least any one of high level syntax (HLS) RBSPs such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, a tile group header, etc. Further, if secondary conversion is applicable, the size of the left-upper sub-block considered in secondary conversion may be indicated in the form of a 1-bit flag in at least any one of the HLS RBSPs. For example, whether an 8×8 size sub-block can be used for secondary conversion considering 4×4 and 8×8 size sub-blocks is indicated by a 1-bit flag in at least any one of the HLS RBSPs.

[0097] If the activation or applicability of the secondary transformation is indicated at a higher level (e.g., HLS), whether the secondary transformation is summarized or not is indicated by a 1-bit flag at the coding unit (CU) level. Also, if the secondary transformation is applied to the current block, an index indicating the transformation kernel used for the secondary transformation at the coding unit level is indicated. The decoder uses the transformation kernel indicated by the corresponding index within a set of transformation kernels preset according to the prediction mode, and performs an inverse secondary transformation on the block to which the secondary transformation is applied. The index indicating the transformation kernel is binary-coded using the truncated unary or fixed-length binary method. The 1-bit flag indicating whether the secondary transformation is applied at the CU level and the index indicating the transformation kernel used for the secondary transformation may be indicated using one syntax element, which is referred to as lfnst_idx[x0][y0] or lfnst_idx in the present invention, but the present invention is not limited thereto. As an example, the first bit of lfnst_idx[x0][y0] indicates whether the secondary transformation can be applied at the CU level. And the remaining bits indicate the index indicating the transformation kernel used for the secondary transformation. That is, lfnst_idx[x0][y0] indicates whether the secondary transformation (LFNST) can be applied and the index indicating the transformation kernel used when the secondary transformation is applied. Such lfnst_idx[x0][y0] is coded via an entropy coder such as CABAC (context-based adaptive binary arithmetic coding) or CAVLC (context-based adaptive variable length coding) that is adaptively coded according to the context. If the current CU is divided into a number of TUs smaller than the CU size, the secondary transformation is not applied, and lfnst_idx[x0][y0], which is a syntax element related to the secondary transformation, is set to 0 without signaling. For example, if lfnst_idx[x0][y0] is 0, it indicates that the secondary transformation is not applied.On the contrary, if lfnst_idx[x0][y0] is greater than 0, it indicates that a second-order transformation is applied, and the transformation kernel used for the second-order transformation is selected based on lfnst_idx[x0][y0].

[0098] As described above, a coding tree unit, a leaf node of a quad tree, and a leaf node of a multi-type tree can be coding units. If the coding unit is not larger than the maximum transformation length, the corresponding coding unit is used as a prediction and / or transformation unit without further division. As an example, if the width or height of the current coding unit is larger than the maximum transformation length, the current coding unit is divided into a plurality of transformation units without explicit signaling regarding division. If the size of the coding unit is larger than the maximum transformation size, it is divided into a plurality of transformation blocks without signaling. In this case, since applying a second-order transformation will result in performance degradation and increased complexity, the maximum coding block (or the maximum size of the coding block) to which the second-order transformation is applied is limited. The size of the maximum coding block is the same as the maximum transformation size. Or, the size of the maximum coding block is defined as the size of a preset coding block. As an example, the preset values may be 64, 32, 16, but the present invention is not limited thereto. At this time, the value compared with the preset value (or the maximum transformation size) is defined as the length of the long side or the number of samples.

[0099] On the one hand, the transformation kernels based on DCT-II, DST-VII, and DCT-VIII basis functions used in the first transformation have separable characteristics. Therefore, two transformations in the vertical / horizontal directions are performed on the samples within the residual block of size N×N, and the size of the transformation kernel is N×N. In contrast, in the case of the second transformation, the transformation kernel has non-separable characteristics. Therefore, if the number of samples considered in the second transformation is n×n, one transformation is performed. At this time, the size of the transformation kernel is (n^2)×(n^2). For example, when performing the second transformation on the left-upper 4×4 coefficient block, a transformation kernel of size 16×16 is applied. And when performing the second transformation on the left-upper 8×8 coefficient block, a transformation kernel of size 64×64 is applied. Since the transformation kernel of size 64×64 involves a large amount of product operations, it can impose a significant burden on the encoder and decoder. Therefore, when the number of samples considered in the second transformation decreases, the amount of computation and the memory required for storing the transformation kernel can be reduced.

[0100] FIG. 12 is a diagram showing a method of applying a second transformation that uses a reduced number of samples according to an embodiment of the present invention. According to an embodiment of the present invention, the second transformation is represented by the product of a second transformation kernel matrix and a first-transformed coefficient vector, and the first-transformed coefficients are interpreted as being mapped to another space. At this time, if the number of coefficients to be second-transformed is reduced, that is, if the number of basis vectors constituting the second transformation kernel is reduced, the amount of computation required for the second transformation and the memory capacity required for storing the transformation kernel can be reduced. For example, when performing the second transformation on the left-upper 8×8 coefficient block, if the number of coefficients to be second-transformed is reduced to 16, a second transformation kernel of size 16 (rows)×64 (columns) (or 16 (rows)×48 (columns)) is applied. The transformation unit of the encoder obtains the second-transformed coefficient vector through the inner product of each row vector constituting the transformation kernel matrix and the first-transformed coefficient vector. The inverse transformation units of the encoder and decoder obtain the first-transformed coefficient vector through the inner product of each column vector constituting the transformation kernel matrix and the second-transformed coefficient vector.

[0101] Referring to FIG. 12, the encoder first performs a forward primary transform on the residual signal block to obtain a forward primary transformed coefficient block. If the size of the forward primary transformed coefficient block is M×N, for an intra-predicted block where the value of min(M,N) is 4, a 4×4 forward secondary transform is performed on the left-upper 4×4 samples of the forward primary transformed coefficient block. For an intra-predicted block where the value of min(M,N) is 8 or more, an 8×8 forward secondary transform is performed on the left-upper 8×8 samples of the forward primary transformed coefficient block. In the case of the 8×8 forward secondary transform, since it involves a large amount of computation and memory, only a part of the 8×8 samples may be utilized. In one embodiment, in order to improve the coding efficiency, for a rectangular block where the value of min(M,N) is 4 and M or N is greater than 8 (for example, a rectangular block of 4×16 or 16×4 size), a 4×4 forward secondary transform may be performed on each of the two left-upper 4×4 sub-blocks within the forward primary transformed coefficient block.

[0102] Since the second transformation is calculated by the product of the second transformation kernel matrix and the input vector, first, the encoder forms the coefficients in the left-upper sub-block of the first-transformed coefficient block in the form of a vector. The method of forming the vector depends on the intra prediction mode. For example, if the intra prediction mode is below the 34th angle mode among the intra prediction modes shown in FIG. 6, the encoder scans the left-upper sub-block of the first-transformed coefficient block horizontally to form the coefficients into a vector. Representing the element in the i-th row and j-th column of the left-upper n×n block of the first-transformed coefficient block as x(i, j), the vectorized coefficients are [X(0,0), X(0,1), …, X(0,n-1), X(1,0), X(1,1), …, X(1,n-1), …, X(n-1,0), X(n-1,1), …, X(n-1,n-1)]. On the contrary, if the intra prediction mode is greater than the 34th angle mode, the encoder scans the left-upper sub-block of the first-transformed coefficient block vertically to form the coefficients into a vector. The vectorized coefficients are [X(0,0), X(1,0), …, X(n-1,0), X(0,1), X(1,1), …, X(n-1,1), …, X(0,n-1), X(1,n-1), …, X(n-1,n-1)]. In order to reduce the amount of computation, when only a part of the 8×8 samples is utilized in the 8×8 second transformation, the coefficient x_ij where i>3 and j>3 may not be included in the above-described vector formation method. In this case, in the 4×4 second transformation, 16 first-transformed coefficients can be the input of the second transformation. In the 8×8 second transformation, 48 first-transformed coefficients can be the input of the second transformation.

[0103] The encoder obtains the coefficients after the second - order transformation through the product of the left - upper sub - block samples of the vectorized first - order transformation coefficient block and the second - order transformation kernel matrix. The second - order transformation kernel applied to the second - order transformation is determined according to the size of the transformation unit or transformation block, the intra - mode, and the syntax element indicating the transformation kernel. As described above, when the number of coefficients to be second - order transformed decreases, the computational amount and the memory required for storing the transformation kernel can be reduced. Therefore, the number of coefficients to be second - order transformed is determined according to the current size of the transformation block. For example, in the case of a 4×4 block, the encoder obtains a coefficient vector of length 8 through the product of a vector of length 16 and an 8 (rows)×16 (columns) transformation kernel matrix. The 8 (rows)×16 (columns) transformation kernel matrix is obtained based on the first to the eighth basis vectors that make up the 16 (rows)×16 (columns) transformation kernel matrix. In the case of a 4×N or M×4 block (N and M are 8 or more), the encoder obtains a coefficient vector of length 16 through the product of a vector of length 16 and a 16 (rows)×16 (columns) transformation kernel matrix. In the case of an 8×8 block, the encoder obtains a coefficient vector of length 8 through the product of a vector of length 48 and an 8 (rows)×48 (columns) transformation kernel matrix. The 8 (rows)×48 (columns) transformation kernel matrix is obtained based on the first to the eighth basis vectors that make up the 16 (rows)×48 (columns) transformation kernel matrix. In the case of an M×N block (M and N are 8 or more) except 8×8, the encoder obtains a coefficient vector of length 16 through the product of a vector of length 48 and a 16 (rows)×48 (columns) transformation kernel matrix.

[0104] According to an embodiment of the present invention, since the coefficients after the second - order transformation are in the form of a vector, they are represented by two - dimensional form data. According to a preset scan order, the coefficients after the second - order transformation are configured into a left - upper coefficient sub - block. In one embodiment, the preset scan order is the upper - right diagonal scan order. The present invention is not limited to this, and the upper - right diagonal scan order is determined based on the methods described in FIGS. 13 and 14 to be described later.

[0105] Also, according to an embodiment of the present invention, the conversion coefficients of the overall conversion unit including the secondarily converted coefficients are included in a bitstream and transmitted after quantization. The bitstream includes syntax elements related to the second conversion. Specifically, the bitstream includes information on whether the second conversion is applied to the current block, and information indicating the conversion kernel applied to the current block.

[0106] The decoder first parses the quantized conversion coefficients from the bitstream and obtains the conversion coefficients through de-quantization. De-quantization is also referred to as scaling. The decoder determines whether the second inverse conversion is to be performed on the current block based on the syntax elements related to the second conversion. If the second inverse conversion is applied to the current conversion unit or conversion block, 8 or 16 conversion coefficients can be input to the second inverse conversion according to the size of the conversion unit or conversion block. The number of coefficients input to the second inverse conversion matches the number of coefficients output by the second conversion of the encoder. For example, if the size of the conversion unit or conversion block is 4×4 or 8×8, 8 conversion coefficients are input to the second inverse conversion; otherwise, 16 conversion coefficients are input to the second inverse conversion. If the size of the conversion unit is M×N, for an intra-predicted block where the value of min(M,N) is 4, 4×4 second inverse conversion is performed on 16 or 8 coefficients of the left-upper 4×4 sub-block of the conversion coefficient block. For an intra-predicted block where min(M,N) is 8 or more, 8×8 second conversion is performed on 16 or 8 coefficients of the left-upper 4×4 sub-block of the conversion coefficient block. In one embodiment, in order to improve the coding efficiency, if min(M,N) is 4 and M or N is greater than 8 (for example, a rectangular block of size 4×16 or 16×4), 4×4 second inverse conversion may be performed on two left-upper 4×4 sub-blocks within the coefficient block respectively.

[0107] According to an embodiment of the present invention, since the secondary inverse transform is calculated by the product of the secondary inverse transform kernel matrix and the input vector, the decoder configures the previously input inverse-quantized transform coefficient block in the form of a vector according to a preset scan order. According to an embodiment, the preset scan order is the upper right diagonal scan order, but the present invention is not limited thereto, and the upper right diagonal scan order is determined based on the methods described in FIGS. 13 and 14 to be described later.

[0108] Also, according to an embodiment of the present invention, the decoder obtains the first-order transformed coefficients through the product of the vectorized transform coefficients and the secondary inverse transform kernel matrix. At this time, the secondary inverse transform kernel is determined according to the size of the transform unit or transform block, the intra mode, and the syntax element indicating the transform kernel. The secondary inverse transform kernel matrix is the transposed matrix of the secondary transform kernel matrix. Considering the implementation complexity, the elements of the kernel matrix are integers represented with an accuracy of 10-bit or 8-bit. Based on the current size of the transform block, the length of the vector that is the output of the secondary inverse transform is determined. For example, in the case of a 4×4 block, a coefficient vector of length 16 is obtained through the product of a vector of length 8 and an 8 (row) × 16 (column) transform kernel matrix. The 8 (row) × 16 (column) transform kernel matrix is obtained based on the first to eighth basis vectors that make up the 16 (row) × 16 (column) transform kernel matrix. In the case of a 4×N or M×N block (N and M are 8 or more), a coefficient vector of length 16 is obtained through the product of a vector of length 16 and a 16 (row) × 16 (column) transform kernel matrix. In the case of an 8×8 block, a coefficient vector of length 48 is obtained through the product of a vector of length 8 and an 8 (row) × 48 (column) transform kernel matrix. The 8 (row) × 48 (column) transform kernel matrix is obtained based on the first to eighth basis vectors that make up the 16 (row) × 48 (column) transform kernel matrix. In the case of an M×N block (M and N are 8 or more) except 8×8, a coefficient vector of length 48 is obtained through the product of a vector of length 16 and a 16 (row) × 48 (column) transform kernel matrix.

[0109] In one embodiment, since the primary conversion coefficients obtained through the secondary inverse conversion are in the form of a vector, the decoder can further represent this in the form of two-dimensional data, which is intra-mode dependent. At this time, the mapping relationship based on the intra-mode applied by the encoder is also applied. As described above, if the intra prediction mode is 34-degree mode or less, the decoder scans the coefficient vector after the secondary inverse conversion in the horizontal direction to obtain a two-dimensional conversion coefficient array. If the intra prediction mode is greater than the 34-degree mode, the decoder scans the coefficient vector after the secondary inverse conversion in the vertical direction to obtain a two-dimensional conversion coefficient array. The decoder performs a primary inverse conversion on all conversion units or conversion coefficient blocks of the conversion block size including the conversion coefficients obtained by performing the secondary inverse conversion to obtain a residual signal.

[0110] Although not shown in FIG. 12, a scaling process using a bit shift operation may be included when applying the conversion or inverse conversion in order to correct the scale increased by the conversion kernel after the change or inverse conversion.

[0111] FIG. 13 is a diagram showing a method for determining the upper right diagonal scan order according to an embodiment of the present invention. According to an embodiment of the present invention, during encoding or decoding, a process of initializing the scan order is performed. Initialization of an array including scan order information is performed according to the block size. Specifically, for the combination of log2BlockWidth and log2BlockHeight, the array initialization process of the upper right diagonal scan order shown in FIG. 13 with 1<<log2BlockWidth and 1<<log2BlockHeight as inputs is called (or performed). The output of the array initialization process of the upper right diagonal scan order is assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight]. Here, log2BlockWidth and log2BlockHeight are variables indicating the values obtained by taking the base-2 logarithm of the width and height of the block, respectively, and are values in the range of [0, 4].

[0112] Through the array initialization process of the upper right diagonal scan order shown in FIG. 13, the encoder / decoder outputs the array diagScan[sPos][sComp] for blkWidth, the width of the input block, and blkHeight, the height of the block. The sPos, which is the index of the array, indicates the scan position (scan index) and is a value in the range of [0, blkWidth*blkHeight-1]. If the sComp, which is the index of the array, is 0, sPos indicates the horizontal component (x), and if sComp is 1, sPos indicates the vertical component (y). The algorithm shown in FIG. 13 is interpreted such that the x-coordinate value and y-coordinate value on the two-dimensional coordinates at the scan position sPos in the upper right diagonal scan order are assigned to diagScan[sPos][0] and diagScan[sPos][1], respectively. That is, the value stored in the DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][sComp] array (or array) means the coordinate value corresponding to sComp at the sPos scan position (scan index) in the upper right diagonal scan order of the block where the width and height of the block are 1<<log2BlockWidth and 1<<log2BlockHeight, respectively.

[0113] FIG. 14 is a diagram showing the upper right diagonal scan order according to an embodiment of the present invention by block size. Referring to FIG. 14(a), if both log2BlockWidth and log2BlockHeight are 2, it means a 4×4 size block. Referring to FIG. 14(b), if both log2BlockWidth and log2BlockHeight are 3, it means an 8×8 size block. In FIG. 14, the numbers represented by the gray shaded areas indicate the scan position (scan index) sPos. The x-coordinate value and y-coordinate value at the sPos position are assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][0] and DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][1], respectively.

[0114] The encoder / decoder codes the transform coefficient information based on the scan order described above. In the present invention, embodiments based on the case where the upper right scan method is used will be mainly described. However, the present invention is not limited to this, and can also be applied to other known scan methods.

[0115] Hereinafter, the decoding process related to the secondary transform will be described in detail. For convenience of explanation, the decoder will be mainly described for the process related to the secondary transform. However, the embodiments described below are applied to the encoder in substantially the same manner.

[0116] FIG. 15 is a diagram showing a method of instructing a second conversion at the coding unit level. The second conversion is instructed at the coding unit level, and the syntax elements related to the second conversion are included in the coding_unit syntax structure. The coding_unit syntax structure includes the syntax elements related to the coding unit. At this time, (x0, y0), which is the coordinate of the left-upper luma sample of the current block based on the left-upper luma sample of the picture, cbWidth which is the width of the block, cbHeight which is the height of the block, and treeType which is a variable indicating the type of coding tree are the inputs of the coding_unit syntax structure. Since there is a correlation between luma and chroma, efficient video compression becomes possible when luma and chroma are coded with the same coding structure. Also, in order to improve the coding efficiency, luma and chroma may be coded with different coding structures from each other. If the variable treeType is SINGLE_TREE, it means that luma and chroma are coded with the same coding tree structure, and the coding unit includes a luma coding block and a chroma coding block according to the color format. If treeType is DUAL_TREE_LUMA, it means that luma and chroma are coded with different coding trees from each other, and it indicates that the currently processed tree is a tree for luma. At this time, the coding unit includes only the luma coding block. If treeType is DUAL_TREE_CHROMA, it means that luma and chroma are coded with different coding trees from each other, and it indicates that the currently processed tree is a tree for chroma. At this time, the coding unit includes a chroma coding block according to the color format.

[0117] In the coding_unit syntax structure, the current prediction method for the coding unit is indicated, and the variable CuPredMode[x0][y0] indicates the prediction method for the current block. If CuPredMode[x0][y0] is MODE_INTRA, it indicates that the intra prediction method is applied to the current block; if it is MODE_INTER, it indicates that the inter prediction method is applied to the current block. Also, if CuPredMode[x0][y0] is MODE_IBC, it indicates that the IBC (Intra Block Copy) prediction, which generates a reference block from the area where the restoration of the current picture has been completed and performs prediction, is applied to the current block. According to the value of the variable CuPredMode[x0][y0], the processing of the syntax elements related to the prediction method is performed. For example, if the variable CuPredMode[x0][y0] indicates intra prediction, the decoder parses the syntax elements including the intra prediction mode, reference line index, and information related to ISP (Intra Sub-Partitions) prediction, or sets the variable related to the intra prediction mode by a preset method.

[0118] After processing the syntax elements related to the prediction method, the processing of the syntax elements related to the residual signal is performed. The transform_tree() syntax structure is the syntax structure for the transform tree. The transform tree is divided into nodes with sizes smaller than the root node with the same size as the coding unit as the root node, and the leaf nodes of the transform tree become the transform units. The transform_tree syntax structure contains information related to the division of the transform tree.

[0119] One of the intra prediction methods is PCM (Pulse Code Modulation) prediction. If PCM prediction is used for the prediction of the current coding unit, conversion and quantization are not performed, so the transform_tree syntax structure does not exist. That is, since the transform_tree syntax structure does not exist, the decoder does not perform operations on the transform_tree syntax structure. PCM prediction is indicated by pcm_flag[x0][y0] when intra prediction is instructed for the current coding unit. That is, if pcm_flag[x0][y0] is 1, the decoder does not perform operations on the transform_tree syntax structure. On the other hand, whether the transform_tree syntax structure exists for the current coding unit is indicated by a 1-bit flag, which is referred to as cu_cbf in the present invention, but is not limited thereto. The decoder parses cu_cbf or, if cu_cbf is not parsed, sets cu_cbf by a preset method. If cu_cbf is 1, the decoder performs operations on the transform_tree syntax structure. If inter prediction or IBC prediction is used for the prediction of the current coding unit, merge prediction can also be used for the prediction of the current coding unit. Whether merge prediction is used is indicated by merge_flag[x0][y0]. If it is indicated that merge prediction is used for the current block (merge_flag[x0][y0]==1), cu_cbf is not parsed and the value of cu_cbf is determined by a preset method. The preset method is a method based on cu_skip_flag[x0][y0] that indicates the skip mode. For example, if cu_skip_flag[x0][y0] is 1, cu_cbf is inferred to be 0, and otherwise cu_cbf is inferred to be 1. If cu_cbf is 1, processing of the transform_tree syntax structure is performed, and the counter value for measuring the number of non-zero quantization coefficients (significant coefficient) is initialized to 0.

[0120] The numSigCoeff variable means a variable indicating the number of non-zero quantization coefficients existing in the transform unit of the current coding unit, and the processing of syntax elements related to the secondary transform may differ depending on the value of numSigCoeff.

[0121] The numZeroOutSigCoeff variable means a variable indicating the number of non-zero quantization coefficients existing at a specific position in the transform unit included in the current coding unit, and the processing of syntax elements related to the secondary transform may differ depending on the value of numZeroOutSigCoeff.

[0122] In the transform_tree, the transform tree is split, and the leaf nodes of the transform tree are transform units. The transform_tree includes a syntax structure, the transform_unit syntax structure, which is related to the transform units that are leaf nodes. The transform_unit processes syntax elements related to the transform unit, and if the corresponding transform unit includes one or more non-zero coefficients, it includes a residual_coding syntax structure. The residual_coding syntax structure includes a syntax structure related to the quantized transform coefficients and processing related thereto. The transform blocks that make up the transform unit may vary depending on the type of tree being processed. If the treeType is SINGLE_TREE, the current transform unit includes a luma transform block and a chroma transform block depending on the color format. If the treeType is DUAL_TREE_LUMA, the current transform unit includes a luma transform block. If the treeType is DUAL_TREE_CHROMA, the current transform unit includes a chroma transform block. The transform_unit syntax structure includes CBF (coded block flag) information, which is information indicating whether the transform block included in the current transform unit includes one or more non-zero coefficients for the transform block, depending on the treeType. The CBF information is information indicated separately for each color component. For example, if the value of the CBF for the luma transform block of the current transform unit indicates that the luma transform block does not include one or more non-zero coefficients, then all the coefficients of the luma transform block are 0, so the residual_coding syntax structure for the luma transform block is not processed. As another example, if the value of the CBF for the chroma Cb transform block of the current transform unit indicates that the chroma Cb transform block includes one or more non-zero coefficients, then there is a residual_coding syntax structure for the Cb transform block of the current transform unit.

[0123] Whether a secondary transformation is applied to the current block is indicated at the CU level. If a secondary transformation is applied, an index indicating the transformation kernel used for the secondary transformation may also be indicated. As described with reference to FIG. 11, the lfnst_idx[x0][y0] syntax element is used to indicate whether a secondary transformation is applied to the current block. The first bit of lfnst_idx[x0][y0] indicates whether a secondary transformation is applied to the current coding unit. If the first bit of lfnst_idx[x0][y0] is 0, i.e., if lfnst_idx[x0][y0] is 0, it indicates that no secondary transformation is applied to the current block. On the other hand, if the first bit of lfnst_idx[x0][y0] is 1, i.e., if lfnst_idx[x0][y0] is greater than 0 (lfnst_idx[x0][y0]>0), it indicates that a secondary transformation is applied to the current block. At this time, additional bits are used to indicate the transformation kernel used for the secondary transformation, and an index indicating the secondary transformation kernel is signaled via the additional bits.

[0124] The lfnst_idx[x0][y0] syntax element is parsed if the conditions described below are satisfied. On the other hand, if the conditions described below are not satisfied, lfnst_idx[x0][y0] does not exist in the current coding unit and lfnst_idx[x0][y0] is set to 0.

[0125] In other words, if the conditions described in the first to fourth embodiments including the lfnst_idx[x0][y0] syntax element parsing condition to be described later are satisfied, the encoder generates a bitstream including the lfnst_idx[x0][y0] syntax element for the current coding unit. On the contrary, if the conditions to be described later are not satisfied, the bitstream generated by the encoder does not include the lfnst_idx[x0][y0] syntax element for the current coding unit, and lfnst_idx[x0][y0] is set to 0. The decoder that receives such a bitstream parses the lfnst_idx[x0][y0] syntax element based on the conditions to be described later.

[0126] Parsing conditions for the lfnst_idx[x0][y0] syntax element

[0127] i) Min(lfnWidth, lfnHeight) >= 4

[0128] First, the first condition relates to the block size. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0129] Specifically, the decoder checks the block size conditions applicable to the second-order transform. The variables SubWidthC and SubHeightC are set according to the color format, and respectively represent the ratio of the width of the luma component of the picture to the width of the chroma component, and the ratio of the height of the luma component to the height of the chroma component. For example, in a 4:2:0 color format video, since it has a structure containing 1 chroma sample for every 4 luma samples, both SubWidthC and SubHeightC are set to 2. As another example, in a 4:4:4 color format video, since it has a structure containing 1 chroma sample for every 1 luma sample, both SubWidthC and SubHeightC are set to 1. The number of samples in the horizontal direction of the current block, lfnWidth, and the number of samples in the vertical direction, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, since the coding unit contains only the chroma component, the number of samples in the horizontal direction of the chroma coding block is the same as the value obtained by dividing the width of the luma coding block, cbwidth, by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma coding block is the same as the value obtained by dividing the height of the luma coding block, cbHeight, by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, since the coding unit contains the luma component, lfnWidth and lfnHeight are set to cbwidth and cbHeight respectively. Since the minimum condition for a block applicable to the second-order transform is 4×4, if Min(lfnWidth, lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is parsed.

[0130] ii) sps_lfnst_enabled_flag == 1

[0131] The second condition relates to a flag value indicating the activation or applicability of the secondary transformation. If the value of the flag (sps_lfnst_enabled_flag) indicating the activation or applicability of the secondary transformation is set to 1, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0132] Specifically, the secondary transformation is indicated in the upper-level syntax RBSP. At least one of SPS, PPS, VPS, tile group header, and slice header includes a flag with a 1-bit size indicating the activation and applicability of the secondary transformation. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element exists within the coding unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element does not exist within the coding unit syntax.

[0133] iii) CuPredMode[x0][y0] == MODE_INTRA

[0134] The third condition relates to the prediction mode. The secondary transformation is applied only to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0135] iv) IntraSubPartitionsSplitType == ISP_NO_SPLIT

[0136] The fourth condition relates to whether the ISP prediction method is applied. If ISP is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0137] Specifically, as described with reference to FIG. 11, when the current CU is currently divided into a large number of transform units smaller than the CU size, no secondary transform is applied to the divided transform units. At this time, lfnst_idx[x0][y0], which is a syntax element related to the secondary transform, is set to 0 without being parsed. When the current CU is divided into a large number of transform units smaller than the transform tree in terms of CU size, it includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that divides the transform tree into a large number of transform units smaller than the CU size by a preset division method when intra prediction is applied to the current coding unit. The ISP prediction mode is indicated at the coding unit level, and based on this, the variable IntraSubPartitionsSplitType is set. At this time, if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. The secondary transform is indicated at the coding unit level, but the actual secondary transform is applied at the transform unit level. Therefore, if the transform tree is divided into a large number of transform units, it is inefficient for the same secondary transform kernel to be applied to all the divided transform units. Also, due to the characteristics of intra prediction that generates prediction samples at the transform unit level, the prediction accuracy is higher when the transform tree is divided into a large number of transform units than when it is not divided. Therefore, if the transform tree is divided into a large number of transform units, even if no secondary transform is applied to the divided large number of transform units, the energy of the residual signal is likely to be efficiently compressed. Also, if the size of the current CU is larger than the size of the luma maximum transform block (MaxTbSizeY) (that is, cbWidth>MaxTbSizeY||cbHeight>MaxTbSizeY), the transform tree is divided into a large number of transform units smaller than the CU size. Although not shown in FIG. 15, no secondary transform is applied even when the size of the current CU is larger than the luma maximum transform block size (MaxTbSizeY).Therefore, the fourth condition may be represented as IntraSubPartitionsSplitType == ISP_NO_SPLIT && cbWidth <= MaxTbSizeY && cbHeight <= MaxTbSize. At this time, MaxTbSizeY is a natural number expressed in the form of a power of 2. MaxTbSizeY is indicated in the upper-level syntax RBSP such as SPS, PPS, slice header, tile group header, etc., or the encoder and decoder may use the same preset value. For example, the preset value may be 64 (2^6).

[0138] v)!intra_mip_flag[x0][y0]

[0139] The fifth condition relates to the intra prediction method. If MIP (Matrix based Intra Prediction) is not applied to the prediction of the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0140] Specifically, MIP is used as one method of intra prediction, and whether MIP can be applied is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and the prediction is performed by the product of the restored samples around the current block and a preset matrix. When MIP is applied, since it shows the nature of the residual signal different from the general intra prediction that performs directional or non-directional prediction, secondary conversion may not be applied to the transform block when MIP is applied.

[0141] vi) numSigCoeff > ((treeType == SIGNLE_TREE)? 2 : 1)

[0142] The sixth condition relates to treeType and coefficients.

[0143] Specifically, if treeType is SINGLE_TREE and the value of the variable numSigCoeff is greater than 2, a second-order transformation is applied to the current block, and the decoder parses the lfnst_idx[x0][y0] syntax element.

[0144] When treeType is DUAL_TREE_LUMA or DUAL_TREE_CHROMA, if the value of the variable numSigCoeff is greater than 1, a second-order transformation is applied to the current block and lfnst_idx[x0][y0] is parsed. At this time, numSigCoeff means a variable indicating the number of valid coefficients existing within the current coding unit. If numSigCoeff is less than the threshold value, there may be no efficient coding even if a second-order transformation is applied to the current block. This is because when the number of valid coefficients is small, the overhead of signaling lfnst_idx[x0][y0] for the bit comparison required for coefficient coding is relatively large. At this time, a valid coefficient means a coefficient that is not 0. Hereinafter, the valid coefficient described in the present invention means a coefficient that is not 0 as described above.

[0145] vii) numZeroOutSigCoeff == 0

[0146] The seventh condition relates to the valid coefficients existing at specific positions.

[0147] Specifically, if a secondary transformation is currently applied to a block, the transform coefficients quantized by the decoder are always 0 at specific positions. Therefore, if there are (quantized) coefficients that are not 0 at specific positions, it means that the secondary transformation is not currently applied to the block. Thus, the purging feasibility of lfnst_idx[x0][y0] is determined according to the number of valid coefficients at specific positions. For example, if numZeroOutSigCoeff is not 0, it means that there are valid coefficients at specific positions. Therefore, lfnst_idx[x0][y0] is set to 0 without being purged. On the contrary, if numZeroOutSigCoeff is 0, it means that there are no valid coefficients at specific positions. Therefore, lfnst_idx[x0][y0] is purged.

[0148] FIG. 16 is a diagram showing a residual_coding syntax structure according to an embodiment of the present invention.

[0149] The residual_coding syntax structure is a syntax structure related to quantized coefficients and receives x0, y0, log2TbWidth, and log2TbHeight as inputs. At this time, x0 and y0 represent the left-upper coordinates (x0, y0) of the transform block, log2TbWidth is the value obtained by taking the base-2 logarithm of the width of the transform block, and log2TbHeight is the value obtained by taking the base-2 logarithm of the height of the transform block. The number of elements within the transform block is coded in sub-block units, and the values of the coefficients within each sub-block are determined based on various syntax elements including sig_coeff_flag. At this time, the coefficients in sub-block units may be expressed as a coefficient group (CG). sig_coeff_flag[xC][yC] indicates whether the coefficient value at the (xC, yC) position within the current block is 0. If sig_coeff_flag[xC][yC] is 1, it indicates that the coefficient value at the corresponding position is not 0, and if sig_coeff_flag[xC][yC] is 0, it indicates that the coefficient value at the corresponding position is 0. In residual_coding, the x-coordinate value and y-coordinate value of the last significant coefficient in the scan order are indicated. Based on the x-coordinate value and y-coordinate value of the last significant coefficient in the scan order, the index (lastSubBlock) of the sub-block including the last significant coefficient in the scan order is determined. The index of the sub-block is also indexed based on the scan order. The scan order is the upper-right diagonal scan order described in FIG. 13. In the coefficient coding in sub-block units, the indices xC and yC indicating the coefficient position (coordinate value) are determined based on the left-upper coordinates (xS<<log2SbW, yS<<log2SbH) of the sub-block and the upper-right diagonal scan order (DiagScanOrder). At this time, xS and yS indicate the horizontal index and vertical index, respectively. log2SbW and log2SbH are the values obtained by taking the base-2 logarithm of the width and height of the sub-block, respectively.

[0150] If the value of sig_coeff_flag[xC][yC] is 1 (i.e., when the number of elements at position (xC, yC) is not zero), and if transform skip is not applied to the current block (i.e.,!transform_skip_flag[x0][y0]), then numSigCoeff is counted. Since secondary transformation may not be applied when transform skip is applied, numSigCoeff used for purging lfnst_idx[x0][y0] counts the number of valid coefficients in the block where transform skip is not applied.

[0151] Also, as described in FIG. 15, if secondary transformation is applied to the transform block, there are no valid coefficients in a specific region within the transform block. Therefore, the numZeroOutSigCoeff counter counts the number of valid coefficients (numZeroOutSigCoeff) present in the specific region, and if numZeroOutSigCoeff is not zero, lfnst_idx[x0][y0] is not purged. Specifically, when secondary transformation is applied to the transform block, the region where no valid coefficients can exist is determined by the size of the transform block.

[0152] For example, for secondary transformation to be applied, when the size of the transform block is 4×4 (i.e., log2TbWidth == 2 && log2TbHeight == 2), the areas of scan order top indices [0,7] and [8,15] within the transform block are divided. Valid coefficients exist in the [0,7] area, and no valid coefficients can exist in the [8,15] area. The 4×4 transform block contains one sub-block. Therefore, when the size of the transform block is 4×4, if the scan position is 8 or more and the index of the sub-block is 0 (i.e., n >= 8 && i == 0), the number of valid coefficients is counted. At this time, the scan order is the upper right diagonal scan order.

[0153] As another example, for the second-order transformation to be applied, when the size of the transformation block is 8×8 (i.e., log2TbWidth == 3 && log2TbHeight == 3), valid coefficients exist only in the first sub-block within the transformation block, and no valid coefficients can exist in the remaining sub-blocks (e.g., the second and third sub-blocks). Even within the first sub-block, valid coefficients exist in the scan order index [0,7] region, but no valid coefficients can exist in the index [8,15] region. Therefore, when the size of the transformation block is 8×8, if the scan position is 8 or more in the first sub-block (i.e., n >= 8 && i == 0), or if the scan position exists in the remaining sub-blocks excluding the first sub-block (e.g., exists in the second and third sub-blocks, i == 1 || i == 2), the number of valid coefficients is counted.

[0154] Finally, when the size of the transformation block is larger than 8×8, valid coefficients exist only in the first sub-block within the transformation block, and no valid coefficients can exist in the remaining sub-blocks (e.g., the second and third sub-blocks). Therefore, if the sub-block is the second or third (i.e., i == 1 || i == 2), the number of valid coefficients is counted. The numZeroOutSigCoeff counter, similar to the numSigCoeff counter, counts the number of valid coefficients only when sig_coeff_flag[xC][yC] is 1 and transform_skip_flag[x0][y0] is 0. At this time, the sub-blocks are indexed in the upper-right diagonal scan order described in FIG. 13.

[0155] In other words, if there are non-zero coefficients in the region (specific region) where no valid coefficients can exist, it means that the second-order transformation has not been performed. Therefore, the valid coefficients are counted to check whether there are non-zero coefficients in the specific region.

[0156] FIG. 17 is a diagram showing a method for instructing second-order transformation at the coding unit level according to an embodiment of the present invention.

[0157] As described with reference to FIGS. 15 and 16, whether or not the second transformation is applied is indicated by the lfnst_idx[x0][y0] syntax element at the coding unit level. In order to parse lfnst_idx[x0][y0], two significant coefficient counters (i.e., the numSigCoeff counter and the numZeroOutSigCoeff counter) are required. In particular, in the case of numSigCoeff, since the numSigCoeff counter should count the number of significant coefficients existing within the entire area of the coding unit, the throughput of the coefficient coding may decrease. Therefore, a method of reducing the number of counters or not using a counter is required.

[0158] The method of indicating the second transformation shown in FIG. 17 is a method of parsing lfnst_idx[x0][y0] regardless of numSigCoeff. In other words, if all of i), ii), iii), iv), and v) among the conditions described in FIG. 15 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. Also, since the value of numSigCoeff is not referenced, the operation of the numSigCoeff counter described in FIG. 16 is not performed.

[0159] Hereinafter, in this specification, a method of indicating the second transformation based on the position information of the last significant coefficient in the scan order will be described. Similar to the case where the number of significant coefficients is small, if the position (scan index) of the last significant coefficient in the scan order is small, the coding efficiency by the second transformation is low. Therefore, it is necessary to efficiently indicate the second transformation based on the position information of the last significant coefficient in the scan order without using a counter.

[0160] (First Embodiment)

[0161] FIG. 18 is a diagram showing a method of indicating the second transformation at the coding unit level according to an embodiment of the present invention.

[0162] FIG. 18 is a diagram showing a method of parsing lfnst_idx[x0][y0] by using the position information of the last valid coefficient in the scan order obtained by residual_coding instead of the numSigCoeff counter.

[0163] According to FIG. 18, since the numSigCoeff counter is not used, the numSigCoeff value does not need to be initialized, and a variable lfnLastScanPos regarding the position of the position information of the last valid coefficient in the scan order is initialized to 1. If the value of lfnLastScanPos is 1, it indicates that the position (scan index) of the last valid coefficient in the scan order is smaller than the threshold value, or all the transform coefficients in the block are 0. On the contrary, if the value of lfnLastScanPos is 0, it indicates that there is one or more valid coefficients in the block, and the position (scan index) of the last valid coefficient in the scan order is greater than or equal to the threshold value. Therefore, if the value of lfnLastScanPos is 1, lfnst_idx[x0][y0] is not parsed, and if the value of lfnLastScanPos is 0, lfnst_idx[x0][y0] is parsed. In addition, lfnst_idx[x0][y0] may also be parsed if the value of lfnLastScanPos is 0 and all of the conditions i), ii), iii), iv), v), vii) described in FIG. 15 are satisfied (if all are true).

[0164] In other words, if there is one or more valid coefficients in the current block and the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold value, then lfnst_idx[x0][y0] is purged. At this time, as will be described later, the threshold value is an integer equal to or greater than 0. For example, assuming that the threshold value is 1, the fact that the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold value means that the valid coefficient exists at a position other than the upper left corner of the block. That is, lfnst_idx[x0][y0] is purged only when the valid coefficient does not exist in the current block or exists only at the upper left corner of the current block, that is, when the valid coefficient exists at a position excluding the upper left corner of the current block. The meaning that the valid coefficient exists at a position excluding the upper left corner of the current block may be represented by "LfnstDConly==0". The upper left corner of the block described in the present invention may mean that the value of the vertical coordinate is (0,0), may mean the first position in a preset scan order (for example, the upper right diagonal order), or may be referred to as DC.

[0165] FIG. 19 is a diagram showing a residual_coding syntax structure according to an embodiment of the present invention.

[0166] FIG. 19 shows the residual_coding syntax structure according to FIG. 18 described above. In residual_coding, the syntax elements regarding the x coordinate and y coordinate of the last significant coefficient in the scan order are parsed, and the LastSignificantCoeffX and LastSignificantCoeffY variables are set. LastSignificantCoeffX indicates the x coordinate of the last significant coefficient in the scan order, and LastSignificantCoeffY indicates the y coordinate of the last significant coefficient in the scan order. Based on LastSignificantCoeffX and LastSignificantCoeffY, the LastScanPos variable, which is the scan index of the last significant coefficient in the scan order, and the index (lastSubBlock) of the sub-block including the last significant coefficient are determined. At this time, as described in FIG. 16, when a secondary transformation is applied to the current block, valid coefficients exist only in the first sub-block. In other words, if valid coefficients exist only in the first sub-block, it means that a secondary transformation is applied.

[0167] For example, in the 4×4 sized block of FIG. 14(a), if LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, then LastScanPos is determined to be 13. Since the 4×4 sized block is composed of one sub-block, the index (lastSubBlock) of the sub-block containing the last significant coefficient is determined to be 0. As another example, the 8×8 sized block of FIG. 14(b) is divided into 4×4 sized sub-blocks. Specifically, in FIG. 14(b), the 4×4 block corresponding to x coordinates 0 to 3 and y coordinates 0 to 3 is the first sub-block, the 4×4 block corresponding to x coordinates 0 to 3 and y coordinates 4 to 37 is the second sub-block, the 4×4 block corresponding to x coordinates 4 to 7 and y coordinates 0 to 34 is the third sub-block, and the 4×4 block corresponding to x coordinates 4 to 7 and y coordinates 4 to 37 is the fourth sub-block. At this time, the first sub-block is indexed as index 0, the second sub-block as index 1, the third sub-block as index 2, and the fourth sub-block as index 3. The sub-blocks are indexed in the upper-right diagonal scan order described in FIG. 13. At this time, if LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, then lastScanPos is determined to be 13. Since lastScanPos is 13, the sub-block containing lastScanPos13 is the first sub-block (that is, sub-block index 0), and the index (lastSubBlock) of the sub-block containing the last significant coefficient is determined to be 0.

[0168] Based on the above-mentioned lastScanPos, lfnstLastScanPos is determined. Specifically, if the width and height of the conversion block are 4 or more and conversion skip is not applied to the conversion block, lfnstLastScanPos is set as shown in the following Equation 1. In other words, if log2TbWidth >= 2, log2TbHeight >= 2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in the following Equation 1. At this time, if transform_skip_flag[x0][y0] is 0, it means that conversion skip is not applied to the current conversion block. Specifically, the flag transform_skip_flag[x0][y0] described in this specification indicates whether primary conversion and secondary conversion are applied to the conversion block. For example, if the value of the transform_skip_flag[x0][y0] is 1, it indicates that primary conversion and secondary conversion are not applied to the conversion block (that is, conversion skip is applied), and if the value of the transform_skip_flag[x0][y0] is 0, it indicates that primary conversion and secondary conversion are applied to the conversion block (that is, conversion skip is not applied).

[0169]

Number

[0170] As described above, the initial value of lfnstLastScanPos is set to 1.

[0171] In Equation 1, cIdx indicates a variable that means the color component of the current conversion block. For example, if cIdx is 0, it indicates that the conversion block processed by residual_coding is the luma Y component. If cIdx is 1, it indicates that the conversion block processed by residual_coding is the chroma Cb component, and if cIdx is 2, it indicates that the conversion block to be processed is the chroma Cr component. The threshold value lfnstLastScanPosTh[cIdx] for lastScanPos is set to different values according to the color component.

[0172] According to Equation 1, if the lfnstLastScanPos of the line is 1 and lastScanPos is smaller than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 1. On the contrary, if the lfnstLastScanPos of the line is 0 or lastScanPos is greater than or equal to lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 0. In other words, if the lastScanPos of all the conversion blocks included in the coding unit is smaller than the threshold value, or if the number of all the conversion blocks is 0, lfnstLastScanPos is determined to be 1, and according to the lfnst_idx[x0][y0] parsing condition in FIG. 18, lfnst_idx[x0][y0] is set to 0 without being parsed. That lfnst_idx[x0][y0] is set to 0 without being parsed indicates that no secondary conversion is applied to the current block. On the contrary, if the LastScanPos of any one of the conversion blocks included in the coding unit is greater than or equal to the threshold value, lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), vii) described in FIG. 15 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a secondary conversion is applied to the current block. If a secondary conversion is applied, it checks / determines the conversion kernel used for the secondary conversion.

[0173] The lfnstLastScanPosTh[cIdx] in Equation 1 is a preset integer value of 0 or more, and both the encoder and the decoder use the same value. Also, the same threshold value may be used for all color components. In this case, lfnstLastScanPos is set as in Equation 2 below. The coding unit described in this specification is composed of a plurality of coding blocks, and there is a conversion block corresponding to each coding block. The conversion block is a conversion block having luminance and color difference components. Specifically, it is a Y conversion block, a Cb conversion block, and a Cr conversion block. At this time, whether to parse lfnst_idx[x0][y0] described in this specification is determined for each conversion block corresponding to each of the coding blocks. That is, if any one of the Y conversion block, the Cb conversion block, and the Cr conversion block satisfies the conditions described in this specification, lfnst_idx[x0][y0] is parsed.

[0174]

Number

[0175] lfnstLastScanPosTh is a preset integer value of 0 or more, and both the encoder and the decoder use the same value. For example, lfnstLastScanPosTh may be 1. That is, if lastScanPos is 1 or more, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is parsed. At this time, since the threshold value (lfnstLastScanPosTh) is an integer value, if lastScanPos is 1 or more, it has the same meaning as the case where lastScanPos is greater than 0. Although the case where the threshold value is 1 has been described as an example of the present invention, the present invention is not limited thereto.

[0176] In other words, based on lastScanPos, it is determined whether lfnst_idx[x0][y0] can be purged. Specifically, as described above, if the secondary transformation is applied, the last valid coefficient in the scan order exists only in the first sub-block of the transformation block. Therefore, the index (lastSubBlock) of the sub-block containing the last valid coefficient in the scan order (where the index indicated by lastScanPos is located) is 0, the width of the transformation block is 4 or more (log2TbWidth>=2), the height of the transformation block is 4 or more (log2TbHeight>=2), transform_skip_flag[x0][y0] is 0 (the transformation skip is not applied), and if LastScanPos is greater than 0 (if LastScanPos is 1 or more), lfnst_idx[x0][y0] is purged. Expressing this as a mathematical formula, it is represented as the following formula 3.

[0177]

Number

[0178] On the other hand, in the above-described first embodiment, since the numSigCoeff counter is not used for purging lfnst_idx[x0][y0], the number of valid coefficients (numSigCoeff) is not counted.

[0179] (Second Embodiment)

[0180] FIG. 20 is a diagram showing a residual_coding syntax structure according to another embodiment of the present invention.

[0181] FIG. 20 is a diagram showing a method in which residual_coding further inputs a treeType variable to FIG. 19 and sets a threshold value for LastScanPos according to treeType.

[0182] If the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as in Equation 4 below. In other words, if log2TbWidth >= 2, log2TbHeight >= 2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as in Equation 4 below. At this time, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.

[0183] [Number]

[0184] In Equation 4, lfnstLastScanPosTh means the threshold value for lastScanPos, and the value is set according to treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, lfnstLastScanPosTh is set to val1, val2, and val3, respectively. If the straight-line lfnstLastScanPos is 1 and lastScanPos is less than lfnstLastScanPosTh, lfnstLastScanPos is updated to 1. On the other hand, if the straight-line lfnstLastScanPos is 0 or lastScanPos is greater than or equal to lfnstLastScanPosTh, lfnstLastScanPos is updated to 0.

[0185] Equation 4 results in lfnstLastScanPos being determined as 1 if the lastScanPos of all the transform blocks included in the coding unit is smaller than the threshold value, or if the number of all the transform blocks is 0. According to the lfnst_idx[x0][y0] parsing condition in FIG. 18, lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that no secondary transform is applied to the current block. On the other hand, if the LastScanPos of any one of the transform blocks included in the coding unit is greater than or equal to the threshold value, lfnstLastScanPos is determined as 0. If all of i), ii), iii), iv), v), and vii) described in FIG. 15 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a secondary transform is applied to the current block. If a secondary transform is applied, the transform kernel used for the secondary transform is checked / determined.

[0186] val1, val2, and val3 are preset integer values of 0 or more, and both the encoder and the decoder use the same values. If treeType is SINGLE_TREE, since both the luma and chroma components are included, the value of val1, which is the value of lfnstLastScanPosTh, may be expressed as the sum of val2 and val3.

[0187] In the second embodiment, since the numSigCoeff counter is not used for parsing lfnst_idx[x0][y0], the number of valid coefficients (numSigCoeff) is not counted.

[0188] (Third Embodiment)

[0189] FIG. 21 is a diagram showing a method of instructing secondary transform at the coding unit level according to another embodiment of the present invention.

[0190] According to FIG. 21, instead of the numSigCoeff counter, the position information of the last valid coefficient in the scan order obtained by residual_coding is utilized to parse lfnst_idx[x0][y0].

[0191] Since the numSigCoeff counter is not used, numSigCoeff does not need to be initialized, and lfnLastScanPos, which is a variable related to the position of the position information of the last valid coefficient in the scan order, is initialized to 0. The lfnstLastScanPos variable in FIG. 21 is the value obtained by adding the lastScanPos of the transform block included in the coding unit. At this time, if lfnLastScanPos is greater than the threshold value and all of the conditions i), ii), iii), iv), v), vii) described in FIG. 15 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a secondary transform is applied to the current block. If a secondary transform is applied, the transform kernel used for the secondary transform is checked / determined. On the contrary, if lfnLastScanPos is less than or equal to the threshold value, lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that no secondary transform is applied.

[0192] The threshold value is set according to treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the threshold values are set to Th1, Th2, and Th3 respectively. Th1, Th2, and Th3 are preset integer values of 0 or more, and both the encoder and the decoder use the same values. If treeType is SINGLE_TREE, since it includes both luma and chroma components, the threshold value Th1 may be expressed as the sum of Th2 and Th3.

[0193] FIG. 22 is a diagram showing the residual_coding syntax structure according to another embodiment of the present invention.

[0194] FIG. 22 shows the residual_coding syntax structure according to FIG. 21 described above. If the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as in Equation 5 below. In other words, if log2TbWidth >= 2, log2TbHeight >= 2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as in Equation 5 below. At this time, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.

[0195]

Equation

[0196] In Equation 5 above, lfnLastScanPos is the value obtained by adding up all the lastScanPos of the transform blocks included in the coding unit. As described in FIG. 21, by comparing lfnLastScanPos with the threshold value, the purging availability of lfnst_idx[x0][y0] is determined.

[0197] In the third embodiment, since the numSigCoeff counter is not used for purging lfnst_idx[x0][y0], the number of valid coefficients (numSigCoeff) is not counted.

[0198] On the one hand, a coding unit includes transform units that are split by a transform tree with a root node of the same size as the coding unit. At this time, each transform unit includes a transform block for each color component. If a second transform is instructed at the coding unit level, after residual coding is performed on all the transform blocks included in the coding unit, lfnst_idx[x0][y0] is parsed based on coefficient information. As another example, the second transform may be instructed at the transform unit level. If the second transform is instructed at the transform unit level, each transform unit included in the coding unit uses a different lfnst_idx[x0][y0]. Therefore, the encoder can search for an optimal lfnst_idx[x0][y0] for each transform unit, and can further improve the coding efficiency. Also, if the second transform is instructed at the coding unit level and the coding unit includes four transform units, residual coding for all the transform blocks included in the four transform units should be processed in order to parse lfnst_idx[x0][y0]. That is, even if the decoder obtains transform coefficients via residual coding for the first transform unit, the decoder cannot perform inverse transform for the first transform unit because it cannot obtain the lfnst_idx[x0][y0] value. This not only increases the buffer size of the decoder, but also may cause excessive delay time in the decoder.

[0199] The first to third embodiments described with reference to FIGS. 18 to 22 are also applicable when the secondary conversion is instructed at the conversion unit level. If the secondary conversion is instructed at the coding unit level, based on the position of the last valid coefficient in the scan order of the conversion block included in the coding unit according to the first to third embodiments, the decision on whether to purge lfnst_idx[x0][y0] is made. Also, if the secondary conversion is instructed at the conversion unit level, based on the position of the last valid coefficient in the scan order of the conversion block included in the conversion unit according to the first to third embodiments, the decision on whether to purge lfnst_idx[x0][y0] is made.

[0200] Hereinafter, a specific method in which the secondary conversion is instructed at the conversion unit level will be described in this specification.

[0201] FIG. 23 is a diagram showing a method of instructing secondary conversion at the conversion unit level according to an embodiment of the present invention.

[0202] According to FIG. 12, instead of the numSigCoeff counter, lfnst_idx[x0][y0] is purged using the position information of the last valid coefficient in the scan order obtained by residual_coding.

[0203] First, before performing residual coding, the variable lfnLastScanPos, which is related to the position of the last valid coefficient in the scan order, is initialized to 1. If the variable lfnLastScanPos is 1, it indicates that for all transform blocks included in the transform unit, the position (scan index) of the last valid coefficient in the scan order is smaller than the threshold value, or all the transform coefficients in the block are 0. If the variable lfnLastScanPos is 0, it indicates that there is one or more valid coefficients in one or more transform blocks included in the transform unit, and the position (scan index) of the last valid coefficient in the scan order is greater than or equal to the threshold value. According to the first embodiment described above, if lfnLastScanPos, which is set based on the position of the last valid coefficient in the scan order of the transform block, is 0, and all of the following conditions i), ii), iii), iv), v), and vi) are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0].

[0204] Parsing conditions for the lfnst_idx[x0][y0] syntax element

[0205] i) Min(lfnWidth, lfnHeight) >= 4

[0206] First, the first condition is related to the size of the block. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0207] Specifically, the decoder checks the block size conditions applicable to the second - order transformation. The variables SubWidthC and SubHeightC are set according to the color format, and respectively represent the ratio of the width of the luma component of the picture to the width of the chroma component and the ratio of the height of the luma component to the height of the chroma component. For example, in a 4:2:0 color format video, since it has a structure that includes 1 chroma sample for every 4 luma samples, both SubWidthC and SubHeightC are set to 2. As another example, in a 4:4:4 color format video, since it has a structure that includes 1 chroma sample for every 1 luma sample, both SubWidthC and SubHeightC are set to 1. lfnWidth, which is the number of samples in the horizontal direction of the current block, and lfnHeight, which is the number of samples in the vertical direction of the current block, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, since the conversion unit only includes the chroma component, the number of samples in the horizontal direction of the chroma conversion block is the same as the value obtained by dividing the width tbwidth of the luma conversion block by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma conversion block is the same as the value obtained by dividing the height tbHeight of the luma conversion block by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, since the conversion unit includes the luma component, lfnWidth and lfnHeight are set to tbwidth and tbHeight respectively. Since the minimum condition for a block applicable to the second - order transformation is 4×4, if Min(lfnWidth, lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is parsed.

[0208] ii) sps_lfn_enabled_flag == 1

[0209] The second condition relates to the flag value indicating the activation or applicability of the secondary transformation. If the value of the flag (sps_lfnst_enabled_flag) indicating the activation or applicability of the secondary transformation is set to 1, the decoder parses lfnst_idx[x0][y0].

[0210] Specifically, the secondary transformation is indicated in the higher-level syntax RBSP. At least one of SPS, PPS, VPS, tile group header, and slice header includes a flag with a 1-bit size indicating the activation and applicability of the secondary transformation. If sps_lfnst_enabled_flag is 1, it indicates that there is an lfnst_idx[x0][y0] syntax element in the transform unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that there is no lfnst_idx[x0][y0] syntax element in the transform unit syntax.

[0211] iii) CuPredMode[x0][y0] == MODE_INTRA

[0212] The third condition relates to the prediction mode. The secondary transformation is applied only to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].

[0213] iv) IntraSubPartitionsSplitType == ISP_NO_SPLIT

[0214] The fourth condition relates to whether the ISP prediction method is applied. If ISP is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0215] Specifically, as described with reference to FIG. 11, when the current CU is currently divided into a large number of transform units smaller than the CU size, no secondary transform is applied to the divided transform units. At this time, lfnst_idx[x0][y0], which is a syntax element related to the secondary transform, is set to 0 without being parsed. When the current CU is divided into a large number of transform units smaller than the transform tree in terms of CU size, it includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method in which, when intra prediction is applied to the current coding unit, the transform tree is divided into a large number of transform units smaller than the CU size by a preset division method. The ISP prediction mode is indicated at the coding unit level, and based on this, the IntraSubPartitionsSplitType variable is set. At this time, if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction that generates prediction samples at the transform unit level, the prediction accuracy is higher when the transform tree is divided into a large number of transform units than when it is not divided. Therefore, even if no secondary transform is applied to the divided large number of transform units, there is a high possibility that the energy of the residual signal can be efficiently compressed.

[0216] v)!intra_mip_flag[x0][y0]

[0217] The fifth condition relates to the intra prediction method. If MIP (Matrix based Intra Prediction) is not applied to the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0218] Specifically, MIP is used as one method of intra prediction, and whether MIP can be applied is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and the prediction is performed by the product of the restored samples around the current block and a preset matrix. When MIP is applied, since it shows the nature of the residual signal different from the general intra prediction that performs directional or non-directional prediction, the secondary transformation may not be applied to the transform block when MIP is applied.

[0219] vi) numZeroOutSigCoeff == 0

[0220] The sixth condition relates to the significant coefficients existing at specific positions.

[0221] Specifically, if the secondary transformation is applied to the current block, the transform coefficients quantized at the decoder are always 0 at specific positions. Therefore, if there are quantized coefficients that are not 0 at specific positions, it means that the secondary transformation has not been applied, so lfnst_idx[x0][y0] is parsed according to the number of significant coefficients at specific positions. For example, if numZeroOutSigCoeff is not 0, it means that there are significant coefficients at specific positions, so lfnst_idx[x0][y0] is set to 0 without being parsed. On the contrary, if numZeroOutSigCoeff is 0, it means that there are no significant coefficients at specific positions, so lfnst_idx[x0][y0] is parsed.

[0222] If it is indicated at the conversion unit level whether the secondary conversion is applied to the current block based on the first embodiment described above, the residual_coding method described in FIG. 19 is followed. If the lastScanPos of all the conversion blocks included in the conversion unit is smaller than the threshold value or the number of all the conversion blocks is 0 according to Equation 1 for determining lfnLastScanPos described in FIG. 19, lfnLastScanPos is determined to be 1 and lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that the secondary conversion is not applied to the current block. On the contrary, if the LastScanPos of any one of the conversion blocks included in the conversion unit is greater than or equal to the threshold value, lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), vi) described in FIG. 23 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether the secondary conversion is applied to the current block. If the secondary conversion is applied, the conversion kernel used for the secondary conversion is checked / determined.

[0223] If it is indicated at the conversion unit level whether the second-order conversion is to be applied based on the second embodiment described above, the conversion unit syntax structure described in FIG. 23 is applied, and the residual_coding method described in FIG. 20 is used. If the lastScanPos of all the conversion blocks included in the conversion unit is smaller than the threshold value or the number of all the conversion blocks is 0 according to Equation 4 for determining lfnLastScanPos described in FIG. 20, lfnLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that the second-order conversion is not applied to the current block. On the other hand, if the LastScanPos of even one of the conversion blocks included in the conversion unit is greater than or equal to the threshold value, lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vi) described in FIG. 23 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether the second-order conversion is applied to the current block. If the second-order conversion is applied, the decoder checks / determines the conversion kernel used for the second-order conversion.

[0224] FIG. 24 is a diagram showing a method of instructing second-order conversion at the conversion unit level according to another embodiment of the present invention.

[0225] According to the third embodiment described above, instead of the numSigCoeff counter, the position information of the last valid coefficient in the scan order obtained by residual_coding is utilized to parse lfnst_idx[x0][y0].

[0226] Before performing residual coding, the variable lfnLastScanPos, which is related to the position of the last valid coefficient in the scan order, is initialized to 0. The variable lfnLastScanPos is the value obtained by adding the lastScanPos of the transform blocks included in the transform unit. At this time, if lfnLastScanPos is greater than the threshold value and all of the conditions i), ii), iii), iv), v), vii) described in Figure 23 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a secondary transform is applied to the current block. If a secondary transform is applied, the transform kernel used for the secondary transform is checked / decided. On the other hand, if lfnLastScanPos is less than or equal to the threshold value, lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that no secondary transform is applied.

[0227] The threshold value is set according to treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the threshold values are set to Th1, Th2, and Th3 respectively. Th1, Th2, and Th3 are preset integer values greater than or equal to 0, and both the encoder and the decoder use the same values. If treeType is SINGLE_TREE, since it includes both luma and chroma components, the threshold value Th1 may be expressed as the sum of Th2 and Th3.

[0228] If it is indicated at the transform unit level whether a secondary transform is applied based on the above-described third embodiment, the residual coding method described in Figure 22 is used. According to Equation 5 for determining lfnLastScanPos described in Figure 22, the variable lfnLastScanPos is set to the value obtained by adding up all the lastScanPos of the transform blocks included in the transform unit. Then, by comparing lfnLastScanPos with the threshold value, it is determined whether lfnst_idx[x0][y0] can be parsed.

[0229] On the one hand, if secondary conversion is instructed at the transform unit level, there may be a high correlation between the transform units included in the coding unit. This is because the prediction method is determined at the coding unit level. Therefore, lfnst_idx[x0][y0] is signaled only in the first transform unit included in the coding unit, and the signaled lfnst_idx[x0][y0] is shared with the remaining transform units. That is, lfnst_idx[x0][y0] may be parsed using the first to third embodiments described above only when subTuIndex indicating the index of the transform unit is 0. If subTuIndex is greater than 0, the corresponding transform unit does not parse lfnst_idx[x0][y0] and uses the value of lfnst_idx[x0][y0] of the first shared transform unit.

[0230] On the other hand, a counter is used to count the valid coefficients, but whether the decoder parses lfnst_idx[x0][y0] is determined by considering only the valid coefficients present in the left-upper sub-block of the transform block. This is to reduce the amount of computation.

[0231] On the one hand, if secondary conversion is instructed at the transform unit level, the decoder delay time is reduced compared to the case where it is instructed at the coding unit level, but other delay times may occur. For example, even if secondary conversion is instructed at the transform unit level, secondary conversion is instructed after the coding of all the luma transform coefficients, Cb transform coefficients, and Cr transform coefficients is completed. Therefore, even after the coding (processing) of the luma transform coefficients is completed, the inverse transform process for the luma transform coefficients is performed after the coding (processing) of the Cb transform coefficients and the Cr transform coefficients is completed. This results in other delay times of the decoder.

[0232] Hereinafter, in this specification, a method of instructing secondary conversion that can minimize the decoder delay time will be described.

[0233] (Fourth Embodiment)

[0234] As an example of a method for instructing a secondary conversion that can minimize the delay time of a decoder, the secondary conversion is instructed at the conversion unit level, and there is a method of parsing lfnst_idx[x0][y0], which is a syntax element related to the secondary conversion, before coding the luma conversion coefficients. Therefore, the decoder can perform the inverse conversion process on the luma conversion coefficients immediately after the coding of the luma conversion coefficients is completed without waiting for the Cb and Cr conversion coefficients. Similarly, the decoder can perform the inverse conversion process on the Cb conversion coefficients immediately after the coding of the Cb conversion coefficients is completed without waiting for the coding of the Cr conversion coefficients. Such a method for instructing a secondary conversion can minimize the delay time of the decoder and solve the pipeline problem.

[0235] FIG. 25 is a diagram showing a coding unit syntax according to an embodiment of the present invention.

[0236] Referring to FIG. 25, since the secondary conversion is instructed at the conversion unit level, the syntax lfnst_idx[x0][y0] related to the secondary conversion is not parsed at the coding unit level but at the conversion unit level divided by transform_tree.

[0237] FIG. 26 is a diagram showing a method for instructing a secondary conversion at the conversion unit level according to another embodiment of the present invention.

[0238] Referring to FIG. 26, the method of instructing the secondary conversion is instructed at the conversion unit level, and the syntax element lfnst_idx[x0][y0] related to the secondary conversion is parsed first before the luma and chroma conversion coefficient coding (residual_coding). For example, if lfnst_idx[x0][y0] is parsed before obtaining the conversion coefficients, then as soon as the coefficient coding for each color component Y, Cb, Cr is completed, the inverse conversion for the Y, Cb, Cr conversion coefficients is processed immediately. For example, as soon as the coefficient coding for the Y component is completed, the inverse conversion for the luma (Y) conversion coefficients is performed immediately. Similarly, as soon as the coefficient coding (residual_coding) for the Cb component is completed, the inverse conversion for the Cb conversion coefficients is performed immediately, and as soon as the coefficient coding (residual_coding) for the Cr component is completed, the inverse conversion for the Cr conversion coefficients is performed immediately.

[0239] If lfnst_idx[x0][y0] is parsed after the coefficient coding (residual_coding) for Y, Cb, Cr, then even if the coefficient coding (residual_coding) for Y is completed, if the coefficient coding (residual_coding) for Cb, Cr is not completed / processed, the inverse conversion for the Y conversion coefficients is not performed / processed. Therefore, even if the coefficient coding (residual_coding) corresponding to Y is completed, the decoder cannot perform the inverse conversion for the Y conversion coefficients until the coefficient coding (residual_coding) for other components (Cb, Cr) is completed, and there is a problem of unnecessary delay time. However, as described above, if lfnst_idx[x0][y0] is parsed first before the coefficient coding (residual_coding), then after the coefficient coding (residual_coding) for each color component (Y, Cb, Cr) is completed, the inverse conversion for the conversion coefficients of each color component is performed immediately, so that the delay time of the decoder is minimized.

[0240] In the transform_unit() syntax structure, elements such as tu_cbf_luma[x0][y0], tu_cbf_cb[x0][y0], tu_cbf_cr[x0][y0], and transform_skip_flag[x0][y0] are parsed.

[0241] Specifically, tu_cbf_luma[x0][y0] is an element indicating whether the current luma transform block contains one or more non-zero transform coefficients. If tu_cbf_luma[x0][y0] is 1, it indicates that the current luma transform block contains one or more non-zero transform coefficients. If tu_cbf_luma[x0][y0] is 0, it indicates that all the transform coefficients of the current luma transform block are 0. tu_cbf_cb[x0][y0] is an element indicating whether the current chroma Cb transform block contains one or more non-zero transform coefficients. If tu_cbf_cb[x0][y0] is 1, it indicates that the current chroma Cb transform block contains one or more non-zero transform coefficients. If tu_cbf_cb[x0][y0] is 0, it indicates that all the transform coefficients of the current chroma Cb transform block are 0. tu_cbf_cr[x0][y0] is an element indicating whether the current chroma Cr transform block contains one or more non-zero transform coefficients. If tu_cbf_cr[x0][y0] is 1, it indicates that the current chroma Cr transform block contains one or more non-zero transform coefficients. If tu_cbf_cr[x0][y0] is 0, it indicates that all the transform coefficients of the current chroma Cr transform block are 0. transform_skip_flag[x0][y0] is a syntax element related to transform skip. If transform_skip_flag[x0][y0] is 1, it indicates that inverse transform is not applied to the current luma transform block. If transform_skip_flag[x0][y0] is 0, it indicates that whether inverse transform is applied to the current luma transform block is determined by other syntax elements.

[0242] As an example of the method for instructing the second conversion according to FIG. 26, instead of being based on the number of non-zero conversion coefficients, the syntax element lfnst_idx[x0][y0] related to the second conversion is parsed based on the position of the last valid coefficient in the scan order.

[0243] First, the variable lfnLastScanPos is initialized and set to 1. As described with reference to FIG. 23, the variable lfnLastScanPos indicates the position information of the last valid coefficient in the scan order of the conversion blocks included in the current conversion unit. Specifically, if lfnLastScanPos is 1, it indicates that for all conversion blocks included in the conversion unit, the position (scan index) of the last valid coefficient in the scan order is less than the threshold value, or all the conversion coefficients in the block are 0. If lfnLastScanPos is 0, it indicates that there is one or more valid coefficients in one or more conversion blocks included in the conversion unit, and the position (scan index) of the last valid coefficient in the scan order is greater than or equal to the threshold value.

[0244] Next, the variable numZeroOutSigCoeff is initialized and set to 0. If the second conversion is applied to the conversion block, there cannot be a last valid coefficient in the scan order. Therefore, the variable numZeroOutSigCoeff indicates whether there is a valid coefficient at a specific position, and based on this, it is confirmed whether the second conversion is applied or not. For example, assume that if the second conversion is applied to the conversion block, only a maximum of 16 valid coefficients are allowed. For 4×4 and 8×8 sized conversion blocks, valid coefficients may exist in the scan order index [0,7] region (allowing a maximum of 8 non-zero conversion coefficients). On the other hand, for conversion blocks with sizes other than 4×4 and 8×8, valid coefficients may exist in the scan order index [0,15] region (allowing a maximum of 16 non-zero conversion coefficients). Therefore, if the position (scan index) of the last valid coefficient in the scan order exists outside the region where the above-mentioned valid coefficients can exist, the decoder can naturally recognize that the second conversion is not applied to the current conversion block.

[0245] Based on the position (scan index) of the last valid coefficient in the scan order, it is determined whether the syntax element lfnst_idx[x0][y0] regarding the secondary transform can be parsed before coefficient coding (residual_coding). Therefore, the decoder processes information regarding the position of the last valid coefficient in the scan order before coefficient coding (residual_coding).

[0246] Specifically, if the current luma transform block contains one or more valid coefficients (tu_cbf_luma[x0][y0]==1) and transform skip is not applied to the current luma transform block (transform_skip_flag[x0][y0]==0), the syntax structure last_significant_pos regarding the position of the last valid coefficient in the luma scan order is processed.

[0247] If the value of tu_cbf_luma[x0][y0] is 0 (tu_cbf_luma[x0][y0]==0), this indicates that all coefficients of the corresponding transform block are 0, which means that coefficient coding (residual_coding) is not performed. Therefore, there is no need to process the position information of the last valid coefficient in the scan order.

[0248] If transform_skip_flag[x0][y0] is 1, it indicates that inverse transform is not applied to the current luma transform block. Therefore, coefficient coding (residual_coding) is performed without relying on the position information of the last valid coefficient in the scan order.

[0249] If the current chroma Cb transform block includes one or more valid coefficients (tu_cbf_cb[x0][y0] == 1), the syntax structure last_significant_pos, which is related to the position of the last non-zero coefficient in the scan order of the current chroma Cb transform block, is processed. The last_significant_pos syntax structure takes as input the left-upper coordinates (x0, y0) of the transform block, the value obtained by taking the base-2 logarithm of the width of the transform block, the value obtained by taking the base-2 logarithm of the height of the transform block, and a variable cIdx indicating which color component the transform block is. For example, if cIdx is 0, it indicates a luma Y transform block; if cIdx is 1, it indicates a chroma Cb transform block; if cIdx is 2, it indicates a chroma Cr transform block. If the value of tu_cbf_cb[x0][y0] is 0 (tu_cbf_cb[x0][y0] == 0), it indicates that all coefficients of the corresponding transform block are 0. Since this means that coefficient coding (residual_coding) is not performed, there is no need to process the position information of the last non-zero coefficient in the scan order.

[0250] On one hand, if the current luma Cr transform block includes one or more significant coefficients (tu_cbf_cr[x0][y0] == 1), the syntax element tu_joint_cbcr_residual[x0][y0], which indicates whether to represent chroma Cb and Cr with one residual signal before processing last_significant_pos, is parsed. For example, if tu_joint_cbcr_residual[x0][y0] is 1, the coefficient coding for Cr is not processed, and the residual signal for Cr is derived from the restored residual signal of Cb. On the contrary, if tu_joint_cbcr_residual[x0][y0] is 0, the coefficient coding for Cr is performed according to the value of tu_cbf_cr[x0][y0]. If the current chroma Cr transform block includes one or more significant coefficients (tu_cbf_cr[x0][y0] == 1), the syntax structure last_significant_pos, which is related to the position of the last significant coefficient in the scan order of chroma Cr, is processed. If the value of tu_cbf_cbr[x0][y0] is 0 (tu_cbf_cr[x0][y0] == 0), it indicates that all coefficients of the chroma Cr transform block are 0. Since this means that no coefficient coding is performed, there is no need to process the position information of the last non-zero coefficient in the scan order.

[0251] By processing last_significant_pos for each color component, the position (scan index) of the last significant coefficient in the scan order for each color component is obtained, and based on this, the lfnLastScanPos and numZeroOutSigCoeff values are updated.

[0252] And if all of the following conditions i), ii), iii), iv), v), vi), vii) are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0] before coefficient coding.

[0253] Parsing conditions for the lfnst_idx[x0][y0] syntax element before coefficient coding (residual_coding)

[0254] i) Min(lfnWidth, lfnHeight) >= 4

[0255] First, the first condition relates to the block size. If the width and height of the block are each 4 pixels or more, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0256] Specifically, the decoder checks the block size conditions applicable to the second - order transformation. The variables SubWidthC and SubHeightC are set according to the color format, indicating the ratio of the width of the luma component of the picture and the width of the chroma component, and the ratio of the height of the luma component of the picture and the height of the chroma component, respectively. For example, in the case of a 4:2:0 color format video, since it has a structure that includes 1 chroma sample for every 4 luma samples, both SubWidthC and SubHeightC are set to 2. As another example, in the case of a 4:4:4 color format video, since it has a structure that includes 1 chroma sample for every 1 luma sample, both SubWidthC and SubHeightC are set to 1. lfnWidth, which is the number of samples in the horizontal direction of the current block, and lfnHeight, which is the number of samples in the vertical direction of the current block, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, since the conversion unit includes only the chroma component, the number of samples in the horizontal direction of the chroma conversion block is the same as the value obtained by dividing the width tbwidth of the luma conversion block by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma conversion block is the same as the value obtained by dividing the height tbHeight of the luma conversion block by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, since the conversion unit includes the luma component, lfnWidth and lfnHeight are set to tbwidth and tbHeight, respectively. Since the minimum condition for a block applicable to the second - order transformation is 4×4, if Min(lfnWidth, lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is parsed.

[0257] ii) sps_lfnst_enabled_flag == 1

[0258] The second condition relates to the flag value indicating the activation or applicability of the second - order transformation. If the value of the flag ( sps_lfnst_enabled_flag ) indicating the activation or applicability of the second - order transformation is set to 1, the decoder parses lfnst_idx[x0][y0].

[0259] Specifically, the secondary transformation is indicated by the upper-level syntax RBSP. Among the SPS, PPS, VPS, tile group header, and slice header, there is a flag with a size of 1 bit that indicates the activation and applicability of the secondary transformation. If sps_lfnst_enabled_flag is 1, it indicates that there is an lfnst_idx[x0][y0] syntax element in the transform unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that there is no lfnst_idx[x0][y0] syntax element in the transform unit syntax.

[0260] iii) CuPredMode[x0][y0] == MODE_INTRA

[0261] The third condition relates to the prediction mode, and the secondary transformation is only applied to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].

[0262] iv) IntraSubPartitionsSplitType == ISP_NO_SPLIT

[0263] The fourth condition relates to whether the ISP prediction method is applied. If ISP is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0264] Specifically, as described with reference to FIG. 11, when the current CU is currently divided into a number of transform units smaller than the CU size, no secondary transform is applied to the divided transform units. At this time, lfnst_idx[x0][y0], which is a syntax element related to the secondary transform, is set to 0 without being parsed. When the current CU is divided into a number of transform units smaller than the CU size from the transform tree, it includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method in which, when intra prediction is applied to the current coding unit, the transform tree is divided into a number of transform units smaller than the CU size by a preset division method. The ISP prediction mode is indicated at the coding unit level, and based on this, the IntraSubPartitionsSplitType variable is set. If IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction that generates prediction samples at the transform unit level, the accuracy of prediction is higher when the transform tree is divided into a number of transform units than when it is not divided. Therefore, even if no secondary transform is applied to the divided number of transform units, the energy of the residual signal is likely to be efficiently compressed.

[0265] v)!intra_mip_flag[x0][y0]

[0266] The fifth condition relates to the intra prediction method. If MIP (Matrix based Intra Prediction) is not applied to the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.

[0267] Specifically, MIP is used as one method of intra prediction, and whether MIP can be applied is indicated by intra_mip_flag[x0][y0] at the coding unit level. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is applied to the prediction of the current coding unit, and the prediction is performed by the product of the restored samples around the current block and a preset matrix. When MIP is applied, since it shows the nature of the residual signal different from that of general intra prediction that performs directional or non-directional prediction, secondary transformation may not be applied to the transform block when MIP is applied.

[0268] vi) lfnLastScanPos == 0 The sixth condition relates to the last valid coefficient in the scan order of the transform block.

[0269] Specifically, if the position information (scan index) of the last valid coefficient in the scan order of the transform block included in the current transform unit is smaller than a preset threshold, there may be little gain in coding efficiency obtained by secondary transformation. Therefore, in such a case, it is highly likely that the encoder does not apply secondary transformation to the transform block (lfnst_idx[x0][y0] is 0), and thus it is considered that signaling lfnst_idx[x0][y0] by the encoder has a large overhead. Therefore, lfnst_idx[x0][y0] is purged only when the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the preset threshold for at least one of the transform blocks included in the transform unit.

[0270] In other words, as described above, the threshold is an integer greater than or equal to 0. For example, assuming the threshold is 1, the fact that the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold means that the valid coefficient exists at a position other than the upper left corner of the block (scan index 0, DC). At this time, the fact that the position of the last valid coefficient in the scan order of the transform block is equal to or greater than the threshold may be represented by "lfnLastScanPos==".

[0271] vii) numZeroOutSigCoeff == 0

[0272] The seventh condition relates to the valid coefficient present at a specific position.

[0273] Specifically, if a second-order transformation is applied to the current block, no valid coefficient can exist at a specific position on the scan position. That is, the numZeroOutSigCoeff variable indicates whether there is a non-zero transformation coefficient at a specific position. For example, assume that if a second-order transformation is applied to the current block, only a maximum of 16 valid coefficients are allowed. For 4×4 and 8×8 sized transformation blocks, valid coefficients can exist in the scan order index [0,7] region (allowing a maximum of 8 non-zero transformation coefficients). On the other hand, for transformation blocks with sizes other than 4×4 and 8×8, valid coefficients can exist in the scan order index [0,15] region (allowing a maximum of 16 non-zero transformation coefficients). Therefore, if the position (scan index) of the last valid coefficient in the scan order exists outside the region where the above-mentioned valid coefficients can exist, the decoder can naturally recognize that a second-order transformation is not applied to the current block. Thus, if numZeroOutSigCoeff > 0, it means that a second-order transformation is not applied to the current block, so lfnst_idx[x0][y0] is set to 0 without being purged.

[0274] In other words, if numZeroOutSigCoeff is not 0, it means that there is a valid coefficient at a specific position, so lfnst_idx[x0][y0] is set to 0 without being purged. On the contrary, if numZeroOutSigCoeff is 0, it means that there is no valid coefficient at a specific position, so lfnst_idx[x0][y0] is purged.

[0275] If all of the above-mentioned conditions i) to vii) are true, lfnst_idx[x0][y0] is purged, otherwise lfnst_idx[x0][y0] is set to 0 without being purged.

[0276] FIG. 27 is a diagram showing a syntax structure regarding the position of the last significant coefficient in the scan order according to an embodiment of the present invention.

[0277] Referring to FIG. 27, the last_significant_pos syntax structure means a syntax structure including position information of the last significant coefficient in the scan order for each color component Y, Cb, Cr conversion block. The last_significant_pos syntax structure receives, as inputs, (x0, y0) which is the left-upper end coordinate of the conversion block, log2TbWidth which is the logarithm to the base 2 of the width of the conversion block, log2TbHeight which is the logarithm to the base 2 of the height of the conversion block, and cIdx which indicates which color component the conversion block represents. If cIdx is 0, it indicates the luma conversion block, if cIdx is 1, it indicates the chroma Cb conversion block, and if cIdx is 2, it indicates the chroma Cr conversion block.

[0278] In the last_significant_pos syntax structure, the syntax elements related to the position information of the last significant coefficient in the scan order are parsed. Specifically, the syntax elements related to the x - coordinate value and y - coordinate value of the last significant coefficient in the scan order are parsed. At this time, each coordinate value is divided into prefix information and suffix information and indicated. The decoder sets the LastSignificantCoeffX variable, which is the x - coordinate of the last significant coefficient in the scan order, based on the prefix information and suffix information for the x - coordinate. Similarly, the decoder sets the LastSignificantCoeffY variable, which is the y - coordinate of the last significant coefficient in the scan order, based on the prefix information and suffix information for the y - coordinate. As shown in Figure 27, the decoder sets the lastScanPos, which is the scan index of the last significant coefficient in the scan order, based on LastSignificantCoeffX, LastSignificantCoeffY, and DiagScanOrder in a do{}while() structure. Also, based on lastScanPos, the decoder updates the numZeroOutSigCoeff and lfnstLastScanPos, which are variables used in the parsing condition of lfnst_idx[x0][y0], a syntax element related to the second - order transformation.

[0279] If a second-order transform is applied to the current block, there cannot be valid coefficients at specific positions in the scan position. The numZeroOutSigCoeff variable indicates whether there are non-zero transform coefficients at such positions. For example, assume that if a second-order transform is applied to the current block, only a maximum of 16 valid coefficients are allowed. For 4×4 and 8×8 sized transform blocks, valid coefficients can exist in the scan order index [0,7] region (allowing a maximum of 8 non-zero transform coefficients). On the other hand, for transform blocks with sizes other than 4×4 and 8×8, valid coefficients can exist in the scan order index [0,15] region (allowing a maximum of 16 non-zero transform coefficients). Therefore, if the position (scan index) of the last valid coefficient in the scan order exists outside the region where the above-mentioned valid coefficients can exist, the decoder can naturally recognize that the second-order transform is not applied to the current block. The minimum size of a block to which the second-order transform can be applied is 4×4. If transform skip is applied (transform_skip_flag[x0][y0]==1), the second-order transform is not applied. Therefore, for a transform block where the width of the transform block is 4 or more (log2TbWidth>=2), the height of the transform block is 4 or more (log2TbHeight>=2), and transform skip is not applied (transform_skip_flag[x0][y0]==0), numZeroOutSigCoeff is updated. If the second-order transform is applied, for 4×4 and 8×8 sized transform blocks, non-zero transform coefficients can exist only in the scan order index [0,7] region. Therefore, if the transform block is 4×4 or 8×8, ((log2TbWidth==2||log2TbHeight==3)&&(log2TbWidth==log2TbHeight)), and lastScanPos is greater than 7 (lastScanPos>7), numZeroOutSigCoeff increases by 1. For transform blocks other than 4×4 and 8×8 sizes to which the second-order transform can be applied, non-zero transform coefficients can exist only in the scan order index [0,15] region. Therefore, if lastScanPos is greater than 15 (lastScanPos>15), numZeroOutSigCoeff increases by 1.

[0280] The decoder determines lfnstLastScanPos based on lastScanPos. Specifically, if the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in Equation 6 below. In other words, if log2TbWidth >= 2, log2TbHeight >= 2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in Equation 1 below. At this time, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.

[0281]

Number

[0282] As described above, the initial value of lfnstLastScanPos is set to 1.

[0283] In Equation 6, cIdx indicates a variable that represents the color component of the current transform block as described above.

[0284] According to Equation 6, if the lfnstLastScanPos of the line is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 1. On the other hand, if the lfnstLastScanPos of the line is 0 or lastScanPos is greater than or equal to lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 0.

[0285] In other words, if the lastScanPos of all the conversion blocks included in the conversion unit is smaller than the threshold value, or if the number of all the conversion blocks is 0, then lfnstLastScanPos is determined to be 1, and according to the lfnst_idx[x0][y0] parsing condition in FIG. 26, lfnst_idx[x0][y0] is set to 0 without being parsed. This indicates that no second-order conversion is applied to the current block. On the contrary, if the LastScanPos of any one of the conversion blocks included in the conversion unit is greater than or equal to the threshold value, then lfnstLastScanPos is determined to be 0, and if any of the conditions i), ii), iii), iv), v), vii) described in FIG. 26 is satisfied (true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a second-order conversion is applied to the current block. If a second-order conversion is applied to the current block, the conversion kernel used for the second-order conversion is checked / determined.

[0286] lfnstLastScanPosTh[cIdx] in Equation 6 is a preset integer value greater than or equal to 0, and both the encoder and the decoder use the same value. Also, all color components may use the same threshold value. In this case, lfnstLastScanPos is set as shown in Equation 7 below.

[0287] [Number]

[0288] lfnstLastScanPosTh is an integer value of 0 or more set in advance, and both the encoder and the decoder use the same value. For example, lfnstLastScanPosTh may be 1. That is, if lastScanPos is 1 or more, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is parsed. At this time, since the threshold value (lfnstLastScanPosTh) is an integer value, if lastScanPos is 1 or more, it has the same meaning as the case where lastScanPos is greater than 0. Although FIG. 27 described the case where all color components have a threshold value of 1, the present invention is not limited to this.

[0289] FIG. 28 is a diagram showing a residual_coding syntax structure according to an embodiment of the present invention.

[0290] Looking at FIG. 28, the position information of the last valid coefficient in the scan order is indicated by the eye of the coefficient coding (residual_coding). Therefore, the syntax structure of the coefficient coding (residual_coding) may not include the syntax structure regarding the position information of the last valid coefficient in the scan order. For example, the position information of the last valid coefficient in the scan order is a prefix for the x coordinate of the last valid coefficient in the scan order, a suffix, a prefix for the y coordinate, and a suffix. When examining the coefficient coding (residual_coding) syntax structure according to FIG. 28, the coefficient coding (residual_coding) is performed based on LastSignificantCoeffX and LastSignificantCoeffY, which are the x coordinate and y coordinate of the last valid coefficient in the scan order determined before the coefficient coding (residual_coding).

[0291] The method for instructing the second conversion according to the fourth embodiment does not use the numSigCoeff counter. Therefore, even if the coefficient at the (xC, yC) position is a valid coefficient (sig_coeff_flag[xC][yC]==1), numSigCoeff is not updated. In other words, the method for instructing the second conversion according to the fourth embodiment is a method that does not use a counter for valid coefficients. Also, according to the method for instructing the second conversion according to the fourth embodiment, since the numZeroOutSigCoeff variable is set based on lastScanPos, a counter based on sig_coeff_flag may not be used in coefficient coding (residual_coding).

[0292] FIG. 29 is a sequence diagram showing a video signal processing method according to an embodiment of the present invention.

[0293] Hereinafter, a video signal processing method and apparatus based on the embodiments described with reference to FIGS. 15 to 28 will be described.

[0294] The video signal decoding apparatus includes a processor that performs the video signal processing method described in FIG. 29.

[0295] First, the processor receives a bitstream including syntax elements related to the second conversion of a coding unit.

[0296] The processor checks whether one or more preset conditions are satisfied. If one or more preset conditions are satisfied, the processor parses the syntax elements related to the second conversion of the coding unit in S2910 and S2920. On the other hand, if one or more preset conditions are not satisfied, the processor does not parse the syntax elements related to the second conversion of the coding unit in S2930. At this time, the value of the syntax element related to the second conversion is set to 0.

[0297] The syntax element related to the second-order transformation of the coding unit described with reference to FIG. 29 is lfnst_idx[x0][y0], which is a syntax element indicating whether the second-order transformation of the transformation block included in the current coding unit described with reference to FIGS. 5 to 28 is applied or not.

[0298] The processor parses the syntax element related to the second-order transformation of the coding unit via step S2920, and based on the parsed syntax element, performs S2940 to check whether the second-order transformation is applied to the transformation block included in the coding unit.

[0299] At this time, if the second-order transformation is applied to the transformation block, the processor performs an inverse second-order transformation based on one or more coefficients of a first sub-block that is one of one or more sub-blocks constituting the transformation block, and performs S2950 to check one or more inverse transformation coefficients for the first sub-block.

[0300] Then, the processor performs a first-order inverse transformation based on the one or more inverse transformation coefficients obtained in step S2950, and performs S2960 to check residual samples for the transformation block.

[0301] The second-order transformation is a low-band non-separable transformation (LFNST). The transformation block is a block to which a first-order transformation that is separately performed for vertical transformation and horizontal transformation is applied. At this time, the first-order inverse transformation is an inverse transformation for the first-order transformation, and the second-order inverse transformation means an inverse transformation for the second-order transformation.

[0302] The syntax element related to the second-order transformation of the coding unit includes information indicating whether the second-order transformation is applied to the coding unit and information indicating a transformation kernel used for the second-order transformation.

[0303] The first sub-block is the first sub-block in a preset scan order. At this time, the index of the first sub-block is 0.

[0304] Among the one or more preset conditions, the first condition is that the index value indicating the position of the first coefficient among the one or more coefficients of the first sub-block is greater than a preset critical value. At this time, the first coefficient is the last valid coefficient in the preset scan order, and the valid coefficient means a coefficient that is not 0. The preset critical value is 0. The preset scan order is the upper right diagonal scan order described in FIGS. 13 and 14.

[0305] Among the one or more preset conditions, the second condition is that the width and height of the conversion block are 4 pixels or more.

[0306] Among the one or more preset conditions, the third condition is that the conversion skip flag value included in the bit stream is not a specific value. At this time, if the value of the conversion skip flag has a specific value, the conversion skip flag indicates that the first conversion and the second conversion are not applied to the conversion block.

[0307] Among the one or more preset conditions, the fourth condition is that at least one of the one or more coefficients of the first sub-block is not 0, and the at least one or more coefficients are present at a position excluding the first position in the preset scan order. At this time, the first position in the preset scan order means a position where the horizontal and vertical coordinate values are (0, 0) as described above, or the first position in the preset scan order (for example, the upper right diagonal order).

[0308] In addition, the coding unit is composed of a plurality of coding blocks. At this time, if at least one of the conversion blocks corresponding to each of the plurality of coding blocks satisfies one or more of the preset conditions, the syntax elements related to the secondary conversion are parsed.

[0309] On the other hand, when the syntax elements related to the secondary conversion are not parsed or set to 0, or when it is confirmed that the secondary conversion is not applied to the conversion block included in the coding unit in step S2930 or S2940, the processor performs an inverse primary conversion based on one or more coefficients of the conversion block to obtain residual samples for the conversion block (S2970).

[0310] At this time, the above-mentioned inverse primary conversion and inverse secondary conversion are the inverse conversions for the primary conversion and secondary conversion respectively.

[0311] A video signal processing method performed by the video signal decoding device described in FIG. 29, or a method similar thereto, is performed by the video signal encoding device.

[0312] The video signal encoding device includes a processor for encoding a video signal.

[0313] At this time, the processor performs a primary conversion on the residual samples of the blocks included in the coding unit to obtain a plurality of primary conversion coefficients for the blocks. A secondary conversion is performed based on one or more of the plurality of primary conversion coefficients to obtain one or more secondary conversion coefficients for a first sub-block which is one of the sub-blocks constituting the block. Information on the one or more secondary conversion coefficients and syntax elements related to the secondary conversion of the coding unit are encoded to obtain a bitstream.

[0314] The secondary conversion is a low-frequency non-separable transform (LFNST), and the primary conversion may be performed separately for horizontal conversion and vertical conversion respectively.

[0315] Also, the syntax elements related to the secondary transformation are coded if they satisfy one or more preset conditions. The syntax elements related to the secondary transformation include information indicating whether the secondary transformation is applied to the coding unit and information indicating the transformation kernel used for the secondary transformation. At this time, the syntax element related to the secondary transformation is lfnst_idx[x0][y0], which is the syntax element described in FIGS. 15 to 28.

[0316] The first sub-block is the first sub-block in a preset scan order. At this time, the index of the first sub-block is 0.

[0317] Among the one or more preset conditions, the first condition is the case where the index value indicating the position of the first coefficient among the one or more secondary transformation coefficients is greater than a preset threshold value. At this time, the first coefficient is the last valid coefficient in a preset scan order, and the valid coefficient means a coefficient that is not 0. The preset threshold value is 0. The preset scan order is the upper right diagonal scan order described in FIGS. 13 and 14.

[0318] Among the one or more preset conditions, the second condition is that the width and height of the primary transformation block are 4 pixels or more.

[0319] Among the one or more preset conditions, the third condition is the case where the transformation skip flag value included in the bitstream is not a specific value. At this time, if the value of the transformation skip flag has a specific value, the transformation skip flag indicates that the primary transformation and the secondary transformation are not applied to the block.

[0320] Among the one or more preset conditions, the fourth condition is that at least one of the one or more secondary conversion coefficients is not zero and the one or more coefficients are present at positions other than the first position according to the preset scan order. At this time, the first position according to the preset scan order means the position where the horizontal and vertical coordinate values are (0, 0) as described above, or the first position according to the preset scan order (for example, the upper right diagonal order).

[0321] Also, the coding unit is composed of a plurality of coding blocks. At this time, if at least one of the (conversion) blocks included in the coding unit corresponding to each of the plurality of coding blocks satisfies the one or more preset conditions, the syntax element related to the secondary conversion is coded.

[0322] Also, the video signal coding device may include a video signal decoding processor that performs the video signal processing method described with reference to FIG. 29.

[0323] As described above, the bitstream includes the syntax element related to the secondary conversion of the coding unit described with reference to FIGS. 15 to 29. At this time, the bitstream is stored in a non-transitory computer-readable medium. On the other hand, if the one or more preset conditions described above are not satisfied, the video signal coding device does not include the syntax element related to the secondary conversion in the bitstream or sets the syntax element related to the secondary conversion to zero. The bitstream is decoded by the video signal decoding device described with reference to FIG. 29 or coded by the video signal coding device described above.

[0324] A method for encoding such a bitstream includes, for example, performing a first-order transform on residual samples of a block included in a coding unit to obtain a plurality of first-order transform coefficients for the block, performing a second-order transform based on one or more of the plurality of first-order transform coefficients to obtain one or more second-order transform coefficients for a first sub-block which is one of sub-blocks constituting the block, and encoding information on the one or more second-order transform coefficients and syntax elements related to the second-order transform of the coding unit.

[0325] Obtaining the coefficients described in this specification means obtaining pixels / blocks related to the coefficients, and obtaining the residual samples means obtaining residual signals / pixels / blocks related to the residual samples.

[0326] The embodiments of the present invention described above are implemented through various means. For example, the embodiments of the present invention are implemented by hardware, firmware, software, or a combination thereof.

[0327] In the case of implementation by hardware, the method according to the embodiments of the present invention is implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSDPs), Programmable Logic Devices (PDLs), Field Programmable Gate Arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0328] In the case of implementation by firmware or software, the method according to an embodiment of the present invention is implemented in the form of modules, procedures, functions, etc. that perform the functions or operations described above. The software code is stored in a memory and implemented by a processor. The memory is located inside or outside the processor and exchanges data with the processor by various means already known.

[0329] Some embodiments are also implemented in the form of a recording medium including computer-executable instruction words such as program modules executed by a computer. A computer-readable medium is any available medium that can be accessed by a computer, including both volatile and non-volatile media, and both removable and non-removable media. Also, a computer-readable medium includes both a storage medium and a communication medium. A computer storage medium includes both volatile and non-volatile media, and both removable and non-removable media implemented by any method or technology for storing information such as computer-readable instruction words, data structures, program modules, or other data. A communication medium typically includes a modulated data signal such as computer-readable instruction words, data structures, or other data of program modules, or other transmission mechanisms, and includes any information transmission medium.

[0330] The above description of the present invention is for illustrative purposes, and those with ordinary knowledge in the technical field to which the present invention pertains should be able to understand that it can be easily changed to other specific forms without changing the technical idea and essential features of the present invention. Therefore, it should be understood that the above-described embodiments are exemplary in all aspects and not restrictive. For example, each component described as a single type may be implemented in a distributed manner, and similarly, components described as being distributed may also be implemented in a combined form.

[0331] The scope of the present invention is defined by the claims that follow, rather than by the detailed description above. All changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be construed as being included within the scope of the present invention.

Explanation of Signs

[0332] 110 Conversion unit 115 Quantization unit 120 Inverse quantization unit 125 Inverse conversion unit 130 Filtering unit 150 Prediction unit 152 Intra prediction unit 154 Inter prediction unit 154a Motion estimation unit 154b Motion compensation unit 160 Entropy coding unit 210 Entropy decoding unit 220 Inverse quantization unit 225 Inverse conversion unit 230 Filtering unit 250 Prediction unit 252 Intra prediction unit 254 Inter prediction unit

Claims

1. In a video decoding method performed by a device, obtaining syntax information regarding a second transformation of a coding unit; determining, based on the obtained syntax information, whether the second transformation is applied to a transform block included in the coding unit; if the second transformation is applied to the transform block, obtaining one or more inverse transform coefficients based on the second transformation; obtaining residual samples for the transform block based on the one or more inverse transform coefficients, wherein the second transformation is a low frequency non-separable transform (LFNST), the syntax information is obtained if one or more conditions are satisfied, among the one or more conditions, a first condition is that an index of a last valid coefficient is greater than a predetermined value according to a preset scan order in a sub-block; among the one or more conditions, a second condition is that a height and a width of the transform block are greater than or equal to 4; the sub-block is a sub-block having a sub-block index 0 according to a preset scan order in the transform block, the video decoding method.

2. The video decoding method according to claim 1, wherein the predetermined value is 0.

3. The video decoding method according to claim 1, wherein the preset scan order is an up-right diagonal scan order.

4. Indexes of the one or more inverse transform coefficients are determined based on the preset scan order, among the one or more inverse transform coefficients, an index of a first coefficient is 0, the video decoding method according to claim 3, wherein the last valid coefficient is a non-zero coefficient.

5. The video decoding method according to claim 1, wherein the residual samples are obtained by performing an inverse transform of a first transform based on the one or more inverse transform coefficients.

6. In a video encoding device, the video encoding device includes at least one processor, the at least one processor is configured to generate a bitstream using an encoding method, the encoding method includes encoding syntax information regarding a second transformation of a coding unit, Obtaining one or more first transformation coefficients based on residual samples for a transformation block included in the coding unit; If the second transformation is applied to the transformation block according to the syntax information, obtaining one or more second transformation coefficients by using the one or more first transformation coefficients based on the second transformation; Encoding the one or more second transformation coefficients, wherein the second transformation is a low frequency non-separable transform (LFNST); the syntax information is encoded if one or more conditions are satisfied; among the one or more conditions, the first condition is that the index of the last valid coefficient is greater than a predetermined value according to a preset scan order in a sub-block; among the one or more conditions, the second condition is that the height and width of the transformation block are greater than or equal to 4; the sub-block is a sub-block having a sub-block index 0 according to a preset scan order in the transformation block, a video encoding device.

7. The video encoding device according to claim 6, wherein the predetermined value is 0.

8. The video encoding device according to claim 6, wherein the preset scan order is an up-right diagonal scan order.

9. The index of the one or more first transformation coefficients is determined based on the preset scan order; among the one or more first transformation coefficients, the index of the first coefficient is 0; The video encoding device according to claim 8, wherein the last valid coefficient is a non-zero coefficient.

10. The video encoding device according to claim 6, wherein the one or more first transformation coefficients are obtained by performing a first transformation based on the residual samples.

11. In a method of storing a bitstream, the bitstream is generated by an encoding method, the encoding method includes: encoding syntax information related to a second transformation of a coding unit; obtaining one or more first transformation coefficients based on residual samples for a transformation block included in the coding unit; If the secondary transformation is applied to the transformation block according to the syntax information, obtaining one or more secondary transformation coefficients using the one or more primary transformation coefficients based on the secondary transformation; encoding the one or more secondary transformation coefficients, but the secondary transformation is a low frequency non-separable transformation (LFNST), the syntax information is encoded if one or more conditions are satisfied, among the one or more conditions, the first condition is that the index of the last valid coefficient is greater than a predetermined value according to a preset scan order in a sub-block, among the one or more conditions, the second condition is that the height and width of the transformation block are greater than or equal to 4, The sub-block is a sub-block having a sub-block index 0 according to a preset scan order in the transformation block, a bitstream storage method.

Citation Information

Patent Citations

  • Binarizing secondary transform index

    US20170324643A1

  • Method and device for coding residual signal in video coding system

    US20180288409A1

  • Transform method in image coding system and apparatus for same

    US20200177889A1