Video signal processing method and apparatus utilizing quadratic conversion
The implementation of secondary conversion with low-frequency non-separable transformations addresses inefficiencies in existing video signal processing, improving coding efficiency by optimizing the encoding and decoding processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-07-17
- Publication Date
- 2026-04-17
AI Technical Summary
Existing video signal processing methods lack efficiency in coding, particularly in handling spatial and temporal correlations, necessitating improved techniques for more effective compression encoding.
Implementing a video signal processing method that utilizes secondary conversion, specifically low-frequency non-separable transformations (LFNST), which includes parsing syntax elements to apply quadratic inverse conversion on coding units based on predefined conditions, enhancing coding efficiency by using quadratic and primary inverse conversions.
The method improves coding efficiency by optimizing the processing of video signals through secondary conversion, specifically by applying low-frequency non-separable transformations, thereby enhancing the overall performance of video encoding and decoding processes.
Smart Images

Figure 0007847703000008 
Figure 0007847703000009 
Figure 0007847703000010
Abstract
Description
Technical Field
[0001] The present invention relates to a video signal processing method and apparatus, and more particularly, to a video signal processing method and apparatus for encoding or decoding a video signal.
Background Art
[0002] Compression encoding means a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc. In particular, the technique of performing compression encoding on video is called video compression. Compression encoding of a video signal is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, due to the recent development of various media and data transmission media, a more efficient video signal processing method and apparatus are required.
Summary of the Invention
Problems to be Solved by the Invention
[0003] An object of the present invention is to improve the coding efficiency of a video signal.
[0004] The present invention has an object to improve the coding efficiency through secondary conversion.
Means for Solving the Problems
[0005] This specification provides a video signal processing method using secondary conversion.
[0006] More specifically, in a video signal decoding device, a processor is included, and if one or more pre-set conditions are met, the processor parses syntax elements related to the quadratic conversion of the coding unit from the bitstream of the video signal, checks whether the quadratic conversion is applied to the conversion block included in the coding unit based on the parsed syntax elements, and if the quadratic conversion is applied to the conversion block, it performs a quadratic inverse conversion based on one or more coefficients of a first subblock, which is one of the one or more subblocks constituting the conversion block, to obtain one or more inverse conversion coefficients for the first subblock, and performs a primary inverse conversion based on the one or more inverse conversion coefficients to obtain residual samples for the conversion block, wherein the quadratic conversion is a Low Frequency Non-Separable conversion. The transform block is a block to which a separable linear transform is applied, which can be performed separately for vertical and horizontal transformations, and the first of the one or more pre-set conditions is that the index value indicating the position of the first coefficient among the one or more coefficients of the first subblock is greater than a pre-set critical value.
[0007] Furthermore, in this specification, the syntax element is characterized by including information indicating whether or not the quadratic transformation is applied to the coding unit, and information indicating the transformation kernel used for the quadratic transformation.
[0008] Furthermore, in this specification, the first coefficient is the last significant coefficient according to a predetermined scan order, and the significant coefficient is a non-zero coefficient.
[0009] Furthermore, in this specification, the first subblock is characterized in that it is the first subblock in a predetermined scan order.
[0010] Furthermore, in this specification, the second of the one or more pre-set conditions is characterized in that the width and height of the conversion block are 4 pixels or more.
[0011] Furthermore, in this specification, the preset critical value is characterized by being 0.
[0012] Furthermore, in this specification, the pre-set scan order is characterized by being an up-right diagonal scan order.
[0013] Furthermore, in this specification, the third condition among the one or more pre-set conditions is that the conversion skip flag value included in the bitstream is not a specific value, and if the conversion skip flag value is the specific value, the conversion skip flag indicates that the primary conversion and the secondary conversion are not applied to the conversion block.
[0014] Furthermore, in this specification, the fourth condition among the one or more pre-set conditions is characterized in that at least one of the one or more coefficients of the first subblock is not 0, and at least one of the coefficients is located in a position other than the first position according to the pre-set scan order.
[0015] Furthermore, in this specification, the coding unit is composed of a number of coding blocks, and if at least one of the transformation blocks corresponding to each of the plurality of coding blocks satisfies one or more of the pre-set conditions, the syntax elements relating to the quadratic transformation are purged.
[0016] Furthermore, in this specification, a video signal decoding device includes a processor, the processor performs a linear transformation on residual samples of a block included in a coding unit to obtain a plurality of linear transformation coefficients for the block, performs a quadratic transformation based on one or more of the plurality of linear transformation coefficients to obtain one or more quadratic transformation coefficients for a first subblock which is one of the subblocks constituting the block, encodes information regarding the one or more quadratic transformation coefficients and syntax elements relating to the quadratic transformation of the coding unit to obtain a bitstream, wherein the quadratic transformation is a low-bandwidth non-separated transformation (LFNST), the linear transformation can be separated into a vertical transformation and a horizontal transformation, the syntax elements relating to the quadratic transformation of the coding unit are encoded if one or more pre-set conditions are satisfied, and the first of the one or more pre-set conditions is that the index value indicating the position of the first coefficient among the one or more quadratic transformation coefficients is greater than a pre-set critical value.
[0017] Furthermore, in this specification, the syntax element is characterized by including information indicating whether or not the quadratic transformation is applied to the coding unit, and information indicating the transformation kernel used for the quadratic transformation.
[0018] Furthermore, in this specification, the first coefficient is the last effective coefficient according to a predetermined scan order, and the effective coefficient is a coefficient that is not zero.
[0019] Furthermore, in this specification, the first subblock is characterized in that it is the first subblock in a predetermined scan order.
[0020] Furthermore, in this specification, the second condition among the one or more pre-set conditions is that the width and height of the primary conversion block are 4 pixels or more.
[0021] Furthermore, in this specification, the preset critical value is characterized by being 0.
[0022] Furthermore, in this specification, the pre-set scan order is characterized by being the diagonal scan order on the upper right side.
[0023] Furthermore, in this specification, the third condition among the one or more pre-set conditions is that the conversion skip flag value included in the bitstream is not a specific value, and if the conversion skip flag value is the specific value, the conversion skip flag indicates that the primary conversion and the secondary conversion are not applied to the block.
[0024] Furthermore, in this specification, the fourth condition among the one or more pre-set conditions is characterized in that at least one of the one or more quadratic conversion coefficients is not 0, and the one or more coefficients are located in positions other than the first position according to the pre-set scan order.
[0025] Also, in this specification, in a non-transitory computer-readable medium storing a bitstream, the bitstream includes: a step of performing a primary transformation on residual samples of a block included in a coding unit to obtain a plurality of primary transformation coefficients for the block; a step of performing a secondary transformation based on any one or more of the plurality of primary transformation coefficients to obtain one or more secondary transformation coefficients for a first sub-block, which is one of the sub-blocks constituting the block; and a step of encoding information regarding the one or more secondary transformation coefficients and syntax elements regarding the secondary transformation of the coding unit. The secondary transformation is a low-frequency non-separable transformation (LFNST), the primary transformation can be separately performed as a vertical transformation and a horizontal transformation, the syntax elements regarding the secondary transformation are encoded if they satisfy one or more preset conditions, and a first condition among the one or more preset conditions is that an index value indicating the position of a first coefficient among the one or more secondary transformation coefficients is greater than a preset threshold value.
Advantages of the Invention
[0026] One embodiment of the present invention provides a video signal processing method using a secondary transformation and an apparatus therefor.
Brief Description of the Drawings
[0027] [Figure 1] It is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention. [Figure 2] It is a schematic block diagram of a video signal decoding apparatus according to an embodiment of the present invention. [Figure 3] It is a diagram showing an example in which a coding tree unit is divided into coding units within a picture. [Figure 4] It is a diagram showing an example of a method of signaling the division of a quad tree and a multi-type tree. [Figure 5] This figure provides a more detailed illustration of the intra-prediction method according to an embodiment of the present invention. [Figure 6] This figure provides a more detailed illustration of the intra-prediction method according to an embodiment of the present invention. [Figure 7] This diagram illustrates in detail how an encoder converts residual signals. [Figure 8] This diagram illustrates in detail how encoders and decoders inversely transform conversion coefficients to obtain a residual signal. [Figure 9] This figure shows basis functions for multiple transformation kernels usable in linear transformations. [Figure 10] This is a block diagram showing the process of restoring a residual signal using a decoder that performs a quadratic conversion according to one embodiment of the present invention. [Figure 11] This figure shows, at a block level, the process of restoring the residual signal using a decoder that performs a quadratic conversion according to one embodiment of the present invention. [Figure 12] This figure shows a method for applying a quadratic transformation to transfer a reduced number of samples according to one embodiment of the present invention. [Figure 13] This figure shows a method for determining the scan order of the upper right diagonal according to one embodiment of the present invention. [Figure 14] This figure shows the upper right diagonal scan order according to one embodiment of the present invention, indicated by the block size. [Figure 15] This diagram shows how to instruct a quadratic transformation at the coding unit level. [Figure 16] This figure shows the residual_coding syntax structure according to one embodiment of the present invention. [Figure 17] This figure shows a method for instructing a quadratic conversion at the coding unit level according to one embodiment of the present invention. [Figure 18] This figure shows a method for instructing a quadratic conversion at the coding unit level according to one embodiment of the present invention. [Figure 19] This figure shows the residual_coding syntax structure according to one embodiment of the present invention. [Figure 20] This figure shows the residual_coding syntax structure according to another embodiment of the present invention. [Figure 21] This figure shows a method for instructing a quadratic transformation at the coding unit level according to another embodiment of the present invention. [Figure 22] This figure shows the residual_coding syntax structure according to another embodiment of the present invention. [Figure 23] This figure shows a method for instructing a secondary conversion at the conversion unit level according to an embodiment of the present invention. [Figure 24] This figure shows a method for instructing a secondary conversion at the conversion unit level according to another embodiment of the present invention. [Figure 25] This figure shows a coding unit syntax according to one embodiment of the present invention. [Figure 26] This figure shows a method for instructing a secondary conversion at the conversion unit level according to another embodiment of the present invention. [Figure 27] This figure shows the syntax structure relating to the position of the last effective coefficient in the scan sequence according to an embodiment of the present invention. [Figure 28] This figure shows the residual_coding syntax structure according to another embodiment of the present invention. [Figure 29] This is a sequence diagram showing a video signal processing method according to an embodiment of the present invention. [Modes for carrying out the invention]
[0028] The terminology used herein has been selected as widely used and general terms as possible, taking into account the function of the present invention; however, this may vary depending on the intent of the articulators, conventions, or the emergence of new technologies. In addition, in certain cases, the applicant has arbitrarily selected some terms, in which case their meaning will be described in the section describing the form of implementation of the invention. Therefore, it is important to clarify that the terminology used herein is not merely a set of names, but should be interpreted based on the substantive meaning of the term and the overall content of this specification.
[0029] In this specification, some terms are interpreted as follows: Coding may be interpreted as encoding or decoding. In this specification, a device that encodes a video signal to generate a video signal bitstream is referred to as an encoding device or encoder, and a device that decodes a video signal bitstream to restore a video signal is referred to as a decoding device or decoder. In this specification, the term video signal processing device is used as a conceptual term that includes both encoders and decoders. Information is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, so the present invention is not limited thereto. "Unit" is used to mean a basic unit of image processing or a specific location in a picture, and refers to an image region that includes at least one of the luminance (luma) component and the chroma component. "Block" refers to an image region that includes a specific component among the luminance component and the chroma component (i.e., Cb and Cr). However, in some embodiments, terms such as "unit," "block," "partition," and "region" may be used interchangeably. Furthermore, in this specification, the term "unit" is used as a concept that includes coding units, prediction units, and transformation units. "Picture" refers to a field or frame, and these terms are used interchangeably depending on the embodiment.
[0030] Figure 1 is a schematic block diagram of a video signal encoding device 100 according to one embodiment of the present invention. Referring to Figure 1, the encoding device 100 according to this specification includes a conversion unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse conversion unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.
[0031] The conversion unit 110 converts the residual signal, which is the difference between the input video signal and the predicted signal generated by the prediction unit 150, to obtain the converted coefficient values. For example, discrete cosine transform (DCT), discrete sine transform (DST), or wavelet transform may be used. Discrete cosine transform and discrete sine transform divide the input picture signal into block form and perform the transformation. In the transformation, the coding efficiency may differ depending on the distribution and characteristics of the values within the transformation domain. The quantization unit 115 quantizes the values of the conversion coefficients output in the conversion unit 110.
[0032] To improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using a pre-coded region via the prediction unit 150, and the restored picture is obtained by adding the residual value between the original picture and the predicted picture to the predicted picture. To prevent mismatches between the encoder and decoder, the encoder should use information that is also available to the decoder when performing prediction. For this purpose, the encoder performs a further process of restoring the currently encoded block. The inverse quantization unit 120 inversely quantizes the conversion coefficient values, and the inverse transformation unit 125 restores the residual value using the inversely quantized conversion coefficient values. Meanwhile, the filtering unit 130 performs filtering operations to improve the quality of the restored picture and enhance coding efficiency. For example, this may include a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter. The filtered picture is stored in the Decoded Picture Buffer (DPB) 156 for output or use as a reference picture.
[0033] To improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using a pre-coded region via the prediction unit 150, and the restored picture is obtained by adding the residual value between the original picture and the predicted picture to the predicted picture. The intra-prediction unit 152 performs intra-prediction within the current picture, and the inter-prediction unit 154 predicts the current picture using the reference buffer stored in the decoded picture buffer 156. The intra-prediction unit 152 performs intra-prediction from the restored region within the current picture and transmits the intra-coded information to the entropy coding unit 160. The inter-prediction unit 154 again includes the motion estimation unit 154a and the motion compensation unit 154b. The motion estimation unit 154a obtains the motion vector value of the current region by referring to a specific restored region. The motion estimation unit 154a transmits position information of the reference region (reference frame, motion vector, etc.) to the entropy coding unit 160 so that it is included in the bitstream. The motion compensation unit 154b performs intermotion compensation using the motion vector values transmitted from the motion estimation unit 154a.
[0034] The prediction unit 150 includes an intra-prediction unit 152 and an inter-prediction unit 154. The intra-prediction unit 152 performs intra-prediction within the current picture, while the inter-prediction unit 154 performs inter-prediction, predicting the current picture using a reference buffer stored in the decoded picture buffer 156. The intra-prediction unit 152 performs intra-prediction from the restored samples in the current picture and transmits intra-coded information to the entropy coding unit 160. The intra-coded information includes at least one of the following: intra-prediction mode, MPM (Most Probable Mode) flag, and MPM index. The intra-coded information includes information about the reference sample. The inter-prediction unit 154 is configured to include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains motion vector values for the current region by referring to a specific region of the restored reference signal picture. The motion estimation unit 154a transmits a set of motion information (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector values transmitted from the motion compensation unit 154a. The inter-prediction unit 154 transmits inter-coded information, including motion information for the reference region, to the entropy coding unit 160.
[0035] In a further embodiment, the prediction unit 150 includes an intrablock copy (BC) prediction unit (not shown). The intraBC prediction unit performs intraBC prediction from the restored samples in the current picture and transmits the intraBC encoded information to the entropy coding unit 160. The intraBC prediction unit obtains a block vector value indicating a reference region used for prediction of the current region by referring to a specific region in the current picture. The intraBC prediction unit performs intraBC prediction using the obtained block vector value. The intraBC prediction unit transmits the intraBC encoded information to the entropy coding unit 160. The intraBC prediction unit includes block vector information.
[0036] Once the picture prediction described above is performed, the conversion unit 110 converts the residual values between the original picture and the predicted picture to obtain conversion coefficient values. In this case, the conversion is performed in units of specific blocks within the picture, but the size of the specific block is variable within a preset range. The quantization unit 115 quantizes the values of the conversion coefficients generated by the conversion unit 110 and transmits them to the entropy coding unit 160.
[0037] The entropy coding unit 160 generates a video signal bitstream by entropy coding information indicating quantized conversion coefficients, intra-coded information, and inter-coded information. The entropy coding unit 160 uses methods such as variable length coding (VLC) and arithmetic coding. Variable length coding (VLC) converts input symbols into a sequence of codewords, but the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and less frequently occurring symbols are represented by long codewords. Context-based Adaptive Variable Length Coding (CAVLC) is used as the variable length coding method. Arithmetic coding converts a sequence of data symbols into a single prime number, but arithmetic coding obtains the optimal number of prime bits necessary to represent each symbol. Context-based Adaptive Binary Arithmetic Coding (CABAC) is used as the arithmetic coding method. For example, the entropy coding unit 160 binary-codes the information indicating the quantized conversion coefficients. The entropy coding unit 160 then arithmetically codes the binary-coded information to generate a bitstream.
[0038] The generated bitstream is encapsulated in Network Abstraction Layer (NAL) units as its basic units. Each NAL unit contains an encoded integer number of coding tree units. To decode the bitstream with a video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded. Meanwhile, the information necessary for decoding the video signal bitstream is transmitted via Raw Byte Sequence Payloads (RBSPs) of higher-level sets such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS).
[0039] On the other hand, the block diagram in Figure 1 shows an encoding device 100 according to one embodiment of the present invention, and the separated blocks show the elements of the encoding device 100 in a logically distinguishable manner. Therefore, the elements of the encoding device 100 described above are mounted on one chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the encoding device 100 described above is performed by a processor (not shown).
[0040] Figure 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to Figure 2, the decoding apparatus 200 according to this specification includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 225, a filtering unit 230, and a prediction unit 250.
[0041] The entropy decoding unit 210 entropy codes the video signal bitstream and extracts conversion coefficient information, intra-encoded information, inter-encoded information, etc., for each region. For example, the entropy decoding unit 210 obtains a binary code for the conversion coefficient information of a specific region from the video signal bitstream. The entropy decoding unit 210 also performs inverse binary coding of the binary code to obtain quantized conversion coefficients. The inverse quantization unit 220 inverse quantizes the quantized conversion coefficients, and the inverse conversion unit 225 uses the inverse quantized conversion coefficients to reconstruct the residual value. The video signal processing device 200 adds the residual value obtained from the inverse conversion unit 225 with the predicted value obtained from the prediction unit 250 to reconstruct the original pixel value.
[0042] Meanwhile, the filtering unit 230 improves image quality by filtering the picture. This includes a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is either output or stored in the decoded picture buffer (DPB) 256 to be used as a reference picture for the next picture.
[0043] The prediction unit 250 includes an intra-prediction unit 252 and an inter-prediction unit 254. The prediction unit 250 generates a prediction picture by utilizing the encoding type decoded via the entropy decoding unit 210 described above, the conversion coefficients for each region, intra / inter-encoded information, etc. To reconstruct the current block being decoded, the current picture containing the current block or the region from which another picture has been decoded is used. A picture (or tile / slice) that uses only the current picture for reconstruction, i.e., performs intra-prediction or intra-BC prediction, is called an intra-picture or I-picture (or tile / slice), and a picture (or tile / slice) that performs intra-prediction, inter-prediction, and intra-BC prediction is called an inter-picture (or tile / slice). A picture (or tile / slice) that uses up to one motion vector and a reference picture index to predict the sample value of each block in an interpicture (or tile / slice) is called a predictive picture or P-picture (or tile / slice), while a picture (or tile / slice) that uses up to two motion vectors and a reference picture index is called a bi-predictive picture or B-picture (or tile / slice). In other words, a P-picture (or tile / slice) uses up to one motion information set to predict each block, and a B-picture (or tile / slice) uses up to two motion information sets to predict each block. Here, a motion information set contains one or more motion vectors and one reference picture index.
[0044] The intra-prediction unit 252 generates a prediction block using intra-encoded information and the recovered sample in the current picture. As described above, the intra-encoded information includes at least one of the intra-prediction mode, the MPM (MOST Probable Mode) flag, and the MPM index. The intra-prediction unit 252 predicts the sample value of the current block using the recovered sample located to the left and / or above the current block as a reference sample. In this disclosure, the recovered sample, the reference sample, and the sample of the current block represent pixels. The sample value also represents the pixel value.
[0045] In one embodiment, the reference sample is a sample included in the surrounding blocks of the current block. For example, the reference sample is a sample adjacent to the left boundary and / or the upper boundary of the current block. Alternatively, the reference sample is a sample located in the surrounding blocks of the current block that lies on a line within a predetermined distance from the left boundary and / or on a line within a predetermined distance from the upper boundary of the current block. In this case, the surrounding blocks of the current block include at least one of the following blocks adjacent to the current block: the left (L) block, the upper (A) block, the lower left (BL) block, the upper right (AR) block, or the upper left (AL) block.
[0046] The interprediction unit 254 generates a prediction block using the reference picture and intercoded information stored in the decoded picture buffer 256. The intercoded information includes a set of motion information for the current block relative to the reference block (reference picture index, motion vector, etc.). Interpretation includes L0 prediction, L1 prediction, and bi-prediction. L0 prediction is a prediction that uses one reference picture included in the L0 picture list, and L1 prediction means a prediction that uses one reference picture included in the L1 picture list. For this, one set of motion information (e.g., motion vector and reference picture index) is required. In the bi-prediction method, up to two reference regions are used, but these two reference regions may reside in the same reference picture or in different pictures. In other words, in the bi-prediction method, up to two sets of motion information (e.g., motion vector and reference picture index) are used, but the two motion vectors may correspond to the same reference picture index or to different reference picture indices. In this case, the reference picture is displayed (or output) either before or after the current picture in terms of time. In one embodiment, the dual prediction method uses two reference regions selected from the L0 picture list and the L1 picture list, respectively.
[0047] The interpretation unit 254 obtains the current reference block using the motion vector and the reference picture index. The reference block resides within the reference picture corresponding to the reference picture index. The sample value of the block identified by the motion vector, or an interpolated value thereof, is used as the predictor for the current block. For motion prediction with sub-pel pixel accuracy, for example, an 8-tab interpolation filter is used for the luminance signal and a 4-tab interpolation filter is used for the chrominance signal. However, the interpolation filter for sub-pel motion prediction is not limited to these. In this way, the interpretation unit 254 performs motion compensation, predicting the texture of the current unit from the previously restored picture. In this process, the interpretation unit utilizes a motion information set.
[0048] In a further embodiment, the prediction unit 250 includes an intra-BC prediction unit (not shown). The intra-BC prediction unit reconstructs the current region by referring to a specific region containing the reconstructed sample in the current picture. The intra-BC prediction unit obtains intra-BC encoded information for the current region from the entropy decoding unit 210. The intra-BC prediction unit obtains a block vector value of the current region that points to a specific region in the current picture. The intra-BC prediction unit performs intra-BC prediction using the obtained block vector value. The intra-BC prediction unit includes block vector information.
[0049] A restored video picture is generated by adding the predicted value output from the intra-prediction unit 252 or the inter-prediction unit 254 and the residual value output from the inverse conversion unit 225. In other words, the video signal decoding device 200 restores the current block using the predicted block generated from the prediction unit 250 and the residual value obtained from the inverse conversion unit 225.
[0050] On the other hand, the block diagram in Figure 2 shows a decoding device 200 according to one embodiment of the present invention, and the separated blocks show the elements of the decoding device 200 in a logically distinguishable manner. Thus, the elements of the decoding device 200 described above are mounted on one chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the decoding device 200 described above is performed by a processor (not shown).
[0051] Figure 3 shows an example in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. In the coding process of a video signal, the picture is divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit consists of two blocks: an NXN block of luminance samples and its corresponding chrominance samples. A Coding Tree Unit is divided into multiple Coding Units. A Coding Tree Unit may also become a leaf node without being divided. In this case, the Coding Tree Unit itself can become a Coding Unit. A Coding Unit refers to a basic unit for processing a picture in the video signal processing processes described above, i.e., intra / inter prediction, transformation, quantization, and / or entropy coding. Within a single picture, the size and pattern of the Coding Units are not constant. Coding Units have a square or rectangular pattern. A rectangular Coding Unit (or rectangular block) includes vertical Coding Units (or vertical blocks) and horizontal Coding Units (or horizontal blocks). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Furthermore, in this specification, a non-square block refers to a rectangular block, but the present invention is not limited to this.
[0052] Referring to Figure 3, the coding tree unit is first divided into a quad tree (QT) structure. That is, in the quad tree structure, one node with a size of 2N × 2N is divided into four nodes with a size of N × N. In this specification, the quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and it is not necessary for all nodes to be divided to the same depth.
[0053] On the other hand, the leaf nodes of the quad tree described above are further divided into a multi-type tree (MTT) structure. According to embodiments of the present invention, in a multi-type tree structure, one node is divided into a binary or ternary tree structure with horizontal or vertical division. In other words, there are four division structures in a multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to embodiments of the present invention, in each of the tree structures, the width and height of the node are both powers of 2. For example, in a binary tree (BT) structure, a node of size 2N×2N is divided into two N×2N nodes by vertical binary division and into two 2N×N nodes by horizontal binary division. Furthermore, in a Ternary Tree (TT) structure, a node of size 2N×2N is divided into (N / 2)×2N, N×2N, and (N / 2)×2N nodes by vertical ternary decomposition, and into 2N×(N / 2), 2N×N, and 2N×(N / 2) nodes by horizontal ternary decomposition. Such multi-type tree decomposition is performed recursively.
[0054] Leaf nodes in a multi-type tree can become coding units. If a coding unit is not larger than the maximum transformation length, it can be used as a unit of prediction and / or transformation without being further subdivided. In one embodiment, if the width or height of a current coding unit is larger than the maximum transformation length, the current coding unit is subdivided into multiple transformation units without explicit signaling regarding subdivision. On the other hand, in the quad tree and multi-type tree described above, at least one of the following parameters is either predefined or transmitted via RBSPs of a higher-level set such as PPS, SPS, VPS, etc. 1) CTU size: size of the root node of the quad tree, 2) MinQtSize: minimum allowed size of QT leaf nodes, 3) MaxBtSize: maximum allowed size of BT root nodes, 4) MaxTtSize: maximum allowed size of TT root nodes, 5) MaxMttDepth: maximum allowed depth of MTT splitting from QT leaf nodes, 6) MinBtSize: minimum allowed size of BT leaf nodes, 7) MinTT size: minimum allowed size of TT leaf nodes.
[0055] Figure 4 shows one example of a method for signaling the splitting of quad trees and multi-type trees. Pre-configured flags are used to signal the splitting of existing quad trees and multi-type trees. Referring to Figure 4, at least one of the following flags is used: "split_cu_flag" which indicates whether a node can be split, "split_qt_flag" which indicates whether a quad tree node can be split, "mtt_split_cu_vertical_flag" which indicates the splitting direction of a multi-type tree node, or "mtt_split_binarycu_flag" which indicates the splitting pattern of a multi-type tree node.
[0056] According to an embodiment of the present invention, the flag "split_cu_flag", which indicates whether the current node can be split or not, is signaled first. If the value of "split_cu_flag" is 0, it indicates that the current node will not be split, and the current node will become a coding unit. If the current node is a coding tree unit, the coding tree unit will contain one coding unit that will not be split. If the current node is a quad tree node "QT node", the current node will be a leaf node "QT leaf node" of the quad tree node and will become a coding unit. If the current node is a multi-type tree node "MTT node", the current node will be a leaf node "MTT leaf node" of the multi-type tree and will become a coding unit.
[0057] If the value of "split_cu_flag" is 1, the current node is split into a quad tree or a multi-type tree node according to the value of "split_qt_flag". A coding tree unit is the root node of a quad tree and is preferentially split into a quad tree structure. In a quad tree structure, "split_qt_flag" is signaled for each node "QT node". If the value of "split_qt_flag" is 1, the node is split into four square nodes, and if the value of "qt_split_flag" is 0, the node becomes a leaf node of the quad tree "QT leaf node", and the node is split into a multi-type node. According to embodiments of the present invention, quad tree splitting can be restricted depending on the type of the current node. If the current node is a coding tree unit (root node of a quad tree) or a quad tree node, quad tree splitting is allowed, but if the current node is a multi-type tree unit, quad tree splitting is not allowed. Each quad tree leaf node, "QT leaf node," is further divided into a multi-type tree structure. As mentioned above, if "split_qt_flag" is 0, the node is currently divided into a multi-type node. "mtt_split_cu_vertical_flag" and "mtt_split_cu_binary_flag" are signaled to indicate the direction and pattern of the division. If the value of "mtt_split_cu_vertical_flag" is 1, a vertical division of the node "MTT node" is indicated, and if the value of "mtt_split_cu_vertical_flag" is 0, a horizontal division of the node "MTT node" is indicated. Also, if the value of "mtt_split_cu_binary_flag" is 1, the node "MTT node" is divided into two rectangular nodes, and if the value of "mtt_split_cu_binary_flag" is 0, the node "MTT node" is divided into three rectangular nodes.
[0058] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes in the coding unit tree). The basic unit for performing such predictions is referred to below as a prediction unit or prediction block.
[0059] Hereinafter, the term "unit" as used herein is used as a substitute for the prediction unit, which is the basic unit for making predictions. However, the present invention is not limited thereto and is understood in a broader sense as a concept that includes the coding unit.
[0060] Figures 5 and 6 illustrate in more detail the intra-prediction method according to an embodiment of the present invention. As described above, the intra-prediction unit predicts the sample value of the current block by using the restored sample located to the left and / or above the current block as a reference sample.
[0061] First, Figure 5 shows an example of a reference sample used to predict the current block in intra-prediction mode. In this example, the reference sample is a sample adjacent to the left boundary and / or the upper boundary of the current block. As shown in Figure 5, if the size of the current block is W×H and a single reference line sample adjacent to the current block is used for intra-prediction, the reference sample is set using up to 2W+2H+1 peripheral samples located to the left and / or above the current block.
[0062] Furthermore, if at least some of the samples to be used as reference samples have not yet been recovered, the intra-prediction unit performs a reference sample padding process to acquire reference samples. The intra-prediction unit also performs a reference sample filtering process to reduce the error of the intra-prediction. That is, it filters the peripheral samples and / or the reference samples acquired through the reference sample padding process to acquire filtered reference samples. The intra-prediction unit uses the reference samples thus acquired to predict the samples of the current block. The intra-prediction unit uses the unfiltered or filtered reference samples to predict the samples of the current block. In this disclosure, peripheral samples include samples on at least one reference line. For example, peripheral samples may include adjacent samples on lines adjacent to the boundary of the current block.
[0063] Next, Figure 6 shows an example of a prediction mode used for intra-prediction. For intra-prediction, intra-prediction mode information indicating the direction of intra-prediction is signaled. The intra-prediction mode indicates one of several intra-prediction modes that make up the intra-prediction mode set. If the current block is an intra-prediction block, the decoder receives the intra-prediction mode information for the current block from the bitstream. The intra-prediction unit of the decoder performs intra-prediction for the current block based on the extracted intra-prediction mode information.
[0064] According to an embodiment of the present invention, the intra-prediction mode set includes all intra-prediction modes used for intra-prediction (e.g., a total of 67 intra-prediction modes). More specifically, the intra-prediction mode set includes planar modes, DC modes, and a plurality of (e.g., 65) angular modes (i.e., directional modes). Each intra-prediction mode is indicated by a preset index (i.e., an intra-prediction mode index). For example, as shown in Figure 6, intra-prediction mode index 0 indicates the planar mode, and intra-prediction mode index 1 indicates the DC mode. Intra-prediction mode indices 2 through 66 each indicate different angular modes. Each angular mode indicates different angles within a preset angular range. For example, an angular mode indicates angles within an angular range between 45 degrees and -135 degrees clockwise (i.e., a first angular range). The angular modes are defined relative to the directional range. In this case, intra-prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra-prediction mode index 18 indicates horizontal (HOR) mode, intra-prediction mode index 34 indicates diagonal (DIA) mode, intra-prediction mode index 50 indicates vertical (VER) mode, and intra-prediction mode index 66 indicates vertical diagonal (VDIA) mode.
[0065] On the other hand, the pre-set angle ranges are set to be different from each other depending on the pattern of the current block. For example, if the current block is a rectangular block, a wide-angle mode is used to specify an angle greater than 45 degrees clockwise or less than -135 degrees. If the current block is a horizontal block, the angle mode specifies an angle within the angle range between (45 + offset1) degrees and (-135 + offset1) degrees clockwise (i.e., the second angle range). In this case, angle modes 67 to 76, which deviate from the first angle range, are also used. Furthermore, if the current block is a waterline block, the angle mode specifies an angle within the angle range between (45 - offset2) degrees and (-135 - offset2) degrees clockwise (i.e., the third angle range). In this case, angle modes -10 to -1, which deviate from the first angle range, are also used. According to embodiments of the present invention, the values of offset1 and offset2 are determined to be different from each other based on the ratio between the width and height of the rectangular block. Also, offset1 and offset2 are positive numbers.
[0066] According to a further embodiment of the present invention, the set of intra-predictive modes comprises a base angle mode and an extended angle mode, wherein the extended angle mode is determined based on the base angle mode.
[0067] According to one embodiment, the basic angle mode corresponds to the angle used in the intra-prediction of the conventional HEVC (High Efficiency Video Coding) standard, and the extended angle mode corresponds to the angle newly added in the intra-prediction of the next-generation video codec standard. More specifically, the basic angle mode is an angle mode corresponding to one of the intra-prediction modes {2, 4, 6, ..., 66}, and the extended angle mode is an angle mode corresponding to one of the intra-prediction modes {3, 5, 6, ..., 65}. In other words, the extended angle mode is an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode is determined based on the angle indicated by the basic angle mode.
[0068] In other embodiments, the basic angle mode is a mode corresponding to an angle within a preset first angle range, and the extended angle mode is a wide-angle mode that deviates from the first angle range. That is, the basic angle mode is an angle mode corresponding to one of the intra-prediction modes {2, 3, 4, ..., 66}, and the extended angle mode is an angle mode corresponding to one of the intra-prediction modes {-10, -9, ..., -1} and {67, 68, ..., 76}. The angle indicated by the extended angle mode is determined to be the angle opposite to the angle indicated by the corresponding basic angle mode. Thus, the angle indicated by the extended angle mode is determined based on the angle indicated by the basic angle mode. On the other hand, the number of extended angle modes is not limited to this, and further extended angles are defined by the size and / or pattern of the block. For example, the extended angle mode may be defined as an angle mode corresponding to one of the intra-prediction modes {-14, -13, ..., -1} and {67, 68, ..., 80}. On the other hand, the total number of intra-prediction modes included in the intra-prediction mode set varies depending on the configuration of the basic angle mode and extended angle mode described above.
[0069] In the above embodiment, the intervals between extended angle modes are set based on the intervals between the corresponding basic angle modes. For example, the intervals between extended angle modes {3, 5, 7, ..., 65} are determined based on the intervals between the corresponding basic angle modes {2, 4, 6, ..., 66}. Also, the intervals between extended angle modes {-10, -9, ..., -1} are determined based on the intervals between the corresponding opposite basic angle modes {56, 57, ..., 65}, and the intervals between extended angle modes {67, 68, ..., 76} are determined based on the intervals between the corresponding opposite basic angle modes {3, 4, ..., 12}. The angle intervals between extended angle modes are set in the same way as the angle intervals between the corresponding basic angle modes. Furthermore, in the intra-prediction mode set, the number of extended angle modes is set to be less than or equal to the number of basic angle modes.
[0070] According to embodiments of the present invention, extended angle modes are signaled based on basic angle modes. For example, a wide-angle mode (i.e., an extended angle mode) substitutes at least one angle mode (i.e., a basic angle mode) within a first angle range. The substitute basic angle mode is the angle mode corresponding to the opposite side of the wide-angle mode. That is, the substitute basic angle mode is an angle mode that corresponds to an angle in the opposite direction to the angle indicated by the wide-angle mode, or an angle that differs from the angle in the opposite direction by a preset offset index. According to embodiments of the present invention, the preset offset index is 1. The intra-predictive mode index corresponding to the substitute basic angle mode is further mapped to the wide-angle mode to signal the wide-angle mode. For example, wide-angle modes {-10, -9, ..., -1} are signaled by intra-predictive mode indices {57, 58, ..., 66}, respectively, and wide-angle modes {67, 68, ..., 76} are signaled by intra-predictive mode indices {2, 3, ..., 11}, respectively. By having the intra-prediction mode index for the basic angle mode signal the extended angle mode in this way, the same set of intra-prediction mode indices can be used to signal the intra-prediction mode even if the configuration of the angle modes used for intra-prediction in each block differs from one another. Therefore, the signaling overhead due to changes in the configuration of the intra-prediction mode is minimized.
[0071] On the other hand, the availability of the extended angle mode is determined based on at least one of the pattern and size of the current block. According to one embodiment, if the size of the current block is larger than a preset size, the extended angle mode is used for intra-prediction of the current block; otherwise, only the basic angle mode is used for intra-prediction of the current block. According to another embodiment, if the current block is not a square block, the extended angle mode is used for intra-prediction of the current block; if the current block is a square block, only the basic angle mode is used for intra-prediction of the current block.
[0072] On the other hand, to improve coding efficiency, instead of directly coding the residual signal as described above, a method is used in which the transformed residual signal is converted, the resulting conversion coefficients are quantized, and the quantized conversion coefficients are then coded. As described above, the conversion unit converts the residual signal to obtain the conversion coefficients. In this process, the residual signal of a particular block may be distributed throughout the entire area of the block. This allows for energy to be concentrated in the low-frequency range through frequency domain conversion of the residual signal, thereby improving coding efficiency. The following describes in detail how the residual signal is converted or inversely converted.
[0073] Figure 7 is a diagram illustrating in detail how the encoder converts the residual signal. As described above, the spatial domain residual signal is converted to the frequency domain. The encoder converts the acquired residual signal to obtain conversion coefficients. First, the encoder acquires at least one residual block containing the residual signal for the current block. The residual block is either the current block or a block divided from the current block. In this disclosure, the residual block is referred to as a residual array or residual matrix containing the residual samples of the current block. Also in this disclosure, the residual block represents a block the same size as the conversion unit or conversion block.
[0074] Next, the encoder uses a conversion kernel to convert the resistive block. The conversion kernel used for the conversion to the resistive block is a conversion kernel having separable properties for vertical and horizontal conversion. In this case, the conversion to the resistive block is performed separately as vertical and horizontal conversion. For example, the encoder applies the conversion kernel to the vertical direction of the resistive block to perform a vertical conversion. The encoder also applies the conversion kernel to the horizontal direction of the resistive block to perform a horizontal conversion. In this disclosure, the term "conversion kernel" is used to refer to a set of parameters used for converting a resistive signal, such as a conversion matrix, conversion array, conversion function, or conversion. In one embodiment, the conversion kernel is one of several available kernels. Furthermore, different conversion types based on each other may be used for the vertical and horizontal conversions, respectively.
[0075] The encoder transmits the transformed block, converted from the residual block, to the quantization unit for quantization. In this case, the transformed block contains multiple transformation coefficients. More specifically, the transformed block consists of multiple transformation coefficients arranged in a two-dimensional array. The size of the transformed block is the same as that of the residual block, either the current block or a block divided from the current block. The transformation coefficients transmitted to the quantization unit are represented by their quantized values.
[0076] Furthermore, the encoder performs a further transformation before the transformation coefficients are quantized. As shown in Figure 7, the above-described transformation method is called the primary transform, and the further transformation is called the secondary transform. The secondary transform is selective for each residual block. In one embodiment, the encoder can improve coding efficiency by performing a secondary transform in areas where it is difficult to concentrate energy in the low-frequency region with only the primary transform. For example, a secondary transform may be added for blocks where the residual value is largely represented in directions other than the horizontal or vertical direction of the residual block. The residual value of an intra-predicted block has a higher probability of changing in directions other than the horizontal or vertical direction compared to the residual value of an inter-predicted block. Therefore, the encoder performs a further secondary transform on the residual signal of the intra-predicted block. Alternatively, the encoder may omit the secondary transform for the residual signal of an inter-predicted block.
[0077] As another example, whether or not to perform a quadratic transformation is determined based on the size of the current block or residual block. Furthermore, differently sized transformation kernels are used depending on the size of the current block or residual block. For example, an 8x8 quadratic transformation is applied to blocks where the length of the shorter side (either width or height) is equal to or greater than a first preset length. A 4x4 quadratic transformation is applied to blocks where the length of the shorter side (either width or height) is equal to or greater than a second preset length, and smaller than the first preset length. In this case, the first preset length may be greater than the second preset length, but this disclosure is not limited to this. Also, unlike a linear transformation, the quadratic transformation does not necessarily have to be separated into vertical and horizontal transformations. Such a quadratic transformation is referred to as a low-bandwidth non-separated transformation (LFNST).
[0078] Furthermore, in the case of video signals in a specific region, high-frequency bandwidth energy does not decrease even if frequency conversion is performed due to rapid changes in brightness. This may lead to a decrease in compression performance due to quantization. Also, when performing conversion on a region where residual values rarely exist, encoding and decoding times may increase unnecessarily. For this reason, conversion of residual signals in a specific region may be omitted. Whether or not to perform conversion on residual signals in a specific region is determined by the syntax element related to the conversion of that region. For example, the syntax element includes transform skip information. Transform skip information is a transform skip flag. If the transform skip information for a residual block indicates a transform skip, the conversion for that residual block is not performed. In this case, the encoder immediately quantizes the residual signal in the region that has not been converted. The operation of the encoder described with reference to Figure 7 is performed via the conversion unit in Figure 1.
[0079] The syntax elements related to the conversion described above are information purged from the video signal bitstream. The decoder entropy-decodes the video signal bitstream to obtain the syntax elements related to the conversion. The encoder then entropy-codes the syntax elements related to the conversion to generate the video signal bitstream.
[0080] Figure 8 is a diagram illustrating in detail how the encoder and decoder obtain the residual signal by inversely transforming the conversion coefficients. For the sake of explanation, it will be explained below that the inverse transformation operation is performed via the inverse transformation unit of the encoder and decoder, respectively. The inverse transformation unit obtains the residual signal by inversely transforming the inversely quantized conversion coefficients. First, the inverse transformation unit detects whether an inverse transformation is performed for a particular region based on the syntax elements related to the transformation of that region. In one embodiment, if the syntax elements related to the transformation of a particular transformation block indicate a transformation skip, the transformation for that transformation block is omitted. In this case, both the first-order inverse transformation and the second-order inverse transformation are omitted for the transformation block. The inversely quantized conversion coefficients are used as the residual signal. For example, the decoder uses the inversely quantized conversion coefficients as the residual signal to restore the current block. The first-order inverse transformation described above refers to the inverse transformation of the first-order transformation and is called an inverse primary transform. The second-order inverse transformation refers to the inverse transformation of the second-order transformation and is called an inverse secondary transform or inverse LFNST. In this invention, a first-order (inverse) transformation is referred to as the first (inverse) transformation, and a second-order (inverse) transformation is referred to as the second (inverse) transformation.
[0081] In other embodiments, the syntax elements for a particular transformation block may not indicate a transformation skip. In this case, the inverse transformation unit decides whether or not to perform a quadratic inverse transformation on the quadratic transformation. For example, if the transformation block is a transformation block of an intra-predicted block, a quadratic inverse transformation is performed on the transformation block. Also, the quadratic transformation kernel used for the transformation block is determined based on the intra-prediction mode for the transformation block. As another example, whether or not to perform a quadratic inverse transformation may be determined based on the size of the transformation block. The quadratic inverse transformation is performed after the inverse quantization process and before the linear inverse transformation is performed.
[0082] The inverse transformer performs a linear inverse transform on the inversely quantized or quadratic inversely transformed transform coefficients. In the case of a linear inverse transform, it is separated into a vertical transform and a horizontal transform, just like the linear transform. For example, the inverse transformer performs a vertical and horizontal inverse transform on the transform block to obtain a residual block. The inverse transformer inverses the transform block based on the transform kernel used to transform the transform block. For example, the encoder signals information indicating which of several available transform kernels is currently applied to the transform block, either explicitly or indexically. The decoder uses the signaled transform kernel information to select the transform kernel to be used for the inverse transform of the transform block from several available transform kernels. The inverse transformer reconstructs the current block using the residual signal obtained through the inverse transform on the inverse transform coefficients.
[0083] On the other hand, the distribution of residual signals in a picture may differ from region to region. For example, the distribution of residual signals within a particular region may differ depending on the prediction method. When performing transformations using the same transformation kernel for multiple different transformation regions, the coding efficiency may differ for each transformation region depending on the distribution and characteristics of the values within the transformation region. As a result, coding efficiency can be further improved by adaptively selecting the transformation kernel used for a particular transformation block from among multiple available transformation kernels. In other words, encoders and decoders are configured to be able to use transformation kernels other than the basic transformation kernel in the transformation of video signals. The method of adaptively selecting a transformation kernel is called adaptive multiple core transform (ATM) or multiple transform selection (MTS). For convenience of explanation, in this disclosure, the transform and inverse transform are collectively referred to as transforms. Also, the transform kernel and inverse transform kernel are collectively referred to as transform kernels.
[0084] The residual signal, which is the difference between the original signal and the predicted signal generated via inter-screen or intra-screen prediction, has its energy distributed throughout the entire pixel domain. If the pixel values of the residual signal themselves are encoded, a problem arises in that the compression efficiency decreases. Therefore, a process is needed to concentrate the energy of the residual signal in the pixel domain into the low-frequency region of the frequency domain through transform coding.
[0085] The HEVC (High Efficiency Video Coding) standard primarily uses DCT-II (discrete cosine transform type-II), which is efficient when the signal is uniformly distributed across the pixel domain (when adjacent pixel values are similar), and uses DST-VII (discrete sine transform type-VII) only for predicted 4x4 blocks within the screen to convert the residual signal in the pixel domain to the frequency domain. The DCT-II transform is suitable for residual signals generated via inter-screen prediction (when energy is uniformly distributed across the pixel domain). However, in the case of residual signals generated via intra-screen prediction, due to the characteristics of intra-screen prediction, which uses reconstructed reference samples around the currently coded unit, the energy of the residual signal tends to increase as it moves further away from the reference sample. Therefore, if only the DCT-II transform is used to convert the residual signal to the frequency domain, high coding efficiency cannot be achieved.
[0086] AMT is a transformation technique that adaptively selects a transformation kernel from a number of pre-configured kernels depending on the prediction method. Because the pattern in the pixel domain of the residual signal (signal characteristics in the horizontal direction, signal characteristics in the vertical direction) differs depending on which prediction method is used, higher coding efficiency can be expected compared to when only DCT-II is used for the transformation of the residual signal. In this invention, AMT may be called MTS (multiple transform selection) in addition to its name.
[0087] Figure 9 shows the basis functions for multiple transformation kernels that can be used for linear transformations.
[0088] For details, Figure 9 shows the basis functions of the transformation kernels used in AMT, and the mathematical formulas for the DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to AMT are shown.
[0089] DCT and DST are expressed as cosine and sine functions, respectively. When the basis function of the transformation kernel for a given number of samples N is represented by Ti(j), index i indicates the index in the frequency domain, and index j indicates the index within the basis function. In other words, a smaller i indicates a lower frequency basis function, and a larger i indicates a higher frequency basis function. If the basis function Ti(j) is represented as a two-dimensional matrix, it represents the j-th element of the i-th row. Since all the transformation kernels shown in Figure 9 have separable characteristics, transformations can be performed on the residual signal X in both the horizontal and vertical directions. That is, if the residual signal block is X and the transformation kernel matrix is T, the transformation on the residual signal X is represented by TXT'. In this case, T' represents the transpose matrix of the transformation kernel matrix T.
[0090] The transformation matrix values defined by the basis functions shown in Figure 9 are in prime number form, not integer form. Prime number values may be difficult to implement in hardware for video encoding and decoding devices. Therefore, an integer-approximated transformation kernel, derived from the original transformation kernel containing prime number values, is used for encoding and decoding video signals. The approximated transformation kernel containing integer values is generated through scaling and rounding of the original transformation kernel. The integer values included in the approximated transformation kernel are within the range representable by a predetermined number of bits. This predetermined number of bits is either 8-bit or 10-bit. Approximation may not maintain the orthogonal properties of DCT and DST. However, since the resulting loss of encoding efficiency is not significant, approximating the transformation kernel to integer form is advantageous from a hardware implementation perspective.
[0091] In the case of linear transformation regions and inverse linear transformations explained in Figures 7 and 8, since the transformation is performed vertically and horizontally by representing a separable transformation kernel with a two-dimensional matrix, it can be considered that two two-dimensional matrix multiplication operations are performed. This involves a large amount of computation and can be problematic from an implementation standpoint. Therefore, from an implementation standpoint, an important issue may arise as to whether the amount of computation can be reduced by using a combination structure of a butterfly structure or half butterfly structure and a half-matrix multiplier, as in DCT-II, or whether the relevant transformation kernel can be decomposed into a transformation kernel with lower implementation complexity (whether the relevant channel can be represented by the product of matrices with lower complexity). Furthermore, since the elements of the transformation kernel (matrix elements of the transformation kernel) should be stored in memory for computation, the memory capacity for storing the kernel matrix should also be considered during implementation. From this perspective, since the implementation complexity of DST-VII and DCT-VIII is relatively high, transformations that exhibit similar characteristics to DST-VII and DCT-VIII but have lower implementation complexity can be used as substitutes for DST-VII and DCT-VIII.
[0092] DST-IV (discrete sine transform type-IV) and DCT-IV (discrete cosine transform type-IV) are considered candidate substitutes for DST-VII and DCT-VIII, respectively. The DCT-II kernel for 2N samples contains the DCT-IV kernel for N samples, and the DST-IV kernel for N samples can be realized from the DCT-IV kernel for N samples by performing a simple operation of sign inversion and reversing the order of the corresponding basis functions. Therefore, DST-IV and DCT-IV for N samples can be easily derived from DCT-II for 2N samples.
[0093] The residual signal, which is the difference between the original signal and the predicted signal, exhibits a characteristic where the energy distribution of the signal changes depending on the prediction method. Therefore, by adaptively selecting the transformation kernel according to the prediction method, as in AMT or MTS, coding efficiency can be increased. Furthermore, as explained in Figures 7 to 8, coding efficiency can be increased by performing additional transformations such as quadratic transformations and inverse quadratic transformations (inverse transformations corresponding to quadratic transformations), in addition to the first-order transformation and inverse first-order transformation (the inverse transformation corresponding to the first-order transformation). Such quadratic transformations improve energy compaction, especially for in-screen predicted residual signal blocks where strong energy is likely to exist in directions other than the horizontal and vertical directions of the residual signal. As mentioned above, such quadratic transformations are called low-bandwidth non-separated transformations (LFNST). The first-order transformation is called the core transform.
[0094] Figure 10 is a block diagram illustrating the process of restoring a residual signal using a decoder that performs a quadratic transformation according to one embodiment of the present invention. First, the entropy coder purges syntax elements related to the residual signal from the bitstream, and quantization coefficients are obtained via inverse binary evolution (de-binarization). The decoder performs inverse quantization on the restored quantization coefficients to obtain transformation coefficients, and then performs an inverse transformation on the transformation coefficients to restore the residual signal block. The inverse transformation is applied to blocks to which transform skip (TS) is not applied. The inverse transformation is performed in the decoder in the order of quadratic inverse transformation followed by linear inverse transformation. In this case, the quadratic inverse transformation may be omitted. For inter-screen predicted blocks, the quadratic inverse transformation may be omitted without being performed. Alternatively, the quadratic inverse transformation may be omitted depending on the block size conditions. The restored residual signal contains quantization errors, and the quadratic transformation can reduce the quantization errors more than when only a linear transformation is performed by changing the energy distribution of the residual signal.
[0095] Figure 11 is a diagram showing the process of restoring residual signals at the block level using a decoder that performs a quadratic transformation according to one embodiment of the present invention. Residual signal restoration is performed at the transform unit (TU) or sub-block level within the TU. Figure 11 shows the restoration process of a residual signal block to which a quadratic transformation is applied, with the inverse quadratic transformation being performed first on the inversely quantized transformation coefficient block. The decoder may perform the inverse quadratic transformation on all W×H (W: width, number of horizontal samples, H: height, number of vertical samples) samples in the TU, but considering complexity, it may also perform the inverse quadratic transformation only on the left-top sub-block of size W'×H', which is the low-frequency region with the greatest influence. In this case, W' is the same as or less than W, and H' is the same as or less than H. The left-top sub-block size W'×H' is set differently depending on the TU size. For example, if min(W,H)=4, both W' and H' are set to 4. If min(W,H)>=8, both W' and H' are set to 8. min(x,y) represents an operation that returns x if x is equal to or less than y, and returns y if x is greater than or equal to y. The decoder performs a quadratic inverse transform, obtains sub-block transformation coefficients of size W'×H' from the left-top end of the TU, and performs a linear inverse transform on the entire W×H size transformation coefficient block to reconstruct the residual signal block.
[0096] The activation or applicability of a quadratic transformation is indicated by a 1-bit flag in at least one of the High Level Syntax (HLS) RBSPs, such as the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, Slice Header, or Tile Group Header. Furthermore, if a quadratic transformation is applicable, the size of the top-left sub-block to be considered in the quadratic transformation may be indicated by a 1-bit flag in at least one of the HLS RBSPs. For example, whether an 8x8 sub-block is available for a quadratic transformation that considers 4x4 and 8x8 sub-blocks is indicated by a 1-bit flag in at least one of the HLS RBSPs.
[0097] If the activation or applicability of a quadratic transformation is indicated at a higher level (e.g., HLS), whether or not a quadratic transformation is applied is indicated at the coding unit (CU) level by a 1-bit flag. Furthermore, if a quadratic transformation is applied to the current block, an index indicating the transformation kernel used for the quadratic transformation is indicated at the coding unit level. The decoder uses the transformation kernel indicated by the corresponding index within a set of transformation kernels pre-configured by the prediction mode to perform an inverse quadratic transformation on the block to which the quadratic transformation is applied. The index indicating the transformation kernel is binary-coded using a truncated unary or a fixed-length binary-coded method. The 1-bit flag indicating whether or not a quadratic transformation is applied at the CU level and the index indicating the transformation kernel used for the quadratic transformation may be indicated using a single syntax element, which in this invention is referred to as lfnst_idx[x0][y0] or lfnst_idx, but the invention is not limited thereto. As one example, the first bit of lfnst_idx[x0][y0] indicates whether a quadratic transformation is applicable at the CU level. The remaining bits indicate an index that points to the transformation kernel used for the quadratic transformation. In other words, lfnst_idx[x0][y0] indicates whether a quadratic transformation (LFNST) is applicable, and an index that points to the transformation kernel used if the quadratic transformation is applied. Such lfnst_idx[x0][y0] is encoded via an entropy coder such as CABAC (context-based adaptive binary arithmetic coding) or CAVLC (context-based adaptive variable length coding), which adaptively encodes based on the context. Currently, if a CU is divided into many TUs smaller than the CU size, the quadratic transformation is not applied, and lfnst_idx[x0][y0], which is the syntax element related to the quadratic transformation, is set to 0 without signaling. For example, if lfnst_idx[x0][y0] is 0, it indicates that the quadratic transformation is not applied.In contrast, if lfnst_idx[x0][y0] is greater than 0, a quadratic transformation is applied, and the transformation kernel used for the quadratic transformation is selected based on lfnst_idx[x0][y0].
[0098] As described above, coding tree units, leaf nodes of quad trees, and leaf nodes of multi-type trees can be coding units. If a coding unit is not larger than the maximum transformation length, it is not further subdivided and is used as a unit of prediction and / or transformation. In one embodiment, if the width or height of a current coding unit is larger than the maximum transformation length, the current coding unit is subdivided into multiple transformation units without explicit signaling regarding subdivision. If the size of a coding unit is larger than the maximum transformation size, it is subdivided into multiple transformation blocks without signaling. In this case, the maximum coding block (or maximum size of a coding block) to which a quadratic transformation is applied is limited because applying a quadratic transformation would degrade performance and increase complexity. The size of the maximum coding block is the same as the maximum transformation size. Alternatively, the size of the maximum coding block is defined as a preset size of a coding block. In one embodiment, the preset value may be 64, 32, or 16, but the present invention is not limited thereto. In this case, the value compared to the preset value (or maximum transformation size) is defined as the length of the long side or the number of samples.
[0099] On the other hand, the transformation kernels based on the DCT-II, DST-VII, and DCT-VIII basis functions used in linear transformations have separable properties. Therefore, two transformations are performed vertically and horizontally for samples within an N×N size residual block, and the size of the transformation kernel is N×N. In contrast, in the case of quadratic transformations, the transformation kernel has non-separable properties. Therefore, if the number of samples considered in a quadratic transformation is n×n, one transformation is performed. In this case, the size of the transformation kernel is (n^2)×(n^2). For example, when performing a quadratic transformation on a 4×4 coefficient block from left to top, a 16×16 size transformation kernel is applied. And when performing a quadratic transformation on an 8×8 coefficient block from left to top, a 64×64 size transformation kernel is applied. A 64×64 size transformation kernel involves a large amount of multiplication, which can place a heavy burden on the encoder and decoder. Therefore, reducing the number of samples considered in a quadratic transformation can reduce the amount of computation and the memory required to store the transformation kernel.
[0100] Figure 12 shows a method for applying a quadratic transformation that transfers a reduced number of samples according to one embodiment of the present invention. According to one embodiment of the present invention, the quadratic transformation is represented by the product of a quadratic transformation kernel matrix and a linearly transformed coefficient vector, and the linearly transformed coefficients are interpreted as being mapped to another space. In this case, reducing the number of coefficients to be quadratic transformed, that is, reducing the number of basis vectors that constitute the quadratic transformation kernel, can reduce the amount of computation required for the quadratic transformation and the memory capacity required to store the transformation kernel. For example, when performing a quadratic transformation on an 8x8 coefficient block from left to top, if the number of coefficients to be quadratic transformed is reduced to 16, a quadratic transformation kernel of size 16 (rows) x 64 (columns) (or 16 (rows) x 48 (columns)) is applied. The transformation unit of the encoder obtains the quadratic transformed coefficient vector via the inner product of each row vector constituting the transformation kernel matrix and the linearly transformed coefficient vector. The inverse transform section of the encoder and decoder obtains a linearly transformed coefficient vector via the inner product of each column vector constituting the transform kernel matrix and the quadratically transformed coefficient vector.
[0101] Referring to Figure 12, the encoder first performs a forward primary transform on the residual signal block to obtain a linearly transformed coefficient block. If the size of the linearly transformed coefficient block is M × N, then for intra-predicted blocks with a min(M,N) value of 4, a 4 × 4 forward secondary transform is performed on the left-top 4 × 4 samples of the linearly transformed coefficient block. For intra-predicted blocks with a min(M,N) value of 8 or greater, an 8 × 8 quadratic transform is performed on the left-top 8 × 8 samples of the linearly transformed coefficient block. In the case of an 8 × 8 quadratic transform, since it involves a large amount of computation and memory, only a portion of the 8 × 8 samples may be used. In one embodiment, in order to improve coding efficiency, for rectangular blocks with a min(M,N) value of 4 and M or N greater than 8 (for example, rectangular blocks of size 4 × 16 or 16 × 4), a 4 × 4 quadratic transform may be performed on the two left-top 4 × 4 subblocks within the linearly transformed coefficient block.
[0102] Since the quadratic transformation is calculated by multiplying the quadratic transformation kernel matrix by the input vector, the encoder first constructs the coefficients in the left-top subblock of the linearly transformed coefficient block into vector form. The method of constructing the vectors is dependent on the intra-prediction mode. For example, if the intra-prediction mode is the 34th angle mode or lower among the intra-prediction modes shown in Figure 6, the encoder scans the left-top subblock of the linearly transformed coefficient block horizontally to construct the coefficients into vectors. If the element in the i-th row and j-th column of the left-top n×n block of the linearly transformed coefficient block is represented as x(i,j), then the vectorized coefficients are represented as [X(0,0), X(0,1), ..., X(0,n-1), X(1,0), X(1,1), ..., X(1,n-1), ..., X(n-1,0), X(n-1,1), ..., X(n-1,n-1)]. In contrast, if the intra-prediction mode is greater than the 34th angle mode, the left-top subblock of the linearly transformed coefficient block is scanned vertically to construct the coefficients into a vector. The vectorized coefficients are represented as [X(0,0), X(1,0), ..., X(n-1,0), X(0,1), X(1,1), ..., X(n-1,1), ..., X(0,n-1), X(1,n-1), ..., X(n-1,n-1)]. To reduce the computational complexity, if only a portion of the 8x8 samples are used in an 8x8 quadratic transformation, coefficients x_ij where i>3 and j>3 do not need to be included in the vector construction method described above. In this case, 16 linearly transformed coefficients can become the input to the quadratic transformation in a 4x4 quadratic transformation. 48 linearly transformed coefficients can become the input to the quadratic transformation in an 8x8 quadratic transformation.
[0103] The encoder obtains quadratically transformed coefficients via the product of a left-top subblock sample of a vectorized linear transformation coefficient block and a quadratically transformed kernel matrix. The quadratically transformed kernel applied to the quadratically transformed coefficients is determined by the size of the transformation unit or transformation block, the intra-mode, and the syntax element that indicates the transformation kernel. As mentioned above, reducing the number of coefficients to be quadratically transformed reduces the computational complexity and the memory required to store the transformation kernel. Therefore, the number of coefficients to be quadratically transformed is currently determined by the size of the transformation block. For example, in the case of a 4x4 block, the encoder obtains a coefficient vector of length 8 via the product of a vector of length 16 and an 8(row) x 16(column) transformation kernel matrix. The 8(row) x 16(column) transformation kernel matrix is obtained based on the first to eighth basis vectors that make up the 16(row) x 16(column) transformation kernel matrix. For 4×N or M×4 blocks (where N and M are 8 or greater), the encoder obtains a coefficient vector of length 16 by multiplying a vector of length 16 by a 16(row)×16(column) transformation kernel matrix. For 8×8 blocks, the encoder obtains a coefficient vector of length 8 by multiplying a vector of length 48 by an 8(row)×48(column) transformation kernel matrix. The 8(row)×48(column) transformation kernel matrix is obtained based on the first to eighth basis vectors that make up the 16(row)×48(column) transformation kernel matrix. For M×N blocks other than 8×8 (where M and N are 8 or greater), the encoder obtains a coefficient vector of length 16 by multiplying a vector of length 48 by a 16(row)×48(column) transformation kernel matrix.
[0104] According to one embodiment of the present invention, the quadratic-transformed coefficients are in vector form and are therefore represented as two-dimensional data. The quadratic-transformed coefficients are organized into left-top coefficient sub-blocks according to a predetermined scan order. In one embodiment, the predetermined scan order is the upper-right diagonal scan order. The present invention is not limited to this, and the upper-right diagonal scan order is determined based on the method described later in Figures 13 and 14.
[0105] Furthermore, according to one embodiment of the present invention, the transformation coefficients of the entire transformation unit, including the quadratic transformation coefficients, are transmitted in a bitstream after quantization. The bitstream includes syntax elements relating to the quadratic transformation. Specifically, the bitstream includes information indicating whether or not a quadratic transformation is applied to the current block, and information indicating the transformation kernel to be applied to the current block.
[0106] The decoder first purges the quantized conversion coefficients from the bitstream and obtains the conversion coefficients via de-quantization. De-quantization is also called scaling. The decoder determines whether a quadratic inverse transformation is performed on the current block based on the syntax elements related to quadratic transformations. If a quadratic inverse transformation is applied to the current transformation unit or block, 8 or 16 conversion coefficients can become inputs to the quadratic inverse transformation, depending on the size of the transformation unit or block. The number of coefficients that become inputs to the quadratic inverse transformation matches the number of coefficients output by the encoder's quadratic transformation. For example, if the size of the transformation unit or block is 4x4 or 8x8, 8 conversion coefficients become inputs to the quadratic inverse transformation; otherwise, 16 conversion coefficients become inputs. If the size of the transformation unit is MxN, then for intra-predicted blocks where min(M,N) is 4, a 4x4 quadratic inverse transformation is performed on the 16 or 8 coefficients of the left-top 4x4 subblock of the conversion coefficient block. For intra-predicted blocks where min(M,N) is 8 or greater, an 8x8 quadratic transformation is performed on 16 or 8 coefficients in the left-top 4x4 subblock of the transformation coefficient block. In one embodiment, to improve coding efficiency, if min(M,N) is 4 and M or N is greater than 8 (for example, a 4x16 or 16x4 rectangular block), a 4x4 inverse quadratic transformation may be performed on each of the two left-top 4x4 subblocks within the coefficient block.
[0107] According to one embodiment of the present invention, the quadratic inverse transform is calculated by the product of the quadratic inverse transform kernel matrix and the input vector. Therefore, the decoder constructs the previously input inverse-quantized transform coefficient blocks into vector form according to a predetermined scan order. In one embodiment, the predetermined scan order is the upper right diagonal scan order, but the present invention is not limited to this, and the upper right diagonal scan order is determined based on the method described later in Figures 13 and 14.
[0108] Furthermore, according to one embodiment of the present invention, the decoder obtains linearly transformed coefficients via the product of vectorized transformation coefficients and a quadratic inverse transformation kernel matrix. In this case, the quadratic inverse transformation kernel is determined according to the size of the transformation unit or transformation block, the intra-mode, and the syntax elements that indicate the transformation kernel. The quadratic inverse transformation kernel matrix is the transpose of the quadratic transformation kernel matrix. Considering the complexity of the implementation, the elements of the kernel matrix are integers expressed with 10-bit or 8-bit precision. Currently, the length of the vector that will be the output of the quadratic inverse transformation is determined based on the size of the transformation block. For example, in the case of a 4x4 block, a coefficient vector of length 16 is obtained via the product of a vector of length 8 and an 8(row) x 16(column) transformation kernel matrix. The 8(row) x 16(column) transformation kernel matrix is obtained based on the first to eighth basis vectors that constitute the 16(row) x 16(column) transformation kernel matrix. For 4×N or M×N blocks (where N and M are 8 or greater), a coefficient vector of length 16 is obtained by multiplying a vector of length 16 by a 16(row)×16(column) transformation kernel matrix. For 8×8 blocks, a coefficient vector of length 48 is obtained by multiplying a vector of length 8 by an 8(row)×48(column) transformation kernel matrix. The 8(row)×48(column) transformation kernel matrix is obtained based on the 8th basis vector from the first basis vector that makes up the 16(row)×48(column) transformation kernel matrix. For M×N blocks other than 8×8 (where M and N are 8 or greater), a coefficient vector of length 48 is obtained by multiplying a vector of length 16 by a 16(row)×48(column) transformation kernel matrix.
[0109] In one embodiment, since the primary conversion coefficients obtained through the secondary inverse conversion are in vector form, the decoder can further represent them as two-dimensional data, which is intra-mode dependent. At this time, the mapping relationship based on the intra-mode applied by the encoder is also applied. As described above, if the intra prediction mode is 34-degree mode or less, the decoder scans the coefficient vector after the secondary inverse conversion horizontally to obtain a two-dimensional conversion coefficient array. If the intra prediction mode is greater than the 34-degree mode, the decoder scans the coefficient vector after the secondary inverse conversion vertically to obtain a two-dimensional conversion coefficient array. The decoder performs a primary inverse conversion on all conversion units or conversion coefficient blocks of the conversion block size including the conversion coefficients obtained by performing the secondary inverse conversion to obtain a residual signal.
[0110] Although not shown in FIG. 12, a scaling process using bit shift operations may be included when applying the conversion or inverse conversion to correct the scale increased by the conversion kernel after the change or inverse conversion.
[0111] FIG. 13 is a diagram showing a method for determining the upper right diagonal scan order according to an embodiment of the present invention. According to an embodiment of the present invention, during encoding or decoding, a process of initializing the scan order is performed. Initialization of an array including scan order information is performed according to the block size. Specifically, for the combination of log2BlockWidth and log2BlockHeight, the array initialization process of the upper right diagonal scan order shown in FIG. 13 with 1<<log2BlockWidth and 1<<log2BlockHeight as inputs is called (or performed). The output of the array initialization process of the upper right diagonal scan order is assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight]. Here, log2BlockWidth and log2BlockHeight are variables indicating the values obtained by taking the base-2 logarithm of the width and height of the block, respectively, and are values in the range [0, 4].
[0112] Through the array initialization process of the upper right diagonal scan order shown in FIG. 13, the encoder / decoder outputs the array diagScan[sPos][sComp] for blkWidth, which is the width of the input block, and blkHeight, which is the height of the block. The sPos, which is the index of the array, indicates the scan position (scan index) and is a value in the range of [0, blkWidth*blkHeight-1]. If the array index sComp is 0, sPos indicates the horizontal component (x), and if sComp is 1, sPos indicates the vertical component (y). The algorithm shown in FIG. 13 is interpreted such that the x coordinate value and y coordinate value on the two-dimensional coordinates at the scan position sPos in the upper right diagonal scan order are assigned to diagScan[sPos][0] and diagScan[sPos][1], respectively. That is, the value stored in the DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][sComp] array (or array) means the coordinate value corresponding to sComp at the sPos scan position (scan index) in the upper right diagonal scan order of the block whose block width and height are 1<<log2BlockWidth and 1<<log2BlockHeight, respectively.
[0113] Figure 14 is a diagram showing the upper right diagonal scan order according to one embodiment of the present invention, defined by block size. Referring to Figure 14(a), if both log2BlockWidth and log2BlockHeight are 2, it means a 4x4 size block. Referring to Figure 14(b), if both log2BlockWidth and log2BlockHeight are 3, it means an 8x8 size block. In Figure 14, the numbers shown in the gray shaded area indicate the scan position (scan index) sPos. The x and y coordinate values at the sPos position are assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][0] and DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][1], respectively.
[0114] The encoder / decoder codes the conversion coefficient information based on the scan order described above. While this invention primarily describes an embodiment based on the use of the upper-right scanning method, the invention is not limited to this and can be applied to other known scanning methods as well.
[0115] The decoding process for quadratic conversion will be explained in detail below. For the sake of explanation, the decoder will be mainly described in relation to the quadratic conversion process, but the embodiments described below are applied to the encoder in essentially the same way.
[0116] Figure 15 shows how to instruct a quadratic transformation at the coding unit level. The quadratic transformation is instructed at the coding unit level, and the syntax elements for the quadratic transformation are included in the coding_unit syntax structure. The coding_unit syntax structure contains the syntax elements for the coding unit. In this case, the inputs to the coding_unit syntax structure are (x0, y0), which is the coordinate of the left-top luma sample of the current block relative to the left-top luma sample of the picture; cbWidth, which is the width of the block; cbHeight, which is the height of the block; and treeType, which is a variable indicating the type of coding tree. Since there is a correlation between luma and chroma, encoding luma and chroma with the same coding structure enables efficient video compression. Alternatively, to increase coding efficiency, luma and chroma may be encoded with different coding structures. If the variable treeType is SINGLE_TREE, it means that luma and chroma are encoded with the same coding tree structure, and the coding unit includes a luma coding block and a chroma coding block depending on the color format. If treeType is DUAL_TREE_LUMA, it means that luma and chroma are encoded with different encoding trees, and that the tree currently being processed is the tree for luma. In this case, the coding unit contains only luma coding blocks. If treeType is DUAL_TREE_CHROMA, it means that luma and chroma are encoded with different encoding trees, and that that tree currently being processed is the tree for chroma. In this case, the coding unit contains chroma coding blocks according to the color format.
[0117] The coding_unit syntax structure indicates the prediction method for the current coding unit, and the variable CuPredMode[x0][y0] indicates the prediction method for the current block. If CuPredMode[x0][y0] is MODE_INTRA, it indicates that the intra prediction method is applied to the current block, and if it is MODE_INTER, it indicates that the inter-prediction method is applied to the current block. Furthermore, if CuPredMode[x0][y0] is MODE_IBC, it indicates that the IBC (Intra Block Copy) prediction, which generates a reference block from the area where the picture restoration is currently complete, is applied to the current block. The syntax elements related to the prediction method are processed according to the value of the variable CuPredMode[x0][y0]. For example, if the variable CuPredMode[x0][y0] indicates intra prediction, the decoder either purges the syntax elements containing information about the intra prediction mode, reference line index, and ISP (Intra Sub-Partitions) prediction, or sets the variable related to the intra prediction mode in a pre-configured manner.
[0118] After processing the syntax elements related to the prediction method, the syntax elements related to the residual signal are processed. The transform_tree() syntax structure is a syntax structure for a transform tree, where the transform tree is divided into nodes smaller than the root node, with the root node being the same size as the coding unit, and the leaf nodes of the transform tree become transform units. The transform_tree syntax structure contains information about the division of the transform tree.
[0119] One of the intra prediction methods is PCM (Pulse Code Modulation) prediction. If PCM prediction is currently used for prediction of a coding unit, no transformation and quantization are performed, and therefore the transform_tree syntax structure does not exist. In other words, because the transform_tree syntax structure does not exist, the decoder does not perform any operations on the transform_tree syntax structure. PCM prediction is indicated by pcm_flag[x0][y0] when intra prediction is currently instructed for a coding unit. In other words, if pcm_flag[x0][y0] is 1, the decoder does not perform any operations on the transform_tree syntax structure. On the other hand, whether or not the transform_tree syntax structure exists for the current coding unit is indicated by a 1-bit flag, which in this invention is referred to as cu_cbf, but is not limited to this. The decoder either parses cu_cbf, or if cu_cbf is not parsed, sets cu_cbf by a pre-configured method. If cu_cbf is 1, the decoder performs operations on the transform_tree syntax structure. If inter prediction or IBC prediction is used for the prediction of the current coding unit, merge prediction is also available for the prediction of the current coding unit. Whether or not merge prediction is used is indicated by merge_flag[x0][y0]. If it is indicated that merge prediction is used for the current block (merge_flag[x0][y0]==1), cu_cbf is not purged, and the value of cu_cbf is determined by a pre-configured method. The pre-configured method is based on cu_skip_flag[x0][y0] which indicates skip mode. For example, if cu_skip_flag[x0][y0] is 1, cu_cbf is inferred to 0, and otherwise cu_cbf is inferred to 1. If cu_cbf is 1, the transform_tree syntax structure is processed, and the counter value for measuring the number of non-zero significant coefficients is initialized to 0.
[0120] The `numSigCoeff` variable represents the number of non-zero quantization coefficients currently present in the transformation unit of the coding unit, and the value of `numSigCoeff` can affect how syntax elements related to quadratic transformations are handled.
[0121] The `numZeroOutSigCoeff` variable represents the number of non-zero quantization coefficients present at a specific location within the transformation unit currently contained in the coding unit. The value of `numZeroOutSigCoeff` can affect how syntax elements related to quadratic transformations are handled.
[0122] In `transform_tree`, the transformation tree is divided, and the leaf nodes of the transformation tree are transformation units. `transform_tree` contains the `transform_unit` syntax structure, which is the syntax structure for the transformation units that are leaf nodes. `transform_unit` processes the syntax elements for the transformation unit and, if the transformation unit contains one or more non-zero coefficients, it contains the `residual_coding` syntax structure. The `residual_coding` syntax structure contains the syntax structure for the quantized transformation coefficients and the processing related thereto. Depending on the type of tree being processed, the transformation blocks that make up a transformation unit may differ. If `treeType` is `SINGLE_TREE`, the current transformation unit contains a lumen transformation block and, depending on the color format, a chroma transformation block. If `treeType` is `DUAL_TREE_LUMA`, the current transformation unit contains a lumen transformation block. If `treeType` is `DUAL_TREE_CHROMA`, the current transformation unit contains a chroma transformation block. The transform_unit syntax structure includes CBF (coded block flag) information, which indicates whether the transformation block currently contained in the transformation unit contains one or more non-zero coefficients, depending on the treeType. The CBF information is indicated for each color component. For example, if the CBF value for the luma transformation block of the current transformation unit indicates that the luma transformation block does not contain one or more non-zero coefficients, then since all the coefficients in the luma transformation block are 0, the residual_coding syntax structure for the luma transformation block is not processed. As another example, if the CBF value for the chroma Cb transformation block of the current transformation unit indicates that the chroma Cb transformation block contains one or more non-zero coefficients, then the residual_coding syntax structure for the Cb transformation block of the current transformation unit exists.
[0123] Whether or not a quadratic transformation is applied to the current block is indicated at the CU level. If a quadratic transformation is applied, an index indicating the transformation kernel used for the quadratic transformation may also be indicated. As explained in Figure 11, the lfnst_idx[x0][y0] syntax element is used to indicate whether or not a quadratic transformation is applied to the current block. The first bit of lfnst_idx[x0][y0] indicates whether or not a quadratic transformation is applied to the current coding unit. If the first bit of lfnst_idx[x0][y0] is 0, that is, if lfnst_idx[x0][y0] is 0, it indicates that a quadratic transformation is not applied to the current block. On the other hand, if the first bit of lfnst_idx[x0][y0] is 1, that is, if lfnst_idx[x0][y0] is greater than 0 (lfnst_idx[x0][y0]>0), it indicates that a quadratic transformation is applied to the current block. In this process, additional bits are used to indicate the conversion kernel used for the quadratic conversion, and an index indicating the quadratic conversion kernel is signaled via these additional bits.
[0124] The lfnst_idx[x0][y0] syntax element will be purged if the conditions described below are met. Conversely, if the conditions described below are not met, lfnst_idx[x0][y0] does not currently exist in the coding unit, and lfnst_idx[x0][y0] will be set to 0.
[0125] In other words, if the conditions described in the first to fourth embodiments, including the lfnst_idx[x0][y0] syntax element purging conditions described later, are satisfied, the encoder generates a bitstream containing the lfnst_idx[x0][y0] syntax element for the current coding unit. Conversely, if the conditions described later are not satisfied, the bitstream generated by the encoder does not contain the lfnst_idx[x0][y0] syntax element for the current coding unit, and lfnst_idx[x0][y0] is set to 0. A decoder that receives such a bitstream purges the lfnst_idx[x0][y0] syntax element based on the conditions described later.
[0126] lfnst_idx[x0][y0] Parsing conditions for syntax elements
[0127] i) Min(lfnWidth,lfnHeight)>=4
[0128] First, regarding the size of the block, if the width and height of the block are each 4 pixels or more, the decoder will parse the lfnst_idx[x0][y0] syntax element.
[0129] More specifically, the decoder checks the block size conditions to which a quadratic transformation can be applied. The variables SubWidthC and SubHeightC are set by the color format and represent the width and height ratio of the chroma component to the lumina component of the picture, respectively. For example, in a 4:2:0 color format video, there is a structure in which there is one chroma sample for every four lumina samples, so both SubWidthC and SubHeightC are set to 2. As another example, in a 4:4:4 color format video, there is a structure in which there is one chroma sample for every one lumina sample, so both SubWidthC and SubHeightC are set to 1. Currently, the horizontal sample count of a block, lfnWidth, and the vertical sample count, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the coding unit contains only chroma components, so the horizontal sample count of a chroma coding block is the same as the width of the lumina coding block, cbwidth, divided by SubWidthC. Similarly, the vertical sample count of a chroma coding block is the same as the height of a luma coding block, cbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the coding unit contains luma components, so lfnWidth and lfnHeight are set to cbwidth and cbHeight, respectively. Since the minimum condition for a block to which a 22nd-order transformation can be applied is 4x4, if Min(lfnWidth,lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is purged.
[0130] ii) sps_lfnst_enabled_flag==1
[0131] The second condition concerns a flag value that indicates whether the quadratic transformation is enabled or applicable. If the value of the flag (sps_lfnst_enabled_flag) that indicates whether the quadratic transformation is enabled or applicable is set to 1, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0132] For details, the quadratic transformation is indicated by the higher-level syntax RBSP. At least one of the following—SPS, PPS, VPS, tile group header, or slice header—contains a 1-bit flag indicating whether the quadratic transformation is activated and applicable. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element exists within the coding unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element does not exist within the coding unit syntax.
[0133] iii)CuPredMode[x0][y0]==MODE_INTRA
[0134] The third condition concerns the prediction mode, and the quadratic transformation is applied only to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0135] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0136] The fourth condition concerns whether or not the ISP prediction method is applied. If the ISP is not currently applied to the block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0137] As explained in detail with reference to Figure 11, if the current CU is split into many transformation units smaller than the CU size, the quadratic transformation is not applied to the split transformation units. In this case, the syntax element lfnst_idx[x0][y0] related to the quadratic transformation is not purged and is set to 0. If the current CU is split into many transformation units smaller than the CU size of the transformation tree, this includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that, when intra prediction is applied to the current coding unit, splits the transformation tree into many transformation units smaller than the CU size according to a pre-configured splitting method. The ISP prediction mode is indicated at the coding unit level, and the variable IntraSubPartitionsSplitType is set based on it. In this case, if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. The quadratic transformation is indicated at the coding unit level, but the actual quadratic transformation is applied at the transformation unit level. Therefore, if the transformation tree is split into many transformation units, it is inefficient to apply the same quadratic transformation kernel to all of the split transformation units. Furthermore, due to the characteristics of intra-prediction, which generates prediction samples at the transformation unit level, the accuracy of predictions increases when the transformation tree is divided into many transformation units compared to when it is not divided. Therefore, if the transformation tree is divided into many transformation units, the energy of the residual signal is likely to be efficiently compressed even if a quadratic transformation is not applied to the divided many transformation units. Also, if the current size of the CU is larger than the size of the Luma maximum transformation block (MaxTbSizeY) (i.e., cbWidth > MaxTbSizeY || cbHeight > MaxTbSizeY), the transformation tree is divided into many transformation units smaller than the CU size. Although not shown in Figure 15, a quadratic transformation is not applied when the current size of the CU is larger than the Luma maximum transformation block size (MaxTbSizeY).Therefore, the fourth condition may be expressed as IntraSubPartitionsSplitType==ISP_NO_SPLIT&&cbWidth<=MaxTbSizeY&&cbHeight<=MaxTbSize, where MaxTbSizeY is a natural number expressed in the form of a power of 2. MaxTbSizeY may be indicated in the higher-level syntax RBSP, such as SPS, PPS, slice headers, and tile group headers, or the encoder and decoder may use the same value that has been pre-set. For example, the pre-set value may be 64 (2^6).
[0138] v)!intra_mip_flag[x0][y0]
[0139] The fifth condition concerns the intra prediction method: currently, if MIP (Matrix-based Intra Prediction) is not applied to the prediction of the coding unit, the decoder purges the lfnst_idx[x0][y0] syntax element.
[0140] More specifically, MIP is used as one method of intra prediction, and whether or not MIP is applicable is indicated at the coding unit level by intra_mip_flag[x0][y0]. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is currently applied to the prediction of the coding unit, and the prediction is made by the product of the reconstructed samples around the current block and a pre-set matrix. When MIP is applied, the residual signal exhibits different properties from general intra predictions that make directional or non-directional predictions, so a quadratic transformation does not need to be applied to the transformation block when MIP is applied.
[0141] vi)numSigCoeff>((treeType==SIGNLE_TREE)?2:1)
[0142] The sixth condition concerns treeType and coefficients.
[0143] For more details, if treeType is SINGLE_TREE, and the value of the variable numSigCoeff is greater than 2, then a quadratic transformation is applied to the current block, and the decoder parses the lfnst_idx[x0][y0] syntax element.
[0144] If treeType is DUAL_TREE_LUMA or DUAL_TREE_CHROMA, and the value of the variable numSigCoeff is greater than 1, a quadratic transformation is applied to the current block and lfnst_idx[x0][y0] is purged. In this case, numSigCoeff is a variable that indicates the number of effective coefficients present in the current coding unit. If numSigCoeff is less than the critical value, even if a quadratic transformation is applied to the current block, efficient coding may not be performed. This is because if the number of effective coefficients is small, the overhead of signaling the bit-to-bit pairs lfnst_idx[x0][y0] required for coefficient coding is relatively large. In this case, the effective coefficient means a coefficient that is not zero. Hereafter, the effective coefficient described in this invention means a coefficient that is not zero as described above.
[0145] vii)numZeroOutSigCoeff==0
[0146] The seventh condition concerns the effectiveness coefficient located at a specific position.
[0147] More specifically, if a quadratic transformation is applied to the current block, the transformation coefficients quantized by the decoder are always 0 at a specific position. Therefore, if there are non-zero (quantized) coefficients at a specific position, it means that a quadratic transformation has not been applied to the current block, and whether lfnst_idx[x0][y0] is purged is determined by the number of effective coefficients at that position. For example, if numZeroOutSigCoeff is not 0, it means that an effective coefficient exists at that position, so lfnst_idx[x0][y0] is set to 0 without being purged. On the other hand, if numZeroOutSigCoeff is 0, it means that no effective coefficient exists at that position, so lfnst_idx[x0][y0] is purged.
[0148] Figure 16 shows the residual_coding syntax structure according to one embodiment of the present invention.
[0149] The residual_coding syntax structure is a syntax structure related to quantization coefficients and receives x0, y0, log2TbWidth, and log2TbHeight as inputs. At this time, x0 and y0 represent the left-upper coordinates (x0, y0) of the transform block, log2TbWidth is the value obtained by taking the base-2 logarithm of the width of the transform block, and log2TbHeight is the value obtained by taking the base-2 logarithm of the height of the transform block. The number within the transform block is coded in sub-block units, and the coefficient values within each sub-block are determined based on various syntax elements including sig_coeff_flag. At this time, the coefficients in sub-block units may be expressed as a Coefficient Group (CG). sig_coeff_flag[xC][yC] indicates whether the coefficient value at the (xC, yC) position within the current block is 0. If sig_coeff_flag[xC][yC] is 1, it indicates that the coefficient value at the corresponding position is a non-zero value, and if sig_coeff_flag[xC][yC] is 0, it indicates that the coefficient value at the corresponding position is 0. In residual_coding, the x-coordinate value and y-coordinate value of the last significant coefficient in the scan order are indicated. Based on the x-coordinate value and y-coordinate value of the last significant coefficient in the scan order, the index (lastSubBlock) of the sub-block including the last significant coefficient in the scan order is determined. The index of the sub-block is also indexed based on the scan order. The scan order is the upper-right diagonal scan order explained in FIG. 13. In the coefficient coding in sub-block units, the indexes xC and yC indicating the coefficient position (coordinate value) are determined based on the left-upper coordinates (xS<<log2SbW, yS<<log2SbH) of the sub-block and the upper-right diagonal scan order (DiagScanOrder). At this time, xS and yS respectively indicate the index in the horizontal direction and the index in the vertical direction. log2SbW and log2SbH are the values obtained by taking the base-2 logarithm of the width and height of the sub-block, respectively.
[0150] If the value of sig_coeff_flag[xC][yC] is 1 (i.e., the number of (xC,yC) positions is not 0), and no transformation skip is currently applied to the block (i.e., !transform_skip_flag[x0][y0]), then numSigCoeff is counted. When a transformation skip is applied, a quadratic transformation may not be applied, so numSigCoeff, which is used for parsing lfnst_idx[x0][y0], counts the number of effective coefficients in blocks to which no transformation skip is applied.
[0151] Furthermore, as explained in Figure 15, if a quadratic transformation is applied to a transformation block, no effective coefficients exist in a specific region within the transformation block. Therefore, the numZeroOutSigCoeff counter counts the number of effective coefficients (numZeroOutSigCoeff) present in that specific region, and if numZeroOutSigCoeff is not 0, lfnst_idx[x0][y0] is not purged. More specifically, when a quadratic transformation is applied to a transformation block, the region where effective coefficients cannot exist is determined by the size of the transformation block.
[0152] For example, for a quadratic transformation to be applied, if the size of the transformation block is 4x4 (i.e., log2TbWidth==2&&log2TbHeight==2), then within the transformation block, the scan order must divide the region into index [0,7] and the region [8,15], and effective coefficients must exist in the [0,7] region, but not in the [8,15] region. The aforementioned 4x4 transformation block contains one subblock. Therefore, if the size of the transformation block is 4x4, and the scan position is 8 or greater, and the index of the subblock is 0 (i.e., n>=8&&i==0), then the number of effective coefficients is counted. In this case, the scan order is the upper right diagonal scan order.
[0153] As another example, for a quadratic transformation to be applied, if the size of the transformation block is 8x8 (i.e., log2TbWidth==3&&log2TbHeight==3), then effective coefficients exist only in the first subblock within the transformation block, and no effective coefficients can exist in the remaining subblocks (e.g., the second and third subblocks). Even within the first subblock, effective coefficients exist in the index [0,7] region in the scan order, but no effective coefficients can exist in the index [8,15] region. Therefore, if the size of the transformation block is 8x8, the number of effective coefficients is counted if the scan position in the first subblock is 8 or greater (i.e., n>=8&&i==0), and if the scan position exists in the remaining subblocks excluding the first subblock (e.g., in the second and third subblocks, i==1||i==2).
[0154] Finally, if the size of the transformation block is larger than 8x8, effective coefficients exist only in the first subblock within the transformation block, and no effective coefficients can exist in the remaining subblocks (e.g., the second and third subblocks). Therefore, if the subblock is the second or third (i.e., i==1||i==2), the number of effective coefficients is counted. The numZeroOutSigCoeff counter, like the numSigCoeff counter, counts the number of effective coefficients only when sig_coeff_flag[xC][yC] is 1 and transform_skip_flag[x0][y0] is 0. In this case, the subblocks are indexed according to the upper right diagonal scan order explained in Figure 13.
[0155] In other words, if a non-zero coefficient exists in a region where effective coefficients cannot exist (a specific region), it means that a quadratic transformation has not been performed. Therefore, effective coefficients are counted to check whether or not a non-zero coefficient exists in a specific region.
[0156] Figure 17 shows a method for instructing a quadratic conversion at the coding unit level according to one embodiment of the present invention.
[0157] As explained in Figures 15 and 16, whether or not a quadratic transformation is applied is indicated at the coding unit level by the lfnst_idx[x0][y0] syntax element, and two significant coefficient counters (i.e., the numSigCoeff counter and the numZeroOutSigCoeff counter) are required for lfnst_idx[x0][y0] to be parsed. In particular, in the case of numSigCoeff, the numSigCoeff counter should count the number of significant coefficients present in the entire region of the coding unit, which may reduce the throughput of coefficient coding. Therefore, it is necessary to either reduce the number of counters or to use a method that does not use counters.
[0158] The quadratic conversion instruction method shown in Figure 17 is a method of purging lfnst_idx[x0][y0] independently of numSigCoeff. In other words, if all of the conditions i), ii), iii), iv), and v) described in Figure 15 are satisfied (if all are true), the decoder will purge lfnst_idx[x0][y0]. Also, since the value of numSigCoeff is not referenced, the operation of the numSigCoeff counter described in Figure 16 does not occur.
[0159] In this specification, a method for instructing a quadratic transformation based on the position information of the last significant coefficient in the scan order is described below. Similar to the case where the number of significant coefficients is small, if the position of the last significant coefficient in the scan order (scan index) is small, the encoding efficiency of the quadratic transformation is low. Therefore, it is necessary to efficiently instruct a quadratic transformation based on the position information of the last significant coefficient in the scan order, without using a counter.
[0160] (First embodiment)
[0161] Figure 18 shows a method for instructing a quadratic transformation at the coding unit level according to one embodiment of the present invention.
[0162] Figure 18 shows a method for purging lfnst_idx[x0][y0] using the position information of the last effective coefficient in the scan order obtained by residual_coding, instead of the numSigCoeff counter.
[0163] As shown in Figure 18, the numSigCoeff counter is not used, so the numSigCoeff value does not need to be initialized, and the variable lfnLastScanPos, which is related to the position information of the last effective coefficient in the scan order, is initialized to 1. If the value of lfnLastScanPos is 1, it indicates that the position (scan index) of the last effective coefficient in the scan order is less than the critical value, or that all the conversion coefficients in the block are 0. On the other hand, if the value of lfnLastScanPos is 0, it indicates that there is one or more effective coefficients in the block and the position (scan index) of the last effective coefficient in the scan order is greater than or equal to the critical value. Therefore, if the value of lfnLastScanPos is 1, lfnst_idx[x0][y0] is not purged, and if the value of lfnLastScanPos is 0, lfnst_idx[x0][y0] is purged. In addition, lfnst_idx[x0][y0] may be purged if the value of lfnLastScanPos is 0 and all of the conditions i), ii), iii), iv), v), and vii) described in Figure 15 are satisfied (if all of them are true).
[0164] In other words, if there is currently one or more effective coefficients in the block, and the position of the last effective coefficient in the scan order (scan index) is greater than or equal to a critical value, then lfnst_idx[x0][y0] is purged. In this case, as will be described later, the critical value is an integer greater than or equal to 0. For example, assuming the critical value is 1, the position of the last effective coefficient in the scan order (scan index) being greater than or equal to the critical value means that the effective coefficient exists at a position other than the upper left corner of the block. That is, lfnst_idx[x0][y0] is purged only in the remaining cases, excluding the cases where there is no effective coefficient in the block or where there is only one effective coefficient at the upper left corner of the block, i.e., when an effective coefficient exists at a position other than the upper left corner of the block. The meaning of an effective coefficient existing at a position other than the upper left corner of the block may also be represented as "LfnstDConly==0". The upper left corner of the block as described in this invention may mean that the value of the vertical coordinate is (0,0), or it may mean the first position according to a predetermined scan order (for example, the upper right diagonal order), or it may be referred to as DC.
[0165] Figure 19 shows the residual_coding syntax structure according to an embodiment of the present invention.
[0166] Figure 19 shows the residual_coding syntax structure as described in Figure 18. In residual_coding, the syntax elements related to the x and y coordinates of the last significant coefficient in the scan order are parsed, and the LastSignificantCoeffX and LastSignificantCoeffY variables are set. LastSignificantCoeffX represents the x coordinate of the last significant coefficient in the scan order, and LastSignificantCoeffY represents the y coordinate of the last significant coefficient in the scan order. Based on LastSignificantCoeffX and LastSignificantCoeffY, the LastScanPos variable, which is the scan index of the last significant coefficient in the scan order, and the index of the subblock containing the last significant coefficient (lastSubBlock) are determined. In this case, as explained in Figure 16, if a quadratic transformation is applied to the current block, the significant coefficient exists only in the first subblock. In other words, if the significant coefficient exists only in the first subblock, then a quadratic transformation is applied.
[0167] For example, in the 4x4 block shown in Figure 14(a), if LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, then LastScanPos is determined to be 13. Since a 4x4 block consists of one subblock, the index of the subblock containing the last significant coefficient (lastSubBlock) is determined to be 0. As another example, the 8x8 block shown in Figure 14(b) is divided into 4x4 subblocks. Specifically, in Figure 14(b), the 4x4 block corresponding to x coordinates 0 to 3 and y coordinates 0 to 3 is set as the first subblock, the 4x4 block corresponding to x coordinates 0 to 3 and y coordinates 4 to 37 is set as the second subblock, the 4x4 block corresponding to x coordinates 4 to 7 and y coordinates 0 to 34 is set as the third subblock, and the 4x4 block corresponding to x coordinates 4 to 7 and y coordinates 4 to 37 is set as the fourth subblock. In this case, the first subblock is indexed to index 0, the second to index 1, the third to index 2, and the fourth to index 3. The subblocks are indexed according to the upper right diagonal scan order explained in Figure 13. In this case, if LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, lastScanPos is determined to be 13. Since lastScanPos is 13, the subblock containing lastScanPos13 is the first subblock (i.e., subblock index 0), and the index of the subblock containing the last significant coefficient (lastSubBlock) is determined to be 0.
[0168] Based on the above-mentioned lastScanPos, lfnstLastScanPos is determined. Specifically, if the width and height of the transform block are 4 or greater, and no transform skip is applied to the transform block, lfnstLastScanPos is set as shown in Formula 1 below. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in Formula 1 below. In this case, transform_skip_flag[x0][y0] being 0 means that no transform skip is currently applied to the transform block. Specifically, the flag transform_skip_flag[x0][y0] described herein indicates whether or not a primary and secondary transform are applied to the transform block. For example, if the value of transform_skip_flag[x0][y0] is 1, it indicates that neither primary nor quadratic transformations are applied to the transformation block (i.e., transformation skipping is applied), and if the value of transform_skip_flag[x0][y0] is 0, it indicates that both primary and quadratic transformations are applied to the transformation block (i.e., transformation skipping is not applied).
[0169]
number
[0170] As mentioned above, the initial value of lfnstLastScanPos is set to 1.
[0171] In Equation 1, cIdx is a variable that represents the color component of the current transformation block. For example, if cIdx is 0, it indicates that the transformation block processed by residual_coding is the luminous Y component. If cIdx is 1, it indicates that the transformation block processed by residual_coding is the chroma Cb component, and if cIdx is 2, it indicates that the transformation block being processed is the chroma Cr component. The critical value for lastScanPos, lfnstLastScanPosTh[cIdx], is set to a different value depending on the color component.
[0172] According to Equation 1, if the lfnstLastScanPos of a line is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx], then lfnstLastScanPos is updated to 1. Conversely, if the lfnstLastScanPos of a line is 0, or if lastScanPos is greater than or equal to lfnstLastScanPosTh[cIdx], then lfnstLastScanPos is updated to 0. In other words, if the lastScanPos of all transformation blocks included in the coding unit is less than the critical value, or if the number of all transformation blocks is 0, then lfnstLastScanPos is determined to be 1, and according to the lfnst_idx[x0][y0] purging condition in Figure 18, lfnst_idx[x0][y0] is not purged and is set to 0. If lfnst_idx[x0][y0] is not parsed and is set to 0, it indicates that no quadratic transformation is currently applied to the block. Conversely, if LastScanPos is greater than or equal to a critical value for any one of the transformation blocks included in the coding unit, lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vii) described in Figure 15 are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if a quadratic transformation is applied, it checks / determines the transformation kernel to be used for the quadratic transformation.
[0173] In Equation 1, lfnstLastScanPosTh[cIdx] is a pre-set integer value of 0 or greater, and both the encoder and decoder use the same value. Alternatively, the same critical value may be used for all color components. In this case, lfnstLastScanPos is set as shown in Equation 2 below. The coding unit described herein consists of multiple coding blocks, and each coding block has a corresponding conversion block. The conversion blocks are conversion blocks that have luminance and chrominance components. Specifically, these are Y conversion blocks, Cb conversion blocks, and Cr conversion blocks. In this case, whether or not to purge lfnst_idx[x0][y0] as described herein is determined for each conversion block corresponding to each coding block. That is, if any one of the Y conversion block, Cb conversion block, or Cr conversion block satisfies the conditions described herein, lfnst_idx[x0][y0] is purged.
[0174]
number
[0175] lfnstLastScanPosTh is a pre-set integer value of 0 or greater, and both the encoder and decoder use the same value. For example, lfnstLastScanPosTh may be 1. In other words, if lastScanPos is 1 or greater, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is purged. In this case, since the critical value (lfnstLastScanPosTh) is an integer value, lastScanPos being 1 or greater has the same meaning as lastScanPos being greater than 0. The present invention has been described as an example in which the critical value is 1, but the present invention is not limited to this.
[0176] In other words, whether lfnst_idx[x0][y0] can be purged is determined based on lastScanPos. More specifically, as mentioned above, if a quadratic transformation is applied, the last effective coefficient in the scan order exists only in the first subblock of the transformation block. Therefore, if the index of the subblock containing the last effective coefficient in the scan order (where the index indicated by lastScanPos is located) (lastSubBlock) is 0, the width of the transformation block is 4 or more (log2TbWidth>=2), the height of the transformation block is 4 or more (log2TbHeight>=2), transform_skip_flag[x0][y0] is 0 (transformation skip is not applied), and LastScanPos is greater than 0 (LastScanPos is 1 or more), then lfnst_idx[x0][y0] will be purged. This can be expressed mathematically as shown in equation 3 below.
[0177]
number
[0178] On the other hand, in the first embodiment described above, the numSigCoeff counter is not used for purging lfnst_idx[x0][y0], so the number of effective coefficients (numSigCoeff) is not counted.
[0179] (Second example)
[0180] Figure 20 shows the residual_coding syntax structure according to another embodiment of the present invention.
[0181] Figure 20 shows how residual_coding is further input to the treeType variable in Figure 19, and how treeType sets the critical value for LastScanPos.
[0182] If the width and height of the transform block are 4 or greater, and no transform skip is applied to the transform block, then lfnstLastScanPos is set as shown in equation 4 below. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, then lfnstLastScanPos is set as shown in equation 4 below. In this case, transform_skip_flag[x0][y0] being 0 means that no transform skip is currently applied to the transform block.
[0183]
number
[0184] In equation 4, lfnstLastScanPosTh represents the critical value for lastScanPos, and its value is set by treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, lfnstLastScanPosTh is set to val1, val2, and val3, respectively. If the lfnstLastScanPos of a line is 1 and lastScanPos is less than lfnstLastScanPosTh, then lfnstLastScanPos is updated to 1. Conversely, if the lfnstLastScanPos of a line is 0, or lastScanPos is greater than or equal to lfnstLastScanPosTh, then lfnstLastScanPos is updated to 0.
[0185] Equation 4 ultimately determines lfnstLastScanPos to be 1 if the lastScanPos of all transformation blocks included in the coding unit is less than the critical value, or if the number of all transformation blocks is 0. In this case, lfnst_idx[x0][y0] is not purged and is set to 0 according to the lfnst_idx[x0][y0] purging condition in Figure 18. This indicates that no quadratic transformation is currently applied to the block. On the other hand, if the LastScanPos of any one of the transformation blocks included in the coding unit is greater than or equal to the critical value, lfnstLastScanPos is determined to be 0. If all of the conditions i), ii), iii), iv), v), and vii) explained in Figure 15 are satisfied (if all are true), the decoder purges lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if so, checks / determines the transformation kernel to be used for the quadratic transformation.
[0186] val1, val2, and val3 are pre-set integer values of 0 or greater, and both the encoder and decoder use the same values. If treeType is SINGLE_TREE, it includes both luma and chroma components, so val1, which is the value of lfnstLastScanPosTh, may be expressed as the sum of val2 and val3.
[0187] In the second embodiment, the numSigCoeff counter is not used for purging lfnst_idx[x0][y0], so the number of effective coefficients (numSigCoeff) is not counted.
[0188] (Third embodiment)
[0189] Figure 21 shows a method for instructing a quadratic transformation at the coding unit level according to another embodiment of the present invention.
[0190] As shown in Figure 21, instead of the numSigCoeff counter, lfnst_idx[x0][y0] is purged by utilizing the position information of the last effective coefficient in the scan order obtained by residual_coding.
[0191] Since the numSigCoeff counter is not used, numSigCoeff does not need to be initialized, and lfnLastScanPos, a variable relating to the position of the last effective coefficient position information in the scan order, is initialized to 0. The lfnstLastScanPos variable in Figure 21 is the sum of the lastScanPos values of the transformation blocks included in the coding unit. In this case, if lfnLastScanPos is greater than the critical value and all of the conditions i), ii), iii), iv), v), and vii) described in Figure 15 are satisfied (if all are true), the decoder purges lfnst_idx[x0][y0]. The decoder purges lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if a quadratic transformation is applied, it checks / determines the transformation kernel to be used for the quadratic transformation. On the other hand, if lfnLastScanPos is less than or equal to the critical value, lfnst_idx[x0][y0] is not purged and is set to 0. This indicates that the quadratic transformation is not applied.
[0192] The critical value is set by the treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the critical values are set to Th1, Th2, and Th3, respectively. Th1, Th2, and Th3 are pre-set non-negative integer values, and both the encoder and decoder use the same values. If treeType is SINGLE_TREE, it includes both luma and chroma components, so the critical value Th1 may be expressed as the sum of Th2 and Th3.
[0193] Figure 22 shows the residual_coding syntax structure according to another embodiment of the present invention.
[0194] Figure 22 shows the residual_coding syntax structure according to Figure 21 described above. If the width and height of the transformation block are 4 or greater, and no transformation skip is applied to the transformation block, then lfnstLastScanPos is set as shown in equation 5 below. In other words, if log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, then lfnstLastScanPos is set as shown in equation 5 below. In this case, if transform_skip_flag[x0][y0] is 0, it means that no transformation skip is currently applied to the transformation block.
[0195]
number
[0196] In the above equation 5, lfnLastScanPos is the sum of all lastScanPos values of the conversion blocks included in the coding unit, and as explained in Figure 21, the possibility of purging lfnst_idx[x0][y0] is determined by comparing lfnLastScanPos with a critical value.
[0197] In the third embodiment, the numSigCoeff counter is not used for purging lfnst_idx[x0][y0], so the number of effective coefficients (numSigCoeff) is not counted.
[0198] On the other hand, a coding unit includes transformation units that are divided by a transformation tree with the same size as the coding unit as the root node. In this case, each transformation unit includes a transformation block for each color component. If a quadratic transformation is specified at the coding unit level, then residual coding is performed on all transformation blocks contained in the coding unit, and then lfnst_idx[x0][y0] is parsed based on coefficient information. In another embodiment, the quadratic transformation may be specified at the transformation unit level. If a quadratic transformation is specified at the transformation unit level, each transformation unit contained in the coding unit will use a different lfnst_idx[x0][y0]. Therefore, the encoder can find the optimal lfnst_idx[x0][y0] for each transformation unit, further improving encoding efficiency. Also, if a quadratic transformation is specified at the coding unit level and the coding unit contains four transformation units, then residual coding should be performed on all transformation blocks contained in the four transformation units in order for lfnst_idx[x0][y0] to be parsed. In other words, even though the decoder obtains the conversion coefficients for the first conversion unit via residual coding, it is unable to obtain the lfnst_idx[x0][y0] value, and therefore cannot perform the inverse conversion for the first conversion unit. This not only increases the decoder's buffer size but can also cause excessive delay time in the decoder.
[0199] The first to third embodiments described in Figures 18 to 22 are also applicable when the quadratic transformation is specified at the transformation unit level. If the quadratic transformation is specified at the coding unit level, the first to third embodiments determine whether lfnst_idx[x0][y0] can be purged based on the position of the last effective coefficient in the scan order of the transformation blocks included in the coding unit.
[0200] Hereinafter, a specific method by which a secondary transformation is instructed at the transformation unit level will be described.
[0201] Figure 23 shows a method for instructing a secondary conversion at the conversion unit level according to an embodiment of the present invention.
[0202] As shown in Figure 12, instead of the numSigCoeff counter, lfnst_idx[x0][y0] is purged using the position information of the last effective coefficient in the scan order obtained by residual_coding.
[0203] First, before performing residual_coding, the variable lfnLastScanPos, which relates to the position of the last effective coefficient in the scan order, is initialized to 1. If the variable lfnLastScanPos is 1, it indicates that for all conversion blocks contained in the conversion unit, the position of the last effective coefficient in the scan order (scan index) is less than a critical value, or that all conversion coefficients within a block are 0. If the variable lfnLastScanPos is 0, it indicates that for one or more conversion blocks contained in the conversion unit, there is one or more effective coefficients within the block, and the position of the last effective coefficient in the scan order (scan index) is greater than or equal to a critical value. According to the first embodiment described above, if lfnLastScanPos, which is set based on the position of the last effective coefficient in the scan order of the conversion block, is 0, and all of the following conditions i), ii), iii), iv), v), and vi) are satisfied (if all are true), the decoder parses lfnst_idx[x0][y0].
[0204] lfnst_idx[x0][y0] Parsing conditions for syntax elements
[0205] i) Min(lfnWidth,lfnHeight)>=4
[0206] First, the first condition concerns the size of the block; if the width and height of the block are each 4 pixels or more, the decoder will parse the lfnst_idx[x0][y0] syntax element.
[0207] More specifically, the decoder checks the block size conditions to determine if a quadratic conversion is applicable. The variables SubWidthC and SubHeightC are set by the color format and represent the width and height ratio of the chroma component to the lumina component of the picture, respectively. For example, a 4:2:0 color format video has a structure where there is one chroma sample for every four lumina samples, so both SubWidthC and SubHeightC are set to 2. As another example, a 4:4:4 color format video has a structure where there is one chroma sample for every one lumina sample, so both SubWidthC and SubHeightC are set to 1. Currently, the horizontal sample count of a block, lfnWidth, and the vertical sample count, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the conversion unit contains only chroma components, so the horizontal sample count of the chroma conversion block is the same as the width of the lumina conversion block, tbwidth, divided by SubWidthC. Similarly, the vertical sample count of a chroma transformation block is the same as the height of a luma transformation block, tbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the transformation unit includes a luma component, so lfnWidth and lfnHeight are set to tbwidth and tbHeight, respectively. Since the minimum condition for a block to which a quadratic transformation can be applied is 4x4, if Min(lfnWidth,lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is purged.
[0208] ii) sps_lfn_enabled_flag==1
[0209] The second condition concerns a flag value that indicates whether the quadratic transformation is activated or applicable. If the value of the flag (sps_lfnst_enabled_flag) that indicates whether the quadratic transformation is activated or applicable is set to 1, the decoder will purge lfnst_idx[x0][y0].
[0210] For details, the quadratic transformation is indicated by the higher-level syntax RBSP. At least one of the following—SPS, PPS, VPS, tile group header, or slice header—includes a 1-bit flag indicating whether the quadratic transformation is enabled and applicable. If sps_lfnst_enabled_flag is 1, it indicates that the lfnst_idx[x0][y0] syntax element exists in the transformation unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element does not exist in the transformation unit syntax.
[0211] iii)CuPredMode[x0][y0]==MODE_INTRA
[0212] The third condition concerns the prediction mode, and the quadratic transformation is applied only to intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].
[0213] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0214] The fourth condition concerns whether or not the ISP prediction method is applied. If the ISP is not currently applied to the block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0215] As explained in detail with reference to Figure 11, when a CU is currently divided into many transformation units smaller than the CU size, the quadratic transformation is not applied to the divided transformation units. In this case, the syntax element lfnst_idx[x0][y0] related to the quadratic transformation is not purged and is set to 0. When a CU is currently divided into many transformation units smaller than the CU size of the transformation tree, this includes cases where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that, when intra prediction is applied to the current coding unit, divides the transformation tree into many transformation units smaller than the CU size according to a pre-configured division method. The ISP prediction mode is indicated at the coding unit level, and the variable IntraSubPartitionsSplitType is set based on this. In this case, if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction, which generates prediction samples at the transformation unit level, the accuracy of the prediction increases when the transformation tree is divided into many transformation units compared to when it is not divided. Therefore, even if quadratic conversion is not applied to the numerous divided conversion units, there is a high probability that the energy of the residual signal will be efficiently compressed.
[0216] v)!intra_mip_flag[x0][y0]
[0217] The fifth condition concerns the intra prediction method: if MIP (Matrix-based Intra Prediction) is not currently applied to the coding unit, the decoder will purge the lfnst_idx[x0][y0] syntax element.
[0218] More specifically, MIP is used as one method of intra prediction, and whether or not MIP is applicable is indicated at the coding unit level by intra_mip_flag[x0][y0]. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is currently applied to the prediction of the coding unit, and the prediction is made by the product of the reconstructed samples around the current block and a pre-set matrix. When MIP is applied, the residual signal exhibits different properties from general intra predictions that make directional or non-directional predictions, so a quadratic transformation does not need to be applied to the transformation block when MIP is applied.
[0219] vi)numZeroOutSigCoeff==0
[0220] The sixth condition concerns the effectiveness coefficient located at a specific position.
[0221] In more detail, if a quadratic transformation is applied to the current block, the transformation coefficients quantized by the decoder are always 0 at a specific position. Therefore, if there are non-zero quantization coefficients at a specific position, it means that the quadratic transformation has not been applied, and lfnst_idx[x0][y0] is purged according to the number of effective coefficients at that position. For example, if numZeroOutSigCoeff is not 0, it means that an effective coefficient exists at that position, so lfnst_idx[x0][y0] is set to 0 without being purged. On the other hand, if numZeroOutSigCoeff is 0, it means that no effective coefficient exists at that position, so lfnst_idx[x0][y0] is purged.
[0222] Based on the first embodiment described above, if the conversion unit level is instructed whether or not a quadratic conversion is applied to the current block, the residual_coding method described in Figure 19 is followed. If the lastScanPos of all conversion blocks included in the conversion unit is less than the critical value, or if the number of conversion blocks is 0, then lfnLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being purged. This indicates that a quadratic conversion is not applied to the current block. On the other hand, if the LastScanPos of any one of the conversion blocks included in the conversion unit is greater than or equal to the critical value, then lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vi) described in Figure 23 are satisfied (if all are true), then the decoder purges lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if so, checks / determines the transformation kernel to be used for the quadratic transformation.
[0223] If the conversion unit level is instructed whether or not a quadratic conversion is applied based on the second embodiment described above, the conversion unit syntax structure described in Figure 23 is applied, and the residual_coding method described in Figure 20 is used. If the lastScanPos of all conversion blocks included in the conversion unit is less than the critical value according to equation 4 which determines lfnLastScanPos described in Figure 20, or if the number of all conversion blocks is 0, then lfnLastScanPos is determined to be 1, and lfnst_idx[x0][y0] is set to 0 without being purged. This indicates that a quadratic conversion is not applied to the current block. On the other hand, if the LastScanPos of even one of the conversion blocks included in the conversion unit is greater than or equal to the critical value, then lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vi) described in Figure 23 are satisfied (if all are true), then the decoder purges lfnst_idx[x0][y0]. The decoder purges lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if so, checks / determines the transformation kernel to be used for the quadratic transformation.
[0224] Figure 24 shows a method for instructing a secondary conversion at the conversion unit level according to another embodiment of the present invention.
[0225] According to the third embodiment described above, instead of the numSigCoeff counter, lfnst_idx[x0][y0] is purged by utilizing the position information of the last effective coefficient in the scan order obtained by residual_coding.
[0226] Before residual coding is performed, the variable lfnLastScanPos, which relates to the position of the last effective coefficient in the scan order, is initialized to 0. The variable lfnLastScanPos is the sum of the lastScanPos values of the transformation blocks contained in the transformation unit. In this case, if lfnLastScanPos is greater than the critical value and all of the conditions i), ii), iii), iv), v), and vii) described in Figure 23 are satisfied (if all are true), the decoder purges lfnst_idx[x0][y0]. The decoder purges lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if a quadratic transformation is applied, it checks / determines the transformation kernel to be used for the quadratic transformation. On the other hand, if lfnLastScanPos is less than or equal to the critical value, lfnst_idx[x0][y0] is not purged and is set to 0. This indicates that a quadratic transformation is not applied.
[0227] The critical value is set by the treeType. If treeType is SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA, the critical values are set to Th1, Th2, and Th3, respectively. Th1, Th2, and Th3 are pre-set non-negative integer values, and both the encoder and decoder use the same values. If treeType is SINGLE_TREE, it includes both luma and chroma components, so the critical value Th1 may be expressed as the sum of Th2 and Th3.
[0228] If the conversion unit level is instructed whether or not a quadratic conversion should be applied based on the third embodiment described above, the residual_coding method described in Figure 22 is used. According to equation 5, which determines lfnLastScanPos as described in Figure 22, the variable lfnLastScanPos is set to the sum of all lastScanPos values of the conversion blocks included in the conversion unit. Then, by comparing lfnLastScanPos with a critical value, it is determined whether or not lfnst_idx[x0][y0] can be purged.
[0229] On the other hand, if a quadratic transformation is instructed at the transformation unit level, there may be a high correlation between the transformation units included in the coding unit. This is because the prediction method is determined at the coding unit level. Therefore, lfnst_idx[x0][y0] is signaled only in the first transformation unit included in the coding unit, and the signaled lfnst_idx[x0][y0] is shared with the remaining transformation units. In other words, lfnst_idx[x0][y0] may be purged using the first to third embodiments described above only if the subTuIndex, which indicates the index of the transformation unit, is 0. If the subTuIndex is greater than 0, the corresponding transformation unit does not purge lfnst_idx[x0][y0] and uses the value of lfnst_idx[x0][y0] of the first transformation unit to be shared.
[0230] On the other hand, while a counter is used to count the effective coefficients, whether or not the decoder purges lfnst_idx[x0][y0] is determined by considering only the effective coefficients present in the left-top subblock of the transformation block. This is to reduce the computational cost.
[0231] On the other hand, if the quadratic conversion is instructed at the conversion unit level, the decoder delay time is reduced compared to when it is instructed at the coding unit level, but other delays may occur. For example, even if the quadratic conversion is instructed at the conversion unit level, the quadratic conversion is instructed only after the coding of the Luma conversion coefficients, Cb conversion coefficients, and Cr conversion coefficients is completed. Therefore, even after the coding (processing) of the Luma conversion coefficients is completed, the inverse conversion process for the Luma conversion coefficients is performed only after the coding (processing) of the Cb conversion coefficients and Cr conversion coefficients is completed. This introduces other delays in the decoder.
[0232] In this specification, a method for instructing a quadratic conversion that can minimize the delay time of the decoder will be described below.
[0233] (Fourth embodiment)
[0234] One example of a quadratic transformation instruction method that can minimize decoder delay is to instruct the quadratic transformation at the transformation unit level, but purge the lfnst_idx[x0][y0] syntax element related to the quadratic transformation before coding the Luma transformation coefficients. Therefore, the decoder can perform the inverse transformation process for the Luma transformation coefficients immediately after coding the Luma transformation coefficients is complete, without waiting for the Cb and Cr transformation coefficients. Similarly, the decoder can perform the inverse transformation process for the Cb transformation coefficients immediately after coding the Cb transformation coefficients is complete, without waiting for the Cr transformation coefficients to be coded. Such a quadratic transformation instruction method can minimize decoder delay and solve pipeline problems.
[0235] Figure 25 shows the coding unit syntax according to one embodiment of the present invention.
[0236] As shown in Figure 25, since the quadratic transformation is instructed at the transformation unit level, the syntax for the quadratic transformation, lfnst_idx[x0][y0], is not purged at the coding unit level, but rather at the transformation unit level, which is divided by transform_tree.
[0237] Figure 26 shows a method for instructing a secondary conversion at the conversion unit level according to another embodiment of the present invention.
[0238] As shown in Figure 26, the method for instructing the quadratic transformation is instructed at the transformation unit level, and the syntax elements related to the quadratic transformation, lfnst_idx[x0][y0], are purged before the lumar and chroma transformation coefficient coding (residual_coding). For example, if lfnst_idx[x0][y0] is purged before the transformation coefficients are obtained, then as soon as the coefficient coding for each color component, Y, Cb, and Cr, is completed, the inverse transformations for the Y, Cb, and Cr transformation coefficients are processed immediately. For example, as soon as the transformation coefficient coding for the Y component is completed, the inverse transformation for the lumar (Y) transformation coefficient is performed immediately. Similarly, as soon as the transformation coefficient coding for the Cb component (residual_coding) is completed, the inverse transformation for the Cb transformation coefficient is performed immediately, and as soon as the transformation coefficient coding for the Cr component (residual_coding) is completed, the inverse transformation for the Cr transformation coefficient is performed immediately.
[0239] If lfnst_idx[x0][y0] is purged after the conversion coefficient coding (residual_coding) for Y, Cb, and Cr, then even if the conversion coefficient coding (residual_coding) for Y is completed, the inverse conversion for the Y conversion coefficient will not be performed / processed until the conversion coefficient coding (residual_coding) for Cb and Cr is completed and processed. Therefore, even if the conversion coefficient coding (residual_coding) corresponding to Y is completed, the decoder cannot perform the inverse conversion for the Y conversion coefficient until the conversion coefficient coding (residual_coding) for the other components (Cb, Cr) is completed, resulting in the problem of unnecessary delay time. However, as mentioned above, if lfnst_idx[x0][y0] is purged before the conversion coefficient coding (residual_coding), then after the conversion coefficient coding (residual_coding) for each color component (Y, Cb, Cr) is completed, the inverse conversion for each color component's conversion coefficient is performed immediately, thus minimizing the decoder's delay time.
[0240] In the transform_unit() syntax structure, elements such as tu_cbf_luma[x0][y0], tu_cbf_cb[x0][y0], tu_cbf_cr[x0][y0], and transform_skip_flag[x0][y0] are purged.
[0241] More specifically, tu_cbf_luma[x0][y0] is an element that indicates whether the current Luma transformation block contains one or more non-zero transformation coefficients. If tu_cbf_luma[x0][y0] is 1, it indicates that the current Luma transformation block contains one or more non-zero transformation coefficients. If tu_cbf_luma[x0][y0] is 0, it indicates that all the transformation coefficients in the current Luma transformation block are 0. tu_cbf_cb[x0][y0] is an element that indicates whether the current Chroma Cb transformation block contains one or more non-zero transformation coefficients. If tu_cbf_cb[x0][y0] is 1, it indicates that the current Chroma Cb transformation block contains one or more non-zero transformation coefficients. If tu_cbf_cb[x0][y0] is 0, it indicates that all the transformation coefficients in the current Chroma Cb transformation block are 0. tu_cbf_cr[x0][y0] is an element that indicates whether the current chroma Cr transformation block contains one or more non-zero transformation coefficients. If tu_cbf_cr[x0][y0] is 1, it indicates that the current chroma Cr transformation block contains one or more non-zero transformation coefficients. If tu_cbf_cr[x0][y0] is 0, it indicates that all of the transformation coefficients in the current chroma Cr transformation block are 0. transform_skip_flag[x0][y0] is a syntax element related to transformation skipping. If transform_skip_flag[x0][y0] is 1, it indicates that the inverse transformation is not currently applied to the luma transformation block. If transform_skip_flag[x0][y0] is 0, it indicates that whether or not the inverse transformation is currently applied to the luma transformation block is determined by other syntax elements.
[0242] As one example of the quadratic transformation instruction method shown in Figure 26, the syntax element lfnst_idx[x0][y0] related to the quadratic transformation is parsed based on the position of the last effective coefficient in the scan order, rather than based on the number of non-zero transformation coefficients.
[0243] First, the lfnLastScanPos variable is initialized to 1. As explained in Figure 23, the variable lfnLastScanPos indicates the position information of the last effective coefficient in the scan order of the conversion blocks currently contained in the conversion unit. Specifically, if lfnLastScanPos is 1, it indicates that for all conversion blocks contained in the conversion unit, the position (scan index) of the last effective coefficient in the scan order is less than the critical value, or that all conversion coefficients within the block are 0. If lfnLastScanPos is 0, it indicates that for one or more conversion blocks contained in the conversion unit, there is one or more effective coefficients within the block, and the position (scan index) of the last effective coefficient in the scan order is greater than or equal to the critical value.
[0244] Next, the variable numZeroOutSigCoeff is initialized to 0. If a quadratic transformation is applied to a transformation block, the last effective coefficient in the scan order cannot exist. Therefore, the variable numZeroOutSigCoeff indicates whether an effective coefficient exists at a specific position, and based on this, it is checked whether or not a quadratic transformation is applied. For example, assume that if a quadratic transformation is applied to a transformation block, a maximum of 16 effective coefficients are allowed. For 4x4 and 8x8 size transformation blocks, an effective coefficient may exist in the index [0,7] region in the scan order (allowing a maximum of 8 non-zero transformation coefficients). On the other hand, for transformation blocks of sizes other than 4x4 and 8x8, an effective coefficient may exist in the index [0,15] region in the scan order (allowing a maximum of 16 non-zero transformation coefficients). Therefore, if the position of the last effective coefficient in the scan order (scan index) is outside the region where effective coefficients can exist as described above, the decoder can automatically recognize that a quadratic transformation is not currently applied to the transformation block.
[0245] Based on the position of the last effective coefficient in the scan order (scan index), the parsing of lfnst_idx[x0][y0], which is a syntax element related to the quadratic transformation, is determined before coefficient coding (residual_coding). Therefore, the decoder processes information regarding the position of the last effective coefficient in the scan order before coefficient coding (residual_coding).
[0246] More specifically, if the current Luma transform block contains one or more non-zero significant coefficients (tu_cbf_luma[x0][y0]==1) and no transform skip has been applied to the current Luma transform block (transform_skip_flag[x0][y0]==0), then last_significant_pos, which is a syntax structure relating to the position of the last significant coefficient in the Luma scan order, is processed.
[0247] If the value of tu_cbf_luma[x0][y0] is 0 (tu_cbf_luma[x0][y0]==0), it indicates that all coefficients in the corresponding transformation block are 0, which means that coefficient coding (residual_coding) will not be performed. Therefore, processing regarding the position information of the last effective coefficient in the scan order is not necessary.
[0248] If transform_skip_flag[x0][y0] is 1, it indicates that the inverse transform is not currently applied to the Luma transform block. Therefore, coefficient coding (residual_coding) is performed without relying on the position information of the last effective coefficient in the scan order.
[0249] Currently, if a chroma Cb transform block contains one or more significant coefficients (tu_cbf_cb[x0][y0]==1), the syntax structure last_significant_pos, which relates to the position of the last non-zero coefficient in the scan order of the chroma Cb transform block, is processed. The last_significant_pos syntax structure takes as input (x0,y0), which is the left-top coordinate of the transform block, the base-2 logarithm of the transform block's width, the base-2 logarithm of the transform block's height, and cIdx, a variable indicating which color component the transform block represents. For example, if cIdx is 0, it represents a luma Y transform block; if cIdx is 1, it represents a chroma Cb transform block; and if cIdx is 2, it represents a chroma Cr transform block. If the value of tu_cbf_cb[x0][y0] is 0 (tu_cbf_cb[x0][y0]==0), it indicates that all coefficients in the corresponding transformation block are 0. This means that coefficient coding (residual_coding) is not performed, so there is no need to process the position information of the last non-zero coefficient in the scan order.
[0250] On the other hand, if the current Luma-Cr conversion block contains one or more significant coefficients (tu_cbf_cr[x0][y0]==1), the syntax element tu_joint_cbcr_residual[x0][y0], which indicates whether or not to represent chroma Cb and Cr with a single residual signal, is parsed before processing last_significant_pos. For example, if tu_joint_cbcr_residual[x0][y0] is 1, coefficient coding for Cr (residual_coding) is not processed, and the residual signal for Cr is derived from the reconstructed residual signal of Cb. In contrast, if tu_joint_cbcr_residual[x0][y0] is 0, coefficient coding for Cr (residual_coding) is performed according to the value of tu_cbf_cr[x0][y0]. Currently, if a chroma Cr conversion block contains one or more significant coefficients (tu_cbf_cr[x0][y0]==1), the syntax structure last_significant_pos, which relates to the position of the last significant coefficient in the scan order of the chroma Cr, is processed. If the value of tu_cbf_cbr[x0][y0] is 0 (tu_cbf_cr[x0][y0]==0), it indicates that all coefficients in the chroma Cr conversion block are 0. This means that coefficient coding (residual_coding) is not performed, so there is no need to process the position information of the last non-zero coefficient in the scan order.
[0251] The `last_significant_pos` process is performed for each color component to obtain the position (scan index) of the last significant coefficient in the scan order for that color component, and the `lfnLastScanPos` and `numZeroOutSigCoeff` values are updated based on this.
[0252] Then, if all of the following conditions i), ii), iii), iv), v), vi), and vii) are satisfied (if all of them are true), the decoder purges lfnst_idx[x0][y0] before coefficient coding (residual_coding).
[0253] Parsing conditions for the lfnst_idx[x0][y0] syntax element before coefficient coding (residual_coding)
[0254] i) Min(lfnWidth,lfnHeight)>=4
[0255] First, the first condition concerns the size of the block; if the width and height of the block are each 4 pixels or more, the decoder will parse the lfnst_idx[x0][y0] syntax element.
[0256] More specifically, the decoder checks the block size conditions to determine if a quadratic conversion is applicable. The variables SubWidthC and SubHeightC are set by the color format and represent the width and height ratio of the chroma component to the lumina component of the picture, respectively. For example, a 4:2:0 color format video has a structure where there is one chroma sample for every four lumina samples, so both SubWidthC and SubHeightC are set to 2. As another example, a 4:4:4 color format video has a structure where there is one chroma sample for every one lumina sample, so both SubWidthC and SubHeightC are set to 1. Currently, the horizontal sample count of a block, lfnWidth, and the vertical sample count, lfnHeight, are set based on SubWidthC and SubHeightC. If treeType is DUAL_TREE_CHROMA, the conversion unit contains only chroma components, so the horizontal sample count of the chroma conversion block is the same as the width of the lumina conversion block, tbwidth, divided by SubWidthC. Similarly, the vertical sample count of a chroma transformation block is the same as the height of a luma transformation block, tbHeight, divided by SubHeightC. If treeType is SINGLE_TREE or DUAL_TREE_LUMA, the transformation unit includes a luma component, so lfnWidth and lfnHeight are set to tbwidth and tbHeight, respectively. Since the minimum condition for a block to which a quadratic transformation can be applied is 4x4, if Min(lfnWidth,lfnHeight)>=4 is satisfied, lfnst_idx[x0][y0] is purged.
[0257] ii) sps_lfnst_enabled_flag==1
[0258] The second condition concerns the flag value that indicates whether the quadratic transformation is activated or applicable, the flag that indicates whether the quadratic transformation is activated or applicable ( sps_lfnst_enabled_flag If the value of ) is set to 1, the decoder will purge lfnst_idx[x0][y0].
[0259] Specifically, the secondary transformation is indicated by the upper-level syntax RBSP. Among the SPS, PPS, VPS, tile group header, and slice header, there is a flag with a size of 1 bit that indicates the activation and applicability of the secondary transformation. If sps_lfnst_enabled_flag is 1, it indicates that there is an lfnst_idx[x0][y0] syntax element in the transform unit syntax. If sps_lfnst_enabled_flag is 0, it indicates that there is no lfnst_idx[x0][y0] syntax element in the transform unit syntax.
[0260] iii)CuPredMode[x0][y0]==MODE_INTRA
[0261] The third condition is related to the prediction mode. The secondary transformation is only applied to the intra-predicted blocks. Therefore, if the current block is an intra-predicted block, the decoder parses lfnst_idx[x0][y0].
[0262] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0263] The fourth condition is related to whether the ISP prediction method is applied. If ISP is not applied to the current block, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0264] Specifically, as described with reference to FIG. 11, when the current CU is currently split into a number of transform units smaller than the CU size, no secondary transform is applied to the split transform units. At this time, the syntax element lfnst_idx[x0][y0] related to the secondary transform is set to 0 without being parsed. When the current CU is split into a number of transform units smaller than the CU size from the transform tree, it includes the case where ISP prediction is applied to the current coding unit. The ISP prediction method is a prediction method that splits the transform tree into a number of transform units smaller than the CU size by a preset splitting method when intra prediction is applied to the current coding unit. The ISP prediction mode is indicated at the coding unit level, and based on this, the variable IntraSubPartitionsSplitType variable is set. If IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristics of intra prediction that generates prediction samples at the transform unit level, the prediction accuracy is higher when the transform tree is split into a number of transform units than when it is not split. Therefore, even if no secondary transform is applied to the split number of transform units, the energy of the residual signal is likely to be efficiently compressed.
[0265] v)!intra_mip_flag[x0][y0]
[0266] The fifth condition relates to the intra prediction method. If MIP (Matrix based Intra Prediction) is not applied to the current coding unit, the decoder parses the lfnst_idx[x0][y0] syntax element.
[0267] More specifically, MIP is used as one method of intra prediction, and whether or not MIP is applicable is indicated at the coding unit level by intra_mip_flag[x0][y0]. If intra_mip_flag[x0][y0] is 1, it indicates that MIP is currently applied to the prediction of the coding unit, and the prediction is made by the product of the reconstructed samples around the current block and a pre-set matrix. When MIP is applied, the residual signal exhibits different properties from general intra predictions that make directional or non-directional predictions, so a quadratic transformation does not need to be applied to the transformation block when MIP is applied.
[0268] vi)lfnLastScanPos==0 The sixth condition concerns the last effective coefficient in the scan order of the transformation blocks.
[0269] In more detail, if the position information (scan index) of the last effective coefficient in the scan order of the conversion block currently included in the conversion unit is smaller than a preset critical value, the gain of encoding efficiency obtained by quadratic conversion may be small. Therefore, in such cases, the encoder is less likely to apply quadratic conversion to the conversion block (lfnst_idx[x0][y0] is 0), and thus, signaling lfnst_idx[x0][y0] by the encoder is considered to have significant overhead. Therefore, lfnst_idx[x0][y0] is purged only if the position (scan index) of the last effective coefficient in the scan order of at least one of the conversion blocks included in the conversion unit is greater than or equal to a preset critical value.
[0270] In other words, as mentioned above, the critical value is a non-negative integer. For example, assuming the critical value is 1, the fact that the position of the last effective coefficient in the scan order (scan index) is greater than or equal to the critical value means that the effective coefficient lies in a position other than the top-left corner of the block (scan index 0, DC). In this case, the fact that the position of the last effective coefficient in the scan order of the transformed block is greater than or equal to the critical value can also be expressed as "lfnLastScanPos==".
[0271] vii)numZeroOutSigCoeff==0
[0272] The seventh condition concerns the effectiveness coefficient present at a specific location.
[0273] More specifically, if a quadratic transformation is applied to the current block, no effective coefficients can exist at a specific location on the scan. In other words, the numZeroOutSigCoeff variable indicates whether or not a non-zero transformation coefficient exists at a specific location. For example, assume that if a quadratic transformation is applied to the current block, a maximum of 16 effective coefficients are allowed. For 4x4 and 8x8 size transformation blocks, effective coefficients can exist in the index [0,7] region in the scan order (allowing a maximum of 8 non-zero transformation coefficients). On the other hand, for transformation blocks of sizes other than 4x4 and 8x8, effective coefficients can exist in the index [0,15] region in the scan order (allowing a maximum of 16 non-zero transformation coefficients). Therefore, if the location of the last effective coefficient in the scan order (scan index) is outside the region where effective coefficients can exist as described above, the decoder can automatically recognize that a quadratic transformation is not applied to the current block. Thus, if numZeroOutSigCoeff > 0, it means that a quadratic transformation is not applied to the current block, and lfnst_idx[x0][y0] is set to 0 without being purged.
[0274] In other words, if numZeroOutSigCoeff is not 0, it means that an effective coefficient exists at that specific location, so lfnst_idx[x0][y0] is set to 0 without being purged. On the other hand, if numZeroOutSigCoeff is 0, it means that an effective coefficient does not exist at that specific location, so lfnst_idx[x0][y0] is purged.
[0275] If any of the above conditions i) through vii) are true, lfnst_idx[x0][y0] is purged; otherwise, lfnst_idx[x0][y0] is not purged and is set to 0.
[0276] Figure 27 shows the syntax structure relating to the position of the last effective coefficient in the scan sequence according to an embodiment of the present invention.
[0277] As shown in Figure 27, the last_significant_pos syntax structure represents a syntax structure that includes the position information of the last significant coefficient in the scan order for each color component Y, Cb, and Cr transformation block. The last_significant_pos syntax structure takes as input (x0, y0), which is the left-top coordinate of the transformation block, log2TbWidth, which is the base-2 logarithm of the transformation block's width, log2TbHeight, which is the base-2 logarithm of the transformation block's height, and cIdx, which color component the transformation block represents. If cIdx is 0, it represents a chroma transformation block; if cIdx is 1, it represents a chroma-Cb transformation block; and if cIdx is 2, it represents a chroma-Cr transformation block.
[0278] The `last_significant_pos` syntax structure parses the syntax elements related to the position information of the last significant coefficient in the scan order. Specifically, the syntax elements related to the x-coordinate and y-coordinate values of the last significant coefficient in the scan order are parsed. In this process, each coordinate value is divided into prefix and suffix information. The decoder sets the `LastSignificantCoeffX` variable, which is the x-coordinate of the last significant coefficient in the scan order, based on the prefix and suffix information for the x-coordinate. Similarly, the decoder sets the `LastSignificantCoeffY` variable, which is the y-coordinate of the last significant coefficient in the scan order, based on the prefix and suffix information for the y-coordinate. As shown in Figure 27, the decoder sets `lastScanPos`, which is the scan index of the last significant coefficient in the scan order, based on `LastSignificantCoeffX`, `LastSignificantCoeffY`, and `DiagScanOrder` in a `do{}while()` structure. Furthermore, the decoder updates numZeroOutSigCoeff and lfnstLastScanPos, which are variables used in the purging condition of lfnst_idx[x0][y0], a syntax element related to quadratic transformations, based on lastScanPos.
[0279] If a quadratic transformation is applied to the current block, no effective coefficients can exist at certain locations on the scan. The numZeroOutSigCoeff variable indicates whether a non-zero transformation coefficient exists at such a location. For example, assume that if a quadratic transformation is applied to the current block, a maximum of 16 effective coefficients are allowed. For 4x4 and 8x8 size transformation blocks, effective coefficients may exist in the index [0,7] region in the scan order (allowing a maximum of 8 non-zero transformation coefficients). On the other hand, for transformation blocks of sizes other than 4x4 and 8x8, effective coefficients may exist in the index [0,15] region in the scan order (allowing a maximum of 16 non-zero transformation coefficients). Therefore, if the location of the last effective coefficient in the scan order (scan index) is outside the region where effective coefficients can exist as described above, the decoder can automatically recognize that a quadratic transformation is not applied to the current block. The minimum size of a block to which a quadratic transformation can be applied is 4x4, and if a transformation skip is applied (transform_skip_flag[x0][y0]==1), the quadratic transformation is not applied. Therefore, numZeroOutSigCoeff is updated for transform blocks where the width of the transform block is 4 or greater (log2TbWidth>=2), the height of the transform block is 4 or greater (log2TbHeight>=2), and transform skipping is not applied (transform_skip_flag[x0][y0]==0). If a quadratic transform is applied, for 4x4 and 8x8 size transform blocks, a non-zero transform coefficient may exist only in the index [0,7] region in the scan order. Therefore, if the transform block is 4x4 or 8x8 ((log2TbWidth==2||log2TbHeight==3)&&(log2TbWidth==log2TbHeight)), and lastScanPos is greater than 7 (lastScanPos>7), numZeroOutSigCoeff increases by 1. For transformation blocks other than the 4x4 and 8x8 sizes to which quadratic transformations can be applied, non-zero transformation coefficients can only exist in the scan-order index [0,15] region. Therefore, if lastScanPos is greater than 15 (lastScanPos>15), numZeroOutSigCoeff increases by 1.
[0280] The decoder determines lfnstLastScanPos based on lastScanPos. Specifically, if the width and height of the transform block are 4 or more and transform skip is not applied to the transform block, lfnstLastScanPos is set as shown in Equation 6 below. In other words, if log2TbWidth >= 2, log2TbHeight >= 2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos is set as shown in Equation 1 below. At this time, if transform_skip_flag[x0][y0] is 0, it means that transform skip is not applied to the current transform block.
[0281]
Number
[0282] As described above, the initial value of lfnstLastScanPos is set to 1.
[0283] In Equation 6, cIdx indicates a variable that represents the color component of the current transform block as described above.
[0284] According to Equation 6, if the lfnstLastScanPos of the line is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 1. On the other hand, if the lfnstLastScanPos of the line is 0 or lastScanPos is greater than or equal to lfnstLastScanPosTh[cIdx], lfnstLastScanPos is updated to 0.
[0285] In other words, if the lastScanPos of all conversion blocks included in the conversion unit is less than the critical value, or if the number of all conversion blocks is 0, then lfnstLastScanPos is determined to be 1, and according to the lfnst_idx[x0][y0] purging condition in Figure 26, lfnst_idx[x0][y0] is not purged and is set to 0. This indicates that a quadratic conversion is not currently applied to the block. On the other hand, if the LastScanPos of any one of the conversion blocks included in the conversion unit is greater than or equal to the critical value, then lfnstLastScanPos is determined to be 0, and if all of the conditions i), ii), iii), iv), v), and vii) described in Figure 26 are satisfied (if true), then the decoder purges lfnst_idx[x0][y0]. The decoder parses lfnst_idx[x0][y0] to check whether a quadratic transformation is applied to the current block, and if so, checks / determines the transformation kernel to be used for the quadratic transformation.
[0286] In equation 6, lfnstLastScanPosTh[cIdx] is a pre-set integer value of 0 or greater, and both the encoder and decoder use the same value. Alternatively, all color components may use the same critical value. In this case, lfnstLastScanPos is set as shown in equation 7 below.
[0287]
number
[0288] lfnstLastScanPosTh is a pre-set integer value of 0 or greater, and both the encoder and decoder use the same value. For example, lfnstLastScanPosTh may be 1. In other words, if lastScanPos is 1 or greater, lfnstLastScanPos is updated to 0, and lfnst_idx[x0][y0] is purged. In this case, since the critical value (lfnstLastScanPosTh) is an integer value, lastScanPos being 1 or greater has the same meaning as lastScanPos being greater than 0. Figure 27 illustrates the case where all color components have the same critical value of 1, but the present invention is not limited to this.
[0289] Figure 28 shows the residual_coding syntax structure according to an embodiment of the present invention.
[0290] As shown in Figure 28, the position information of the last significant coefficient in the scan order is indicated in the residual coding. Therefore, the syntax structure of residual coding does not necessarily have to include the syntax structure related to the position information of the last significant coefficient in the scan order. For example, the position information of the last significant coefficient in the scan order is the prefix and suffix for the x coordinate and the prefix and suffix for the y coordinate of the last significant coefficient in the scan order. Examining the residual coding syntax structure in Figure 28, residual coding is performed based on the x and y coordinates of the last significant coefficient in the scan order, LastSignificantCoeffX and LastSignificantCoeffY, which were determined before residual coding.
[0291] The quadratic transformation instruction method according to the fourth embodiment does not use the numSigCoeff counter. Therefore, even if the coefficient at the (xC,yC) position is a significant coefficient (sig_coeff_flag[xC][yC]==1), numSigCoeff is not updated. In other words, the quadratic transformation instruction method according to the fourth embodiment is a method that does not use a counter for significant coefficients. Furthermore, according to the quadratic transformation instruction method according to the fourth embodiment, the numZeroOutSigCoeff variable is set based on lastScanPos, so a counter based on sig_coeff_flag does not need to be used in coefficient coding (residual_coding).
[0292] Figure 29 is a sequence diagram showing a video signal processing method according to an embodiment of the present invention.
[0293] The following describes a video signal processing method and apparatus based on the embodiments explained with reference to Figures 15 to 28.
[0294] The video signal decoding device includes a processor that performs the video signal processing method described in Figure 29.
[0295] First, the processor receives a bitstream containing syntax elements related to the quadratic transformation of the coding unit.
[0296] The processor checks whether one or more pre-set conditions are satisfied, and if one or more of the pre-set conditions are satisfied, it purges the syntax elements related to the quadratic transformation of the coding unit in S2910, S2920. On the other hand, if one or more of the pre-set conditions are not satisfied, the processor does not purge the syntax elements related to the quadratic transformation of the coding unit in S2930. In this case, the value of the syntax elements related to the quadratic transformation is set to 0.
[0297] The syntax element for the quadratic transformation of the coding unit, as explained in Figure 29, is lfnst_idx[x0][y0], which indicates whether or not the quadratic transformation of the transformation block currently included in the coding unit, as explained in Figures 5 to 28, is applied.
[0298] The processor purges syntax elements related to the quadratic transformation of the coding unit via step S2920, and S2940 confirms whether the quadratic transformation is applied to the transformation block included in the coding unit based on the purged syntax elements.
[0299] In this case, if the quadratic transformation has been applied to the transformation block, the processor performs a quadratic inverse transformation based on one or more coefficients of the first subblock, which is one of the subblocks that make up the transformation block, and confirms one or more inverse transformation coefficients for the first subblock S2950.
[0300] Then, in step S2960, the processor performs a linear inverse transform based on the one or more inverse transform coefficients obtained in step S2950, and checks the residual sample for the transform block.
[0301] The aforementioned quadratic transformation is a low-bandwidth non-separated transformation (LFNST). The transformation block is a block to which a primary transformation is applied, which is separated into a vertical transformation and a horizontal transformation, respectively. In this case, the primary inverse transformation is the inverse transformation of the primary transformation, and the secondary inverse transformation is the inverse transformation of the secondary transformation.
[0302] The syntax element relating to the quadratic transformation of the coding unit includes information indicating whether or not the quadratic transformation is applied to the coding unit, and information indicating the transformation kernel used for the quadratic transformation.
[0303] The first subblock is the first subblock in the pre-configured scan order, and in this case, the index of the first subblock is 0.
[0304] The first of the one or more pre-set conditions is that the index value indicating the position of the first coefficient among the one or more coefficients of the first subblock is greater than a pre-set critical value. In this case, the first coefficient is the last effective coefficient in the pre-set scan order, and the effective coefficient means a coefficient that is not 0. The pre-set critical value is 0. The pre-set scan order is the upper right diagonal scan order explained in Figures 13 and 14.
[0305] The second of the one or more pre-set conditions is that the width and height of the conversion block are 4 pixels or more.
[0306] The third of the one or more pre-set conditions is that the conversion skip flag value included in the bitstream is not a specific value. In this case, if the value of the conversion skip flag is a specific value, the conversion skip flag indicates that the primary conversion and the secondary conversion are not applied to the conversion block.
[0307] The fourth condition among the one or more pre-set conditions is that at least one of the one or more coefficients of the first subblock is not 0, and at least one of the coefficients is located in a position other than the first position according to the pre-set scan order. In this case, the first position according to the pre-set scan order means either the position where the horizontal and vertical coordinate values are (0,0) as described above, or the first position according to the pre-set scan order (for example, the upper right diagonal order).
[0308] Furthermore, the coding unit is composed of multiple coding blocks. In this case, if at least one of the transformation blocks corresponding to each of the multiple coding blocks satisfies one or more of the pre-set conditions, the syntax elements related to the quadratic transformation are purged.
[0309] On the other hand, if the syntax elements relating to the quadratic transformation are not parsed or are set to 0 in step S2930, or if it is confirmed in step S2940 that the quadratic transformation is not applied to the transformation block included in the coding unit, the processor performs a linear inverse transformation based on one or more coefficients of the transformation block to obtain a residual sample for the transformation block in S2970.
[0310] In this context, the linear inverse transform and quadratic inverse transform mentioned above are the inverse transforms of the linear and quadratic transforms, respectively.
[0311] The video signal processing method used in the video signal decoding device described in Figure 29, or a similar method, is performed in the video signal encoding device.
[0312] A video signal encoding device includes a processor that encodes video signals.
[0313] In this process, the processor performs a linear transformation on the residual samples of the blocks included in the coding unit to obtain a plurality of linear transformation coefficients for the blocks. A quadratic transformation is performed based on one or more of the plurality of linear transformation coefficients to obtain one or more quadratic transformation coefficients for a first subblock, which is one of the subblocks constituting the block. Information on the one or more quadratic transformation coefficients and syntax elements related to the quadratic transformation of the coding unit are encoded to obtain a bitstream.
[0314] The aforementioned secondary conversion is a low-bandwidth non-separated conversion (LFNST), and the primary conversion may be separated into horizontal and vertical conversions, respectively.
[0315] Furthermore, the syntax element relating to the quadratic transformation is encoded if one or more pre-set conditions are satisfied. The syntax element relating to the quadratic transformation includes information indicating whether or not the quadratic transformation is applied to the coding unit, and information indicating the transformation kernel used for the quadratic transformation. In this case, the syntax element relating to the quadratic transformation is lfnst_idx[x0][y0], which is the syntax element described in Figures 15 to 28.
[0316] The aforementioned first subblock is the first subblock in the pre-configured scan order. In this case, the index of the first subblock is 0.
[0317] The first of the one or more pre-set conditions is that the index value indicating the position of the first coefficient among the one or more quadratic transformation coefficients is greater than a pre-set critical value. In this case, the first coefficient is the last effective coefficient in the pre-set scan order, and the effective coefficient means a coefficient that is not 0. The pre-set critical value is 0. The pre-set scan order is the upper right diagonal scan order explained in Figures 13 and 14.
[0318] The second of the one or more pre-set conditions is that the width and height of the primary conversion block are 4 pixels or more.
[0319] The third of the one or more pre-set conditions is that the conversion skip flag value included in the bitstream is not a specific value. In this case, if the value of the conversion skip flag is a specific value, the conversion skip flag indicates that the primary conversion and the secondary conversion will not be applied to the block.
[0320] The fourth condition among the one or more pre-set conditions is that at least one of the one or more quadratic transformation coefficients is not 0, and the one or more coefficients are located at a position other than the first position according to the pre-set scan order. In this case, the first position according to the pre-set scan order means either the position where the horizontal and vertical coordinate values are (0,0) as described above, or the first position according to the pre-set scan order (for example, the upper right diagonal order).
[0321] Furthermore, the coding unit is composed of multiple coding blocks. In this case, if at least one of the (transformation) blocks included in the coding unit corresponding to each of the multiple coding blocks satisfies one or more of the pre-set conditions, the syntax elements related to the quadratic transformation are encoded.
[0322] Furthermore, the video signal encoding device may include a video signal decoding processor that performs the video signal processing method described in Figure 29.
[0323] As described above, the bitstream includes syntax elements related to the quadratic transformation of the coding unit described in Figures 15 to 29. In this case, the bitstream is stored on a non-temporary computer-readable medium. On the other hand, the video signal encoder either omits the syntax elements related to the quadratic transformation from the bitstream or sets the syntax elements related to the quadratic transformation to 0 unless one or more of the above-described pre-set conditions are satisfied. The bitstream is either decoded by the video signal decoder described via Figure 29 or encoded by the video signal encoder described above.
[0324] A method for encoding such a bitstream includes, for example, the process of performing a linear transformation on the residual samples of a block included in a coding unit to obtain a plurality of linear transformation coefficients for the block, performing a quadratic transformation based on one or more of the plurality of linear transformation coefficients to obtain one or more quadratic transformation coefficients for a first subblock which is one of the subblocks constituting the block, and encoding information for the one or more quadratic transformation coefficients and syntax elements relating to the quadratic transformation of the coding unit.
[0325] In this specification, obtaining a coefficient means obtaining pixels / blocks related to the coefficient, and obtaining a residual sample means obtaining residual signals / pixels / blocks related to the residual sample.
[0326] The embodiments of the present invention described above can be embodied through a variety of means. For example, embodiments of the present invention can be embodied through hardware, firmware, software, or a combination thereof.
[0327] In the case of hardware implementation, the method according to the embodiment of the present invention is implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSDPs (Digital Signal Processing Devices), PDLs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.
[0328] In the case of implementation by firmware or software, the method according to the embodiment of the present invention is implemented in the form of a module, procedure, or function that performs the functions or operations described above. The software code is stored in memory and implemented by a processor. The memory is located inside or outside the processor and exchanges data with the processor by various already known means.
[0329] Some embodiments also embody the form of recording media containing computer-executable instructions, such as program modules executed by a computer. Computer-readable media are any available media that can be accessed by a computer, and include both volatile and non-volatile media, and isolated and non-isolated media. Computer-readable media also include both storage media and communication media. Computer storage media include both volatile and non-volatile media, isolated and non-isolated media, and embodied in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Communication media typically include computer-readable instructions, data structures, or other data such as modulated data signals or program modules, or other transmission mechanisms, and include any information transmission media.
[0330] The above description of the present invention is illustrative, and a person with ordinary skill in the art to which the invention pertains should understand that it can be easily modified into other specific forms without altering the technical idea or essential features of the invention. Therefore, the above embodiments should be understood to be illustrative and not limiting in all respects. For example, each component described as a single type may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.
[0331] The scope of the present invention is indicated by the claims described below rather than by the detailed description above, and all modifications or altered forms derived from the meaning and scope of the claims and the concept of equivalents thereto should be interpreted as being included within the scope of the present invention. [Explanation of symbols]
[0332] 110 Conversion Unit 115 Quantization section 120 Inverse quantization section 125 Inverse Transform Section 130 Filtering section 150 Prediction Section 152 Intra Prediction Unit 154 Interpretation Unit 154a Motion Estimation Unit 154b Motion compensation unit 160 Entropy coding section 210 Entropy Decoding Unit 220 Inverse quantization section 225 Inverse Transform Section 230 Filtering section 250 Prediction Section 252 Intra Prediction Unit 254 Interpretation Unit
Claims
1. In a video decoding method performed by a device, The stage of obtaining syntax information regarding the secondary transformation of the coding unit, Based on the acquired syntax information, the step of determining whether the secondary transformation is applied to the transformation block included in the coding unit, If the quadratic transformation is applied to the transformation block, the steps include obtaining one or more inverse transformation coefficients based on the quadratic transformation, The step includes obtaining a residual sample for the transformation block based on one or more inverse transformation coefficients, The aforementioned secondary conversion is a low-frequency non-separable transform (LFNST), The syntax information mentioned above is obtained if one or more conditions are met. Of the one or more conditions mentioned above, the first condition is that in a subblock, the index of the last effective coefficient is greater than the default value due to the previously set scan order. Of the one or more conditions mentioned above, the second condition is that intra prediction is applied to the coding unit. A video decoding method in which the subblock is a subblock having subblock index 0 according to the scan order set in the conversion block.
2. The video decoding method according to claim 1, wherein the default value is 0.
3. The video decoding method according to claim 1, wherein the previously set scan order is the up-right diagonal scan order.
4. The index of the one or more inverse transformation coefficients is determined based on the previously set scan order. Of the one or more inverse transformation coefficients mentioned above, the index of the first coefficient is 0. The video decoding method according to claim 3, wherein the last effective coefficient is a non-zero coefficient.
5. The video decoding method according to claim 1, wherein the residual sample is obtained by performing an inverse transform of a linear transform based on one or more inverse transform coefficients.
6. In a video encoding device, the video encoding device is It includes at least one processor, The aforementioned at least one processor is It is configured to generate a bitstream using an encoding method, The aforementioned encoding method is The stage of encoding syntax information related to the secondary transformation of the coding unit, The steps include obtaining one or more linear transformation coefficients based on residual samples for the transformation blocks included in the coding unit, If the quadratic transformation is applied to the transformation block based on the syntax information, the steps include obtaining one or more quadratic transformation coefficients using one or more linear transformation coefficients based on the quadratic transformation, The step includes encoding one or more quadratic transformation coefficients, The aforementioned secondary conversion is a low-frequency non-separable transform (LFNST), The syntax information is encoded if one or more conditions are met. Of the one or more conditions mentioned above, the first condition is that in a subblock, the index of the last effective coefficient is greater than the default value due to the previously set scan order. Of the one or more conditions mentioned above, the second condition is that intra prediction is applied to the coding unit. A video encoding device in which the subblock is a subblock having subblock index 0 according to the scan order set in the conversion block.
7. The video encoding apparatus according to claim 6, wherein the default value is 0.
8. The video encoding apparatus according to claim 6, wherein the previously set scan order is an up-right diagonal scan order.
9. The index of the one or more linear transformation coefficients is determined based on the previously set scan order. Of the one or more linear transformation coefficients mentioned above, the index of the first coefficient is 0. The video encoding apparatus according to claim 8, wherein the last effective coefficient is a non-zero coefficient.
10. The video encoding apparatus according to claim 6, wherein the one or more linear transformation coefficients are obtained by performing a linear transformation based on the residual samples.
11. In methods for saving bitstreams, The bitstream is generated by the encoding method, The aforementioned encoding method is The stage of encoding syntax information related to the secondary transformation of the coding unit, The steps include obtaining one or more linear transformation coefficients based on residual samples for the transformation blocks included in the coding unit, If the quadratic transformation is applied to the transformation block based on the syntax information, the steps include obtaining one or more quadratic transformation coefficients using one or more linear transformation coefficients based on the quadratic transformation, The step includes encoding one or more quadratic transformation coefficients, The aforementioned secondary conversion is a low-frequency non-separable transform (LFNST), The syntax information is encoded if one or more conditions are met. Of the one or more conditions mentioned above, the first condition is that in a subblock, the index of the last effective coefficient is greater than the default value due to the previously set scan order. Of the one or more conditions mentioned above, the second condition is that intra prediction is applied to the coding unit. A bitstream storage method in which the subblock is a subblock having subblock index 0 according to the scan order set in the conversion block.
Citation Information
Patent Citations
Binarizing secondary transform index
US20170324643A1
Method and device for coding residual signal in video coding system
US20180288409A1
Transform method in image coding system and apparatus for same
US20200177889A1