Video signal processing method and apparatus using quadratic transformation
By introducing a quadratic transformation technique, particularly the low-frequency non-separable transform (LFNST) and the upper-right diagonal scanning sequence, the problem of insufficient coding efficiency in existing technologies has been solved, achieving more efficient video signal encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2020-06-25
- Publication Date
- 2026-07-17
AI Technical Summary
Existing video signal processing methods are insufficient in terms of encoding efficiency, and more effective processing methods are needed to improve compilation efficiency.
The video signal is processed using a quadratic transform technique, which involves vertical and horizontal transforms of the transform blocks. Under specific conditions, the upper right diagonal scanning order is applied to parse and encode syntax elements, thereby improving the efficiency of encoding and decoding.
By employing a double-transformation technique, the encoding and decoding efficiency of video signals is significantly improved, thereby enhancing compilation efficiency.
Smart Images

Figure CN115941975B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 202080002631.0 (PCT / KR2020 / 008301), filed with the China Patent Office on November 4, 2020, with an international application date of June 25, 2020, entitled "Video Signal Processing Method and Apparatus Using Secondary Transformation". Technical Field
[0002] The present invention relates to video signal processing methods and apparatus, and more specifically, to video signal processing methods and apparatus for encoding or decoding video signals. Background Technology
[0003] Compression encoding refers to a series of signal processing techniques used to transmit digitized information over communication lines or to store information in a form suitable for storage media. The objects of compression encoding include objects such as speech, video, and text, and in particular, techniques used to perform compression encoding on images are called video compression. Compression encoding of video signals is performed by removing excess information, taking into account spatial, temporal, and random correlations. However, with the recent development of various media and data transmission media, there is a need for more efficient video signal processing methods and devices. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this invention is to improve the compilation efficiency of video signals.
[0006] The present invention aims to increase compilation efficiency through a secondary transformation.
[0007] Technical solution
[0008] This manual provides a method for video signal processing using quadratic transformation.
[0009] Specifically, a video signal decoding apparatus includes a processor configured to: parse syntax elements related to a secondary transform of a compilation unit from a bitstream of a video signal when one or more preset conditions are met; check, based on the parsed syntax elements, whether a secondary transform is applied to a transform block included in the compilation unit; when a secondary transform is applied to a transform block, obtain one or more inverse transform coefficients for the first sub-block by performing an inverse secondary transform based on one or more coefficients of the first sub-block, which is one or more sub-blocks constituting the transform block; and obtain residual samples for the transform block by performing an inverse primary transform based on the one or more inverse transform coefficients. The secondary transform is a low-frequency inseparable transform (LFNST), the transform block is a block to which a primary transform, separable into a vertical transform and a horizontal transform, is applied, and the first condition of the one or more preset conditions is that the index value indicating the position of the first coefficient among the one or more coefficients of the first sub-block is greater than a preset threshold.
[0010] Furthermore, according to this specification, the syntax elements include information indicating whether a quadratic transformation is applied to a compilation unit and information indicating the transformation kernel used for the quadratic transformation.
[0011] Furthermore, according to this specification, the first coefficient is the last valid coefficient according to the preset scanning order, and the valid coefficient is a non-zero coefficient.
[0012] Furthermore, according to this specification, the first sub-block is the first sub-block according to the preset scanning order.
[0013] Furthermore, according to this specification, a second condition of one or more preset conditions is that the width and height of the transform block are 4 pixels or more.
[0014] Furthermore, according to this instruction manual, the preset threshold is 0.
[0015] Furthermore, according to this instruction manual, the default scanning order is the upper right diagonal scanning order.
[0016] Furthermore, according to this specification, a third condition among one or more preset conditions is that the value of the transform skip flag included in the bitstream is not a specific value, and the transform skip flag indicates that the primary and secondary transforms were not applied to the transform block when the value of the transform skip flag has a specific value.
[0017] Furthermore, according to this specification, a fourth condition of one or more preset conditions is that at least one coefficient of one or more coefficients of the first sub-block is not zero, and at least one coefficient exists in a location other than the first position according to the preset scanning order.
[0018] Furthermore, according to this specification, the compilation unit consists of multiple compilation blocks, and when at least one transformation block corresponding to each of the multiple compilation blocks satisfies one or more preset conditions, the syntax elements related to the secondary transformation are parsed.
[0019] Furthermore, according to this specification, a video signal encoding apparatus includes a processor, wherein the processor is configured to: obtain a plurality of primary transform coefficients for a block by performing a primary transform on residual samples of a block included in a compilation unit; obtain one or more secondary transform coefficients for a first sub-block, which is one of the sub-blocks constituting the block, by performing a secondary transform based on one or more of the plurality of primary transform coefficients; and obtain a bitstream by encoding information for the one or more secondary transform coefficients and syntax elements related to the secondary transform of the compilation unit. The secondary transform is a low-frequency inseparable transform (LFNST), the primary transform can be separated into a vertical transform and a horizontal transform, and syntax elements related to the secondary transform of the compilation unit are encoded when one or more preset conditions are met, and the first condition of the one or more preset conditions is that the index value indicating the position of the first coefficient among the one or more secondary transform coefficients is greater than a preset threshold.
[0020] Furthermore, according to this specification, the syntax elements include information indicating whether a quadratic transformation is applied to a compilation unit and information indicating the transformation kernel used for the quadratic transformation.
[0021] Furthermore, according to this specification, the first coefficient is the last valid coefficient according to the preset scanning order, and the valid coefficient is a non-zero coefficient.
[0022] Furthermore, according to this specification, the first sub-block is the first sub-block according to the preset scanning order.
[0023] Furthermore, according to this specification, a second condition of one or more preset conditions is that the width and height of the initial transformation block are 4 pixels or more.
[0024] Furthermore, according to this instruction manual, the preset threshold is 0.
[0025] Furthermore, according to this instruction manual, the default scanning order is the upper right diagonal scanning order.
[0026] Furthermore, according to this specification, a third condition of one or more preset conditions is that the value of the transform skip flag included in the bitstream is not a specific value, and the transform skip flag indicates that when the transform skip flag value has a specific value, the primary and secondary transforms are not applied to the block.
[0027] Furthermore, according to this specification, a fourth condition of one or more preset conditions is that at least one of the one or more quadratic transformation coefficients is not zero, and at least one coefficient exists in a location other than the first position according to the preset scanning order.
[0028] Furthermore, according to this specification, a non-transitory computer-readable medium stores a bitstream. The bitstream is encoded using an encoding method comprising: obtaining a plurality of primary transform coefficients for a block by performing a primary transform on residual samples of a block included in a compilation unit; obtaining one or more secondary transform coefficients for a first sub-block, which is one of the sub-blocks constituting the block, by performing a secondary transform based on one or more of the plurality of primary transform coefficients; and encoding information for the one or more secondary transform coefficients and syntax elements related to the secondary transform of the compilation unit. The secondary transform is a low-frequency inseparable transform (LFNST), the primary transform can be separated into a vertical transform and a horizontal transform, and syntax elements related to the secondary transform are encoded when one or more preset conditions are met, and the first condition of the one or more preset conditions is that the index value indicating the position of the first coefficient among the one or more secondary transform coefficients is greater than a preset threshold.
[0029] Beneficial effects
[0030] Embodiments of the present invention provide a video signal processing method and apparatus using quadratic transformation. Attached Figure Description
[0031] Figure 1 This is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention.
[0032] Figure 2 This is a schematic block diagram of a video signal decoding apparatus according to an embodiment of the present invention.
[0033] Figure 3 An example is shown in which the compiler tree unit is divided into compiler units in the image.
[0034] Figure 4 An embodiment of a method for sending partitions of quadtrees and multi-type trees using signals is shown.
[0035] Figure 5 and Figure 6 The intra-frame prediction method according to an embodiment of the present invention is illustrated in more detail.
[0036] Figure 7 It is a diagram that specifically illustrates a method for transforming residual signals using an encoder.
[0037] Figure 8This diagram illustrates a method for obtaining residual signals by performing an inverse transform on the transform coefficients using an encoder and a decoder.
[0038] Figure 9 This is a graph illustrating the basis functions of multiple transformation kernels that can be used in the initial transformation.
[0039] Figure 10 This is a block diagram illustrating the process of reconstructing the residual signal in a decoding unit that performs a quadratic transformation according to an embodiment of the present invention.
[0040] Figure 11 This is a block-level diagram illustrating the process of reconstructing the residual signal in a decoding unit that performs a second transformation according to an embodiment of the present invention.
[0041] Figure 12 This is a diagram illustrating a method for applying a quadratic transformation using a reduced number of samples according to an embodiment of the present invention.
[0042] Figure 13 This is a diagram illustrating a method for determining the upper right diagonal scanning order according to an embodiment of the present invention.
[0043] Figure 14 This is a diagram illustrating the upper right diagonal scanning sequence according to an embodiment of the present invention, based on the block size diagram.
[0044] Figure 15 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level.
[0045] Figure 16 This is a diagram illustrating the residual_coding syntax structure according to an embodiment of the present invention.
[0046] Figure 17 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to an embodiment of the present invention.
[0047] Figure 18 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to an embodiment of the present invention.
[0048] Figure 19 This is a diagram illustrating the residual_coding syntax structure according to an embodiment of the present invention.
[0049] Figure 20 This is a diagram illustrating the residual_coding syntax structure according to another embodiment of the present invention.
[0050] Figure 21 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to another embodiment of the present invention.
[0051] Figure 22 This is a diagram illustrating the residual_coding syntax structure according to another embodiment of the present invention.
[0052] Figure 23 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to an embodiment of the present invention.
[0053] Figure 24 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to another embodiment of the present invention.
[0054] Figure 25 This is a diagram illustrating the syntax of a compilation unit according to an embodiment of the present invention.
[0055] Figure 26 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to another embodiment of the present invention.
[0056] Figure 27 The illustration shows a syntax structure related to the position of the last valid coefficient in the scanning order according to an embodiment of the present invention.
[0057] Figure 28 This is a diagram illustrating the residual_coding syntax structure according to another embodiment of the present invention.
[0058] Figure 29 This is a flowchart illustrating a video signal processing method according to an embodiment of the present invention. Detailed Implementation
[0059] Considering the functions of this invention, the terminology used in this specification may be currently widely used general terms, but may be changed according to the intent, customs, or emergence of new technologies of those skilled in the art. Furthermore, in some cases, terms may be arbitrarily chosen by the applicant, and in such cases, their meanings are described in the corresponding descriptive sections of the invention. Therefore, the terminology used in this specification should be interpreted based on the substantive meaning of the terms and content throughout the specification.
[0060] In this specification, some terms may be interpreted as follows. In some cases, compilation can be interpreted as encoding or decoding. In this specification, an apparatus that generates a video signal bitstream by performing encoding (compilation) of a video signal is called an encoding apparatus or encoder, and an apparatus that performs decoding (decoding) of a video signal bitstream to reconstruct a video signal is called a decoding apparatus or decoder. Additionally, in this specification, video signal processing apparatus is used as a term encompassing both the concepts of encoder and decoder. Information is a term that includes all values, parameters, coefficients, elements, etc. In some cases, the meaning is interpreted differently, therefore the invention is not limited thereto. "Unit" is used to refer to a basic unit of image processing or a specific location in an image, and refers to an image region including at least one of the luminance and chrominance components. Additionally, "block" refers to an image region including a specific component of the luminance and chrominance components (i.e., Cb and Cr). However, depending on the embodiment, terms such as "unit," "block," "partition," and "region" may be used interchangeably. Furthermore, in this specification, "unit" can be used as a concept encompassing all compilation units, prediction units, and transformation units. The image refers to a field or frame, and according to embodiments, these terms may be used interchangeably.
[0061] Figure 1 This is a schematic block diagram of a video signal encoding apparatus 100 according to an embodiment of the present invention. (See reference) Figure 1 The encoding device 100 of the present invention includes a transformation unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transformation unit 125, a filtering unit 130, a prediction unit 150, and an entropy compilation unit 160.
[0062] Transform unit 110 obtains the values of transform coefficients by transforming the residual signal, which is the difference between the input video signal and the prediction signal generated by prediction unit 150. For example, discrete cosine transform (DCT), discrete sine transform (DST), or wavelet transform can be used. DCT and DST perform the transform by dividing the input image signal into multiple blocks. During the transform, the compilation efficiency can vary depending on the distribution and characteristics of the values in the transform region. Quantization unit 115 quantizes the values of the transform coefficients output from transform unit 110.
[0063] To improve compilation efficiency, instead of compiling the image signal as is, a method is used that predicts the image using the region already compiled by prediction unit 150, and obtains the reconstructed image by adding the residual value between the original image and the predicted image to the predicted image. To prevent mismatches in the encoder and decoder, information that can be used in the decoder should be used when performing prediction in the encoder. For this purpose, the encoder performs the processing of the current block of reconstruction encoding again. Inverse quantization unit 120 inverse quantizes the values of the transform coefficients, and inverse transform unit 125 uses the inverse quantized transform coefficient values to reconstruct the residual values. Simultaneously, filtering unit 130 performs filtering operations to improve the quality of the reconstructed image and improve compilation efficiency. For example, it may include a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter. The filtered image is output or stored in decoded image buffer (DPB) 156 for use as a reference image.
[0064] To increase compilation efficiency, instead of compiling the image signal as is, a method for obtaining a reconstructed image is used. This method uses regions already compiled by prediction unit 150 to predict the image, and adds the residual value between the original image and the predicted image to the predicted image. Intra-frame prediction unit 152 performs intra-frame prediction within the current image, and inter-frame prediction unit 154 predicts the current image using a reference image stored in decoded image buffer 156. Intra-frame prediction unit 152 performs intra-frame prediction from the reconstructed region in the current image and sends intra-frame coding information to entropy compilation unit 160. Again, inter-frame prediction unit 154 may include motion estimation unit 154a and motion compensation unit 154b. Motion estimation unit 154a obtains the motion vector value of the current region by referencing a specific reconstructed region. Motion estimation unit 154a can send the location information of the reference region (reference frame, motion vector, etc.) to entropy compilation unit 160 to be included in the bitstream. Motion compensation unit 154b performs inter-frame motion compensation using the motion vector value sent from motion estimation unit 154a.
[0065] Prediction unit 150 includes intra-frame prediction unit 152 and inter-frame prediction unit 154. Intra-frame prediction unit 152 performs intra-frame prediction in the current image, and inter-frame prediction unit 154 performs inter-frame prediction to predict the current image using a reference image stored in DPB 156. Intra-frame prediction unit 152 performs intra-frame prediction based on reconstructed samples in the current image and sends intra-frame encoding information to entropy encoding unit 160. Intra-frame encoding information may include at least one of intra-frame prediction mode, most probable mode (MPM) flag, and MPM index. Intra-frame encoding information may include information about the reference samples. Inter-frame prediction unit 154 may include motion estimation unit 154a and motion compensation unit 154b. Motion estimation unit 154a references a specific region of the reconstructed reference image to obtain motion vector values for the current region. Motion estimation unit 154a sends a set of motion information about the reference region (reference image index, motion vector information, etc.) to entropy encoding unit 160. Motion compensation unit 154b uses the motion vector values sent from motion estimation unit 154a to perform motion compensation. Inter-frame prediction unit 154 sends inter-frame coding information, including a set of motion information about the reference region, to entropy compilation unit 160.
[0066] According to another embodiment, prediction unit 150 may include an intra-block copy (BC) prediction unit (not shown). The intra-BC prediction unit performs intra-BC prediction from reconstructed samples in the current image and sends intra-BC coding information to entropy compilation unit 160. The intra-BC prediction unit references a specific region in the current image and obtains block vector values indicating the reference region to be used for prediction of the current region. The intra-BC prediction unit can use the obtained block vector values to perform intra-BC prediction. The intra-BC prediction unit sends intra-BC coding information to entropy compilation unit 160. The intra-BC coding information may include block vector information.
[0067] When performing the image prediction described above, the transformation unit 110 transforms the residual values between the original image and the predicted image to obtain transformation coefficient values. In this case, the transformation can be performed on a unit of specific blocks within the image, and the size of the specific blocks can be changed within a preset range. The quantization unit 115 quantizes the transformation coefficient values generated in the transformation unit 110 and sends them to the entropy compilation unit 160.
[0068] Entropy compilation unit 160 performs entropy compilation on quantized transform coefficient information, intra-frame compilation information, and inter-frame compilation information to generate a video signal bitstream. In entropy compilation unit 160, variable-length compilation (VLC) methods, arithmetic compilation methods, etc., can be used. The VLC method transforms input symbols into continuous codewords, and the length of the codewords can be variable. For example, frequently occurring symbols are expressed as short codewords, while less frequently occurring symbols are expressed as long codewords. As a VLC method, context-based adaptive variable-length compilation (CAVLC) can be used. Arithmetic compilation transforms continuous data symbols into single decimals, and arithmetic compilation can obtain the optimal number of decimal places required to represent each symbol. As an arithmetic compilation method, context-based adaptive arithmetic compilation (CABAC) can be used. For example, entropy compilation unit 160 can binarize the information representing quantized transform coefficients. Additionally, entropy compilation unit 160 can generate a bitstream by performing arithmetic compilation on binary information.
[0069] The generated bitstream is encapsulated using Network Abstraction Layer (NAL) units as the basic unit. Each NAL unit comprises an integer number of compiled tree units. To decode the bitstream in the video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded. Simultaneously, the information required for decoding the video signal bitstream can be sent via Raw Byte Sequence Payloads (RBSPs) of higher-level sets such as Picture Parameter Sets (PPS), Sequence Parameter Sets (SPS), and Video Parameter Sets (VPS).
[0070] at the same time, Figure 1 The block diagram illustrates an encoding device 100 according to an embodiment of the present invention, and the separately displayed blocks logically distinguish and illustrate the elements of the encoding device 100. Therefore, depending on the device design, the elements of the encoding device 100 can be mounted as one or more chips. According to an embodiment, the operation of each element of the encoding device 100 can be performed by a processor (not shown).
[0071] Figure 2 This is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. (See reference) Figure 2 The decoding device 200 of the present invention includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.
[0072] Entropy decoding unit 210 performs entropy decoding on the video signal bitstream and extracts transform coefficient information, intra-frame coding information, and inter-frame coding information for each region. For example, entropy decoding unit 210 can obtain binary code for transform coefficient information for a specific region from the video signal bitstream. Furthermore, entropy decoding unit 210 obtains quantized transform coefficients by inverse binarizing the binarized code. Inverse quantization unit 220 inverse quantizes the quantized transform coefficients, and inverse transform unit 225 reconstructs the residual value using the inverse quantized transform coefficients. Video signal processing apparatus 200 reconstructs the original pixel value by adding the residual value obtained in inverse transform unit 225 to the predicted value obtained in prediction unit 250.
[0073] Simultaneously, the filtering unit 230 performs filtering on the image to improve image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire image. The filtered image is output or stored in the DPB 256 as a reference image for the next image.
[0074] Prediction unit 250 includes intra-frame prediction unit 252 and inter-frame prediction unit 254. Prediction unit 250 generates prediction images using the coding type decoded by the entropy decoding unit 210 described above, the transform coefficients of each region, and intra-frame / inter-frame coding information. To reconstruct the current block in which decoding is performed, the current image or the decoded regions of other images including the current block can be used. Images (or tiles / slices) that are used only for reconstruction (i.e., performing intra-frame prediction or intra-frame BC prediction) are called intra-frame images or I-images (or tiles / slices), and images (or tiles / slices) that perform all intra-frame prediction, inter-frame prediction, and intra-frame BC prediction are called inter-frame images (or tiles / slices). To predict the sample values of each block in an inter-frame image (or tile / slice), an image (or tile / slice) using at most one motion vector and a reference image index is called a prediction image or P-image (or tile / slice), and an image (or tile / slice) using at most two motion vectors and a reference image index is called a bidirectional prediction image or B-image (or tile / slice). In other words, a P-image (or tile / slice) uses at most one set of motion information to predict each block, and a B-image (or tile / slice) uses at most two sets of motion information to predict each block. Here, the set of motion information includes one or more motion vectors and a reference image index.
[0075] Intra-prediction unit 252 generates prediction blocks using intra-coding information and recovered samples from the current image. As described above, the intra-coding information may include at least one of intra-prediction mode, most probable mode (MPM) flag, and MPM index. Intra-prediction unit 252 predicts sample values for the current block by using recovered samples located to the left and / or above the current block as reference samples. In this disclosure, the recovered samples, reference samples, and samples of the current block can represent pixels. Moreover, sample values can represent pixel values.
[0076] According to an embodiment, the reference sample may be a sample included in the neighboring blocks of the current block. For example, the reference sample may be a sample adjacent to the left boundary of the current block and / or a sample adjacent to the top boundary. Alternatively, the reference sample may be a sample in the samples of the neighboring blocks of the current block located on a line within a predetermined distance from the left boundary of the current block and / or a sample located on a line within a predetermined distance from the top boundary of the current block. In this case, the neighboring blocks of the current block may include the left (L) block, the top (A) block, the bottom left (BL) block, the top right (AR) block, or the top left (AL) block.
[0077] Inter-frame prediction unit 254 generates prediction blocks using reference images and inter-frame coding information stored in DPB 256. The inter-frame coding information may include a set of motion information (reference image index, motion vector information, etc.) for the current block used as the reference block. Inter-frame prediction may include L0 prediction, L1 prediction, and bidirectional prediction. L0 prediction means making a prediction using a reference image included in the L0 image list, and L1 prediction means making a prediction using a reference image included in the L1 image list. For this purpose, a set of motion information (e.g., motion vectors and reference image indexes) may be required. In the bidirectional prediction method, up to two reference regions can be used, and the two reference regions may exist in the same reference image or in different images. That is, in the bidirectional prediction method, up to two sets of motion information (e.g., motion vectors and reference image indexes) can be used, and the two motion vectors may correspond to the same reference image index or different reference image indices. In this case, the reference images may be displayed (or output) before and after the current image in terms of time. According to an embodiment, the two reference regions used in the bidirectional prediction scheme may be regions selected from each of the L0 image list and the L1 image list.
[0078] The inter-frame prediction unit 254 can obtain a reference block for the current block using motion vectors and a reference image index. The reference block is located in the reference image corresponding to the reference image index. Furthermore, the sample value of the block specified by the motion vector, or its interpolated value, can be used as a predictor for the current block. For motion prediction with sub-pellet unit pixel precision, for example, an 8-tap interpolation filter for the luma signal and a 4-tap interpolation filter for the chroma signal can be used. However, the interpolation filter for motion prediction at the sub-pellet level is not limited to this. In this way, the inter-frame prediction unit 254 performs motion compensation to predict the texture of the current unit based on a motion image previously reconstructed using motion information. In this case, the inter-frame prediction unit can use a set of motion information.
[0079] According to another embodiment, prediction unit 250 may include an intra-frame BC prediction unit (not shown). The intra-frame BC prediction unit can reconstruct a current region by referencing a specific region including reconstructed samples in the current image. The intra-frame BC prediction unit obtains intra-frame BC coding information about the current region from entropy decoding unit 210. The intra-frame BC prediction unit obtains block vector values indicating a specific region in the current image. The intra-frame BC prediction unit can use the obtained block vector values to perform intra-frame BC prediction. The intra-frame BC coding information may include block vector information.
[0080] The reconstructed video image is generated by adding the predicted value output from the intra-frame prediction unit 252 or the inter-frame prediction unit 254 to the residual value output from the inverse transform unit 225. That is, the video signal decoding device 200 uses the predicted block generated by the prediction unit 250 and the residual obtained from the inverse transform unit 225 to reconstruct the current block.
[0081] at the same time, Figure 2 The block diagram illustrates a decoding device 200 according to an embodiment of the present invention, and the separately displayed blocks logically distinguish and illustrate the elements of the decoding device 200. Therefore, depending on the device design, the elements of the decoding device 200 can be mounted as one or more chips. According to an embodiment, the operation of each element of the decoding device 200 can be performed by a processor (not shown).
[0082] Figure 3The illustration shows an embodiment where a Compiler Tree Unit (CTU) in an image is segmented into Compiler Units (CUs). During the compilation of a video signal, an image can be segmented into a series of Compiler Tree Units (CTUs). A Compiler Tree Unit consists of N×N blocks of luminance samples and two blocks of corresponding chrominance samples. A Compiler Tree Unit can be segmented into multiple Compiler Units. A Compiler Tree Unit may not be segmented and may be a leaf node. In this case, the Compiler Tree Unit itself can be a Compiler Unit. A Compiler Unit refers to the basic unit used to process images during the aforementioned video signal processing, i.e., intra / inter-frame prediction, transform, quantization, and / or entropy compilation. The size and shape of a Compiler Unit in an image may not be constant. A Compiler Unit can have a square or rectangular shape. A rectangular Compiler Unit (or rectangular block) includes vertical Compiler Units (or vertical blocks) and horizontal Compiler Units (or horizontal blocks). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Furthermore, in this specification, a non-square block may refer to a rectangular block, but the invention is not limited thereto.
[0083] refer to Figure 3 First, the compiler tree unit is divided into a quadtree (QT) structure. That is, in a quadtree structure, a node of size 2×2N can be divided into four nodes of size N×N. In this specification, a quadtree can also be referred to as a quaternion tree. Quadtree partitioning can be performed recursively, and not all nodes need to be partitioned at the same depth.
[0084] Simultaneously, the leaf nodes of the aforementioned quadtree can be further divided into a multi-type tree (MTT) structure. According to embodiments of the present invention, in the MTT structure, a node can be divided into a horizontally or vertically partitioned binary or ternary tree structure. That is, in the MTT structure, there are four partitioning structures, such as vertical binary partitioning, horizontal binary partitioning, vertical ternary partitioning, and horizontal ternary partitioning. According to embodiments of the present invention, in each tree structure, the width and height of the node can both be powers of 2. For example, in a binary tree (BT) structure, a node of size 2×2N can be partitioned into two NX2N nodes through vertical binary partitioning, and further partitioned into two 2NXN nodes through horizontal binary partitioning. Additionally, in a ternary tree (TT) structure, a node of size 2×2N is partitioned into (N / 2)×2N, NX2N, and (N / 2)×2N nodes through vertical ternary partitioning, and further partitioned into 2NX(N / 2), 2NXN, and 2NX(N / 2) nodes through horizontal ternary partitioning. This multi-type tree split can be performed recursively.
[0085] Leaf nodes of multi-type trees can be compilation units. When a compilation unit is not greater than the maximum transformation length, it can be used as a unit for prediction and / or transformation without further segmentation. As an example, if the width or height of the current compilation unit is greater than the maximum transformation length, the current compilation unit can be segmented into multiple transformation units without explicit signaling regarding the segmentation. Alternatively, at least one of the following parameters for the aforementioned quadtrees and multi-type trees can be predefined or sent via an RBSP of a high-level set such as PPS, SPS, VPS, etc. 1) CTU size: Size of the root node of the quadtree; 2) Minimum QT size MinQtSize: Minimum allowed QT leaf node size; 3) Maximum BT size MaxBtSize: Maximum allowed BT root node size; 4) Maximum TT size MaxTtSize: Maximum allowed TT root node size; 5) Maximum MTT depth MaxMttDepth: Maximum allowed MTT depth derived from the leaf nodes of the QT; 6) Minimum BT size MinBtSize: Minimum allowed BT leaf node size; 7) Minimum TT size MinTtSize: Minimum allowed TT leaf node size.
[0086] Figure 4 The illustration shows an embodiment of a method for signaling quadtree and multi-type tree segmentations. Preset markers can be used to signal the aforementioned quadtree and multi-type tree segmentations. (Reference) Figure 4 At least one of the following flags can be used: "split_cu_flag" indicating whether a node has been split, "split_qt_flag" indicating whether a quadtree node has been split, "mtt_split_cu_vertical_flag" indicating the splitting direction of a multi-type tree node, or "mtt_split_cu_binary_flag" indicating the splitting shape of a multi-type tree node.
[0087] According to an embodiment of the present invention, a "split_cu_flag" signal can be sent first, which is a flag indicating whether the current node has been split. When the value of "split_cu_flag" is 0, it indicates that the current node has not been split, and the current node becomes a compilation unit. When the current node is a compilation tree unit, the compilation tree unit includes a non-splitting compilation unit. When the current node is a quadtree node "QT node", the current node is a leaf node of the quadtree "QT leaf node" and becomes a compilation unit. When the current node is a multi-type tree node "MTT node", the current node is a leaf node of the multi-type tree "MTT leaf node" and becomes a compilation unit.
[0088] When the value of "split_cu_flag" is 1, the current node can be split into nodes of a quadtree or a polytree based on the value of "split_qt_flag". The compiler tree unit is the root node of the quadtree and can be initially split into a quadtree structure. In the quadtree structure, "split_qt_flag" is sent as a signal for each node, the "QT node". When the value of "split_qt_flag" is 1, the node is split into four square nodes, while when the value of "qt_split_flag" is 0, the node becomes a leaf node of the quadtree's "QT leaf node" and is split into polytree nodes. According to embodiments of the present invention, quadtree splitting can be restricted based on the type of the current node. When the current node is a compiler tree unit (the root node of the quadtree) or a quadtree node, quadtree splitting is allowed, while when the current node is a polytree node, quadtree splitting may not be allowed. Each quadtree leaf node, the "QT leaf node", can be further split into a polytree structure. As described above, when "split_qt_flag" is 0, the current node can be split into multiple types of nodes. To indicate the splitting direction and shape, "mtt_split_cu_vertical_flag" and "mtt_split_cu_binary_flag" can be sent using signals. When the value of "mtt_split_cu_vertical_flag" is 1, it indicates a vertical split for the "MTT node," and when the value of "mtt_split_cu_vertical_flag" is 0, it indicates a horizontal split for the "MTT node." Additionally, when the value of "mtt_split_cu_binary_flag" is 1, the "MTT node" is split into two rectangular nodes, and when the value of "mtt_split_cu_binary_flag" is 0, the "MTT node" is split into three rectangular nodes.
[0089] Image prediction (motion compensation) for compilation is performed on the compilation units that are no longer segmented (i.e., the leaf nodes of the coding unit tree). The basic unit that performs this prediction will be referred to as a prediction unit or prediction block in the following text.
[0090] In the following description, the term "unit" may be used in place of "prediction unit," which is the basic unit used to perform prediction. However, the invention is not limited thereto and can be more broadly understood to include the concept of a compilation unit.
[0091] Figure 5 and Figure 6The intra-frame prediction method according to an embodiment of the present invention is illustrated in more detail. As described above, the intra-frame prediction unit predicts the sample value of the current block by using the recovered sample located to the left and / or above the current block as a reference sample.
[0092] first, Figure 5 An embodiment of reference samples for prediction of the current block in intra-frame prediction mode is shown. According to the embodiment, the reference samples may be samples adjacent to the left boundary and / or the top boundary of the current block. Figure 5 As shown, when the size of the current block is WXH and a single reference line adjacent to the current block is used for intra-frame prediction, the reference sample can be configured using the maximum 2W+2H+1 neighboring samples located to the left and top of the current block.
[0093] Additionally, if at least some of the samples to be used as reference samples have not yet been recovered, the intra-prediction unit can obtain reference samples by performing a reference sample padding process. Furthermore, the intra-prediction unit can perform reference sample filtering to reduce errors in intra-prediction. That is, filtering can be performed on the surrounding samples and / or reference samples obtained through the reference sample padding process to obtain filtered reference samples. The intra-prediction unit uses the thus obtained reference samples to predict samples for the current block. The intra-prediction unit predicts samples for the current block using either unfiltered or filtered reference samples. In this disclosure, surrounding samples may include samples on at least one reference line. For example, surrounding samples may include adjacent samples on lines adjacent to the boundary of the current block.
[0094] Next, Figure 6 An embodiment of a prediction mode for intra-frame prediction is illustrated. For intra-frame prediction, intra-frame prediction mode information indicating the direction of intra-frame prediction can be transmitted via signaling. The intra-frame prediction mode information indicates one of a plurality of intra-frame prediction modes included in the set of intra-frame prediction modes. When the current block is an intra-frame prediction block, the decoder receives the intra-frame prediction mode information of the current block from the bitstream. The decoder's intra-frame prediction unit performs intra-frame prediction on the current block based on the extracted intra-frame prediction mode information.
[0095] According to embodiments of the present invention, the intra-prediction mode set may include all intra-prediction modes used in intra-prediction (e.g., a total of 67 intra-prediction modes). More specifically, the intra-prediction mode set may include planar modes, DC modes, and multiple (e.g., 65) angle modes (i.e., orientation modes). Each intra-prediction mode can be indicated by a preset index (i.e., an intra-prediction mode index). For example, as... Figure 6As shown, intra-prediction mode index 0 indicates a planar mode, and intra-prediction mode index 1 indicates a DC mode. Furthermore, intra-prediction mode indices 2 through 66 can each indicate different angle modes. Each angle mode indicates an angle that differs from the others within a preset angle range. For example, an angle mode can indicate an angle within a clockwise angle range of 45 degrees and -135 degrees (i.e., a first angle range). An angle mode can be defined based on the 12 o'clock direction. In this case, intra-prediction mode index 2 indicates a horizontal diagonal (HDIA) mode, intra-prediction mode index 18 indicates a horizontal (horizontal, HOR) mode, intra-prediction mode index 34 indicates a diagonal (DIA) mode, intra-prediction mode index 50 indicates a vertical (VER) mode, and intra-prediction mode index 66 indicates a vertical diagonal (VDIA) mode.
[0096] Simultaneously, preset angle ranges can be set differently depending on the shape of the current block. For example, when the current block is a rectangular block, a wide-angle mode indicating an angle greater than 45 degrees or less than -135 degrees in the clockwise direction can be used. When the current block is a horizontal block, the angle mode can indicate an angle within a clockwise range of (45 + offset 1) degrees and (-135 + offset 1) degrees (i.e., the second angle range). In this case, angle modes 67 to 76, outside the first angle range, can be used. Furthermore, when the current block is a vertical block, the angle mode can indicate an angle within a clockwise range of (45 - offset 2) degrees and (-135 - offset 2) degrees (i.e., the third angle range). In this case, angle modes -10 to -1, outside the first angle range, can be used. According to embodiments of the present invention, the values of offset 1 and offset 2 can be determined differently based on the ratio between the width and height of the rectangular block. Furthermore, offset 1 and offset 2 can be positive numbers.
[0097] According to another embodiment of the present invention, the multiple angle modes included in the intra-frame prediction mode set may include a basic angle mode and an extended angle mode. In this case, the extended angle mode can be determined based on the basic angle mode.
[0098] According to an embodiment, the basic angle pattern is a pattern corresponding to an angle used in intra-frame prediction in the existing High Efficiency Video Coding (HEVC) standard, and the extended angle pattern can be a pattern corresponding to an angle newly added in intra-frame prediction in the next-generation video codec standard. More specifically, the basic angle pattern is an angle pattern corresponding to any one of the intra-frame prediction patterns {2,4,6,…,66}, while the extended angle pattern is an angle pattern corresponding to any one of the intra-frame prediction patterns {3,5,7,…,65}. That is, the extended angle pattern can be an angle pattern among the basic angle patterns within a first angle range. Therefore, the angle indicated by the extended angle pattern can be determined based on the angle indicated by the basic angle pattern.
[0099] According to another embodiment, the basic angle mode can be a mode corresponding to an angle within a preset first angle range, while the extended angle mode can be a wide-angle mode outside the first angle range. That is, the basic angle mode is an angle mode corresponding to any one of the intra-prediction modes {2, 3, 4, ..., 66}, and the extended angle mode is an angle mode corresponding to any one of the intra-prediction modes {-10, -9, ..., -1} and {67, 68, ..., 76}. The angle indicated by the extended angle mode can be determined as the opposite angle to the angle indicated by the corresponding basic angle mode. Therefore, the angle indicated by the extended angle mode can be determined based on the angle indicated by the basic angle mode. Meanwhile, the number of extended angle modes is not limited to this, and additional extended angles can be defined according to the size and / or shape of the current block. For example, the extended angle mode can be defined as an angle mode corresponding to any one of the intra-prediction modes {-14, -13, ..., -1} and {67, 68, ..., 80}. Meanwhile, the total number of intra-prediction modes included in the intra-prediction mode set can vary depending on the configuration of the basic angle mode and the extended angle mode mentioned above.
[0100] In the above embodiments, the interval between extended angle modes can be set based on the interval between corresponding basic angle modes. For example, the interval between extended angle modes {3, 5, 7, ..., 65} can be determined based on the interval between corresponding basic angle modes {2, 4, 6, ..., 66}. Similarly, the interval between extended angle modes {-10, -9, ..., -1} can be determined based on the interval between corresponding opposite-side basic angle modes {56, 57, ..., 65}, and the interval between extended angle modes {67, 68, ..., 76} can be determined based on the interval between corresponding opposite-side basic angle modes {3, 4, ..., 12}. The angular interval between extended angle modes can be configured to be the same as the angular interval between corresponding basic angle modes. Furthermore, the number of extended angle modes in the intra-frame prediction mode set can be configured to be less than or equal to the number of basic angle modes.
[0101] According to embodiments of the present invention, extended angle modes can be signaled based on a basic angle mode. For example, a wide-angle mode (i.e., an extended angle mode) can replace at least one angle mode (i.e., a basic angle mode) within a first angle range. The basic angle mode to be replaced can be an angle mode corresponding to the opposite side of the wide-angle mode. That is, the basic angle mode to be replaced is an angle mode corresponding to an angle in the opposite direction of the angle indicated by the wide-angle mode or an angle differing from the angle in the opposite direction by a preset offset index. According to embodiments of the present invention, the preset offset index is 1. The intra-frame prediction mode index corresponding to the replaced basic angle mode can be mapped back to the wide-angle mode to signal the wide-angle mode. For example, the wide-angle mode {-10, -9, ..., -1} can be signaled using the intra-frame prediction mode index {57, 58, ..., 66}, and the wide-angle mode {67, 68, ..., 76} can be signaled using the intra-frame prediction mode index {2, 3, ..., 11}. Thus, because the intra-prediction mode index used for the basic angle mode is signaled to the extended angle mode, even if the configuration of the angle mode used for intra-prediction in each block is different, the same set of intra-prediction mode indices can be used for signaling of the intra-prediction mode. Therefore, the signaling overhead caused by changes in the intra-prediction mode configuration can be minimized.
[0102] Simultaneously, the use of the extended angle mode can be determined based on at least one of the shape and size of the current block. According to one embodiment, when the size of the current block is greater than a preset size, the extended angle mode can be used for intra-frame prediction of the current block; otherwise, only the basic angle mode can be used for intra-frame prediction of the current block. According to another embodiment, when the current block is a block other than a square, the extended angle mode can be used for intra-frame prediction of the current block, and when the current block is a square block, only the basic angle mode can be used for intra-frame prediction of the current block.
[0103] On the other hand, to increase compilation efficiency, instead of compiling the residual signal as is, a method can be used to quantize the transform coefficient values obtained by transforming the residual signal and then compile the quantized transform coefficients. As described above, the transform unit can obtain transform coefficient values by transforming the residual signal. In this case, the residual signal of a specific block can be distributed across the entire region of the current block. Therefore, compilation efficiency can be improved by concentrating energy in the low-frequency domain through frequency domain transformation of the residual signal. The methods for transforming or inversely transforming the residual signal will be described in detail below.
[0104] Figure 7 This diagram specifically illustrates a method for transforming residual signals using an encoder. As described above, residual signals in the spatial domain can be transformed to the frequency domain. The encoder can obtain transform coefficients by transforming the acquired residual signals. First, the encoder can acquire at least one residual block including the residual signals of the current block. The residual block can be the current block or any of the blocks into which the current block is divided. In this disclosure, the residual block can be referred to as a residual array or residual matrix including residual samples of the current block. Furthermore, in this disclosure, the residual block can represent a transform unit or a block having the same size as the transform block.
[0105] Next, the encoder can use a transform kernel to transform the residual block. The transform kernel used to transform the residual block can be a transform kernel with characteristics that can be separated into vertical and horizontal transforms. In this case, the transformation of the residual block can be separated into vertical and horizontal transforms. For example, the encoder can perform a vertical transform by applying the transform kernel in the vertical direction of the residual block. Additionally, the encoder can perform a horizontal transform by applying the transform kernel in the horizontal direction of the residual block. In this disclosure, the term "transform kernel" can be used to refer to a set of parameters used to transform the residual signal, such as a transform matrix, transform array, and transform function. According to embodiments, the transform kernel can be any one of a plurality of available kernels. Furthermore, transform kernels based on different transform types can be used for each of the vertical and horizontal transforms.
[0106] The encoder can send a transform block, transformed from the residual block, to the quantization unit for quantization. In this case, the transform block can include multiple transform coefficients. Specifically, the transform block can consist of multiple transform coefficients arranged in two dimensions. Similar to the residual block, the size of the transform block can be the same as the size of the current block or any of the blocks into which the current block is divided. The transform coefficients sent to the quantization unit can be expressed as quantized values.
[0107] Additionally, the encoder can perform an additional transformation before quantizing the transform coefficients. For example... Figure 7As illustrated in the diagram, the above transformation method can be referred to as the primary transformation, and the additional transformation can be referred to as the secondary transformation. The secondary transformation can be selective for each residual block. According to an embodiment, the encoder can improve compilation efficiency by performing a secondary transformation on regions where it is difficult to concentrate energy in the low-frequency domain using only the primary transformation. For example, the secondary transformation can be applied to blocks in which the residual values are relatively large in directions other than the horizontal or vertical direction of the residual block. Compared to the residual values of inter-frame predicted blocks, the residual values of intra-frame predicted blocks can have a relatively higher probability of varying in directions other than the horizontal or vertical direction. Therefore, the encoder can additionally perform a secondary transformation on the residual signals of intra-frame predicted blocks. Alternatively, the encoder can omit the secondary transformation on the residual signals of inter-frame predicted blocks.
[0108] For another example, whether to perform a quadratic transformation can be determined based on the size of the current block or residual block. Additionally, transformation kernels of different sizes can be used depending on the size of the current block or residual block. For example, an 8×8 quadratic transformation can be applied to blocks in which the shorter side of the width or height is equal to or greater than a first preset length. Alternatively, a 4×4 quadratic transformation can be applied to blocks in which the shorter side of the width or height is equal to or greater than a second preset length and less than a first preset length. In this case, the first preset length can be a value greater than the second preset length; however, this disclosure is not limited thereto. Furthermore, unlike the primary transformation, the quadratic transformation may not be separable into a vertical transformation and a horizontal transformation. This quadratic transformation can be referred to as a low-frequency inseparable transformation (LFNST).
[0109] Furthermore, in the case of video signals in specific regions, sudden changes in brightness may not reduce energy in high-frequency bands even with frequency transformation. Therefore, the compression performance due to quantization may degrade. Additionally, when performing transformation on regions where residual values are scarce, encoding and decoding times may unnecessarily increase. Therefore, transformation of the residual signal in specific regions can be omitted. Whether to perform transformation on the residual signal in a specific region can be determined by syntax elements related to the transformation of that region. For example, syntax elements can include transform skip information. Transform skip information can be a transform skip flag. When the transform skip information for a residual block indicates a transform skip, the transformation of the residual block is not performed. In this case, the encoder can immediately quantize the residual signal that has not yet undergone region transformation. (Reference) Figure 7 The operation of the encoder described can be achieved through Figure 1 The transformation unit is used to perform this.
[0110] The aforementioned syntax elements related to the transform can be information parsed from the video signal bitstream. The decoder can perform entropy decoding on the video signal bitstream to obtain the transform-related syntax elements. Alternatively, the encoder can generate the video signal bitstream by entropy encoding the transform-related syntax elements.
[0111] Figure 8 This diagram specifically illustrates a method for obtaining a residual signal by performing an inverse transform on the transform coefficients using an encoder and a decoder. In the following description, for ease of explanation, the inverse transform operation will be described by the inverse transform unit of each of the encoder and decoder. The inverse transform unit can obtain the residual signal by performing an inverse transform on the inverse-quantized transform coefficients. First, the inverse transform unit can detect whether an inverse transform of a specific region has been performed from the transform-related syntax elements of that specific region. According to an embodiment, when a transform-related syntax element on a specific transform block indicates that the transform is skipped, the transform of the transform block can be omitted. In this case, both the inverse primary transform and the inverse secondary transform can be omitted for the transform block. Furthermore, the inverse-quantized transform coefficients can be used as the residual signal. For example, the decoder can reconstruct the current block using the inverse-quantized transform coefficients as the residual signal. The aforementioned inverse primary transform represents the inverse transform used for the primary transform and can be referred to as the primary inverse transform. The inverse secondary transform represents the inverse transform used for the secondary transform and can be referred to as the secondary inverse transform or inverse LFNST. In this invention, the (inverse) initial transformation can be called the first-order (inverse) transformation, and the (inverse) second-order transformation can be called the second-order (inverse) transformation.
[0112] According to another embodiment, the transform-related syntax elements for a specific transform block may not indicate transform skipping. In this case, the inverse transform unit can determine whether to perform an inverse quadratic transform on the quadratic transform. For example, when the transform block is a transform block of an intra-predicted block, an inverse quadratic transform can be performed on the transform block. Additionally, the quadratic transform kernel for the transform block can be determined based on the intra-prediction mode corresponding to the transform block. For another example, the determination of whether to perform an inverse quadratic transform can be based on the size of the transform block. The inverse quadratic transform can be performed after the inverse quantization process and before the inverse primary transform is performed.
[0113] The inverse transform unit can perform an inverse primary transform on the inversely quantized transform coefficients or the coefficients of the inverse quadratic transform. Like the primary transform, the inverse primary transform can be separated into a vertical transform and a horizontal transform. For example, the inverse transform unit can perform a vertical inverse transform and a horizontal inverse transform on a transform block to obtain a residual block. The inverse transform unit can perform an inverse transform on a transform block based on a transform kernel used to transform the transform block. For example, the encoder can explicitly or implicitly signal information indicating which of a plurality of available transform kernels should be applied to the current transform block. The decoder can select the transform kernel to be used for the inverse transform of the transform block from among the plurality of available transform kernels by using the information indicating the transform kernel signaled. The inverse transform unit can reconstruct the current block using the residual signal obtained by performing an inverse transform on the transform coefficients.
[0114] On the other hand, the distribution of the residual signal in an image can be different for each region. For example, the distribution of residual signal values in a specific region can vary depending on the prediction method. When the same transform kernel is used to transform multiple different transform regions, the compilation efficiency varies for each transform region depending on the distribution and characteristics of the values in the transform region. Therefore, compilation efficiency can be further improved by adaptively selecting a transform kernel from multiple available transform kernels for transforming a specific transform block. That is, the encoder and decoder can be configured to use transform kernels other than the basic transform kernel when transforming the video signal. The method for adaptively selecting the transform kernel can be called Adaptive Multi-Kernel Transform (AMT) or Multi-Transform Selection (MTS). In this disclosure, for ease of description, transform and inverse transform are collectively referred to as transform. Furthermore, transform kernel and inverse transform kernel are collectively referred to as transform kernel.
[0115] The residual signal, which is the difference between the original signal and the predicted signal generated through inter-frame prediction or intra-frame prediction, has energy distributed across the entire pixel domain, and therefore compression efficiency may be poor when the pixel values of the residual signal itself are encoded. Therefore, a process is needed to concentrate the energy in the low-frequency region of the frequency domain by transcoding the residual signal in the pixel domain.
[0116] In the High Efficiency Video Compilation (HEVC) standard, when the signal is uniformly distributed in the pixel domain (when neighboring pixel values are similar), the residual signal in the pixel domain is primarily transformed to the frequency domain by using an efficient Discrete Cosine Transform Type II (DCT-II) and by restricting the Discrete Sine Transform Type VII (DST-VII) to be used only in intra-frame prediction 4×4 blocks. The DCT-II transform may be suitable for residual signals generated by inter-frame prediction (when the energy is uniformly distributed in the pixel domain). However, for residual signals generated by intra-frame prediction, due to the nature of intra-frame prediction using reconstructed reference samples around the current compilation unit, the energy of the residual signal may tend to increase with increasing distance from the reference sample. Therefore, high compilation efficiency cannot be achieved when only the DCT-II transform is used to transform the residual signal to the frequency domain.
[0117] AMT is a transform technique that adaptively selects a transform kernel from several preset transform kernels based on a prediction method. Because the patterns (signal characteristics in the horizontal direction, signal characteristics in the vertical direction) in the pixel domain of the residual signal differ depending on the prediction method used, higher compilation efficiency can be expected compared to when only DCT-II is used for the transform of the residual signal. In this invention, the name AMT is not limited to what is described herein and may be referred to as Multiple Transform Selection (MTS).
[0118] Figure 9 This is a graph illustrating the basis functions of multiple transformation kernels that can be used in the initial transformation.
[0119] Specifically Figure 9 This is a graph illustrating the basis functions of the transform kernels used in AMT, and shows the kernel equations for DCT-II (Discrete Cosine Transform Type II), DCT-V (Discrete Cosine Transform Type V), DCT-VIII (Discrete Cosine Transform Type VIII), DST-I (Discrete Sine Transform Type I), and DST-VII (Discrete Sine Transform Type VII) applied to AMT.
[0120] DCT and DST can be expressed as cosine and sine functions, respectively. When the basis function of the transform kernel for a sample size N is represented as Ti(j), index i represents the index in the frequency domain, and index j represents the index in the basis function. That is, smaller i indicates a low-frequency basis function, and larger i indicates a high-frequency basis function. When expressed as a two-dimensional matrix, the basis function Ti(j) can represent the j-th element in the i-th row, and because... Figure 9All the transform kernels illustrated in the diagram are separable, allowing transformations of the residual signal X to be performed separately in the horizontal and vertical directions. That is, when the residual signal block is represented by X and the transform kernel matrix is represented by T, the transformation of the residual signal X can be represented by TXT'. In this case, T' represents the transpose of the transform kernel matrix T.
[0121] Depend on Figure 9 The values of the transformation matrix defined by the basis functions illustrated in the diagram can be in decimal form, rather than integer form. Decimal values may be difficult to implement in the hardware of video encoding and decoding devices. Therefore, an approximate integer transformation kernel from the original transformation kernel, which includes decimal values, can be used for encoding and decoding video signals. An approximate transformation kernel including integer values can be generated by scaling and rounding the original transformation kernel. The integer values included in the approximate transformation kernel can be values within a range that can be represented by a preset number of bits. The preset number of bits can be 8 bits or 10 bits. This approximation may not maintain the orthogonality of DCT and DST. However, because the resulting loss in compilation efficiency is small, approximating the transformation kernel to integer form in hardware implementation may be advantageous.
[0122] for Figure 7 and Figure 8 The primary and inverse primary transformations described herein involve two matrix multiplication operations, as the separable transformation kernel is represented as a two-dimensional matrix and the transformations are performed in both the vertical and horizontal directions. This can be seen as a problem from an implementation perspective, given the large amount of computation involved. Therefore, from an implementation standpoint, it is important to consider whether the computational load can be reduced by using a butterfly structure like DCT-II, or a combination of a semi-butterfly structure and a semi-matrix multiplier, or whether it is possible to decompose the transformation kernel into a less complex transformation kernel (whether the kernel can be expressed as a product of less complex matrices). Furthermore, since the elements of the transformation kernel (the matrix elements of the transformation kernel) must be stored in memory for computation, the memory capacity for storing the kernel matrix must also be considered in the implementation. From this perspective, because DST-VII and DCT-VIII have relatively high implementation complexity, transformations with similar characteristics and lower implementation complexity can replace DST-VII and DCT-VIII.
[0123] DST-IV (Discrete Sine Transform Type IV) and DCT-IV (Discrete Cosine Transform Type IV) can be candidates to replace DST-VII and DCT-VIII, respectively. A DCT-II kernel for 2N samples can contain a DCT-IV kernel for N samples, and a DST-IV kernel for N samples can be implemented from a DCT-IV kernel for N samples by performing a sign transformation and reversing the order of the basis functions (a simple operation). Therefore, DST-IV and DCT-IV for N samples can be easily derived from a DCT-II for 2N samples.
[0124] Because the residual signal, which is the difference between the original signal and the predicted signal, shows the characteristic that the energy distribution of the signal varies according to the prediction method, compilation efficiency can be improved when the transform kernel is adaptively selected according to the prediction method such as AMT or MTS. Additionally, as... Figure 7 and Figure 8 As described, compilation efficiency can be improved by performing a second transformation and an inverse second transformation (corresponding to the inverse of the second transformation) as additional transformations besides the primary transformation and the inverse primary transformation (corresponding to the inverse of the primary transformation). Specifically, the second transformation can improve energy compression of intra-prediction residual blocks where strong energy is likely to exist in directions other than the horizontal or vertical directions of the residual block. As mentioned above, the second transformation can be referred to as the Low-Frequency Inseparable Transform (LFNST). Furthermore, the primary transformation can be referred to as the core transformation.
[0125] Figure 10 This is a block diagram illustrating the process of reconstructing a residual signal in a decoding unit performing a quadratic transform according to an embodiment of the present invention. First, the entropy encoder can parse the syntax elements related to the residual signal from the bitstream and obtain quantization coefficients through debinarization. The decoder can perform inverse quantization on the reconstructed quantization coefficients to obtain transform coefficients, and can perform an inverse transform on the transform coefficients to reconstruct a residual signal block. The inverse transform can be applied to blocks where transform skip (TS) has not been applied. The inverse transform can be performed in the decoding unit in the order of inverse quadratic transform and inverse primary transform. In this case, the inverse quadratic transform can be omitted. The inverse quadratic transform can be performed on inter-frame prediction blocks and can be omitted. Alternatively, depending on the block size condition, the inverse quadratic transform can be omitted. The reconstructed residual signal includes quantization error, and the quadratic transform can reduce quantization error by changing the energy distribution of the residual signal compared to performing only the primary transform.
[0126] Figure 11 This diagram illustrates, at the block level, the process of reconstructing the residual signal in a decoding unit performing a secondary transformation according to an embodiment of the present invention. The reconstruction of the residual signal can be performed on a unit-by-unit (TU) or a sub-block within a TU. Figure 11 The diagram illustrates the process of reconstructing the residual signal block to which a quadratic transform has been applied, and the inverse quadratic transform can be performed first on the inverse-quantized transform coefficient block. The decoder can perform the inverse quadratic transform on all samples of size W×H (W: width, number of horizontal samples, H: height, number of vertical samples) in the TU; however, for complexity reasons, the inverse quadratic transform can be performed only on the top-left sub-block with size W'×H', which is the low-frequency region with the greatest influence. In this case, W' is less than or equal to W, and H' is less than or equal to H. The size W'×H' of the top-left sub-block can be set differently depending on the size of the TU. For example, when min(W, H) = 4, both W' and H' can be set to 4. When min(W, H) >= 8, both W' and H' can be set to 8. min(x, y) represents the operation of returning x when x is less than or equal to y, and returning y when x is greater than y. After performing the inverse quadratic transform, the decoder can obtain the transform coefficients of the upper left sub-block of size W'×H' in the TU, and can perform the inverse primary transform on the entire transform coefficient block of size W×H to reconstruct the residual signal block.
[0127] By including as a 1-bit marker in at least one of the High-Level Syntax (HLS) RBSPs such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, Slice Header, or Tile Group Header, it can be indicated whether a secondary transformation can be enabled or applied. Additionally, when a secondary transformation is applicable, the size of the top-left sub-block considered in the secondary transformation can be indicated as a 1-bit marker in at least one HLS RBSP. For example, a 1-bit marker in at least one HLS RBSP can indicate whether an 8×8 sub-block can be used for a secondary transformation considering a 4×4 or 8×8 sub-block.
[0128] When the enablement or applicability of a quadratic transform is indicated at a higher level (e.g., HLS), a 1-bit flag can be used to indicate whether a quadratic transform is applied at the compiler unit (CU) level. Additionally, when a quadratic transform is applied to the current block, an index indicating the transform kernel used for the quadratic transform can be indicated at the CU level. The decoder can perform an inverse quadratic transform on the block to which the quadratic transform has been applied, based on the prediction pattern, using the transform kernel indicated by the index within a preset transform kernel set. The index representing the transform kernel can be binarized using a truncated unary or fixed-length binarization method. A 1-bit flag indicating whether a quadratic transform is applied at the CU level and an index indicating the transform kernel used for the quadratic transform can be used as a syntax element, and in this specification, this is referred to as lfnst_idx[x0][y0] or lfnst_idx, but the invention is not limited to these names. As an example, the first bit of lfnst_idx[x0][y0] can indicate whether a quadratic transform is applied at the CU level. The remaining bits can represent the index indicating the transform kernel used for the quadratic transform. That is, lfnst_idx[x0][y0] can represent whether a quadratic transform (LFNST) is applied, and the index of the transform kernel to be used when the quadratic transform is applied. Such lfnst_idx[x0][y0] can be encoded by an entropy encoder such as context-based adaptive binary arithmetic compiler (CABAC) and context-based adaptive variable-length compiler (CAVLC). When the current CU is divided into multiple TUs smaller than the CU size, the quadratic transform may not be applied, and the syntax element lfnst_idx[x0][y0] related to the quadratic transform can be set to 0 without signaling. For example, when lfnst_idx[x0][y0] is 0, it may indicate that the quadratic transform is not applied. On the other hand, when lfnst_idx[x0][y0] is greater than 0, it may indicate that the quadratic transform is applied, and the transform kernel used for the quadratic transform can be selected based on lfnst_idx[x0][y0].
[0129] As described above, compilation tree units, leaf nodes of quadtrees, and leaf nodes of multi-type trees can be compilation units. When a compilation unit is not larger than the maximum transform length, it can be used as a unit for prediction and / or transformation without further segmentation. As an example, when the width or height of the current compilation unit is greater than the maximum transform length, the current compilation unit can be divided into multiple transform units without explicit signaling regarding segmentation. When the size of the compilation unit is greater than the maximum transform size, it can be divided into multiple transform blocks without signaling. In this case, performance degradation and complexity may increase when applying a quadratic transform, and therefore, the maximum compilation block (or the maximum size of the compilation block) for which a quadratic transform is applied can be limited. The size of the maximum compilation block can be the same as the maximum transform size. Alternatively, the size of the maximum compilation block can be defined as a preset compilation block size. As an example, the preset value can be 64, 32, or 16; however, the invention is not limited thereto. In this case, the value to be compared with the preset value (or the maximum transform size) can be defined as the length of the long side or the total number of samples.
[0130] On the other hand, the transform kernels used in the primary transform, based on DCT-II, DST-VII, and DCT-VIII basis functions, are separable. Therefore, two transforms in the vertical / horizontal directions can be performed on samples in a residual block of size N×N, and the transform kernel size can be N×N. Conversely, for the secondary transform, the transform kernel is not separable. Therefore, when the number of samples to be considered in the secondary transform is n×n, only one transform can be performed. In this case, the transform kernel size can be (n^2)×(n^2). For example, when performing a secondary transform on a 4×4 coefficient block in the upper left corner, a 16×16 transform kernel can be applied. Additionally, when performing a secondary transform on an 8×8 coefficient block in the upper left corner, a 64×64 transform kernel can be applied. A 64×64 transform kernel involves a large number of multiplication operations, which can place a heavy burden on the encoder and decoder. Therefore, when the number of samples to be considered in the secondary transform is reduced, the computational load and memory required for storing the transform kernel can be reduced.
[0131] Figure 12This diagram illustrates a method for applying a secondary transformation using a reduced number of samples according to an embodiment of the present invention. According to an embodiment of the present invention, the secondary transformation can be expressed by multiplying the secondary transformation kernel matrix by the primary transformation coefficient vector, and can be interpreted as mapping the primary transformation coefficients to another space. In this case, when the number of coefficients to be transformed is reduced, i.e., when the number of basis vectors constituting the secondary transformation kernel is reduced, the computational load required for the secondary transformation and the memory capacity required to store the transformation kernel may be reduced. For example, when performing a secondary transformation on the top-left 8×8 coefficient block, when the number of coefficients to be transformed is reduced to 16, a secondary transformation kernel of size 16 (rows) × 64 (columns) (or size 16 (rows) × 48 (columns)) can be applied. The encoder's transformation unit can obtain the secondary transformation coefficient vector by the inner product of each row vector constituting the transformation kernel matrix with the primary transformation coefficient vector. The inverse transformation units of the encoder and decoder can obtain the primary transformation coefficient vector by the inner product of each column vector constituting the transformation kernel matrix with the secondary transformation coefficient vector.
[0132] refer to Figure 12 The encoder can first perform an initial transform (forward initial transform) on the residual signal block to obtain the initial transform coefficient block. When the size of the initial transform coefficient block is M×N, for an intra-prediction block with a min(M,N) value of 4, a 4×4 quadratic transform (forward quadratic transform) can be performed on the top-left 4×4 samples of the initial transform coefficient block. For an intra-prediction block with a min(M,N) value equal to or greater than 8, an 8×8 quadratic transform can be performed on the top-left 8×8 samples of the initial transform coefficient block. Because the 8×8 quadratic transform involves a large amount of computation and memory, only some of the 8×8 samples can be utilized. In one embodiment, to improve compilation efficiency, for a rectangular block where the min(M,N) value is 4 and M or N is greater than 8 (e.g., a rectangular block of size 4×16 or 16×4), a 4×4 quadratic transform can be performed on each of the two top-left 4×4 sub-blocks in the initial transform coefficient block.
[0133] Because the second transformation can be computed by multiplying the second transformation kernel matrix by the input vector, the encoder can first construct the coefficients in the upper-left sub-block of the first transform coefficient block in vector form. The method used to construct the coefficients in vector form can depend on the intra-prediction mode. For example, when the intra-prediction mode is less than or equal to... Figure 6In the 34th angle mode of the intra-prediction mode shown, the encoder can construct coefficients in vector form by scanning the top-left sub-block of the coefficient block of the initial transform in the horizontal direction. When the element in the i-th row and j-th column of the top-left n×n block in the initial transform coefficient block is expressed as x(i,j), the vectorized coefficients can be expressed as [x(0,0),x(0,1),...,x(0,n-1),x(1,0),x(1,1),...,x(1,n-1),...,x(n-1,0),x(n-1,1),...,x(n-1,n-1)]. On the other hand, if the intra-prediction mode is greater than the 34th angle mode, the coefficients can be constructed in vector form by scanning the top-left sub-block of the coefficient block of the initial transform in the vertical direction. Vectorized coefficients can be expressed as [x(0,0),x(1,0),...,x(n-1,0),x(0,1),x(1,1),...,x(n-1,1),...,x(0,n-1),x(1,n-1),...,x(n-1,n-1)]. When only some samples from the 8×8 sample are used in the 8×8 quadratic transformation to reduce computation, the coefficients x_ij where i>3 and j>3 may not be included in the above method for constructing coefficients as vectors. In this case, in a 4×4 quadratic transformation, 16 primary transformation coefficients can be the inputs to the quadratic transformation. In an 8×8 quadratic transformation, 48 primary transformation coefficients can be the inputs to the quadratic transformation.
[0134] The encoder obtains the secondary transform coefficients by multiplying the top-left sub-block sample in the vectorized primary transform coefficient block with the secondary transform kernel matrix. The secondary transform kernel applied to the secondary transform can be determined by the size of the transform unit or transform block, the intra-frame mode, and the syntax element indicating the transform kernel. As mentioned above, reducing the number of coefficients to be secondary transformed reduces the amount of computation and memory required to store the transform kernel. Therefore, the number of coefficients to be secondary transformed can be determined by the size of the current transform block. For example, for a 4×4 block, the encoder can obtain a coefficient vector of length 8 by multiplying a vector of length 16 with an 8(row)×16(column) transform kernel matrix. The 8(row)×16(column) transform kernel matrix can be obtained based on the first to eighth basis vectors constituting the 16(row)×16(column) transform kernel matrix. For a 4×N block or M×4 (where N and M are 8 or larger), the encoder can obtain a coefficient vector of length 16 by multiplying a vector of length 16 with a 16(row)×16(column) transform kernel matrix. For an 8×8 block, the encoder can obtain a coefficient vector of length 8 by multiplying a vector of length 48 with an 8 (row) × 48 (column) transform kernel matrix. An 8 (row) × 48 (column) transform kernel matrix can be obtained based on the first to eighth basis vectors that constitute the 16 (row) × 48 (column) transform kernel matrix. For M×N blocks other than 8×8 (where M and N are 8 or greater), the encoder can obtain a coefficient vector of length 16 by multiplying a vector of length 48 with a 16 (row) × 48 (column) transform kernel matrix.
[0135] According to embodiments of the present invention, since the quadratic transformation coefficients are in vector form, they can be expressed as two-dimensional data. Coefficients that have undergone quadratic transformation according to a preset scan order can form a coefficient sub-block in the upper left corner. In an embodiment, the preset scan order can be a diagonal scan order from the upper right. The invention is not limited thereto and can be based on what will be described later. Figure 13 and Figure 14 The method described in the text determines the top-right diagonal scanning order.
[0136] Furthermore, according to embodiments of the present invention, the transform coefficients, including the total transform unit size of the quadratic transform coefficients, can be included in the bitstream and transmitted after quantization. The bitstream may include syntax elements related to the quadratic transform. Specifically, the bitstream may include information about whether the quadratic transform is applied to the current block and information indicating the transform kernel applied to the current block.
[0137] The decoder can first parse the quantized transform coefficients from the bitstream and obtain the transform coefficients through dequantization. Dequantization can be referred to as scaling. The decoder can determine whether to perform an inverse quadratic transform on the current block based on syntax elements related to the quadratic transform. When the inverse quadratic transform is applied to the current transform unit or transform block, depending on the size of the transform unit or transform block, 8 or 16 transform coefficients can be the inputs to the inverse quadratic transform. The number of coefficients to be used as inputs to the inverse quadratic transform can be matched with the number of coefficients output from the encoder's quadratic transform. For example, when the size of the transform unit or transform block is 4×4 or 8×8, 8 transform coefficients can be the inputs to the inverse quadratic transform, and otherwise, 16 transform coefficients can be the inputs to the inverse quadratic transform. When the size of the transform unit is M×N, for an intra-prediction block with a min(M,N) value of 4, a 4×4 inverse quadratic transform can be performed on 16 or 8 coefficients of the top-left 4×4 sub-block in the transform coefficient block. For intra-prediction blocks where min(M, N) is 8 or greater, an 8×8 inverse quadratic transform can be performed on the 16 or 8 coefficients of the top-left 4×4 sub-block in the transform coefficient block. In an embodiment, to improve compilation efficiency, if min(M, N) is 4 and M or N is greater than 8 (e.g., a rectangular block of size 4×16 or 16×4), an 8×4 inverse quadratic transform can be performed on each of the two top-left 4×4 sub-blocks in the transform coefficient block.
[0138] According to embodiments of the present invention, since the inverse quadratic transform can be calculated by multiplying the inverse quadratic transform kernel matrix and the input vector, the decoder can construct in vector form the dequantized transform coefficients that have been input according to a preset scan order. In embodiments, the preset scan order may be a top-right diagonal scan order, and the invention is not limited thereto, and may be based on what will be described later. Figure 13 and Figure 14 The method described in the text determines the top-right diagonal scanning order.
[0139] Furthermore, according to embodiments of the present invention, the decoder can obtain the primary transform coefficients by multiplying the vectorized transform coefficients with the inverse quadratic transform kernel matrix. In this case, the inverse quadratic transform kernel can be determined by the size of the transform unit or transform block, the intra-frame mode, and the syntax elements indicating the transform kernel. The inverse quadratic transform kernel matrix can be the transpose of the quadratic transform kernel matrix. Considering the complexity of implementation, the elements of the kernel matrix can be integers expressed with 10-bit or 8-bit precision. The length of the vector as the output of the inverse quadratic transform can be determined based on the size of the current transform block. For example, for a 4×4 block, a coefficient vector of length 16 can be obtained by multiplying a vector of length 8 with an 8(row)×16(column) transform kernel matrix. The 8(row)×16(column) transform kernel matrix can be obtained based on the first to eighth basis vectors constituting the 16(row)×16(column) transform kernel matrix. For a 4×N block or M×N (N and M are 8 or greater), a coefficient vector of length 16 can be obtained by multiplying a vector of length 16 with a 16(row)×16(column) transform kernel matrix. For an 8×8 block, a coefficient vector of length 48 can be obtained by multiplying a vector of length 8 with an 8 (rows) × 48 (columns) transformation kernel matrix. An 8 (rows) × 48 (columns) transformation kernel matrix can be obtained based on the first to eighth basis vectors that constitute the 16 (rows) × 48 (columns) transformation kernel matrix. For M×N blocks other than 8×8 (where M and N are 8 or greater), a coefficient vector of length 48 can be obtained by multiplying a vector of length 16 with a 16 (rows) × 48 (columns) transformation kernel matrix.
[0140] In this embodiment, because the primary transform coefficients obtained through the inverse quadratic transform are in vector form, the decoder can again represent them as two-dimensional data, which may depend on the intra-frame mode. In this case, the mapping relationship based on the intra-frame mode applied by the encoder can be applied equally. As described above, when the intra-frame prediction mode is less than or equal to the 34th angle mode, the decoder can obtain a two-dimensional transform coefficient array by scanning the inverse quadratic transform coefficient vector in the horizontal direction. When the intra-frame prediction mode is greater than the 34th angle mode, the decoder can obtain a two-dimensional transform coefficient array by scanning the inverse quadratic transform coefficient vector in the vertical direction. The decoder can obtain the residual signal by performing an inverse primary transform on the entire transform unit, which includes transform coefficients obtained by performing the inverse quadratic transform or a transform coefficient block of transform block size.
[0141] Although not in Figure 12 As illustrated in the diagram, in order to correct for the scaling increase caused by the transform kernel after the transform or inverse transform, a scaling process using shift operations can be included in the application of the transform or inverse transform.
[0142] Figure 13FIG. is a diagram illustrating a method for determining a top - right diagonal scan order according to an embodiment of the present invention. According to an embodiment of the present invention, a process of initializing a scan order during encoding or decoding can be performed. An array including scan order information can be initialized according to a block size. Specifically, Figure 13 the initialization process of the top - right diagonal scan order arrangement illustrated in Figure 13 is called (or executed), where the combined input for log2BlockWidth and log2BlockHeight is 1 << log2BlockWidth and 1 << log2BlockHeight. The output of the initialization process of the top - right diagonal scan order arrangement can be assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight]. Here, log2BlockWidth and log2BlockHeight are variables representing values obtained by taking the logarithm to the base 2 of the width and height of the block, respectively, and can be values in the range [0, 4].
[0143] By Figure 13 the initialization process of the top - right diagonal scan order arrangement illustrated in Figure 13 , the encoder / decoder can output an array diagScan[sPos][sComp] for blkWidth as the width of the block and blkHeight as the height of the block (all of which are received). The array index sPos can represent a scan position (scan index) and can be a value within the range [0, blkWidth * blkHeight - 1]. When sComp as the array index is 0, sPos can represent the horizontal component (x), and when sComp is 1, sPos can represent the vertical component (y). In Figure 13 the algorithm illustrated in Figure 13 , the x - coordinate and y - coordinate values on the two - dimensional coordinates at the scan position sPos can be respectively interpreted as being assigned to diagScan[sPos][0] and diagScan[sPos][1] in the top - right diagonal scan order. That is, the value stored in the DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][sComp] arrangement (array) can refer to the coordinate value corresponding to sComp at the sPos scan position (scan index) in the top - right diagonal scan order of the block, whose width and height are 1 << log2BlockWidth and 1 << log2BlockHeight, respectively.
[0144] Figure 14 FIG. is a diagram illustrating the top - right diagonal scan order according to an embodiment of the present invention according to a block size. Refer to Figure 14(a) When both log2BlockWidth and log2BlockHeight are 2, it can refer to a 4×4 block. (See reference) Figure 14 (b) When both log2BlockWidth and log2BlockHeight are 3, it can refer to an 8×8 block. Figure 14 In the diagram, the number displayed in the gray shaded area indicates the scan position (scan index) sPos. The x and y coordinate values at the sPos position can be assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][0] and DiagScanOrder[log2BlockWidth][log2BlockHeight][sPos][1], respectively.
[0145] The encoder / decoder can compile the transform coefficient information based on the scan order described above. This invention primarily describes an embodiment based on the use of the upper-right scan method; however, the invention is not limited thereto, and other known scan methods can also be applied.
[0146] The decoding process related to the quadratic transform will be described in detail below. For ease of description, the process related to the quadratic transform will be described primarily in terms of the decoder, but the embodiments described later can be applied to the encoder in essentially the same way.
[0147] Figure 15This diagram illustrates a method for indicating secondary transformations at the compilation unit level. Secondary transformations can be indicated at the compilation unit level, and syntax elements related to these transformations can be included in the `coding_unit` syntax structure. The `coding_unit` syntax structure can include syntax elements related to the compilation unit. In this case, the inputs to the `coding_unit` syntax structure are: the top-left luminance sample (x0, y0) of the image (which are the coordinates of the top-left luminance sample of the current block), `cbWidth` as the block width, `cbHeight` as the block height, and `treeType` as a variable representing the type of the compilation tree. Because there is a correlation between luminance and chrominance, efficient image compression can be achieved if luminance and chrominance are encoded using the same compilation structure. Alternatively, to improve compilation efficiency, luminance and chrominance can be encoded using different compilation structures. When the variable `treeType` is `SINGLE_TREE`, it can mean that luminance and chrominance are encoded using the same compilation tree structure, and the compilation unit can include a luminance compilation block and a chrominance compilation block according to the color format. When `treeType` is `DUAL_TREE_LUMA`, it means that luma and chroma are encoded with different compilation trees, and the currently processed tree may indicate the tree used for luma. In this case, a compilation unit may include only a luma compilation block. When `treeType` is `DUAL_TREE_CHROMA`, it means that luma and chroma are encoded with different compilation trees, and the currently processed tree may indicate the tree used for chroma. In this case, a compilation unit may include a chroma compilation block depending on the color format.
[0148] In the `coding_unit` syntax structure, the prediction method used for the current compilation unit can be indicated, and the variable `CuPredMode[x0][y0]` can indicate the prediction method used for the current block. When `CuPredMode[x0][y0]` is `MODE_INTRA`, it indicates that the intra-frame prediction method is applied to the current block, and when `CuPredMode[x0][y0]` is `MODE_INTER`, it indicates that the inter-frame prediction method is applied to the current block. Additionally, when `CuPredMode[x0][y0]` is `MODE_IBC`, it indicates that intra-block copy (IBC) prediction, which performs prediction by generating a reference block from the region where the current image has been reconstructed, is applied to the current block. Depending on the value of the variable `CuPredMode[x0][y0]`, syntax elements related to the prediction method can be processed. For example, when `CuPredMode[x0][y0]` indicates intra-frame prediction, the decoder can parse syntax elements that include information related to the intra-frame prediction mode, reference line index, and intra-segmentation sub-partition (ISP) prediction, or it can set variables related to the intra-frame prediction mode according to a preset method.
[0149] After processing the syntax elements related to the prediction method, syntax elements related to the residual signal can be processed. The `transform_tree()` syntax structure is the syntax structure for transforming the tree, and by setting a node of the same size as the compilation unit as the root node, the transform tree can be divided into nodes whose size is smaller than the root node, and the leaf nodes of the transform tree can be transform units. The `transform_tree` syntax structure can include information related to the division of the transform tree.
[0150] One method for intra-frame prediction is Pulse Compilation Modulation (PCM) prediction. When PCM prediction is used for prediction in the current compilation unit, the `transform_tree` syntax structure may not exist because no transformation or quantization is performed. That is, because the `transform_tree` syntax structure does not exist, the decoder can choose not to perform any operations on it. When intra-frame prediction is indicated in the current compilation unit, PCM prediction can be indicated by `pcm_flag[x0][y0]`. That is, when `pcm_flag[x0][y0]` is 1, the decoder can choose not to perform any operations on the `transform_tree` syntax structure. Simultaneously, the existence of a `transform_tree` syntax structure in the current compilation unit can be indicated by a 1-bit flag, referred to in this specification as `cu_cbf`, but not limited to this. The decoder can set `cu_cbf` according to a preset method when parsing `cu_cbf`, or when not parsing `cu_cbf`. When `cu_cbf` is 1, the decoder can perform operations on the `transform_tree` syntax structure. Merge prediction can also be used for prediction in the current compilation unit when inter-frame prediction or IBC prediction is used. Whether to use merge prediction can be indicated by `merge_flag[x0][y0]`. When merge prediction is indicated to be used in the current block (`merge_flag[x0][y0] == 1`), `cu_cbf` may not be resolved, and its value can be determined according to a preset method. The preset method can be based on `cu_skip_flag[x0][y0]` indicating the skip mode. For example, when `cu_skip_flag[x0][y0]` is 1, `cu_cbf` is inferred as 0; otherwise, it can be inferred as 1. When `cu_cbf` is 1, the `transform_tree` syntax structure can be processed, and the counter value used to measure the number of non-zero quantization coefficients (effective coefficients) can be initialized to 0.
[0151] The variable numSigCoeff can refer to the number of non-zero quantization coefficients (effective coefficients) present in the transformation unit of the current compilation unit, and can handle syntactic elements related to the quadratic transformation differently depending on the value of numSigCoeff.
[0152] The variable numZeroOutSigCoeff can refer to a variable representing the number of non-zero quantized coefficients (effective coefficients) at a specific location in the transformation unit included in the current compilation unit, and syntactic elements related to the quadratic transformation can be treated differently depending on the value of numZeroOutSigCoeff.
[0153] In a `transform_tree`, the transform tree can be segmented, and the leaf nodes of the transform tree can be transform units. The `transform_tree` can include a `transform_unit` syntax structure, which is a syntax structure associated with the transform units that are leaf nodes. The `transform_unit` can process syntax elements associated with the transform unit, and when the transform unit includes one or more non-zero transform coefficients, it can include a `residual_coding` syntax structure. This `residual_coding` syntax structure can include syntax structures associated with the quantization transform coefficients and their related processing. The transform blocks constituting a transform unit can vary depending on the type of tree currently being processed. When `treeType` is `SINGLE_TREE`, the current transform unit can include a luma transform block and, depending on the color format, a chroma transform block. When `treeType` is `DUAL_TREE_LUMA`, the current transform unit can include a luma transform block. When `treeType` is `DUAL_TREE_CHROMA`, the current transform unit can include a chroma transform block. The `transform_unit` syntax structure can include Compiler Block Flag (CBF) information, which, based on `treeType`, indicates whether the transform block included in the current transform unit includes one or more non-zero coefficients. CBF information can be information indicated for each color component. For example, if the CBF value of the luma transform block used for the current transform unit indicates that the luma transform block does not include one or more non-zero coefficients, then all coefficients of the luma transform block are 0, and therefore, the residual_coding syntax structure for the luma transform block may not be processed. As another example, if the CBF value of the chroma Cb transform block used for the current transform unit indicates that the chroma Cb transform block includes one or more non-zero coefficients, then the residual_coding syntax structure for the Cb transform block used for the current transform unit may exist.
[0154] At the CU level, it can be indicated whether a quadratic transform is applied to the current block. When a quadratic transform is applied, the index of the transform kernel used for the quadratic transform can be additionally indicated. (See reference...) Figure 11As described, the application of a quadratic transformation to the current block can be indicated using the `lfnst_idx[x0][y0]` syntax element. The first bit of `lfnst_idx[x0][y0]` indicates whether a quadratic transformation is applied to the current compilation unit. When the first bit of `lfnst_idx[x0][y0]` is 0, it indicates that the quadratic transformation is not applied to the current block. Conversely, when the first bit of `lfnst_idx[x0][y0]` is 1, it indicates that the quadratic transformation is applied to the current block (`lfnst_idx[x0][y0]>0`). In this case, additional bits can be used to indicate the transformation kernel used for the quadratic transformation, and the index of the quadratic transformation kernel can be signaled using additional bits.
[0155] The lfnst_idx[x0][y0] syntax element can be parsed when the conditions described later are met. On the other hand, if the conditions described later are not met, lfnst_idx[x0][y0] does not exist in the current compilation unit, and lfnst_idx[x0][y0] can be set to 0.
[0156] In other words, if the conditions described in the first to fourth embodiments, including the conditions for parsing the lfnst_idx[x0][y0] syntax element (described later), are met, the encoder can generate a bitstream including the lfnst_idx[x0][y0] syntax element for the current compilation unit. On the other hand, if the conditions described later are not met, the lfnst_idx[x0][y0] syntax element for the current compilation unit is not included in the bitstream generated by the encoder, and lfnst_idx[x0][y0] can be set to 0. The decoder receiving such a bitstream can parse the lfnst_idx[x0][y0] syntax element based on the conditions described later.
[0157] lfnst_idx[x0][y0] Syntax element parsing condition
[0158] i)Min(lfnstWidth, lfnstHeight)>=4
[0159] First, the first condition is related to the block size. When the width and height of the block are 4 pixels or more, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0160] Specifically, the decoder can check the block size condition to which a quadratic transformation can be applied. The variables SubWidthC and SubHeightC are set according to the color format and can represent the ratio of the width of the chroma component to the width of the luma component, and the ratio of the height of the chroma component to the height of the luma component, respectively. For example, because a 4:2:0 color format image has a structure where every four luma samples include one chroma sample, both SubWidthC and SubHeightC can be set to 2. For another example, because a 4:4:4 color format image has a structure where every luma sample includes one chroma sample, both SubWidthC and SubHeightC can be set to 1. lfnstWidth, as the number of samples in the horizontal direction of the current block, and lfnstHeight, as the number of samples in the vertical direction, can be set based on SubWidthC and SubHeightC. When treeType is DUAL_TREE_CHROMA, because the compilation unit only includes the chroma component, the number of samples in the horizontal direction of the chroma compilation block is equal to the value obtained by dividing cbWidth, which is the width of the luma compilation block, by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma compilation block is equal to the value obtained by dividing cbHeight, which is the height of the luma compilation block, by SubHeightC. When treeType is SINGLE_TREE or DUAL_TREE_LUMA, lnfnstWidth and lfnstHeight can be set to cbWidth and cbHeight respectively because the compilation unit includes the luma component. Since the minimum condition for a block to which a quadratic transformation can be applied is 4×4, lfnst_idx[x0][y0] can be resolved if Min(lfnstWidth, lfnstHeight)>=4.
[0161] ii) sps_lfnst_enabled_flag==1
[0162] The second condition involves a flag value indicating whether a second transformation can be enabled or applied, and when the flag value indicating whether a second transformation can be enabled or applied (sps_lfnst_enabled_flag) is set to 1, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0163] Specifically, secondary transformations can be indicated using high-level syntax RBSP. A 1-bit flag indicating whether secondary transformations can be enabled and applied can be included in at least one of the SPS, PPS, VPS, tile group header, and slice header. When sps_lfnst_enabled_flag is 1, it indicates the presence of the lfnst_idx[x0][y0] syntax element in the compiler unit syntax. When sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element does not exist in the compiler unit syntax.
[0164] iii)CuPredMode[x0][y0]==MODE_INTRA
[0165] The third condition relates to the prediction mode, and the quadratic transform can be applied only to intra-predictive blocks. Therefore, when the current block is an intra-predictive block, the decoder can parse the lfnst_idx[x0][y0] syntax elements.
[0166] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0167] The fourth condition concerns whether the ISP prediction method is applied. When ISP is not applied to the current block, the decoder can parse the lfnst_idx[x0][y0] syntax elements.
[0168] Specifically, as referenced Figure 11When the current CU is partitioned into multiple transform units smaller than the CU size, the secondary transform may not be applied to the partitioned transform units. In this case, lfnst_idx[x0][y0], as a syntax element related to the secondary transform, can be set to 0 and not resolved. When the transform tree for the current CU is partitioned into multiple transform units smaller than the CU size, ISP prediction may be applied to the current compilation unit. When intra-frame prediction is applied to the current compilation unit, the ISP prediction method can be a prediction method used to partition the transform tree into multiple transform units smaller than the CU size according to a preset partitioning method. The ISP prediction mode can be indicated at the compilation unit level, and the variable IntraSubPartitionsSplitType can be set based on it. In this case, when IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. The secondary transform is indicated at the compilation unit level, but the actual secondary transform can be applied at the transform unit level. Therefore, when the transform tree is partitioned into multiple transform units, applying the same secondary transform kernel to all partitioned transform units may be inefficient. Furthermore, due to the intra-frame prediction characteristic of generating prediction samples at the transform unit level, the prediction accuracy may be higher when the transform tree is divided into multiple transform units compared to when it is not divided. Therefore, if the transform tree is divided into multiple transform units, it is likely to effectively compress the energy of the residual signal even without applying the quadratic transform to the divided transform units. Additionally, when the current CU size is greater than the maximum luminance transform block size (MaxTbSizeY) (i.e., cbWidth > MaxTbSizeY || cbHeight > MaxTbSizeY), the transform tree can be divided into multiple transform units smaller than the CU size. Although in Figure 15 Not illustrated, even if the current CU size is larger than the maximum luminance transform block size (MaxTbSizeY), a quadratic transform may not be applied. Therefore, the fourth condition can be expressed as IntraSubPartitionsSplitType == ISP_NO_SPLIT && cbWidth <= MaxTbSizeY && cbHeight <= MaxTbSizeY. In this case, MaxTbSizeY can be a natural number expressed as an exponent of 2. MaxTbSizeY can be indicated by being included in a high-level RBSP (such as SPS, PPS, slice header, and tile group header), or the encoder and decoder can use the same preset value. For example, the preset value could be 64 (2^6).
[0169] v)!intra_mip_flag[x0][y0]
[0170] The fifth condition relates to the intra-frame prediction method. When matrix-based intra-frame prediction (MIP) is not applied to the prediction of the current compilation unit, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0171] Specifically, MIP can be used as a method for intra-frame prediction, and whether MIP is applied can be indicated at the compilation unit level via `intra_mip_flag[x0][y0]`. When `Intra_mip_flag[x0][y0]` is 1, it indicates that MIP is applied to the prediction of the current compilation unit, and prediction can be performed by multiplying the reconstructed samples around the current block with a preset matrix. Because the residual signal properties exhibited when MIP is applied differ from those of general intra-frame prediction performing directional or non-directional prediction, the quadratic transform may not be applied to the transform block when MIP is applied.
[0172] vi)numSigCoeff>((treeType==SINGLE_TREE)?2:1)
[0173] The sixth condition is related to treeType and coefficients.
[0174] Specifically, when treeType is SINGLE_TREE, when the value of variable numSigCoeff is greater than 2, the quadratic transformation can be applied to the current block, and the decoder can parse the lfnst_idx[x0][y0] syntax elements.
[0175] When `treeType` is `DUAL_TREE_LUMA` or `DUAL_TREE_CHROMA`, a quadratic transformation can be applied to the current block and `lfnst_idx[x0][y0]` can be resolved when the value of the variable `numSigCoeff` is greater than 1. In this case, `numSigCoeff` refers to a variable representing the number of valid coefficients present in the current compilation unit. When `numSigCoeff` is less than the threshold, valid encoding may not be performed even if a quadratic transformation is applied to the current block. When the number of valid coefficients is small, the overhead of signaling `lfnst_idx[x0][y0]` may be relatively large compared to the bits required for coefficient compilation. In this case, valid coefficients can refer to non-zero coefficients. In the following description, valid coefficients can refer to non-zero coefficients as described above.
[0176] vii)numZeroOutSigCoeff==0
[0177] The seventh condition involves the effective coefficients that exist at a specific location.
[0178] Specifically, when the quadratic transform is applied to the current block, the transform coefficients quantized in the decoder can always be 0 at a specific location. Therefore, since the quadratic transform is not applied to the current block when there are non-zero (quantized) coefficients at a specific location, whether to resolve lfnst_idx[x0][y0] depends on the number of valid coefficients at that location. For example, when numZeroOutSigCoeff is not 0, this means there are valid coefficients at that location, and therefore lfnst_idx[x0][y0] can be set to 0 without resolution. On the other hand, when numZeroOutSigCoeff is 0, this means there are no valid coefficients at that location, and therefore lfnst_idx[x0][y0] can be resolved.
[0179] Figure 16 This is a diagram illustrating the residual_coding syntax structure according to an embodiment of the present invention.
[0180] The `residual_coding` syntax structure can be a syntax structure related to quantization coefficients and can accept `x0`, `y0`, `log2TbWidth`, and `log2TbHeight` as inputs. In this case, `x0` and `y0` can refer to (x0, y0), which is the top-left coordinate of the transform block. `log2TbWidth` can be a value obtained by taking the logarithm of the transform block's width to the base 2, and `log2TbHeight` can be a value obtained by taking the logarithm of the transform block's height to the base 2. Coefficients in a transform block can be compiled on a sub-block basis, and the coefficient values in each sub-block can be determined based on several syntax elements including `sig_coeff_flag`. In this case, the coefficients of a sub-block can be expressed as a coefficient group (CG). `sig_coeff_flag[xC][yC]` can indicate whether the coefficient value at position (xC, yC) in the current transform block is 0. If `sig_coeff_flag[xC][yC]` is 1, it indicates that the coefficient value at that position is not 0; if `sig_coeff_flag[xC][yC]` is 0, it indicates that the coefficient value at that position is 0. In `residual_coding`, the x and y coordinates of the last valid coefficient in the scan order can be indicated. The index `lastSubBlock` of the sub-block containing the last valid coefficient in the scan order can be determined based on the x and y coordinates of the last valid coefficient in the scan order. The sub-block index can also be based on the scan order. The scan order can be referenced... Figure 13The described upper-right diagonal scan order. In sub-block unit coefficient compilation, the indices xC and yC representing the positions (coordinate values) of the coefficients can be determined based on the upper-left coordinates (xS << log2SbW, yS << log2SbH) of the sub-block and the upper-right diagonal scan order (DiagScanOrder). In this case, xS and yS represent the indices in the horizontal direction and the vertical direction, respectively. log2SbW and log2SbH can be values obtained by taking the base-2 logarithm of the width and height of the sub-block.
[0181] When the value of sig_coeff_flag[xC][yC] is 1 (i.e., the coefficient at position (xC, yC) is not 0) and transform skip is not applied to the current block (i.e.,!transform_skip_flag[x0][y0]), numSigCoeff can be counted. When transform skip is applied, since the secondary transform may not be applied, the numSigCoeff used to parse lfnst_idx[x0][y0] can count the number of valid coefficients of the block to which transform skip is not applied.
[0182] In addition, as described in Figure 15 When the secondary transform is applied to the transform block, there may be no valid coefficients in a specific region within the transform block. Therefore, the numZeroOutSigCoeff counter counts the number of valid coefficients (numZeroOutSigCoeff) present in the specific region, and when numZeroOutSigCoeff is not 0, lfnst_idx[x0][y0] may not be parsed. Specifically, when the secondary transform is applied to the transform block, the region where there cannot be valid coefficients can be determined according to the size of the transform block.
[0183] For example, to apply the secondary transform, when the size of the transform block is 4×4 (i.e., log2TbWidth == 2 && log2TbHeight == 2), the index [0, 7] region and the index [8, 15] region can be divided in the scan order within the transform block such that valid coefficients can exist in the [0, 7] region and cannot exist in the [8, 15] region. The 4×4 transform block can include one sub-block. Therefore, when the size of the transform block is 4×4, when the scan position is 8 or greater and the index of the sub-block is 0 (i.e., n >= 8 && i == 0), the number of valid coefficients can be calculated. In this case, the scan order can be the upper-right diagonal scan order.
[0184] For another example, to apply the quadratic transformation, when the size of the transform block is 8×8 (i.e., log2TbWidth == 3 && log2TbHeight == 3), valid coefficients may exist only in the first sub-block of the transform block, and may not exist in the remaining sub-blocks (e.g., the second and third sub-blocks). Even within the first sub-block, valid coefficients may appear in the index [0, 7] region in scan order, but valid coefficients may not exist in the index [8, 15] region. Therefore, when the size of the transform block is 8×8, the number of valid coefficients can be counted when the scan position in the first sub-block is 8 or greater (i.e., n >= 8 && i == 0) or the scan position exists in the remaining sub-blocks other than the first sub-block (e.g., in the second and third sub-blocks, i == 1 || i == 2).
[0185] Finally, when the size of the transform is greater than 8×8, valid coefficients may exist only in the first sub-block of the transform block, and may not exist in the remaining sub-blocks (e.g., the second and third sub-blocks). Therefore, the number of valid coefficients can be counted when the sub-block is the second or third block (i.e., i==1||i==2). Similar to the numSigCoeff counter, the numZeroOutSigCoeff counter can only count the number of valid coefficients when sig_coeff_flag[xC][yC] is 1 and transform_skip_flag[x0][y0] is 0. In this case, it can be determined according to the reference. Figure 13 The description describes the top-right diagonal scan order for indexing sub-blocks.
[0186] In other words, since the existence of non-zero coefficients in regions (specific regions) where there may be no valid coefficients indicates that a quadratic transformation should not be performed, valid coefficients are counted in order to check whether there are non-zero coefficients in specific regions.
[0187] Figure 17 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to an embodiment of the present invention.
[0188] like Figure 15 and Figure 16As described above, the application of a quadratic transformation can be indicated at the compilation unit level via the syntax element `lfnst_idx[x0][y0]`, and resolving `lfnst_idx[x0][y0]` may require two valid coefficient counters (i.e., the `numSigCoeff` counter and the `numZeroOutSigCoeff` counter). In particular, in the case of `numSigCoeff`, the throughput of coefficient compilation may be reduced because the `numSigCoeff` counter has to count the number of valid coefficients present throughout the entire compilation unit region. Therefore, a method is needed to reduce the number of counters or eliminate the use of counters altogether.
[0189] Figure 17 The quadratic transformation indicator method illustrated in the figure is a method that can parse lfnst_idx[x0][y0] regardless of numSigCoeff. In other words, if in Figure 15 If conditions i), ii), iii), iv), v), and vii) described in the code are all satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. Furthermore, since the value of numSigCoeff is not referenced, execution is unnecessary. Figure 16 The operation of the numSigCoeff counter described in [the document].
[0190] The following specification describes a method for indicating a secondary transformation based on the position information of the last valid coefficient in the scan order. Similar to a small number of valid coefficients, when the position (scan index) of the last valid coefficient in the scan order is small, the compilation efficiency due to the secondary transformation may be low. Therefore, it is necessary to efficiently indicate the secondary transformation based on the position information of the last valid coefficient in the scan order without using a counter.
[0191] (First Embodiment)
[0192] Figure 18 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to an embodiment of the present invention.
[0193] Figure 18 This is a diagram illustrating a method for resolving lfnst_idx[x0][y0] by using the position information of the last valid coefficient in the scan order obtained from residual_coding instead of the numSigCoeff counter.
[0194] refer to Figure 18Because the numSigCoeff counter is not used, it is not necessary to initialize the numSigCoeff value, and the variable lfnstLastScanPos, which relates to the position of the last valid coefficient in the scan order, can be initialized to 1. When lfnstLastScanPos is 1, it indicates that the position (scan index) of the last valid coefficient in the scan order is less than a threshold, or that all transform coefficients in the block are 0. On the other hand, when lfnstLastScanPos is 0, it indicates that there is at least one valid coefficient in the block, and the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold. Therefore, if lfnstLastScanPos is 1, lfnst_idx[x0][y0] may not be parsed, and if lfnstLastScanPos is 0, lfnst_idx[x0][y0] can be parsed. Additionally, if lfnstLastScanPos is 0, and... Figure 15 If all of the conditions i), ii), iii), iv), v), and vii) described in the text are satisfied (if all are true), then lfnst_idx[x0][y0] can be parsed.
[0195] In other words, lfnst_idx[x0][y0] can be resolved when there is at least one valid coefficient in the current block and the position (scan index) of the last valid coefficient in the scan order is equal to or greater than a threshold. In this case, as described later, the threshold can be an integer equal to or greater than 0. For example, assuming the threshold is 1, the fact that the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold may mean that the valid coefficient exists in a position other than the top left of the block. That is, lfnst_idx[x0][y0] can be resolved only when the valid coefficient does not exist in the current block or exists only in the top left of the current block. The meaning of the existence of a valid coefficient in a position other than the top left of the current block can be expressed as "LfnstDConly == 0". The top left of the block described in this specification can mean that the value of the horizontal and vertical coordinates is (0, 0), can refer to the first position in a preset scan order (e.g., top right diagonal order), or can be referred to as DC.
[0196] Figure 19 This is a diagram illustrating the residual_coding syntax structure according to an embodiment of the present invention.
[0197] Figure 19 The diagram is as described above. Figure 18The described residual_coding syntax structure allows for the parsing of syntax elements related to the x and y coordinates of the last valid coefficient in the scan order, enabling the setting of the variables LastSignificantCoeffX and LastSignificantCoeffY. LastSignificantCoeffX represents the x-coordinate of the last valid coefficient in the scan order, and LastSignificantCoeffY represents the y-coordinate. Based on LastSignificantCoeffX and LastSignificantCoeffY, the lastScanPos variable, which serves as the scan index of the last valid coefficient in the scan order, and the index lastSubBlock, which includes the last valid coefficient, can be determined. In this case, as referenced... Figure 16 As described, when a quadratic transformation is applied to the current block, only the first sub-block can have valid coefficients. In other words, a quadratic transformation can be applied when valid coefficients exist only in the first sub-block.
[0198] For example, when in Figure 14 In (a), when LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3 in a 4×4 block, we can determine that lastScanPos is 13. Since a 4×4 block can consist of a sub-block, we can determine that the index lastSubBlock of the sub-block containing the last valid coefficient is 0. For example, Figure 14 (b) The 8×8 block can be divided into 4×4 sub-blocks. Specifically, in Figure 14 In (b), the 4×4 blocks corresponding to x-coordinates 0 to 3 and y-coordinates 0 to 3 can be set as the first sub-block, the 4×4 blocks corresponding to x-coordinates 0 to 3 and y-coordinates 4 can be set as the second sub-block, the 4×4 blocks corresponding to x-coordinates 4 to 7 and y-coordinates 0 to 3 can be set as the third sub-block, and the 4×4 blocks corresponding to x-coordinates 4 to 7 and y-coordinates 4 to 7 can be set as the fourth sub-block. In this case, the first sub-block can be indexed as index 0, the second sub-block as index 1, the third sub-block as index 2, and the fourth sub-block as index 3. This can be done according to the reference... Figure 13The top-right diagonal scan order is described to index subblocks. In this case, when LastSignificantCoeffX is 2 and LastSignificantCoeffY is 3, lastScanPos is determined to be 13. Because lastScanPos is 13, the subblock including lastScanPos 13 is the first subblock (i.e., subblock index 0), and therefore the index (lastSubBlock) of the subblock including the last valid coefficient can be determined as 0.
[0199] Based on the above `lastScanPos`, `lfnstLastScanPos` can be determined. Specifically, when the width and height of the transform block are 4 or greater and transform skipping is not applied to the transform block, `lfnstLastScanPos` can be set as shown in Equation 1 below. In other words, when `log2TbWidth>=2`, `log2TbHeight>=2`, and `transform_skip_flag[x0][y0]` is 0, `lfnstLastScanPos` can be set as shown in Equation 1 below. In this case, when `transform_skip_flag[x0][y0]` is 0, it may mean that the transform skipping is not applied to the current transform block. Specifically, the flag `transform_skip_flag[x0][y0]` described in this specification can indicate whether the primary and secondary transforms are applied to the transform block. For example, when the value of transform_skip_flag[x0][y0] is 1, it can indicate that the primary and secondary transformations are not applied to the transform block (i.e., transformation skip is applied), and when the value of transform_skip_flag[x0][y0] is 0, it can indicate that the primary and secondary transformations can be applied to the transform block (i.e., transformation skip is not applied).
[0200] [Equation 1]
[0201] lfnstLastScanPos=lfnstLastScanPos&&(lastScanPos <lfnstLastScanPosTh[cldx])
[0202] As mentioned above, the initial value of lfnstLastScanPos can be set to 1.
[0203] In Equation 1, cIdx can represent a variable indicating the color component of the current transform block. For example, when cIdx is 0, it can indicate that the transform block to be processed in residual_coding is the luminance Y component. When cIdx is 1, it can indicate that the transform block to be processed in residual_coding is the chrominance Cb component, and when cIdx is 2, it can indicate that the transform block to be processed is the chrominance Cr component. The threshold lfnstLastScanPosTh[cIdx] of lastScanPos can be set to different values depending on the color component.
[0204] According to Equation 1, lfnstLastScanPos can be updated to 1 when the immediately preceding lfnstLastScanPos is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx]. Conversely, lfnstLastScanPos can be updated to 0 when the immediately preceding lfnstLastScanPos is 0 or lastScanPos is equal to or greater than lfnstScanPosTh[cIdx]. In other words, if the lastScanPos of all transform blocks included in the compilation unit is less than the threshold or the coefficients of all transform blocks are 0, it can be determined that lfnstLastScanPos is 1, and lfnst_idx[x0][y0] can be set to 0 without needing to consider... Figure 18 The parsing condition of lfnst_idx[x0][y0] is used for parsing. The fact that lfnst_idx[x0][y0] is not parsed and is set to 0 indicates that the second transformation is not applied to the current block. On the other hand, if any transformation block included in the compilation unit has a lastScanPos equal to or greater than the threshold, it can be determined that lfnstLastScanPos is 0, and if in Figure 15 If all conditions i), ii), iii), iv), v), and vii) described in the code are satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform was applied to the current block, and if a quadratic transform was applied, the transform kernel used for the quadratic transform can be checked / determined.
[0205] The lfnstLastScanPosTh[cIdx] in Equation 1 is a preset integer value equal to or greater than 0, and both the encoder and the decoder can use the same value. Additionally, the same threshold can be used for all color components. In this case, lfnstLastScanPos can be set as in Equation 2 below. The compilation units described in this specification can include multiple compilation blocks, and there can be transform blocks corresponding to each compilation block. The transform blocks can be transform blocks with luminance and chrominance components. Specifically, the transform blocks can be Y transform blocks, Cb transform blocks, or Cr transform blocks. In this case, it can be determined whether to parse lfnst_idx[x0][y0] described in this specification for each transform block corresponding to each compilation block. That is, when any one of the Y transform block, Cb transform block, and Cr transform block satisfies the conditions described in this specification, lfnst_idx[x0][y0] can be parsed.
[0206] [Equation 2]
[0207] lfnstLastScanPos = lfnstLastScanPos && (lastScanPos < lfnstLastScanPosTh) The lfnstLastScanPosTh is a preset integer value equal to or greater than 0, and both the encoder and the decoder can use the same value. For example, lfnstLastScanPosTh can be 1. That is, when lastScanPos is 1 or greater, lfnstLastScanPos can be updated to 0, and lfnst_idx[x0][y0] can be parsed. In this case, since the threshold lfnstLastScanPosTh is an integer value, the case where lastScanPos is 1 or greater can have the same meaning as the case where lastScanPos is greater than 0. As an example of the present invention, the case where the threshold is 1 has been described, however, the present invention is not limited thereto.
[0208] In other words, whether to parse lfnst_idx[x0][y0] can be determined based on lastScanPos. Specifically, as mentioned above, when a quadratic transform is applied, the last valid coefficient in the scan order may only exist in the first sub-block of the transform block. Therefore, lfnst_idx[x0][y0] can be parsed when the index lastSubBlock (the position of the index indicated by lastScanPos) of the sub-block containing the last valid coefficient in the scan order is 0, the width of the transform block is 4 or greater (log2TbWidth>=2), the height of the transform block is 4 or greater (log2TbHeight>=2), transform_skip_flag[x0][y0] is 0 (no transform skipped), and lastScanPos is greater than 0 (lastScanPos is 1 or greater). This can be expressed as shown in Equation 3 below.
[0209] [Equation 3]
[0210] lastSubBlock==0&&log2TbWidth>=2&&log2TbHeight>=2&&
[0211] ! transform_skip_flag[x0][y0][cldx]&&lastScanPos>0
[0212] Meanwhile, in the first embodiment above, since the numSigCoeff counter is not used to parse lfnst_idx[x0][y0], the number of valid coefficients numSigCoeff does not need to be calculated.
[0213] (Second Embodiment)
[0214] Figure 20 This is a diagram illustrating the residual_coding syntax structure according to another embodiment of the present invention.
[0215] Figure 20 The diagram shows that for residual_coding, besides Figure 19 In addition, it also receives the treeType variable and sets the threshold of lastScanPos based on treeType.
[0216] When the width and height of the transform block are 4 or greater and transform skip is not applied to the transform block, lfnstLastScanPos can be set as shown in Equation 4 below. In other words, when log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos can be set as shown in Equation 4 below. In this case, when transform_skip_flag[x0][y0] is 0, it may mean that the transform skip is not applied to the current transform block.
[0217] [Equation 4]
[0218] lfnstLastScanPosTh=(treeType==SlNGLE_TREE)? val1: ((treeType==DUAL_TREE_LUMA)? val2: val3)
[0219] lfnstLastScanPos=lfnstLastScanPos&&(lastScanPos <lfnstLastScanPosTh)
[0220] In Equation 4, lfnstLastScanPosTh refers to the threshold of lastScanPos, and this value can be set according to treeType. When treeType is SINGLE_TREE, DUAL_TREE_LUMA, and DUAL_TREE_CHROMA, lfnstLastScanPosTh can be set to val1, val2, and val3, respectively. When the immediately preceding lfnstLastScanPos is 1 and lastScanPos is less than lfnstLastScanPosTh, lfnstLastScanPos can be updated to 1. On the other hand, when the immediately preceding lfnstLastScanPos is 0 or lastScanPos is equal to or greater than lfnstScanPosTh, lfnstLastScanPos can be updated to 0.
[0221] As a result, in Equation 4, when the lastScanPos of all transform blocks included in the compilation unit is less than the threshold or the coefficients of all transform blocks are 0, it can be determined that lfnstLastScanPos is 1, and lfnst_idx[x0][y0] can be set to 0 without considering... Figure 18The resolution condition of lfnst_idx[x0][y0] is resolved. This indicates that the quadratic transformation was not applied to the current block. On the other hand, if the lastScanPos of any of the transformation blocks included in the compilation unit is equal to or greater than the threshold, it can be determined that lfnstLastScanPos is 0, and if in Figure 15 If all of the conditions described in i), ii), iii), iv), v), and vii) are satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform was applied to the current block, and if a quadratic transform was applied, the transform kernel used for the quadratic transform can be checked / determined.
[0222] val1, val2, and val3 are preset integer values equal to or greater than 0, and both the encoder and decoder can use the same values. When treeType is SINGLE_TREE, both luminance and chrominance components are included, and therefore the value val1 of lfnstLastScanPosTh can be expressed as the sum of val2 and val3.
[0223] In the second embodiment described above, since the numSigCoeff counter is not used to parse lfnst_idx[x0][y0], the number of valid coefficients numSigCoeff can be left uncounted.
[0224] (Third Embodiment)
[0225] Figure 21 This is a diagram illustrating a method for indicating a secondary transformation at the compiler unit level according to another embodiment of the present invention.
[0226] refer to Figure 21 lfnst_idx[x0][y0] can be resolved by using the position information of the last valid coefficient in the scan order obtained from residual_coding instead of the numSigCoeff counter.
[0227] Since the numSigCoeff counter is not used, it is not necessary to initialize numSigCoeff, and the variable lfnstLastScanPos, which relates to the position of the last valid coefficient in the scan order, can be initialized to 0. Figure 21 The variable lfnstLastScanPos can be a value obtained by adding the lastScanPos values of the transform blocks included in the compilation unit. In this case, if lfnstLastScanPos is greater than a threshold and Figure 15If all conditions i), ii), iii), iv), v), and vii) described in the code are satisfied (if all are true), the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform was applied to the current block, and if so, to check / determine the transform kernel used for the quadratic transform. On the other hand, when lfnstLastScanPos is less than or equal to the threshold, lfnst_idx[x0][y0] can be set to 0 and not parsed. This indicates that no quadratic transform was applied.
[0228] The threshold can be set according to `treeType`. When `treeType` is `SINGLE_TREE`, `DUAL_TREE_LUMA`, or `DUAL_TREE_CHROMA`, the threshold can be set to `Th1`, `Th2`, and `Th3`, respectively. `Th1`, `Th2`, and `Th3` are preset integer values equal to or greater than 0, and both the encoder and decoder can use the same values. When `treeType` is `SINGLE_TREE`, both luminance and chrominance components are included, and therefore `Th1` as the threshold can be expressed as the sum of `Th2` and `Th3` as thresholds.
[0229] Figure 22 This is a diagram illustrating the residual_coding syntax structure according to another embodiment of the present invention.
[0230] Figure 22 Illustrated reference Figure 21 The residual_coding syntax structure is described, and when the width and height of the transform block are 4 or greater and transform skip is not applied to the transform block, lfnstLastScanPos can be set as shown in Equation 5 below. In other words, when log2TbWidth>=2, log2TbHeight>=2, and transform_skip_flag[x0][y0] is 0, lfnstLastScanPos can be set as shown in Equation 5 below. In this case, when transform_skip_flag[x0][y0] is 0, it may mean that the transform skip is not applied to the current transform block.
[0231] [Equation 5]
[0232] lfnstLastScanPos=lfnstLastScanPos+lastScanPos
[0233] In Equation 5 above, lfnstLastScanPos is the value obtained by adding up all lastScanPos of the transform blocks included in the compilation unit. For example... Figure 21 As described in the document, it can be determined whether to parse lfnst_idx[x0][y0] by comparing lfnstLastScanPos with a threshold.
[0234] In the third embodiment described above, since the numSigCoeff counter is not used to parse lfnst_idx[x0][y0], the number of valid coefficients numSigCoeff can be left uncounted.
[0235] On the other hand, a compilation unit may include a transformation unit divided by a transformation tree, with the same size as the compilation block as the root node. In this case, the transformation unit may include a transformation block for each color component. When a quadratic transformation is indicated at the compilation unit level, lfnst_idx[x0][y0] can be resolved based on coefficient information after residual compilation is performed on all transformation blocks included in the compilation unit. In another embodiment, a quadratic transformation may be indicated at the transformation unit level. When a quadratic transformation is indicated at the transformation unit level, each transformation unit included in the compilation unit can use a different lfnst_idx[x0][y0]. Therefore, the encoder can find an optimized lfnst_idx[x0][y0] for each transformation unit and can further improve encoding efficiency. In addition, when a quadratic transformation is indicated at the compilation unit level and the compilation unit includes four transformation units, residual compilation for all transformation blocks included in the four transformation units is processed so that lfnst_idx[x0][y0] can be resolved. That is, even if the decoder obtains the transform coefficients through residual compilation for the first transform unit, it cannot perform the inverse transform on the first transform unit because it does not obtain the lfnst_idx[x0][y0] value. This could not only increase the decoder's buffer size but also lead to excessive latency in the decoder.
[0236] Figures 18 to 22 The first to third embodiments described herein can be applied even when a secondary transformation is indicated at the transformation unit level. When a secondary transformation is indicated at the compilation unit level, according to the first to third embodiments, whether to parse lfnst_idx[x0][y0] can be determined based on the position of the last valid coefficient in the scan order of the transformation blocks included in the compilation unit. Furthermore, when a secondary transformation is indicated at the transformation unit level, according to the first to third embodiments, whether to parse lfnst_idx[x0][y0] can be determined based on the position of the last valid coefficient in the scan order of the transformation blocks included in the transformation unit.
[0237] The following section will describe the specific method for indicating the secondary transformation at the transformation unit level.
[0238] Figure 23 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to an embodiment of the present invention.
[0239] refer to Figure 23 lfnst_idx[x0][y0] can be resolved by using the position information of the last valid coefficient in the scan order obtained from residual_coding instead of the numSigCoeff counter.
[0240] First, before performing residual_coding, the variable lfnstLastScanPos, which relates to the position of the last valid coefficient in the scan order, can be initialized to 1. When the variable lfnstLastScanPos is 1, it indicates that the position (scan index) of the last valid coefficient in the scan order of all transform blocks included in the transform unit is less than a threshold or that all transform coefficients in the block are 0. When the variable lfnstLastScanPos is 0, it indicates that for one or more transform blocks included in the transform unit, there are one or more valid coefficients in the block, and the position (scan index) of the last valid coefficient in the scan order is equal to or greater than a threshold. According to the first embodiment described above, if lfnstLastScanPos, which is set to 0 based on the position of the last valid coefficient in the scan order of the transform block, is 0, and conditions i), ii), iii), iv), v), and vi), which will be described later, are all satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0].
[0241] lfnst_idx[x0][y0] Syntax element parsing condition
[0242] i)Min(lfnstWidth, lfnstHeight)>=4
[0243] First, the first condition is related to the block size. When the width and height of the block are 4 pixels or more, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0244] Specifically, the decoder can check the block size condition to which a quadratic transform can be applied. The variables SubWidthC and SubHeightC are set according to the color format and can represent the ratio of the width of the chroma component to the width of the luminance component, and the ratio of the height of the chroma component to the height of the luminance component, respectively. For example, because a 4:2:0 color format image has a structure where every four luminance samples include one chroma sample, both SubWidthC and SubHeightC can be set to 2. For another example, because a 4:4:4 color format image has a structure where every luminance sample includes one chroma sample, both SubWidthC and SubHeightC can be set to 1. lfnstWidth, the number of samples in the horizontal direction of the current block, and lfnstHeight, the number of samples in the vertical direction, can be set based on SubWidthC and SubHeightC. When treeType is DUAL_TREE_CHROMA, because the transform unit only contains the chroma component, the number of samples in the horizontal direction of the chroma transform block is equal to the value obtained by dividing tbWidth, which is the width of the luminance transform block, by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma transform block is equal to the value obtained by dividing tbHeight, which is the height of the luminance transform block, by SubHeightC. When treeType is SINGLE_TREE or DUAL_TREE_LUMA, lnfnstWidth and lfnstHeight can be set to tbWidth and tbHeight respectively because the transform unit includes the luminance component. Since the minimum condition for a block to which a quadratic transform can be applied is 4×4, lfnst_idx[x0][y0] can be resolved if Min(lfnstWidth, lfnstHeight)>=4.
[0245] ii) sps_lfnst_enabled_flag==1
[0246] The second condition involves a flag value indicating whether a second transformation can be enabled or applied, and when the flag value indicating whether a second transformation can be enabled or applied (sps_lfnst_enabled_flag) is set to 1, the decoder can parse lfnst_idx[x0][y0].
[0247] Specifically, secondary transformations can be indicated using high-level syntax RBSP. A 1-bit flag indicating whether secondary transformations can be enabled or applied can be included in at least one of the SPS, PPS, VPS, tile group header, and slice header. When sps_lfnst_enabled_flag is 1, it indicates the presence of the lfnst_idx[x0][y0] syntax element in the transform unit syntax. When sps_lfnst_enabled_flag is 0, it indicates that the lfnst_idx[x0][y0] syntax element does not exist in the transform unit syntax.
[0248] iii)CuPredMode[x0][y0]==MODE_INTRA
[0249] The third condition relates to the prediction mode, and the quadratic transform can be applied only to intra-prediction blocks. Therefore, when the current block is an intra-prediction block, the decoder can resolve lfnst_idx[x0][y0].
[0250] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0251] The fourth condition concerns whether the ISP prediction method is applied. When ISP is not applied to the current block, the decoder can parse the lfnst_idx[x0][y0] syntax elements.
[0252] Specifically, as referenced Figure 11As described, when the current CU is partitioned into multiple transform units smaller than the CU size, the quadratic transform may not be applied to the partitioned transform units. In this case, lfnst_idx[x0][y0], as a syntax element related to the quadratic transform, can be set to 0 and not parsed. When the transform tree for the current CU is partitioned into multiple transform units smaller than the CU size, ISP prediction can be applied to the current compilation unit. When intra-frame prediction is applied to the current compilation unit, the ISP prediction method can be a prediction method used to partition the transform tree into multiple transform units smaller than the CU size according to a preset partitioning method. The ISP prediction mode can be indicated at the compilation unit level, and the variable IntraSubPartitionsSplitType can be set based on it. In this case, when IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristics of intra-frame prediction that generates prediction samples at the transform unit level, the prediction accuracy may be higher when the transform tree is partitioned into multiple transform units compared to when the transform tree is not partitioned. Therefore, even if the quadratic transform is not applied to the partitioned multiple transform units, it is very likely that the energy of the residual signal can be effectively compressed.
[0253] v)!intra_mip_flag[x0][y0]
[0254] The fifth condition relates to the intra-frame prediction method. When matrix-based intra-frame prediction (MIP) is not applied to the prediction of the current compilation unit, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0255] Specifically, matrix-based intra-frame prediction (MIP) can be used as a method for intra-frame prediction, and whether MIP is applied can be indicated at the compilation unit level via `intra_mip_flag[x0][y0]`. When `intra_mip_flag[x0][y0]` is 1, it indicates that MIP is applied to the prediction of the current compilation unit, and prediction can be performed by multiplying the reconstructed samples around the current block with a preset matrix. Because the residual signal properties exhibited when MIP is applied differ from those of general intra-frame prediction performing directional or non-directional prediction, the quadratic transform may not be applied to the transform block when MIP is applied.
[0256] vi)numZeroOutSigCoeff==0
[0257] The sixth condition involves the effective coefficients that exist at a specific location.
[0258] Specifically, when the quadratic transform is applied to the current block, the transform coefficients quantized in the decoder can always be 0 at a specific location. Therefore, since the quadratic transform is not applied when there are non-zero quantized coefficients at a specific location, lfnst_idx[x0][y0] can be resolved based on the number of valid coefficients at that location. For example, when numZeroOutSigCoeff is not 0, it means there are valid coefficients at that location, and therefore lfnst_idx[x0][y0] can be set to 0 without resolution. On the other hand, when numZeroOutSigCoeff is 0, it means there are no valid coefficients at that location, and therefore lfnst_idx[x0][y0] can be resolved.
[0259] When indicating whether a secondary transformation is applied to the current block at the transformation unit level based on the first embodiment described above, it can be followed Figure 19 The residual_coding method described in [the document] will be applied if the lastScanPos of all transform blocks included in the transform unit is less than [the value specified in the document]. Figure 19 If the threshold of Equation 1 used to determine lfnstLastScanPos, as described in [the document], or if all coefficients of the transform block are 0, then lfnstLastScanPos can be determined to be 1, and lfnst_idx[x0][y0] can be set to 0 without parsing. This indicates that the quadratic transform was not applied to the current block. On the other hand, if any of the transform blocks included in the transform unit has a lastScanPos equal to or greater than the threshold, then lfnstLastScanPos can be determined to be 0, and if [the threshold value is missing in the original text]. Figure 23 If all conditions i), ii), iii), iv), v), and vii) described in the code are satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform has been applied to the current block, and if a quadratic transform has been applied, the transform kernel used for the quadratic transform can be confirmed / determined.
[0260] When, based on the second embodiment described above, an indication is given at the transformation unit level as to whether a secondary transformation is applied. Figure 23 The transformation unit syntax structure described in [the document] can be applied, and Figure 20 The `residual_coding` method described in [the document] can be used. When based on the [method used to determine...] Figure 20Equation 4 of lfnstLastScanPos, as described in the section, states that if the lastScanPos of all transform blocks in the transform unit is less than the threshold or if all coefficients in the transform blocks are 0, then lfnstLastScanPos can be determined to be 1, and lfnst_idx[x0][y0] can be set to 0 without resolution. This indicates that the quadratic transform was not applied to the current block. On the other hand, if any of the transform blocks included in the transform unit has a lastScanPos equal to or greater than the threshold, then lfnstLastScanPos can be determined to be 0, and if... Figure 23 If all conditions i), ii), iii), iv), v), and vi) described in the code are satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform is applied to the current block, and when a quadratic transform is applied, it can check / determine the transform kernel used for the quadratic transform.
[0261] Figure 24 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to another embodiment of the present invention.
[0262] According to the third embodiment described above, lfnst_idx[x0][y0] can be resolved by using the position information of the last valid coefficient in the scan order obtained from residual_coding instead of the numSigCoeff counter.
[0263] Before performing residual coding, the variable lfnstLastScanPos, which relates to the position of the last valid coefficient in the scan order, can be initialized to 0. The variable lfnstLastScanPos can be a value obtained by adding the lastScanPos of the transform blocks included in the transform unit. In this case, if lfnstLastScanPos is greater than a threshold and Figure 23 If all conditions i), ii), iii), iv), v), and vi) described in the code are satisfied (if all are true), the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform was applied to the current block, and if so, it can check / determine the transform kernel used for the quadratic transform. On the other hand, when lfnstLastScanPos is less than or equal to the threshold, lfnst_idx[x0][y0] can be set to 0 and no parsing can be performed. This indicates that no quadratic transform was applied.
[0264] The threshold can be set according to `treeType`. When `treeType` is `SINGLE_TREE`, `DUAL_TREE_LUMA`, or `DUAL_TREE_CHROMA`, the threshold can be set to `Th1`, `Th2`, and `Th3`, respectively. `Th1`, `Th2`, and `Th3` are preset integer values equal to or greater than 0, and both the encoder and decoder can use the same values. When `treeType` is `SINGLE_TREE`, both luminance and chrominance components are included, and therefore `Th1` as the threshold can be expressed as the sum of `Th2` and `Th3` as thresholds.
[0265] When, based on the third embodiment described above, an indication is made at the transformation unit level as to whether a secondary transformation is applied. Figure 22 The `residual_coding` method described in [the document] can be used. According to... Figure 22 Equation 5, described in [the document], is used to determine lfnstLastScanPos. The variable lfnstLastScanPos can be set to a value obtained by summing all lastScanPos of the transform block included in the transform unit. Alternatively, whether to parse lfnst_idx[x0][y0] can be determined by comparing lfnstLastScanPos with a threshold.
[0266] On the other hand, when a secondary transformation is indicated at the transformation unit level, the correlation between transformation units included in the compilation unit can be high. This is because the method used for prediction is determined at the compilation unit level. Therefore, lfnst_idx[x0][y0] is signaled only in the first transformation unit included in the compilation unit, and the signaled lfnst_idx[x0][y0] can be shared with the remaining transformation units. That is, lfnst_idx[x0][y0] can only be resolved by using the first to third embodiments described above when subTuIndex, which indicates the index of the transformation unit, is 0. If subTuIndex is greater than 0, the corresponding transformation unit does not resolve lfnst_idx[x0][y0], and the value of lfnst_idx[x0][y0] of the shared first transformation unit can be used.
[0267] On the other hand, a counter can be used to count the valid coefficients, but it is possible to determine whether the decoder resolves lfnst_idx[x0][y0] by only considering the valid coefficients present in the sub-block of the top-left transform block. This is to reduce the amount of computation.
[0268] On the other hand, while indicating the secondary transform at the transform unit level can reduce decoder latency compared to indicating it at the compilation unit level, it may introduce another latency. For example, even if the secondary transform is indicated at the transform unit level, it may only be indicated after the compilation of the luminance transform coefficients, Cb transform coefficients, and Cr transform coefficients is complete. Therefore, even if the compilation (processing) of all luminance transform coefficients is finished, the inverse transform processing for the luminance transform coefficients may be performed after the compilation (processing) of the Cb and Cr transform coefficients is completed. This results in another latency for the decoder.
[0269] The following section will describe a quadratic transform indication method for minimizing the delay time of the decoder.
[0270] (Fourth Embodiment)
[0271] Taking a secondary transform indication method for minimizing decoder latency as an example, the secondary transform is indicated at the transform unit level. However, there might be a method for resolving the syntax element lfnst_idx[x0][y0] related to the secondary transform before the luma transform coefficients are compiled. Therefore, the decoder can perform the inverse transform on the luma transform coefficients immediately after compilation, without waiting for the Cb and Cr transform coefficients to be compiled. Similarly, the decoder can perform the inverse transform on the Cb transform coefficients immediately after compilation, without waiting for the Cr transform coefficients to be compiled. This secondary transform indication method minimizes decoder latency and can resolve pipeline problems.
[0272] Figure 25 This is a diagram illustrating the syntax of a compilation unit according to an embodiment of the present invention.
[0273] refer to Figure 25 Because the second transformation is indicated at the transformation unit level, the syntax related to the second transformation, lfnst_idx[x0][y0], is not parsed at the compilation unit level, but can be parsed at the transformation unit level that is divided by transform_tree.
[0274] Figure 26 This is a diagram illustrating a method for indicating a secondary transformation at the transformation unit level according to another embodiment of the present invention.
[0275] refer to Figure 26The method for indicating the secondary transformation can be specified at the transformation unit level, and lfnst_idx[x0][y0], which is a syntax element related to the secondary transformation, can be parsed first before the luma and chroma transformation coefficients are compiled (residual_coding). For example, when lfnst_idx[x0][y0] is parsed first before obtaining the transformation coefficients, the inverse transformations of the Y, Cb, and Cr transformation coefficients can be processed once the coefficients for each of the color components Y, Cb, and Cr are compiled. For example, once the transformation coefficients for the Y component are compiled, the inverse transformation of the luma (Y) transformation coefficients can be performed. Similarly, once the transformation coefficients for the Cb component are compiled (residual_coding), the inverse transformation of the Cb transformation coefficients can be performed, and once the transformation coefficients for the Cr component are compiled (residual_coding), the inverse transformation of the Cr transformation coefficients can be performed.
[0276] When parsing lfnst_idx[x0][y0] after compiling (residual_coding) the transform coefficients for Y, Cb, and Cr, even if the compilation (residual_coding) of the transform coefficients for Y is complete, the inverse transform of the Y transform coefficients cannot be performed / processed if the compilation (residual_coding) of the transform coefficients for Cb and Cr is not completed / processed. Therefore, even if the compilation (residual_coding) of the transform coefficients for Y is complete, the decoder may not perform the inverse transform of the Y transform coefficients until the compilation (residual_coding) of the transform coefficients for the other components Cb and Cr is completed, which may cause unnecessary delay time. However, as mentioned above, if lfnst_idx[x0][y0] is parsed first before the compilation (residual_coding) of the transform coefficients, the inverse transform of each of the color components Y, Cb, and Cr can be performed immediately after the compilation (residual_coding) of the transform coefficients for each of the color components is completed, so there is an effect of minimizing the delay time of the decoder.
[0277] The `transform_unit()` syntax structure can parse `tu_cbf_luma[x0][y0]`, `tu_cbf_cb[x0][y0]`, `tu_cbf_cr[x0][y0]`, `transform_skip_flag[x0][y0]`, etc.
[0278] Specifically, `tu_cbf_luma[x0][y0]` is an element indicating whether the current luminance transform block includes one or more non-zero transform coefficients. If `tu_cbf_luma[x0][y0]` is 1, it indicates that the current luminance transform block includes one or more non-zero transform coefficients. If `tu_cbf_luma[x0][y0]` is 0, it indicates that all transform coefficients of the current luminance transform block are 0. `tu_cbf_cb[x0][y0]` is an element indicating whether the current chrominance Cb transform block includes one or more non-zero transform coefficients. If `tu_cbf_cb[x0][y0]` is 1, it indicates that the current chrominance Cb transform block includes one or more non-zero transform coefficients. If `tu_cbf_cb[x0][y0]` is 0, it indicates that all transform coefficients of the current Cb transform block are 0. `tu_cbf_cr[x0][y0]` is an element indicating whether the current chrominance Cr transform block includes one or more non-zero transform coefficients. If `tu_cbf_cr[x0][y0]` is 1, it indicates that the current chroma (Cr) transform block includes one or more non-zero transform coefficients. If `tu_cbf_cr[x0][y0]` is 0, it indicates that all transform coefficients in the current chroma (Cr) transform block are 0. `transform_skip_flag[x0][y0]` is a syntax element related to transform skipping. If `transform_skip_flag[x0][y0]` is 1, it indicates that the inverse transform is not applied to the luma (luminance) transform block. If `transform_skip_flag[x0][y0]` is 0, it indicates that whether the inverse transform is applied to the luma transform block is determined by another syntax element.
[0279] For reference Figure 26 An embodiment of the quadratic transform indication method can parse the syntax element lfnst_idx[x0][y0] related to the quadratic transform based on the position of the last valid coefficient in the scan order, rather than based on the number of non-zero transform coefficients (valid coefficients).
[0280] First, the lfnstLastScanPos variable can be set by initializing it to 1. (See reference...) Figure 23The variable lfnstLastScanPos indicates the position information of the last valid coefficient in the scan order of the transform blocks included in the current transform unit. Specifically, when lfnstLastScanPos is 1, it indicates that the position (scan index) of the last valid coefficient in the scan order of all transform blocks included in the transform unit is less than a threshold, or it indicates that all transform coefficients in the block are 0. When lfnstLastScanPos is 0, it indicates that there are one or more valid coefficients in the block of one or more transform blocks included in the transform unit, and the position (scan index) of the last valid coefficient in the scan order is equal to or greater than a threshold.
[0281] Next, the variable `numZeroOutSigCoeff` can be set by initializing it to 0. When a quadratic transform is applied to a transform block, valid coefficients may not exist at specific positions in the scan order. Therefore, the variable `numZeroOutSigCoeff` can indicate whether valid coefficients exist at specific positions, and based on this, it can be checked whether a quadratic transform has been applied. For example, when a quadratic transform is applied to a transform block, it is assumed that only a maximum of 16 valid coefficients are allowed. In 4×4 and 8×8 transform blocks, valid coefficients may exist in the index [0, 7] region in the scan order (maximum of 8 non-zero transform coefficients allowed). On the other hand, in transform blocks that are not 4×4 or 8×8, valid coefficients may exist in the index [0, 15] region in the scan order (maximum of 16 non-zero transform coefficients allowed). Therefore, if the position of the last valid coefficient in the scan order (scan index) is outside the aforementioned region where valid coefficients may exist, the decoder can clearly identify that a quadratic transform has not been applied to the current transform block.
[0282] Whether to parse syntax elements related to the quadratic transform lfnst_idx[x0][y0] before coefficient compilation (residual_coding) can be determined based on the position (scan index) of the last valid coefficient in the scan order. Therefore, the decoder can process information related to the position of the last valid coefficient in the scan order before coefficient compilation (residual_coding).
[0283] Specifically, when the current luminance transform block includes one or more valid coefficients (tu_cbf_luma[x0][y0] == 1) and transform skip is not applied to the current luminance transform block (transform_skip_flag[x0][y0] == 0), last_significant_pos, which is a syntax structure related to the position of the last valid coefficient in the luminance scan order, can be processed.
[0284] When the value of tu_cbf_luma[x0][y0] is 0 (tu_cbf_luma[x0][y0] == 0), it indicates that all coefficients of the corresponding transform block are 0, which in turn indicates that coefficient residual coding has not been performed. Therefore, it is unnecessary to process the position information of the last valid coefficient in the scan order.
[0285] When transform_skip_flag[x0][y0] is 1, it indicates that the inverse transform has not been applied to the current brightness transform block. Therefore, coefficient compilation (residual_coding) can be performed without relying on the position information of the last valid coefficient in the scan order.
[0286] When the current chroma Cb transform block includes one or more valid coefficients (tu_cbf_cb[x0][y0] == 1), the last_significant_pos syntax structure, which relates to the position of the last valid coefficient (excluding 0) in the scan order of the chroma Cb transform block, can be processed. The last_significant_pos syntax structure can receive (x0, y0) as the top-left coordinate of the transform block, a value obtained by taking the logarithm of the transform block's width to the base 2, a value obtained by taking the logarithm of the transform block's height to the base 2, and cIdx as a variable indicating which color component the transform block represents. For example, cIdx of 0 can represent a luminance Y transform block, cIdx of 1 can represent a chroma Cb transform block, and cIdx of 2 can represent a chroma Cr transform block. When the value of tu_cbf_cb[x0][y0] is 0 (tu_cbf_cb[x0][y0] == 0), it indicates that all coefficients of the corresponding transform block are 0. This means that coefficient residual coding is not performed, and therefore, it is unnecessary to process the position information of the last valid coefficients in the scan order other than 0.
[0287] On the other hand, if the current chroma Cr transform block includes one or more valid coefficients (tu_cbf_cr[x0][y0] == 1), then tu_joint_cbcr_residual[x0][y0], which is a syntax element indicating whether chroma Cb and Cr are expressed as a residual signal, can be parsed before last_significant_pos processing. For example, when tu_joint_cbcr_residual[x0][y0] is 1, coefficient compilation (residual_coding) for Cr is not processed, and the residual signal for Cr can be derived from the reconstructed Cb residual signal. On the other hand, when tu_joint_cbcr_residual[x0][y0] is 0, coefficient compilation (residual_coding) for Cr can be performed based on the value of tu_cbf_cr[x0][y0]. If the current chroma Cr transform block includes one or more valid coefficients (tu_cbf_cr[x0][y0] == 1), the syntax structure last_significant_pos, which relates to the position of the last valid coefficient in the chroma Cr scan order, can be processed. When the value of tu_cbf_cr[x0][y0] is 0 (tu_cbf_cr[x0][y0] == 0), it indicates that all coefficients in the chroma Cr transform block are 0. This means that coefficient compilation (residual_coding) is not performed, and therefore the position information of the last valid coefficient in the scan order, except for 0, does not need to be executed.
[0288] When the processing of last_significant_pos for each color component is performed, the position (scan index) of the last valid coefficient in the scan order for each color component can be obtained, and based on this, the values of lfnstLastScanPos and numZeroOutSigCoeff can be updated.
[0289] Additionally, if all of the conditions i), ii), iii), iv), v), vi), and vii (if all are true) are met, the decoder can parse lfnst_idx[x0][y0] before coefficient compilation (residual_coding).
[0290] Solving the syntax elements of lfnst_idx[x0][y0] before coefficient compilation (residual_coding) Analysis conditions
[0291] i)Min(lfnstWidth, lfnstHeight)>=4
[0292] First, the first condition is related to the block size. When the width and height of the block are 4 pixels or more, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0293] Specifically, the decoder can check the block size condition to which a quadratic transform can be applied. The variables SubWidthC and SubHeightC are set according to the color format and can represent the ratio of the width of the chroma component to the width of the luminance component, and the ratio of the height of the chroma component to the height of the luminance component, respectively. For example, because a 4:2:0 color format image has a structure where every four luminance samples include one chroma sample, both SubWidthC and SubHeightC can be set to 2. For another example, because a 4:4:4 color format image has a structure where every luminance sample includes one chroma sample, both SubWidthC and SubHeightC can be set to 1. lfnstWidth, as the number of samples in the horizontal direction of the current block, and lfnstHeight, as the number of samples in the vertical direction, can be set based on SubWidthC and SubHeightC. When treeType is DUAL_TREE_CHROMA, because the transform unit only includes the chroma component, the number of samples in the horizontal direction of the chroma transform block is equal to the value obtained by dividing tbWidth, which is the width of the luminance transform block, by SubWidthC. Similarly, the number of samples in the vertical direction of the chroma transform block is equal to the value obtained by dividing tbHeight, which is the height of the luminance transform block, by SubHeightC. When treeType is SINGLE_TREE or DUAL_TREE_LUMA, because the transform unit includes the luminance component, lnfnstWidth and lfnstHeight can be set to tbWidth and tbHeight, respectively. Since the minimum condition for a block to which a quadratic transform can be applied is 4×4, lfnst_idx[x0][y0] can be resolved if Min(lfnstWidth, lfnstHeight)>=4.
[0294] ii) sps_lfnst_enabled_flag==1
[0295] The second condition involves a flag value indicating whether a second transformation can be enabled or applied, and when the flag value indicating whether a second transformation can be enabled or applied (sps_lfnst_enabled_flag) is set to 1, the decoder can parse lfnst_idx[x0][y0].
[0296] Specifically, secondary transformations can be indicated using high-level syntax RBSP. A 1-bit flag indicating whether secondary transformations can be enabled or applied can be included in at least one of the SPS, PPS, VPS, tile group header, and slice header, and when sps_lfnst_enabled_flag is 1, it can indicate the presence of the lfnst_idx[x0][y0] syntax element in the transform unit syntax. When sps_lfnst_enabled_flag is 0, it may indicate that the lfnst_idx[x0][y0] syntax element does not exist in the transform unit syntax.
[0297] iii)CuPredMode[x0][y0]==MODE_INTRA
[0298] The third condition relates to the prediction mode, and the quadratic transform can be applied only to intra-prediction blocks. Therefore, when the current block is an intra-prediction block, the decoder can resolve lfnst_idx[x0][y0].
[0299] iv)IntraSubPartitionsSplitType==ISP_NO_SPLIT
[0300] The fourth condition concerns whether the ISP prediction method is applied. When ISP is not applied to the current block, the decoder can parse the lfnst_idx[x0][y0] syntax elements.
[0301] Specifically, as referenced Figure 11As described, when the current CU is partitioned into multiple transform units smaller than the CU size, the quadratic transform may not be applied to the partitioned transform units. In this case, lfnst_idx[x0][y0], as a syntax element related to the quadratic transform, can be set to 0 and not parsed. When the transform tree for the current CU is partitioned into multiple transform units smaller than the CU size, ISP prediction can be applied to the current compilation unit. When intra-frame prediction is applied to the current compilation unit, the ISP prediction method can be a prediction method used to partition the transform tree into multiple transform units smaller than the CU size according to a preset partitioning method. The ISP prediction mode can be indicated at the compilation unit level, and the variable IntraSubPartitionsSplitType can be set based on it. When IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that ISP is not applied to the current block. Due to the characteristic of intra-frame prediction that generates prediction samples at the transform unit level, the prediction accuracy may be higher when the transform tree is partitioned into multiple transform units compared to when the transform tree is not partitioned. Therefore, even if the quadratic transform is not applied to the partitioned multiple transform units, it is very likely that the energy of the residual signal can be effectively compressed.
[0302] v)!intra_mip_flag[x0][y0]
[0303] The fifth condition relates to the intra-frame prediction method. When matrix-based intra-frame prediction (MIP) is not applied to the prediction of the current compilation unit, the decoder can parse the lfnst_idx[x0][y0] syntax element.
[0304] Specifically, matrix-based intra-frame prediction (MIP) can be used as a method for intra-frame prediction, and whether MIP is applied can be indicated at the compilation unit level via `intra_mip_flag[x0][y0]`. When `intra_mip_flag[x0][y0]` is 1, it indicates that MIP is applied to the prediction of the current compilation unit, and prediction can be performed by multiplying the reconstructed samples around the current block with a preset matrix. Because the residual signal properties exhibited when MIP is applied differ from those of general intra-frame prediction when performing directional or non-directional prediction, a quadratic transform cannot be applied to the transform block when MIP is applied.
[0305] vi)lfnstLastScanPos==0
[0306] The sixth condition relates to the last valid coefficient in the scan order of the transform block.
[0307] Specifically, when the position information (scan index) of the last valid coefficient in the scan order of the transform blocks included in the current transform unit is less than a preset threshold, the compilation efficiency gain that can be obtained through the secondary transform is likely to be small. Therefore, in this case, the encoder is likely not to apply the secondary transform to the transform block (lfnst_idx[x0][y0] is 0), and thus, the encoder may be considered to have a large overhead for signaling lfnst_idx[x0][y0]. Therefore, lfnst_idx[x0][y0] can only be parsed if the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the preset threshold for at least one transform block included in the transform unit.
[0308] In other words, as mentioned above, the threshold can be an integer equal to or greater than 0. For example, assuming the threshold is 1, the fact that the position (scan index) of the last valid coefficient in the scan order is equal to or greater than the threshold may mean that the valid coefficient exists outside the top left of the block (scan index 0, DC). In this case, the fact that the position of the last valid coefficient in the scan order of the transform block is equal to or greater than the threshold can be expressed as "lfnstLastScanPos==0".
[0309] vii)numZeroOutSigCoeff==0
[0310] The seventh condition involves the effective coefficients that exist in a specific location.
[0311] Specifically, when a quadratic transform is applied to the current block, valid coefficients may not exist at a specific position in the scan sequence. That is, the `numZeroOutSigCoeff` variable indicates whether a non-zero transform coefficient exists at a specific position. For example, when a quadratic transform is applied to the current block, assume that only a maximum of 16 valid coefficients are allowed. In transform blocks of 4×4 and 8×8 size, valid coefficients can exist in the index [0, 7] region of the scan sequence (maximum of 8 non-zero transform coefficients allowed). On the other hand, in transform blocks of sizes other than 4×4 and 8×8, valid coefficients can exist in the index [0, 15] region of the scan sequence (maximum of 16 non-zero transform coefficients allowed). Therefore, if the position of the last valid coefficient in the scan sequence (scan index) exists outside the aforementioned region where valid coefficients may exist, the decoder can clearly identify that the quadratic transform was not applied to the current block. Therefore, since the quadratic transform is not applied to the current block when `numZeroOutSigCoeff > 0`, `lfnst_idx[x0][y0]` can be set to 0 without resolution.
[0312] In other words, when numZeroOutSigCoeff is not 0, it means that there is a valid coefficient at a specific position, and therefore lfnst_idx[x0][y0] can be set to 0 without parsing. On the other hand, when numZeroOutSigCoeff is 0, it means that there is no valid coefficient at a specific position, and therefore lfnst_idx[x0][y0] can be parsed.
[0313] If all conditions i) to vii) above are true, then lfnst_idx[x0][y0] can be parsed; otherwise, lfnst_idx[x0][y0] can be set to 0 and no parsing can be performed.
[0314] Figure 27 The illustration shows a syntax structure related to the position of the last valid coefficient in the scanning order according to an embodiment of the present invention.
[0315] refer to Figure 27 The `last_significant_pos` syntax structure refers to a syntax structure that includes the position information of the last valid coefficient in the scan order of the transformations used for each of the color components Y, Cb, and Cr. Additionally, the `last_significant_pos` syntax structure can accept the top-left coordinates (x0, y0) of the transform block, log2TbWidth (obtained by taking the logarithm of the transform block's width to the base 2), log2TbHeight (obtained by taking the logarithm of the transform block's height to the base 2), and cIdx representing the color components of the transform block as inputs. When cIdx is 0, it represents a luminance transform block; when cIdx is 1, it represents a chrominance Cb transform block; and when cIdx is 2, it represents a chrominance Cr transform block.
[0316] In the `last_significant_pos` syntax structure, syntax elements related to the position information of the last valid coefficient in the scan order can be parsed. Specifically, syntax elements related to the x-coordinate and y-coordinate values of the last valid coefficient in the scan order can be parsed. In this case, this can be indicated by dividing each coordinate value into prefix and suffix information. The decoder can set the `LastSignificantCoeffX` variable, which is the x-coordinate of the last valid coefficient in the scan order, based on the prefix and suffix information of the x-coordinate. Similarly, the decoder can set the `LastSignificantCoeffY` variable, which is the y-coordinate of the last valid coefficient in the scan order, based on the prefix and suffix information of the y-coordinate. Figure 27As illustrated in the diagram, within the `do{}while()` structure, the decoder can set `lastScanPos` based on `LastSignificantCoeffX`, `LastSignificantCoeffY`, and `DiagScanOrder`, where `lastScanPos` is the scan index of the last valid coefficient in the scan order. Additionally, the decoder can update `numZeroOutSigCoeff` and `lfnstLastScanPos`, which are variables used as parse conditions in `lfnst_idx[x0][y0]`, and `lfnst_idx[x0][y0]` are syntactic elements related to the quadratic transformation, based on `lastScanPos`.
[0317] If a quadratic transform is applied to the current block, valid coefficients cannot exist at a specific location on the scan position. The `numZeroOutSigCoeff` variable indicates whether a non-zero transform coefficient exists at that location. For example, when a quadratic transform is applied to the current block, assume that only a maximum of 16 valid coefficients are allowed. In transform blocks of 4×4 and 8×8 size, valid coefficients may exist in the index [0, 7] region of the scan order (maximum of 8 non-zero transform coefficients allowed). On the other hand, in transform blocks of sizes other than 4×4 and 8×8, valid coefficients may exist in the index [0, 15] region of the scan order (maximum of 16 non-zero transform coefficients allowed). Therefore, if the location of the last valid coefficient in the scan order (scan index) exists outside the aforementioned region where valid coefficients may exist, the decoder can clearly identify that a quadratic transform was not applied to the current block. The minimum size of the block to which a quadratic transform can be applied is 4×4, and a quadratic transform may not be applied when a transform skip is applied (transform_skip_flag[x0][y0] == 1). Therefore, for a transform block with a width of 4 or greater (log2TbWidth>=2), a height of 4 or greater (log2TbHeight>=2), and where transform skipping is not applied (transform_skip_flag[x0])[y0]==0), numZeroOutSigCoeff can be updated. When a quadratic transform is applied, for a transform block of size 4×4 or 8×8, non-zero transform coefficients (valid coefficients) may only exist in the index [0, 7] region in the scan order. Therefore, when the transform block is 4×4 or 8×8 ((log2TbWidth==2||log2TbHeight==3)&&(log2TbWidth==log2TbHeight)) and lastScanPos is greater than 7 (lastScanPos>7), numZeroOutSigCoeff can be incremented by 1. For blocks other than those with a size of 4×4 or 8×8 where the quadratic transformation can be applied, non-zero transformation coefficients can exist only in the index [0, 15] region of the scan order. Therefore, numZeroOutSigCoeff can be incremented by 1 when lastScanPos is greater than 15 (lastScanPos>15).
[0318] The decoder can determine `lfnstLastScanPos` based on `lastScanPos`. Specifically, when the width and height of the transform block are 4 or greater and transform skipping is not applied to the transform block, `lfnstLastScanPos` can be set as shown in Equation 6 below. In other words, when `log2TbWidth>=2`, `log2TbHeight>=2`, and `transform_skip_flag[x0][y0]`...
[0319] When it is 0, lfnstLastScanPos can be set as in Equation 1 below. In this case, when transform_skip_flag[x0][y0] is 0, it may mean that the transformation skips the transformation that has not been applied to the current transformation block.
[0320] [Equation 6]
[0321] lfnstLastScanPos=lfnstLastScanPos&&(IastScanPos <lfnstLastScanPosTh[cldx])
[0322] As mentioned above, the initial value of lfnstLastScanPos can be set to 1.
[0323] As mentioned above, in Equation 6, cIdx can represent a variable that indicates the color component of the current transform block.
[0324] According to Equation 6, when the immediately preceding lfnstLastScanPos is 1 and lastScanPos is less than lfnstLastScanPosTh[cIdx], lfnstLastScanPos can be updated to 1. On the other hand, when the immediately preceding lfnstLastScanPos is 0 or lastScanPos is equal to or greater than lfnstScanPosTh[cIdx], lfnstLastScanPos can be updated to 0.
[0325] In other words, when the lastScanPos of all transform blocks included in the transform unit is less than the threshold or the coefficients of all transform blocks are 0, it can be determined that lfnstLastScanPos is 1, and without considering... Figure 26In the case of conditional parsing of lfnst_idx[x0][y0], lfnst_idx[x0][y0] can be set to 0. This indicates that the quadratic transformation was not applied to the current block. On the other hand, if any of the transformation blocks included in the transformation unit has a lastScanPos equal to or greater than the threshold, then lfnstLastScanPos can be determined to be 0, and if Figure 26 If all conditions i), ii), iii), iv), v), and vii) are satisfied (if all are true), then the decoder can parse lfnst_idx[x0][y0]. The decoder can parse lfnst_idx[x0][y0] to check if a quadratic transform is applied to the current block, and when a quadratic transform is applied to the current block, it can check / determine the transform kernel used for the quadratic transform.
[0326] In Equation 6, lfnstLastScanPosTh[cIdx] is a preset integer value equal to or greater than 0, and both the encoder and decoder can use the same value. Furthermore, all color components can use the same threshold. In this case, lfnstLastScanPos can be set as shown in Equation 7 below.
[0327] [Equation 7]
[0328] lfnstLastScanPos=lfnstLastScanPos&&(lastScanPos <lfnstLastScanPosTh)
[0329] LfnstLastScanPosTh is a preset integer value equal to or greater than 0, and both the encoder and decoder can use the same value. For example, lfnstLastScanPosTh can be 1. That is, when lastScanPos is 1 or greater, lfnstLastScanPos can be updated to 0, and lfnst_idx[x0][y0] can be parsed. In this case, because the threshold lfnstLastScanPosTh is an integer value, the case where lastScanPos is 1 or greater can have the same meaning as the case where lastScanPos is greater than 0. Figure 27 The case where all color components have the same threshold 1 has already been described; however, the present invention is not limited thereto.
[0330] Figure 28 This is a diagram illustrating the residual_coding syntax structure according to an embodiment of the present invention.
[0331] refer to Figure 28The position information of the last valid coefficient in the scan order can be indicated before coefficient compilation (residual_coding). Therefore, the coefficient compilation (residual_coding) syntax structure may not include syntax structures related to the position information of the last valid coefficient in the scan order. For example, the position information of the last valid coefficient in the scan order could be a prefix and suffix of the x-coordinate and a prefix or suffix of the y-coordinate of the last valid coefficient in the scan order. (See reference...) Figure 28 The coefficient compilation (residual_coding) syntax structure can be performed based on LastSignificantCoeffX and LastSignificantCoeffY, which are the x and y coordinates of the last valid coefficient in the scan order determined before coefficient compilation (residual_coding).
[0332] The quadratic transformation indication method according to the fourth embodiment does not use the numSigCoeff counter. Therefore, even if the coefficient at position (xC, yC) is a valid coefficient (sig_coeff_flag[xC][yC] == 1), numSigCoeff may not be updated. In other words, the quadratic transformation indication method according to the fourth embodiment can be a method that does not use a counter for valid coefficients. Furthermore, in the quadratic transformation indication method according to the fourth embodiment, since the numZeroOutSigCoeff variable can be set based on lastScanPos, a counter based on sig_coeff_flag can be omitted during coefficient compilation (residual_coding).
[0333] Figure 29 This is a flowchart illustrating a video signal processing method according to an embodiment of the present invention.
[0334] The following description is based on a reference. Figures 15 to 28 The video signal processing method and apparatus of the described embodiments.
[0335] Video signal decoding apparatus may include performing Figure 29 The processor of the video signal processing method described in the document.
[0336] First, the processor can receive a bitstream that includes syntax elements related to the secondary transformations of the compilation unit.
[0337] The processor can check whether one or more preset conditions are met, and when one or more preset conditions are met, the processor can parse the syntax elements related to the secondary transformation of the compilation unit (S2910 and S2920). On the other hand, when one or more preset conditions are not met, the processor may not parse the syntax elements related to the secondary transformation of the compilation unit (S2930). In this case, the value of the syntax elements related to the secondary transformation can be set to 0.
[0338] and Figure 29 The syntax element related to the quadratic transformation of the compilation unit described in the document can be lfnst_idx[x0][y0], which indicates whether the quadratic transformation is applied to the compiler unit included in the document. Figures 15 to 28 The syntax elements of the transformation block in the current compilation unit described in the document.
[0339] The processor can parse the syntax elements related to the secondary transformation of the compilation unit through step S2920, and can check whether the secondary transformation is applied to the transformation block included in the compilation unit based on the parsed syntax elements (S2940).
[0340] In this case, when the quadratic transformation is applied to the transform block, the processor can obtain one or more inverse transform coefficients for the first sub-block by performing an inverse quadratic transformation based on one or more coefficients of the first sub-block, which is one or more sub-blocks constituting the transform block (S2950).
[0341] Then, the processor can obtain the residual samples of the transform block by performing an inverse primary transform based on one or more inverse transform coefficients obtained in S2950 (S2960).
[0342] The second-order transform can be a low-frequency non-separable transform (LFNST). Alternatively, the transform block can be a block to which the first-order transform, which can be separable into a vertical transform and a horizontal transform, is applied. In this case, the inverse first-order transform can refer to the inverse transform used for the first-order transform, and the inverse second-order transform can refer to the inverse transform used for the second-order transform.
[0343] Syntax elements related to the second transformation of a compilation unit may include information indicating whether the second transformation is applied to the compilation unit and information indicating the transformation kernel used for the second transformation.
[0344] The first sub-block can be the first sub-block according to a preset scan order, and in this case, the index of the first sub-block can be 0.
[0345] The first condition in one or more preset conditions can be that the index value of the position of the first coefficient among one or more coefficients indicating the first sub-block is greater than a preset threshold. In this case, the first coefficient can be the last valid coefficient according to a preset scan order, and a valid coefficient can refer to a non-zero coefficient. The preset threshold can be 0. The preset scan order can be in... Figure 13 and Figure 14 The upper right diagonal scanning order described in the text.
[0346] The second condition in one or more preset conditions can be a case where the width and height of the transform block are 4 pixels or more.
[0347] The third condition in one or more preset conditions may be a case where the value of the transform skip flag included in the bitstream is not a specific value. In this case, when the transform skip flag value has a specific value, the transform skip flag can indicate that the primary and secondary transforms are to be applied to the transform block.
[0348] The fourth condition in one or more preset conditions can be that at least one coefficient of one or more coefficients of the sub-block is not 0 and that at least one coefficient exists in a place other than the first position value according to the preset scan order. In this case, the first position in the preset scan order can mean the position where the horizontal and vertical coordinates are (0, 0) as described above, or it can be the first position according to the preset scan order (e.g., the top right diagonal order).
[0349] Furthermore, a compilation unit may include multiple compilation blocks. In this case, when at least one of the transformation blocks corresponding to the multiple compilation blocks satisfies one or more preset conditions, the syntax elements related to the secondary transformation can be parsed.
[0350] On the other hand, when the syntax elements related to the second transformation are not parsed or are set to 0 (S2930), or when it is confirmed in step S2940 that the second transformation is not applied to the transformation block included in the compilation unit, the processor can obtain the residual sample of the transformation block by performing the inverse first transformation based on one or more coefficients of the transformation block (S2970).
[0351] In this case, the inverse primary transformation and the inverse secondary transformation described above can be the inverse transformations of the primary transformation and the secondary transformation, respectively.
[0352] Depend on Figure 29 The video signal processing method or similar method performed by the video signal decoding device described herein can be performed by the video signal encoding device.
[0353] Video signal encoding apparatus may include a processor for encoding video signals.
[0354] In this scenario, the processor can obtain multiple primary transform coefficients for a block by performing a primary transform on residual samples of the block included in the compilation unit. The processor can obtain one or more secondary transform coefficients for a first sub-block, which is one of the sub-blocks constituting the block, by performing a secondary transform based on one or more of these primary transform coefficients. The processor can obtain a bitstream by encoding information about the one or more secondary transform coefficients and syntax elements related to the secondary transform of the compilation unit.
[0355] The second transformation can be called the low-frequency inseparable transformation (LFNST), and the first transformation can be separated into a vertical transformation and a horizontal transformation.
[0356] Additionally, syntax elements related to the quadratic transformation can be encoded when one or more preset conditions are met. These syntax elements may include information indicating whether the quadratic transformation is applied to a compilation unit and information indicating the transformation kernel used for the quadratic transformation. In this case, the syntax element related to the quadratic transformation could be lfnst_idx[x0][y0], which is... Figures 15 to 28 The grammatical elements described in the document.
[0357] The first sub-block can be the first sub-block according to a preset scan order. In this case, the index of the first sub-block can be 0.
[0358] The first condition in one or more preset conditions can be a case where the index value indicating the position of the first coefficient among one or more quadratic transform coefficients is greater than a preset threshold. In this case, the first coefficient can be the last valid coefficient according to a preset scan order, and a valid coefficient can refer to a non-zero coefficient. The preset threshold can be 0. The preset scan order can be... Figure 13 and Figure 14 The upper right diagonal scanning order described in the text.
[0359] The second condition in one or more preset conditions may be the case where the width and height of the initial transformation block are 4 pixels or more.
[0360] The third condition in one or more preset conditions can be a case where the value of the transform skip flag included in the bitstream is not a specific value. In this case, when the transform skip flag value has a specific value, the transform skip flag can indicate that the primary and secondary transforms were not applied to the block.
[0361] The fourth condition in one or more preset conditions can be that at least one of the quadratic transformation coefficients is not 0 and that at least one coefficient exists in a location other than the first position according to the preset scan order. In this case, the first position in the preset scan order can mean the position where the horizontal and vertical coordinates are (0, 0) as described above, or the first position according to the preset scan order (e.g., the top right diagonal order).
[0362] Additionally, a compilation unit may include multiple compilation blocks. In this case, syntax elements related to the secondary transformation can be encoded when at least one of the (transformation) blocks included in the compilation unit corresponding to each of the multiple compilation blocks satisfies one or more preset conditions.
[0363] Additionally, the video signal encoding device may include an execution Figure 29 The video signal decoding processor described in the video signal processing method.
[0364] As described above, a bitstream can include... Figures 15 to 29 The syntax elements related to the second transformation of the compilation unit described in the document. In this case, the bitstream can be stored in a non-transitory computer-readable medium. Meanwhile, when one or more of the above preset conditions are not met, the video signal encoding apparatus may omit the syntax elements related to the second transformation from the bitstream, or may set the syntax elements related to the second transformation to 0. The bitstream can be obtained from a reference... Figure 29 The video signal decoding device described herein can decode the signal, or the video signal encoding device described above can encode the signal.
[0365] Methods for encoding bitstreams may include, for example, performing an initial transformation on residual samples of a block included in a compilation unit to obtain multiple initial transformation coefficients for the block, performing a secondary transformation based on one or more of the multiple initial transformation coefficients to obtain one or more secondary transformation coefficients for a first sub-block that is one of the sub-blocks constituting the block, and encoding information about the one or more secondary transformation coefficients and syntactic elements related to the secondary transformation of the compilation unit.
[0366] In this specification, acquiring a coefficient may mean acquiring a pixel / block related to that coefficient, and acquiring a residual sample may mean acquiring a residual signal / pixel / block related to the residual sample.
[0367] The above embodiments of the present invention can be implemented by various means. For example, the embodiments of the present invention can be implemented by hardware, firmware, software, or a combination thereof.
[0368] In the case of hardware implementation, the method according to the embodiments of the present invention can be implemented by one or more of the following: application-specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field-programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, etc.
[0369] When implemented via firmware or software, the method according to embodiments of the invention can be implemented in the form of modules, processes, or functions that perform the above-described functions or operations. Software code can be stored in memory and driven by a processor. The memory can be located inside or outside the processor and can exchange data with the processor in various known ways.
[0370] Some embodiments may also be implemented in the form of a recording medium including computer-executable instructions, such as a computer-executable program module. A computer-readable medium can be any available medium accessible to a computer and includes volatile and non-volatile media, removable and non-removable media. Furthermore, a computer-readable medium can include computer storage media and communication media. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented using any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Communication media typically include computer-readable instructions, data structures, other data in modulated data signals such as program modules, or other transmission mechanisms, and includes any information delivery medium.
[0371] The above description of the present invention is for illustrative purposes only, and it will be understood that those skilled in the art to which this invention pertains can make changes to the invention without altering its technical concept or essential characteristics, and that the invention can be readily modified in other specific forms. Therefore, the above embodiments are illustrative and not limiting in any way. For example, each component described as a single entity can be distributed and implemented, and similarly, components described as distributed can also be implemented in an associated manner.
[0372] The scope of this invention is defined by the appended claims rather than the foregoing detailed description, and all changes or modifications derived from the meaning and scope of the appended claims and their equivalents shall be construed as being included within the scope of this invention.
Claims
1. A video signal decoding device, comprising a processor, in, The processor is configured to: Parse the syntax elements related to the secondary transformations of the compilation unit; When the quadratic transform is applied to a transform block, one or more inverse transform coefficients are obtained by performing the inverse transform of the quadratic transform; and Residual samples for the transform block are obtained based on the one or more inverse transform coefficients. The second transformation is the low-frequency inseparable transform LFNST. Specifically, the syntax element is parsed when one or more conditions are met. Among these conditions, the first condition is that the index of the last valid coefficient in the sub-block according to the preset scan order is greater than 0. The second condition in one or more of the conditions is that the flag information related to enabling the secondary transformation has a value indicating that the syntax element related to the secondary transformation exists in the compiler unit syntax. The sub-block is a sub-block with sub-block index 0 in the transform block according to the preset scan order. The final effective coefficient is a non-zero coefficient.
2. The video signal decoding device according to claim 1, in, The preset scanning order is the upper right diagonal scanning order.
3. The video signal decoding device according to claim 1, wherein, The syntax element indicates whether the quadratic transformation is applied to the compilation unit and which transformation kernel is used for the quadratic transformation.
4. The video signal decoding device according to claim 1, wherein, The third condition in one or more of the conditions is that the width and height of the transform block are 4 or greater.
5. The video signal decoding device according to claim 1, wherein, The transformation block is a block to which an initial transformation, which can be separated into vertical and horizontal transformations, is applied.
6. The video signal decoding device according to claim 5, wherein, The fourth condition in one or more of the conditions is that the value of the transform skip flag is not a specific value, and Specifically, when the value of the transform skip flag is the specific value, the transform skip flag indicates that the primary transform and the secondary transform were not applied to the transform block.
7. A video signal encoding device, comprising a processor, in, The processor is configured to: Obtain the bitstream decoded by the decoder using the decoding method. The decoding method includes: Parse the syntax elements related to the secondary transformations of the compilation unit; When the quadratic transform is applied to a transform block, one or more inverse transform coefficients are obtained by performing the inverse transform of the quadratic transform; and Residual samples for the transform block are obtained based on the one or more inverse transform coefficients. The second transformation is the low-frequency inseparable transform LFNST. Specifically, the syntax element is parsed when one or more conditions are met. Among these conditions, the first condition is that the index of the last valid coefficient in the sub-block according to the preset scan order is greater than 0. The second condition in one or more of the conditions is that the flag information related to enabling the secondary transformation has a value indicating that the syntax element related to the secondary transformation exists in the compiler unit syntax. The sub-block is a sub-block with sub-block index 0 in the transform block according to the preset scan order. The final effective coefficient is a non-zero coefficient.
8. The video signal encoding device according to claim 7, in, The preset scanning order is the upper right diagonal scanning order.
9. The video signal encoding apparatus according to claim 7, wherein, The syntax element indicates whether the quadratic transformation is applied to the compilation unit and which transformation kernel is used for the quadratic transformation.
10. The video signal encoding apparatus according to claim 7, wherein, The third condition in one or more of the conditions is that the width and height of the transform block are 4 or greater.
11. The video signal encoding apparatus according to claim 7, wherein, The transformation block is a block to which an initial transformation, which can be separated into vertical and horizontal transformations, is applied.
12. The video signal encoding apparatus according to claim 11, wherein, The fourth condition in one or more of the conditions is that the value of the transform skip flag is not a specific value, and Specifically, when the value of the transform skip flag is the specific value, the transform skip flag indicates that the primary transform and the secondary transform were not applied to the transform block.
13. A method for transmitting a bit stream, the method comprising: The bit stream is generated by executing an encoding method; and Send the bit stream, The encoding method includes: One or more first transformation coefficients are obtained based on the residual samples of the transform block; When the quadratic transform is applied to the transform block, one or more second transform coefficients are obtained by performing the quadratic transform; and The syntax elements related to the secondary transformation of the compilation unit are encoded. The second transformation is the low-frequency inseparable transform LFNST. Specifically, the syntax element is encoded when one or more conditions are met. Among these conditions, the first condition is that the index of the last valid coefficient in the sub-block according to the preset scan order is greater than 0. The second condition in one or more of the conditions is that the flag information related to enabling the secondary transformation has a value indicating that the syntax element related to the secondary transformation exists in the compiler unit syntax. The sub-block is a sub-block with sub-block index 0 in the transform block according to the preset scan order. The final effective coefficient is a non-zero coefficient.
14. The method according to claim 13, in, The preset scanning order is the upper right diagonal scanning order.
15. The method according to claim 13, wherein, The syntax element indicates whether the quadratic transformation is applied to the compilation unit and which transformation kernel is used for the quadratic transformation.
16. The method according to claim 13, wherein, The third condition in one or more of the conditions is that the width and height of the transform block are 4 or greater.
17. The method according to claim 13, wherein, The transformation block is a block to which an initial transformation, which can be separated into vertical and horizontal transformations, is applied.
18. The method according to claim 17, wherein, The fourth condition in one or more of the conditions is that the value of the transform skip flag is not a specific value, and Specifically, when the value of the transform skip flag is the specific value, the transform skip flag indicates that the primary transform and the secondary transform were not applied to the transform block.
19. A method for decoding a video signal, the method comprising: Parse the syntax elements related to the secondary transformations of the compilation unit; When the quadratic transform is applied to a transform block, one or more inverse transform coefficients are obtained by performing the inverse transform of the quadratic transform; and Residual samples for the transform block are obtained based on the one or more inverse transform coefficients. The second transformation is the low-frequency inseparable transform LFNST. Specifically, the syntax element is parsed when one or more conditions are met. Among these conditions, the first condition is that the index of the last valid coefficient in the sub-block according to the preset scan order is greater than 0. The second condition in one or more of the conditions is that the flag information related to enabling the secondary transformation has a value indicating that the syntax element related to the secondary transformation exists in the compiler unit syntax. The sub-block is a sub-block with sub-block index 0 in the transform block according to the preset scan order. The final effective coefficient is a non-zero coefficient.