Video coding method and device based on inseparable primary transformation

By adopting inseparable primary transformation (NSPT) technology based on intra prediction mode, transform block size and transform coefficient characteristics in video encoding, the problems of insufficient video encoding efficiency and image quality in the prior art are solved, and more efficient video data processing is achieved.

CN120226360APending Publication Date: 2025-06-27HYUNDAI MOTOR CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380068874.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-11
Filing Date
2023-08-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing video encoding technologies process high-resolution and high frame rate video data, the encoding efficiency and image enhancement effect are insufficient, making it difficult to meet the growing demand for video data.

Method used

An inseparable primary transformation (NSPT) is performed based on the intra prediction mode of the current block, the size of the transform block, and the characteristics of the transform coefficients, and an inseparable primary transformation is performed on the large transform block by implicit division.

Benefits of technology

Improves video encoding efficiency, enhances video quality, and can more effectively process high resolution and high frame rate video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226360A_ABST
    Figure CN120226360A_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding based on an inseparable primary transform are disclosed. In a disclosed embodiment, a video decoding device acquires inverse quantized transform coefficients for a transform block of a current block, and decodes a non-separable primary transform (NSPT) flag from a bitstream. The video decoding device checks an NSPT flag, and when the NSPT flag is true, determines an inseparable primary inverse transform kernel based on a size of a transform block, an intra prediction mode of a current block, and a characteristic of an inverse quantization transform coefficient. A video decoding device generates a residual signal by applying an inseparable primary inverse transform kernel to a transform coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video coding method and apparatus based on a non-separable primary transform. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Since video data has a large amount of data compared to audio data or still image data, a large amount of hardware resources (including memory) are required to store or transmit video data without compression and processing.

[0004] Therefore, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) (which has an encoding efficiency improvement of about 30% or more compared to HEVC).

[0005] However, due to the gradual increase in image size, resolution, and frame rate, the amount of data to be encoded also increases. Therefore, there is a need to provide a new compression technology with higher encoding efficiency and improved image enhancement effects than existing compression technologies.

[0006] The non-separable primary transform (NSPT) technology uses a pre-trained transform kernel to perform a non-separable primary transform on the residual signal of a block predicted in the intra prediction mode, rather than performing vertical and horizontal transforms. If the size of the residual block on which the transform is performed is W tb ×H tb , then the size of the non-separable primary transform kernel K is (W tb ×H tb )×(W tb ×H tb ). In other words, (W tb ×H tb ) residual signals are transformed to generate (W tb ×H tb ) transform coefficients.

[0007] For the non-separable primary transform, the transform kernel can be defined based on the block shape and the intra prediction mode. For example, using the symmetry of a square block, 35 transform kernels are defined for square blocks according to the intra prediction mode. In addition, 67 transform kernels are defined for rectangular blocks.

[0008] The non-separable primary transform can be applied to the luminance component. In addition, when the non-separable primary transform is applied, the secondary transform (i.e., the low-frequency non-separable transform (LFNST)) may not be performed. Therefore, in order to improve video coding efficiency and enhance video quality, it is necessary to improve the non-separable primary transform and adopt an integrated operation method. Summary of the Invention

[0009] [Technical Problem]

[0010] The present disclosure seeks to provide a video coding method and apparatus for performing a non-separable primary transform (NSPT) based on the intra prediction mode of a current block, the size of a transform block, and the characteristics of transform coefficients. In addition, the video coding method and apparatus according to the present disclosure perform a non-separable primary transform on a large transform block that cannot apply a non-separable transform using implicit partitioning.

[0011] [Technical Solution]

[0012] At least one aspect of the present disclosure provides a method for reconstructing a current block performed by a video decoding apparatus. The method includes: obtaining inverse quantization transform coefficients for a transform block of the current block. The method further includes: decoding a non-separable primary transform (NSPT) flag from a bitstream. The NSPT flag indicates whether the non-separable primary transform is applied. The method further includes: checking the NSPT flag. When the NSPT flag is true, the method further includes: determining a non-separable primary inverse transform kernel based on the size of the transform block, the intra prediction mode of the current block, and the characteristics of the inverse quantization transform coefficients. The method further includes: performing a primary inverse transform by generating a residual signal by applying the non-separable primary inverse transform kernel to the transform coefficients.

[0013] Another aspect of the present disclosure provides a method for encoding a current block performed by a video coding apparatus. The method includes: obtaining a residual signal for a transform block of the current block. The method further includes: determining a non-separable primary transform kernel based on the size of the transform block, the intra prediction mode of the current block, and the characteristics of the quantization transform coefficients. The method further includes: generating first primary transform coefficients by applying the non-separable primary transform kernel to the residual signal. The method further includes: explicitly or implicitly determining a pair of primary transform kernels for the vertical direction and the horizontal direction of the transform block. The method further includes: generating second primary transform coefficients by applying the pair of primary transform kernels to the residual signal.

[0014] Another aspect of the present disclosure provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes: obtaining a residual signal for a transform block of a current block. The video encoding method further includes: determining a non-separable first-order transform kernel based on the size of the transform block, the intra prediction mode of the current block, and the characteristics of the quantized transform coefficients. The video encoding method further includes: generating first-order transform coefficients by applying the non-separable first-order transform kernel to the residual signal. The video encoding method further includes: explicitly or implicitly determining a pair of first-order transform kernels for the vertical and horizontal directions of the transform block. The video encoding method further includes: generating second-order transform coefficients by applying the pair of first-order transform kernels to the residual signal.

[0015] [Advantageous Effects]

[0016] As described above, the present disclosure provides a video encoding method and apparatus that perform non-separable first-order transforms based on the intra prediction mode of a current block, the size of a transform block, and the characteristics of transform coefficients. Accordingly, the video encoding method and apparatus improve video encoding efficiency and enhance video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of a video encoding apparatus capable of implementing the technology of the present disclosure.

[0018] Figure 2 illustrates a method of partitioning blocks using a quadtree-binary tree-trinary tree (QTBTTT) structure.

[0019] Figure 3A and 3B illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.

[0020] Figure 4 illustrates adjacent blocks of a current block.

[0021] Figure 5 is a block diagram of a video decoding apparatus capable of implementing the technology of the present disclosure.

[0022] Figure 6 is a block diagram showing in detail a part of a video decoding apparatus according to an embodiment of the present disclosure.

[0023] Figure 7 illustrates a method for determining an inverse transform kernel according to an embodiment of the present disclosure.

[0024] Figure 8 illustrates a method for determining an inverse transform kernel according to another embodiment of the present disclosure.

[0025] Figure 9 illustrates the partitioning of a transform block according to an embodiment of the present disclosure.

[0026] Figure 10 Shows the partitioning of a transform block according to another embodiment of the present disclosure.

[0027] Figure 11 Shows the reconstruction order of sub - blocks according to an embodiment of the present disclosure.

[0028] Figure 12A and 12B Shows the inverse transform process of reconstructing transform coefficients according to an embodiment of the present disclosure.

[0029] Figure 13 Shows the scan order for generating a first - stage transform coefficient vector according to an embodiment of the present disclosure.

[0030] Figure 14 Shows the scan order for generating a first - stage transform coefficient vector according to another embodiment of the present disclosure.

[0031] Figure 15 Shows the scan order for generating a first - stage transform coefficient vector according to an embodiment of the present disclosure.

[0032] Figure 16A and 16B Is a flowchart showing a method for transforming a transform block by a video encoding device according to an embodiment of the present disclosure.

[0033] Figure 17 Is a flowchart showing a method for inverse - transforming a transform block by a video decoding device according to an embodiment of the present disclosure. Detailed Description of the Invention

[0034] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, the same reference numerals denote the same elements, although these elements are shown in different drawings. Further, in the following description of some embodiments, for the sake of clarity and conciseness, detailed descriptions of related known components and functions may be omitted when it is considered that such detailed descriptions would obscure the subject matter of the present disclosure.

[0035] Figure 1 Is a block diagram of a video encoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 1 the description of

[0036] The encoding device may include a picture splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0037] Each component of the encoding device can be implemented as hardware or software, or as a combination of hardware and software. In addition, the function of each component can be implemented as software, and the microprocessor can also be implemented to execute the functions of the software corresponding to each component.

[0038] A video consists of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed on each region. For example, a picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile and / or slice is divided into one or more coding tree units (CTUs). In addition, each CTU is divided into one or more coding units (CUs) in a tree structure. The information applied to each coding unit (CU) is encoded as the syntax of the CU, and the information commonly applied to the CUs included in a CTU is encoded as the syntax of the CTU. In addition, the information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and the information applied to all blocks constituting one or more pictures is encoded as a picture parameter set (PPS) or a picture header. In addition, the information commonly referred to by a plurality of pictures is encoded as a sequence parameter set (SPS). In addition, the information commonly referred to by one or more SPSs is encoded as a video parameter set (VPS). In addition, the information commonly applied to a tile or a tile group can also be encoded as the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntax.

[0039] The picture splitter 110 determines the size of the coding tree unit (CTU). The information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and transmitted to the video decoding device.

[0040] The picture splitter 110 divides each picture constituting the video into a plurality of coding tree units (CTUs) of a predetermined size, and then recursively divides the CTUs using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units of encoding.

[0041] The tree structure can be a quadtree (QT), in which a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), in which a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), in which a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree-binary tree (QTBT) structure can be used, or a quadtree-binary tree-ternary tree (QTBTTT) structure can be used. Here, a binary tree-ternary tree (BTTT) is added to a tree structure called a multi-type tree (MTT).

[0042] Figure 2 is a diagram for describing a method of dividing a block by using a QTBTTT structure.

[0043] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four nodes of a layer is encoded by the entropy encoder 155 and signaled to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or TT structure. For example, there can be two directions, namely the direction in which the block of the corresponding node is horizontally divided and the direction in which the block of the corresponding node is vertically divided. As Figure 2 shown, when the MTT division starts, a second flag (mtt_split_flag) indicating whether the node is divided, and (if the node is divided) a flag indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) are encoded by the entropy encoder 155 and signaled to the video decoding device.

[0044] Alternatively, before encoding the first flag (QT_split_flag) of four nodes indicating whether each node is split into layers, the CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (which is the basic unit of encoding). When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag by the above scheme.

[0045] When using the QTBT as another example of the tree structure, there may be two types, that is, the type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetric horizontal split) and the type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetric vertical split). The split flag (split_flag) of the block indicating whether each node of the BT structure is split into layers and the split type information indicating the split type are encoded by the entropy encoder 155 and transmitted to the video decoding device. On the other hand, there may additionally be a type in which the block of the corresponding node is split into two blocks that are asymmetric to each other. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in the diagonal direction.

[0046] According to the QTBT or QTBTTT split from the CTU, the CU may have various sizes. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". Due to the QTBTTT split, the shape of the current block may be rectangular in addition to square.

[0047] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0048] Generally, each current block in a picture can be predictively encoded. Usually, the prediction of the current block can be performed by using intra prediction techniques (using data from the picture including the current block) or inter prediction techniques (using data from the pictures encoded before the picture including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.

[0049] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) around the current block in the current picture including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3AAs shown, multiple intra prediction modes may include two non-directional modes, including a planar mode and a DC mode, and may include 65 directional modes. For each prediction mode, the surrounding pixels and equations to be used are defined differently.

[0050] For efficient directional prediction of a current block having a rectangular shape, directional modes as shown by the dashed arrows in Figure 3B (modes #67 to #80, intra prediction modes #-1 to #-14) may be additionally used. These directional modes may be referred to as "wide-angle intra prediction modes". In Figure 3B the arrows indicate the corresponding reference samples for prediction and do not represent the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode that performs prediction in a direction opposite to a specific direction mode without additional bit transmission. In this case, among the wide-angle intra prediction modes, some wide-angle intra prediction modes available for the current block may be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide-angle intra prediction modes (intra prediction modes #67 to #80) having an angle less than 45 degrees are available. When the current block has a rectangular shape with a width greater than the height, wide-angle intra prediction modes having an angle greater than -135 degrees are available.

[0051] The intra predictor 122 may determine the intra prediction to be used for encoding the current block. In some examples, the intra predictor 122 may encode the current block by using multiple intra prediction modes and may also select an appropriate intra prediction mode to be used from among the tested modes. For example, the intra predictor 122 may calculate rate-distortion values by using rate-distortion analysis of multiple tested intra prediction modes and may also select the intra prediction mode having the best rate-distortion characteristics from among the tested modes.

[0052] The intra predictor 122 selects one intra prediction mode from among multiple intra prediction modes and predicts the current block by using the surrounding pixels (reference pixels) and equations determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0053] The inter - frame predictor 124 generates a predicted block of the current block by using a motion compensation process. The inter - frame predictor 124 searches for the block most similar to the current block in a reference picture that was encoded and decoded earlier than the current picture, and generates a predicted block of the current block by using the searched - for block. In addition, a motion vector (MV) corresponding to the displacement between the current block in the current picture and the predicted block in the reference picture is generated. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance components. Motion information including information on the reference picture and information on the motion vector for predicting the current block is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0054] The inter - frame predictor 124 can also perform interpolation on the reference picture or reference block in order to increase the prediction accuracy. In other words, interpolation of sub - samples between two consecutive integer samples is performed by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference picture, fractional - unit precision rather than integer - sample - unit precision can be expressed for the motion vector. The precision or resolution of the motion vector can be set differently for each target region to be encoded (e.g., units such as slices, tiles, CTUs, CUs, etc.). When such an adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. The information on the motion vector resolution can be information indicating the precision of the motion vector difference described below.

[0055] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference pictures and two motion vectors representing the positions of the blocks most similar to the current block in each reference picture are used. The inter-frame predictor 124 selects the first reference picture and the second reference picture from the reference picture list 0 (RefPicList0) and the reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference picture to generate the first reference block and the second reference block. In addition, the predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. In addition, the motion information including information about the two reference pictures used for predicting the current block and including information about the two motion vectors is transmitted to the entropy encoder 155. Here, the reference picture list 0 may be composed of pictures among the pre-reconstructed pictures that are before the current picture in the display order, and the reference picture list 1 may be composed of pictures among the pre-reconstructed pictures that are after the current picture in the display order. However, although not particularly limited thereto, the pre-reconstructed pictures after the current picture in the display order may be additionally included in the reference picture list 0. Conversely, the pre-reconstructed pictures before the current picture may also be additionally included in the reference picture list 1.

[0056] To minimize the amount of bits consumed for encoding the motion information, various methods can be used.

[0057] For example, when the reference picture and the motion vector of the current block are the same as those of an adjacent block, the information capable of identifying the adjacent block is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.

[0058] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.

[0059] As the adjacent blocks for deriving the merge candidates, all or some of the left-end block A0, the lower-left block A1, the upper-end block B0, the upper-right block B1, and the upper-left block B2 adjacent to the current block in the current picture can be used, as Figure 4 shown. In addition, blocks located in a reference picture (which may be the same as or different from the reference picture used for predicting the current block) other than the current picture in which the current block is located can also be used as merge candidates. For example, the block at the same position as the current block in the reference picture or the block adjacent to the block at the same position can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, a zero vector is added to the merge candidates.

[0060] The inter - frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as the motion information of the current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0061] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only the neighboring block selection information is sent without sending the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for pictures with little motion, still pictures, screen content pictures, etc.

[0062] Hereinafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.

[0063] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0064] In the AMVP mode, the inter - frame predictor 124 derives motion vector prediction value candidates for the motion vector of the current block by using neighboring blocks of the current block. As neighboring blocks for deriving motion vector prediction value candidates, all or some of the left - end block A0, left - lower - end block A1, upper - end block B0, right - upper - end block B1, and left - upper - end block B2 adjacent to the current block in the current picture can be used, as Figure 4 shown. In addition, blocks within a reference picture (which may be the same as or different from the reference picture used for predicting the current block) other than the current picture in which the current block is located can also be used as neighboring blocks for deriving motion vector prediction value candidates. For example, a block at the same position as the current block in the reference picture or a block adjacent to the block at the same position can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0065] The inter - frame predictor 124 derives motion vector prediction value candidates by using the motion vectors of neighboring blocks, and determines the motion vector prediction value for the motion vector of the current block by using the motion vector prediction value candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction value from the motion vector of the current block.

[0066] A motion vector prediction value can be obtained by applying a predefined function (e.g., central value and average value calculation, etc.) to motion vector prediction value candidates. In this case, the video decoding device also knows the predefined function. Additionally, since the neighboring blocks used to derive the motion vector prediction value candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction value candidates. Thus, in this case, the information regarding the motion vector difference and the information regarding the reference picture for predicting the current block are encoded.

[0067] On the other hand, the motion vector prediction value can also be determined by selecting any one of the motion vector prediction value candidates. In this case, additionally, the information for identifying the selected motion vector prediction value candidate is encoded together with the information regarding the motion vector difference and the information regarding the reference picture for predicting the current block.

[0068] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0069] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the total size of the residual block as the transform unit, or the residual block can also be divided into multiple sub - blocks, and the transform can be performed by using the sub - blocks as the transform unit. Alternatively, the residual block is divided into two sub - blocks (a transform region and a non - transform region), and the residual signal is transformed by using only the transform region sub - block as the transform unit. Here, the transform region sub - block can be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates only the transformed sub - block, and the directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. Additionally, with respect to the horizontal axis (or vertical axis), the size of the transform region sub - block can have a size ratio of 1:3. In this case, a flag (cu_sbt_quad_flag) for distinguishing the corresponding segmentation is additionally encoded by the entropy encoder 155 and signaled to the video decoding device.

[0070] On the other hand, the transformer 140 can perform transformations on the residual blocks in the horizontal direction and the vertical direction respectively. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multi-transformation set (MTS). The transformer 140 can select a pair of transformation functions in the MTS with the highest transformation efficiency, and can transform the residual blocks in each of the horizontal direction and the vertical direction. Information (mts_idx) about the pair of transformation functions in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.

[0071] The quantizer 145 uses quantization parameters to quantize the transform coefficients output from the transformer 140, and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual blocks without performing any block transformation or frame transformation. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the two-dimensional arrangement of the quantized transform coefficients can be encoded and signaled to the video decoding device.

[0072] The rearrangement unit 150 can perform re-alignment of the coefficient values on the quantized residual values.

[0073] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the DC coefficient to the high-frequency domain coefficients by using zigzag scanning or diagonal scanning to output a 1D coefficient sequence. According to the size of the transform unit and the intra-frame prediction mode, vertical scanning that scans the 2D coefficient array in the column direction and horizontal scanning that scans the 2D block type coefficients in the row direction can also be used instead of zigzag scanning. In other words, according to the size of the transform unit and the intra-frame prediction mode, the scanning method to be used can be determined among zigzag scanning, diagonal scanning, vertical scanning, and horizontal scanning.

[0074] The entropy encoder 155 encodes the 1D quantized transform coefficient sequence output from the rearrangement unit 150 by using various coding schemes, including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc., to generate a bitstream.

[0075] In addition, the entropy encoder 155 encodes information related to block partitioning, such as CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, MTT partitioning direction, etc., to allow the video decoding device to evenly partition blocks to the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (the merge index in the case of the merge mode and information regarding the reference picture index and motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization, i.e., information regarding the quantization parameter and information regarding the quantization matrix.

[0076] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0077] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next sequential block, the pixels in the reconstructed current block can be used as reference pixels.

[0078] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blur artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180 as the loop filter may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0079] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove block artifacts that occur due to block unit encoding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The SAO filter 184 applies an offset as a CTU unit to enhance the subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block unit filtering and applies different filters to compensate for distortion by distinguishing the boundaries of the corresponding blocks and the degree of variation. Information regarding the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.

[0080] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in a picture are reconstructed, the reconstructed picture can then be used as a reference picture for inter prediction of blocks within a picture to be encoded.

[0081] The video encoding device may store the bitstream of the encoded video data in a non-transitory storage medium or transmit the bitstream to a video decoding device via a communication network.

[0082] Figure 5 is a block diagram of a video decoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 5 , the video decoding device and its components will be described.

[0083] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0084] Similar to Figure 1 the video encoding device, each component of the video decoding device may be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0085] The entropy decoder 510 decodes the bitstream generated by the video encoding device to extract information related to block partitioning to determine the current block to be decoded, and extracts prediction information and information about the residual signal required to reconstruct the current block.

[0086] The entropy decoder 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), and divides the picture into CTUs having the determined size. In addition, the CTU is determined as the highest layer of the tree structure, i.e., the root node, and the partitioning information of the CTU can be extracted to partition the CTU using the tree structure.

[0087] For example, when the CTU is partitioned using the QTBTTT structure, first, the first flag (QT_split_flag) related to the partitioning of the QT is extracted to divide each node into four lower-layer nodes. In addition, for the node corresponding to the leaf node of the QT, the second flag (mtt_split_flag) related to the partitioning of the MTT, the partitioning direction (vertical / horizontal), and / or the partitioning type (binary / trinary) are extracted to divide the corresponding leaf node into the MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a BT structure or a TT structure.

[0088] As another example, when using the QTBTTT structure to split a CTU, a CU split flag (split_cu_flag) indicating whether to split the CU is extracted. When splitting the corresponding block, a first flag (QT_split_flag) can also be extracted. During the splitting process, for each node, zero or more recursive MTT splits can occur after zero or more recursive QT splits. For example, for a CTU, the MTT split can occur immediately, or conversely, only multiple QT splits can occur.

[0089] As another example, when using the QTBT structure to split a CTU, a first flag (QT_split_flag) related to the split of the QT is extracted to split each node into four lower-layer nodes. In addition, a split flag (split_flag) indicating whether to further split the node corresponding to the leaf node of the QT into a BT and split direction information are extracted.

[0090] On the other hand, when the entropy decoder 510 uses a tree-structured split to determine the current block to be decoded, the entropy decoder 510 extracts information on the prediction type indicating whether the current block is intra-frame prediction or inter-frame prediction. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts the syntax element for the intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information representing the syntax element for the inter-frame prediction information, that is, the motion vector and the reference picture referred to by the motion vector.

[0091] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on the quantized transform coefficients of the current block as information on the residual signal.

[0092] The rearrangement unit 515 can change the 1D sequence of quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in an order opposite to the coefficient scan order performed by the video coding device.

[0093] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the 2D-arranged quantized transform coefficients. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.

[0094] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain, thereby generating a residual block of the current block.

[0095] In addition, when the inverse transformer 530 performs inverse transformation on a partial region (sub-block) of a transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the non-inverse-transformed region with zero values as the residual signal to generate the final residual block of the current block.

[0096] In addition, when applying MTS, the inverse transformer 530 determines a transform index or a transform matrix to be applied in each of the horizontal direction and the vertical direction by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal direction and the vertical direction by using the determined transform function.

[0097] The predictor 540 may include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.

[0098] The intra predictor 542 determines the intra prediction mode of the current block among a plurality of intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using the neighboring reference pixels of the current block according to the intra prediction mode.

[0099] The inter predictor 544 determines the motion vector of the current block and the reference picture referred to by the motion vector by using the syntax element of the inter prediction mode extracted from the entropy decoder 510.

[0100] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 and the prediction block output from the inter predictor 544 or the intra predictor 542. When performing intra prediction on a block to be decoded later, the pixels in the reconstructed current block are used as reference pixels.

[0101] The loop filter unit 560 serving as a loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between reconstructed blocks to remove block artifacts that occur due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using information about the filter coefficients decoded from the bitstream.

[0102] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks within a picture to be encoded later.

[0103] In some embodiments, the present disclosure relates to encoding and decoding video pictures as described above. More specifically, the present disclosure provides a video coding method and apparatus for performing a non-separable single transform (NSPT) based on the intra prediction mode of a current block, the size of a transform block, and the characteristics of transform coefficients. In addition, the video coding method and apparatus use an implicit partitioning of large transform blocks for which non-separable transforms cannot be applied to perform the non-separable single transform.

[0104] The following embodiments may be performed by the transformer 140 and the inverse transformer 165 in a video coding apparatus. The following embodiments may also be performed by the inverse transformer 530 in a video decoding apparatus.

[0105] When encoding a current block, the video coding apparatus may generate signal information associated with the present embodiment in terms of optimizing rate distortion. The video coding apparatus may use an entropy encoder 155 to encode the signal information and send the encoded signal information to the video decoding apparatus. The video decoding apparatus may use an entropy decoder 510 to decode the signal information associated with the decoding of the current block from the bitstream.

[0106] In the following description, the term "target block" may be used interchangeably with the current block or the coding unit (CU), or may refer to a certain region of the coding unit.

[0107] In addition, a value of true for a flag indicates that the flag is set to 1. In addition, a value of false for a flag indicates that the flag is set to 0.

[0108] I. Transform Technology - Single Transform Technology

[0109] As described above, for efficient video compression, after prediction based on various prediction techniques, quantization or scaling can be additionally applied to the remaining residual signal. At this time, after applying a transform technique based on the importance of the cognitive visual information inherent in the residual signal and clustering the residual signal into specific regions according to its frequency components, scaling can be performed. However, in the case of screen content (which is an unnatural signal), the frequency-based transform technique may be inefficient. In this case, the transform technique can be omitted; instead, only scaling can be performed, or encoding / decoding can be performed without applying scaling.

[0110] When applying a transform in HEVC, DCT-II is used as the transform kernel (hereinafter, used interchangeably with the transform type) to transform the residual signal. However, in order to apply a more appropriate transform technique according to the diversity of residual signal characteristics, multi-transform selection (MTS) can be used. MTS determines one or two optimal types among multiple transform types, and then transforms the block according to the determined transform type. For example, in VVC, as shown in Table 1, in addition to DCT-II, two other transform types, DCT-VIII and DST-VII, are added so that the residual signal can be transformed in various ways.

[0111]

Table 1

[0112]

[0113] Here, the basis functions constitute the transform matrix that defines each transform type. Hereinafter, DCT-II, DCT-VIII, and DST-VII are used interchangeably with DCT2, DCT8, and DST7, respectively.

[0114] On the other hand, a flag that determines whether to use MTS can be controlled at the block level. In addition, an activation flag at a higher SPS level can be used to control the use of MTS.

[0115] If MTS is activated in the SPS, a CU-level flag indicating the application of MTS can be displayed. Here, MTS can be applied to the luminance component. When both the width and height of the transform block (TB) are less than or equal to 32 pixels and the coding block flag (CBF) indicating the existence of non-zero values among the transform coefficient levels is true, the CU-level flag can be expressed.

[0116] If the CU-level flag is 0, DCT2 can be used as the kernel in both the horizontal and vertical directions. On the other hand, if the CU-level flag is not 0, MTS can be applied. In this case, explicit MTS and implicit MTS can be supported.

[0117] In the case of explicit MTS, the kernel for TB can be sent explicitly. Generally, the index of the transform kernel can be sent. For example, as shown in Table 2, a kernel index called mts_idx can be defined.

[0118] [Table 2]

[0119] mts_idx 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2

[0120] Here, trTypeHor and trTypeVer represent the transform types in the horizontal and vertical directions respectively. In addition, 0 represents DCT2; 1 represents DST7; and 2 represents DCT8.

[0121] On the other hand, in implicit MTS, for example, in the case of an intra block, even if the MTS is not signaled explicitly, the transform type can be determined implicitly. In VVC, as shown in Equation 1, the transform types in the horizontal and vertical directions can be determined implicitly.

[0122] [Equation 1]

[0123] trTypeHor = (nTbW >= 4 && nTbW <= 16)? DST7 : DCT2

[0124] trTypeVer = (nTbH >= 4 && nTbH <= 16)? DST7 : DCT2

[0125] Here, nTbW and nTbH represent the lengths of the transform block in the horizontal and vertical directions respectively.

[0126] For example, when applying a specific coding technique, explicit MTS or implicit MTS can be applied. For example, in the case of matrix weighted intra prediction (MIP), explicit intra MTS can be used. In the case of the intra sub - partition (ISP) mode, implicit intra MTS can be used, and DST7 or DCT2 can be used as the transform type.

[0127] When the transform block includes at least one non - DC coefficient, mts_idx is signaled. In other words, if the position of the last valid coefficient according to the scan order is greater than 0, mts_idx is signaled. On the other hand, if the transform block includes only one non - DC coefficient, signaling of mts_idx is omitted, mts_idx is derived as 0, and DCT2 is applied as the transform kernel.

[0128] The first bit of the signaled mts_idx indicates whether mts_idx is greater than 0. If mts_idx is greater than 0 (i.e., pointing to one of the values from 1 to 4), then additionally a two-bit fixed-length code is signaled to indicate the signaled mts_idx among the four candidates.

[0129] On the other hand, in the enhanced compression model (ECM) software as a next-generation technology, the number and type of MTS cores are increased, and DST7, DCT8, DCT, DST4, DST1, and the identity transform are added.

[0130] II. Low-Frequency Non-Separable Transform (LFNST)

[0131] The Low-Frequency Non-Separable Transform (LFNST) technology performs a secondary transform on the low-frequency region among the transform coefficients generated by a single transform of a transform unit (TU) during intra-frame prediction. From a coding perspective, the LFNST technology applies a secondary transform to L low-frequency primary transform coefficients among W×H primary transform coefficients to generate K secondary transform coefficients (where K ≤ L). Here, the size of the LFNST transform kernel is L×K. In other words, the LFNST technology represents L low-frequency primary transform coefficients among W×H primary transform coefficients in the form of a 1×L vector, and applies an L×K transform kernel to generate a 1×K vector. After that, the LFNST technology rearranges the 1×K vector into a two-dimensional array in the low-frequency region to use the vector for subsequent processes such as quantization.

[0132] Compared with the primary transform that applies separate transform kernels in the horizontal and vertical directions, the LFNST technology performs a non-separable transform of a one-dimensional vector.

[0133] On the other hand, the type of the transform kernel can be determined based on the intra-frame prediction mode of the current TU, the TU size, and the LFNST index (lfnst_idx). For example, a set of transform kernels can be determined according to the intra-frame prediction mode (IntraPredMode) of the current TU, as shown in Table 3.

[0134]

Table 3

[0135] IntraPredMode lfnstTrSetIdx IntraPredMode<0 1 0 ≤ IntraPredMode ≤ 1 0 2 ≤ IntraPredMode ≤ 12 1 13 ≤ IntraPredMode ≤ 23 2 24 ≤ IntraPredMode ≤ 44 3 45 ≤ IntraPredMode ≤ 55 2 56 ≤ IntraPredMode ≤ 80 1 81 ≤ IntraPredMode ≤ 83 0

[0136] Here, the intra-frame prediction mode (IntraPredMode) follows the example shown in Figure 3B . In addition, lfnstTrSetIdx represents an index indicating the kernel set. In Table 3, the intra-frame prediction modes 81, 82, and 83 correspond to the cross-component linear model (CCLM) prediction mode.

[0137] For each kernel set (lfnstTrSetIdx), two types of kernels are defined. Which kernel to select between the two types of kernels can be determined by the LFNST index. An LFNST index of 0 indicates that LFNST is not performed, while an LFNST index of 1 or 2 indicates the use of different LFNST kernels within the same kernel set. Since there are additional kernel sets according to the TU size, a total of 4 × 2 × 2 = 16 transform kernels are available. The kernel sizes of LFNST are defined as 16×16 and 16×48. In addition, the kernel size can be adjusted according to the TU size, as shown in Table 4.

[0138]

Table 4

[0139] Block size Transformation size (K × L) 4×4 8×16 4 × N, N × 4 (N > 8) 16×16 8×8 8×48 Greater than 8 × 8 16×48

[0140] On the other hand, if the DCT2 / DCT2 transform kernel is applied as a primary transform to the TU of intra prediction, the LFNST technique can be applied as a secondary transform.

[0141] Although the following embodiments are described with reference to a video decoding apparatus, these embodiments can also be implemented in a video encoding apparatus in the same or similar manner as in the video decoding apparatus.

[0142] III. Embodiments according to the present disclosure

[0143] Figure 6 is a block diagram showing a part of a video decoding apparatus according to an embodiment of the present disclosure.

[0144] The video decoding apparatus according to the present embodiment can determine a prediction unit and a transform unit, perform prediction and inverse transform on a current block corresponding to the determined unit using a specified prediction technique and prediction mode, and finally generate a reconstructed block of the current block. Figure 6 The block diagram shown can be executed by an entropy decoder 510, an inverse quantizer 520, an inverse transformer 530, a predictor 540, and an adder 550 of the video decoding apparatus. On the other hand, as Figure 6 shown, an inverse quantizer 160, an inverse transformer 165, a picture splitter 110, a predictor 120, and an adder 170 of the video encoding apparatus can perform the same operations. At this time, the video decoding apparatus can use the encoded information parsed from the bitstream, but the video encoding apparatus can use the encoded information set from a high-level module to minimize rate distortion. Hereinafter, for convenience of description, this embodiment is described with respect to the video decoding apparatus.

[0145] As Figure 5 shown, the predictor 540 includes an intra predictor 542 and an inter predictor 544 based on the adopted prediction technique; however, as Figure 6 shown, the predictor 540 can include a prediction mode determiner 602 and a prediction executor 604.

[0146] In Figure 6 the example of, the sub-block splitter 606 can be part of the entropy decoder 510, the inverse transformer 530, or the predictor 540. From the perspective of the video coding device, the operation of the sub-block splitter 606 can be performed by the inverse transformer 165, the picture splitter 110, or the predictor 120.

[0147] When the color format of the input video is the YUV format (e.g., YUV420, YUV411, YUV422, YUV444), the video decoding device can first perform the prediction and reconstruction of the luminance component, and then can perform the prediction and reconstruction of the chrominance component. In other words, the luminance and chrominance components can be Figure 6 reconstructed in sequence by the components shown. Here, in the case of the YUV format, the color format represents the correspondence between the luminance component pixels and the chrominance component pixels.

[0148] The prediction mode determiner 602 determines the prediction technique (e.g., intra prediction, inter prediction, intra block copy (IBC) mode, or palette mode) for the current block. In addition, the prediction mode determiner 602 determines the specific prediction mode for the selected prediction technique. The prediction executor 604 generates a prediction block for the current block based on the determined prediction technique and prediction mode.

[0149] The inverse quantizer 520 inverse quantizes the quantized transform coefficients decoded for the current transform block to generate an inverse quantized signal. At this time, the inverse quantizer 520 can use one or more inverse quantizers to perform the inverse quantization. When using multiple inverse quantizers and the number of inverse quantizers is N q , the video coding and decoding device can select the inverse quantizer based on a state machine having identical states. At this time, the inverse quantizer can be selected based on the current state and the least significant bit (LSB) of the previous transform coefficient value. If k is the LSB of the previous transform coefficient value and N q = 2, the state transition table can be expressed as shown in Table 5.

[0150]

Table 5

[0151] Current state k=0 k=1 0 0 2 1 2 0 2 1 3 3 3 1

[0152] In addition, if k is the LSB of the previous transform coefficient and N q = 3, the state transition table can be expressed as shown in Table 6.

[0153]

Table 6

[0154]

[0155]

[0156] The sub - block splitter 606 divides the current block into sub - blocks based on the sub - block split enable flag, split method flag (or index), and / or split direction (vertical or horizontal) flag, as well as the aspect ratio / width / height of the current block. According to an embodiment, parsing of the split method flag and / or split direction flag (or index) may be omitted based on the aspect ratio / width / height of the current block, and the split method and / or split direction may be implicitly derived according to the protocol between the video encoding device and the video decoding device. At this time, flags and indices for deriving sub - block splitting may be defined based on the prediction technique (inter - frame or intra - frame) of the current block. In addition, prediction and / or transformation may be performed in units of the divided sub - blocks.

[0157] The inverse transformer 530 generates a residual signal by performing an inverse transform on the TU expressed by the inverse - quantized signal.

[0158] The adder 550 generates a reconstructed block by adding the prediction block and the residual signal. The reconstructed block is stored in the memory and is used later for predicting other blocks.

[0159] Hereinafter, the TU and the transform block may be used interchangeably.

[0160] As Figure 6 shown, the inverse transformer 530 may include all or some of an inverse transform kernel determiner 610, a transform unit determiner 612, and an inverse transform executor 614. The inverse transformer 530 may perform a non - separable primary inverse transform (NSPIT) on the transform coefficients using the above - mentioned components.

[0161] The inverse transform kernel determiner 610 determines the type of the inverse transform kernel of the current transform block according to Figure 7 or Figure 8 an example of.

[0162] Figure 7 shows a method for determining an inverse transform kernel according to an embodiment of the present disclosure.

[0163] For example, the inverse transform kernel determiner 610 parses the explicit_transform_flag (hereinafter referred to as the "explicit transform flag"), which indicates whether to use an explicit inverse transform kernel. If the parsed explicit_transform_flag is 0, the horizontal inverse transform kernel and the vertical inverse transform kernel are implicitly determined to be DCT-2. On the other hand, if the explicit_transform_flag is 1, the inverse transform kernel determiner 610 further parses the NSPT_flag (hereinafter referred to as the "non-separable primary transform flag" or "NSPT flag"), which indicates whether to apply a non-separable primary inverse transform. If the NSPT_flag is 0, the horizontal inverse transform kernel and the vertical inverse transform kernel can be explicitly determined according to the MTS. The inverse transform kernel determiner 610 parses the mts_idx to obtain the horizontal and vertical transform kernel pairs indicated by the mts_idx. At this time, the vertical / horizontal kernel pairs can be defined according to the protocol between the video encoding device and the video decoding device. If the NSPT_flag is 1, the inverse transform kernel determiner 610 further parses the NSPT_idx (hereinafter referred to as the "non-separable primary transform index" or "NSPT index") to determine the type of the non-separable primary inverse transform kernel. In other words, if the kernel set includes two or more inverse transform kernels, the inverse transform kernel determiner 610 can parse the NSPT_idx and select the primary inverse transform kernel indicated by the NSPT_idx from the kernel set. At this time, the kernel set can be selected based on factors such as the intra prediction mode of the current block and the size of the transform block. Alternatively, if the kernel set includes only one inverse transform kernel, the parsing of the NSPT_idx can be omitted, and the non-separable primary inverse transform kernel can be set to this one inverse transform kernel.

[0164] Figure 8 FIG. shows a method for determining an inverse transform kernel according to another embodiment of the present disclosure.

[0165] In another example, the inverse transform kernel determiner 610 first parses the NSPT_flag to determine whether to apply a non-separable inverse transform. If the NSPT_flag is 1, the inverse transform kernel determiner 610 may further parse the NSPT_idx to determine the type of the non-separable primary inverse transform kernel. If the NSPT_flag is 0, the inverse transform kernel determiner 610 parses the explicit_transform_flag. If the parsed explicit_transform_flag is 0, the horizontal inverse transform kernel and the vertical inverse transform kernel are implicitly determined to be DCT-2. On the other hand, if the explicit_transform_flag is 1, the horizontal inverse transform kernel and the vertical inverse transform kernel can be explicitly determined. The inverse transform kernel determiner 610 may parse the mts_idx to obtain the horizontal and vertical transform kernel pairs indicated by the mts_idx.

[0166] For example, if the transform kernel of the current transform block is determined according to the mts_idx and the current block is predicted based on the intra prediction mode and / or the weighted sum of the intra prediction mode and the inter prediction, the MTS kernel candidate list may be determined based on the intra prediction mode of the current block and the size of the current transform block. In other words, the kernel indicated by the mts_idx may vary according to the intra prediction mode, the size of the transform block, etc.

[0167] For example, the MTS list for the current transform block may be determined based on the sum of the absolute values of the quantized transform coefficients decoded by the entropy decoder 510, the sum of the absolute values of the inverse quantized transform coefficients by the inverse quantizer 520, or the position of the first non-zero transform coefficient (lastScanPos). Here, lastScanPos is determined according to the scan order defined by the protocol between the video encoding device and the video decoding device.

[0168] For example, the number of the MTS list may be determined as follows based on the value of lastScanPos. According to an embodiment, the number of sets determined based on a threshold (the number of the MTS list) and the number of transform kernels included in each set may vary.

[0169] MTS transform kernel candidates: {K0, K1, K2, K3, K4, K5}

[0170] List candidate set 0: {K0}, when lastScanPos ≤ th0

[0171] List candidate set 1: {K0, K1, K2, K3}, when th0 < lastScanPos ≤ th1

[0172] List candidate set 2: {K0, K1, K2, K3, K4, K5}, when lastScanPos > th1

[0173] At this time, the threshold can be defined based on the protocol between the video encoding device and the video decoding device. If the size of the list candidate set determined by the threshold is 1, the mts_idx signal for determining the kernel can be omitted.

[0174] In another example, the inverse transform kernel determiner 610 can parse MTS_ver_idx and MTS_hor_idx to determine the kernels in the vertical direction and the horizontal direction, respectively. MTS_ver_idx and MTS_hor_idx indicate the kernels in the vertical direction and the horizontal direction, respectively.

[0175] On the other hand, when applying the non-separable one-time inverse transform, the kernel set for the non-separable one-time inverse transform can be determined based on the intra prediction mode of the current block and / or the size of the current transform block. Hereinafter, it is assumed that based on satisfying "minNSPT ≤ Tb W , Tb H ≤ maxNSPT" and the T (Tb W × Tb H ) transform blocks of the intra prediction mode, there are M kernel sets, and there are C inverse transform kernel candidates for each kernel set. At this time, minNSPT and maxNSPT can be defined based on the protocol between the video encoding device and the video decoding device.

[0176] For example, Figure 3B the intra prediction modes shown in [] can be divided into six mode sets as shown in Table 7.

[0177]

Table 7

[0178] IntraPredMode Mode set IntraPredMode<0 1 0 ≤ IntraPredMode ≤ 1 0 2 ≤ IntraPredMode ≤ 12 1 13 ≤ IntraPredMode ≤ 23 2 24 ≤ IntraPredMode ≤ 44 3 45 ≤ IntraPredMode ≤ 55 4 56 ≤ IntraPredMode ≤ 80 5

[0179] Based on Table 7, the number of kernel sets M can be defined as 6 × T. In other words, based on the size of the transform block and the intra prediction mode (whether the mode is a directional mode or a non-directional mode and whether it is a prediction direction in the case of a directional mode), the kernel sets can be divided into M sets.

[0180] For example, by utilizing the symmetry of the square block, the mode sets of the square blocks in Table 7 can be divided into the mode sets shown in Table 8.

[0181]

Table 8

[0182] IntraPredMode Mode set IntraPredMode<0 1 0 ≤ IntraPredMode ≤ 1 0 2 ≤ IntraPredMode ≤ 12 1 13 ≤ IntraPredMode ≤ 23 2 24 ≤ IntraPredMode ≤ 44 3 45 ≤ IntraPredMode ≤ 55 2 56 ≤ IntraPredMode ≤ 80 1

[0183] If the mode set shown in Table 8 is adopted, mode m and mode 68 - m are included in the same mode set. Therefore, in the case of mode m and mode 68 - m, the inverse transform kernel determiner 610 can use the same kernel set for the non-separable one-time transform.

[0184] In Tables 7 and 8, Mode-1 to Mode-14 and Mode 67 to Mode 80 correspond to wide-angle prediction modes. The wide-angle prediction modes can be classified into a separate set of prediction modes or included in the same set of prediction modes as the nearest directional modes.

[0185] In addition, matrix-based intra prediction modes can be included in the non-directional mode set (Mode Set 0) or classified into a separate set.

[0186] In another example, in the case of a rectangular block, M sets of kernels can be identified based on the mode sets defined in Table 7. In this case, for a non-separable transform, a transform block having a size of A×B and predicted in mode m can use the same kernel as a block having a size of B×A and predicted in mode 68 - m.

[0187] For example, the number of inverse transform kernel candidates C for each set of kernels for the current transform block can be adaptively determined based on the sum of the absolute values of the quantized transform coefficients decoded by the entropy decoder 510, the sum of the absolute values of the inverse-quantized transform coefficients by the inverse quantizer 520, or the position of the first non-zero transform coefficient (lastScanPos). Based on one or more thresholds, the video decoding apparatus can compare the sum of the absolute values or lastScanPos with the threshold to determine the number of inverse transform kernel candidates.

[0188] For example, the number of inverse transform kernel candidates can be determined by comparing lastScanPos with a threshold th. If lastScanPos is less than or equal to the threshold, the inverse transform kernel determiner 610 can set the number of inverse transform kernel candidates to N0 (a non-negative integer). On the other hand, if lastScanPos exceeds the threshold, the number of inverse transform kernel candidates can be set to N1 (an integer greater than or equal to 1).

[0189] After that, the inverse transform kernel determiner 610 can parse NSPT_idx to determine the candidate indicated by the parsed NSPT_idx among the inverse transform kernel candidates as the non-separable one-time inverse transform kernel. If there is only one inverse transform kernel candidate, that candidate can be selected as the non-separable one-time inverse transform kernel without additional parsing of NSPT_idx.

[0190] If the width and / or height of the current transform block is greater than maxNSPT (which is the maximum size of a transform block to which non-separable one-time inverse transform can be applied), the transform unit determiner 612 can implicitly split the current transform block until its width and height become less than maxNSPT.

[0191] For example, if the width or height of the current transform block (or transform sub-block) exceeds maxNSPT, the transform unit determiner 612 may split the current transform block (or transform sub-block) as follows.

[0192] If the width T of the current transform block W (or the width sbT of the transform sub-block W ) is greater than maxNSPT, the transform unit determiner 612 may recursively perform SPLIT_BT_VER (vertical binary tree splitting) until the width of the split transform block becomes less than or equal to maxNSPT. Alternatively, if the height T of the current transform block H (or the height sbT of the transform sub-block H ) is greater than maxNSPT, the transform unit determiner 612 may recursively perform SPLIT_BT_HOR (horizontal binary tree splitting) until the height of the split transform block becomes less than or equal to maxNSPT.

[0193] In another example, if both the width and height of the current transform block (or transform sub-block) are greater than maxNSPT, the transform unit determiner 612 may split the current transform block into four sub-blocks using SPLIT_QT (quad tree splitting).

[0194] For example, when maxNSPT is 16, the transform unit determiner 612 may implicitly split a transform block with T W = 64 and T H = 16, as shown in Figure 9 . Additionally, when maxNSPT is 16, the transform unit determiner 612 may implicitly split a transform block with T W = 32 and T H = 64, as shown in Figure 10 .

[0195] On the other hand, the inverse transform and reconstruction of the split sub-blocks may be performed by following the z-scan order. At this time, the inverse transform order for the sub-blocks may be determined based on the intra prediction mode of the current block. If the intra prediction mode of the current block is the vertical mode (mode 50) or greater than a specific mode k (where k > 50), the z-scan order starting from the upper-right sub-block may be used to perform the inverse transform on the sub-blocks, as shown in the left example of Figure 11 . Additionally, if the intra prediction mode is the horizontal mode (mode 18) or less than a specific mode k (where k < 18), the z-scan order starting from the lower-left sub-block may be used to perform the inverse transform on the sub-blocks, as shown in the right example of Figure 11 .

[0196] For example, if the width and / or height of the current transform block is greater than a specific size maxNTsize predetermined according to a protocol between a video encoding device and a video decoding device, parsing of NSPT_flag and NSPT_idx may be omitted, and non-separable transform / inverse transform may not be performed. Further, if the width and / or height of the current transform block is less than minNSPT, parsing of NSPT_flag and NSPT_idx may be omitted, and non-separable transform / inverse transform may not be performed. Here, maxNTsize represents the maximum size of a transform block for parsing NSPT_flag, and maxNSPT represents the maximum size of a transform block size to which non-separable transform / inverse transform can be actually applied. Further, minNSPT represents the minimum size of a transform block size to which non-separable transform / inverse transform can be actually applied.

[0197] Hereinafter, maxNTsize is referred to as the "predefined NSPT maximum size". Further, maxNSPT represents the "predefined NSPT applicable maximum size", and minNSPT represents the "predefined NSPT applicable minimum size".

[0198] The inverse transform executor 614 performs an inverse transform based on an inverse transform kernel and the size of a transform block. The inverse transform executor 614 may parse a second transform flag or a second transform index to determine whether to perform a second inverse transform (i.e., LFNST). If the second transform index is parsed and the parsed second transform index is 0, the inverse transform executor 614 does not perform the second inverse transform. When performing the second inverse transform, the inverse transform executor 614 reconstructs first transform coefficients by performing a second inverse transform on inverse quantized second transform coefficients, and then performs a first inverse transform on the first transform coefficients to reconstruct a residual signal. If the second inverse transform is not performed, the inverse transform executor 614 reconstructs a residual signal by performing a first inverse transform on inverse quantized first transform coefficients.

[0199] According to an embodiment, the second inverse transform and the non-separable first inverse transform may not be performed simultaneously. In other words, if it is determined to perform the second inverse transform based on parsing of the second transform flag or the second transform index, parsing of NSPT_flag and NSPT_idx may be omitted. Alternatively, if NSPT_flag is 1, the second transform flag or the second transform index may be implicitly derived as 0.

[0200] If a non-separable first inverse transform is performed on a current transform block of width T W and height T H the inverse transform executor 614 may multiply the reconstructed transform coefficients of size P by a matrix having dimensions P×(T W ×T HThe inverse transformation matrix (i.e., the inverse transformation kernel) invT of ( ) is used to perform the inverse transformation as expressed in Equation 2.

[0201]

Equation 2

[0202]

[0203] Here, represents the vector of the residual signal reconstructed through the inverse transformation. represents the vector of the primary transformation coefficients reconfigured from the reconstruction transformation coefficients of size P according to the scan order. Here, the scan order can be predefined according to the protocol between the video encoding device and the video decoding device. In addition, the scan order including the z-scan order can vary according to the embodiment.

[0204] On the other hand, the relationship can hold such that P ≤ T W ×T H . For example, the size P can be determined as a multiple of the coefficient group (CG). The size W of the CG CG ×H CG can vary and, according to the embodiment, can be 4×4, 8×8, etc. In addition, the size of P can be determined based on the size of the current transformation block.

[0205] For example, if the relationship holds such that P ≤ T W ×T H , the inverse transformation process for reconstructing the transformation coefficients can be described as shown in Figure 12A and 12B . In this case, the size of P, the vectorized scan order of the reconstruction transformation coefficients, and the packing method (i.e., the scan order) for reconstructing the residual signal can vary according to the embodiment.

[0206] When applying the non-separable primary inverse transformation to a square transformation block, the inverse transformation executor 614 can use the same inverse transformation kernel to perform the inverse transformation for both the block predicted by the intra prediction mode m and the block predicted by the mode 68 - m. At this time, when the reconstruction transformation coefficients are reconfigured into the vector of the primary transformation coefficients , the scan order can be adaptively changed according to the intra prediction mode.

[0207] For example, for the intra prediction mode L (where 2 ≤ L ≤ 34) that is closer to the horizontal direction mode (mode 18) than the vertical direction mode, the inverse transformation executor 614 can use the scan order 1 (or scan order 2) to generate the vector as shown in Figure 13 . In addition, for the intra prediction mode L (where 34 < L ≤ 66) that is closer to the vertical direction mode (mode 50), the inverse transformation executor 614 can use the scan order 2 (or scan order 1) to generate the vector

[0208] When an inseparable single inverse transform is applied to a rectangular transform block, the inverse transform executor 614 can use the same kernel to perform an inseparable single transform on a block of size A×B predicted in mode m (2 ≤ m ≤ 34) and a block of size B×A predicted in mode 68 - m. At this time, in the case of the A×B block predicted in mode m (2 ≤ m ≤ 34), the inverse transform executor 614 can use Figure 14 the scan order 1 (or scan order 2) shown in to generate a vector. In addition, in the case of the B×A block predicted in mode 68 - m, the inverse transform executor 614 can use Figure 15 the scan order 2 (or scan order 1) shown in

[0209] When the transform unit determiner 612 implicitly divides a transform block into sub - blocks, the prediction executor 604 can perform intra - prediction of the current sub - block using the reconstruction reference samples of the sub - blocks reconstructed first according to the z - scan order.

[0210] Hereinafter, a method for performing a transform and an inverse transform on a transform block of a current block will be described.

[0211] Figure 16A And 16B are flowcharts showing a method for varying a transform block by a video coding device according to an embodiment of the present disclosure.

[0212] The video coding device obtains a residual signal for the transform block of the current block (S1600).

[0213] The video coding device determines an inseparable single transform kernel based on the size of the transform block, the intra - prediction mode of the current block, and the characteristics of the quantized transform coefficients (S1602).

[0214] The video coding device selects a mode set including the intra - prediction mode of the current block among a predefined set of modes. The video coding device selects a set of kernels based on the size of the transform block and the selected mode set.

[0215] If the set of kernels includes a single transform kernel candidate, the video coding device determines the single kernel candidate as the inseparable single transform kernel.

[0216] If the set of kernels includes multiple transform kernel candidates, the video coding device can determine one of the multiple transform kernel candidates as the inseparable single transform kernel in terms of rate - distortion optimization. Thereafter, the video coding device can encode the index indicating the determined candidate.

[0217] The video encoding device applies an inseparable first - order transform kernel to the residual signal to generate first - order transform coefficients (S1604).

[0218] The video encoding device generates a one - dimensional vector of the residual signal by arranging all or part of the residual signal in a predefined first scanning order based on the size and type of the inseparable first - order transform kernel. The video encoding device performs matrix multiplication between the vector of the residual signal and the inseparable first - order transform kernel to generate a first - order transform coefficient vector. The video encoding device distributes the first - order transform coefficient vector to transform blocks according to the predefined scanning order to generate first - order transform coefficients.

[0219] The video encoding device can adaptively determine the number of inverse transform kernel candidates for the kernel set based on the sum of the absolute values of the transform coefficients or the position of the first non - zero transform coefficient.

[0220] The video encoding device determines a pair of first - order transform kernels in the vertical and horizontal directions of the transform block (S1606).

[0221] The video encoding device applies the pair of first - order transform kernels to the residual signal in the vertical and horizontal directions to generate second - order transform coefficients (S1608).

[0222] The video encoding device applies predefined first - order transform kernels to the residual signal in the vertical and horizontal directions to generate third - order transform coefficients (S1610).

[0223] The video encoding device determines the NSPT flag based on the first, second, and third first - order transform coefficients (S1612). Here, the NSPT flag indicates whether the inseparable first - order transform is applied.

[0224] From the perspective of rate - distortion optimization, the video encoding device can determine the NSPT flag. For example, if the first first - order transform coefficients are optimal, the video encoding device can set the NSPT flag to true. On the other hand, if the second or third first - order transform coefficients are optimal, the video encoding device can set the NSPT flag to false.

[0225] The video encoding device encodes the NSPT flag (S1614).

[0226] The video encoding device checks the NSPT flag (S1616).

[0227] If the NSPT flag is true (the "yes" in S1616), the video encoding device encodes the first first - order transform coefficients (S1618). The video encoding device can quantize and entropy - encode the first first - order transform coefficients to generate a bitstream of the first first - order transform coefficients.

[0228] On the other hand, if the NSPT flag is false ("No" in S1616), the video encoding device performs the following steps.

[0229] The video encoding device determines an explicit transform flag (S1630) based on the second and third first transform coefficients. Here, the explicit transform flag indicates whether to use a pair of first transform kernels in the vertical and horizontal directions.

[0230] From the perspective of rate - distortion optimization, the video encoding device can determine the explicit transform flag. For example, if the second first transform coefficient is optimal, the video encoding device can set the explicit transform flag to true. On the other hand, if the third first transform coefficient is optimal, the video encoding device can set the explicit transform flag to false.

[0231] The video encoding device encodes the explicit transform flag (S1632).

[0232] The video encoding device checks the explicit transform flag (S1634).

[0233] If the explicit transform flag is true ("Yes" in S1634), the video encoding device performs the following steps.

[0234] The video encoding device encodes an index indicating a pair of first transform kernels in the vertical and horizontal directions (S1636).

[0235] The video encoding device encodes the second first transform coefficient (S1638). The video encoding device can quantize and entropy - encode the second first transform coefficient to generate a bitstream of the second first transform coefficient.

[0236] On the other hand, if the explicit transform flag is false ("No" in S1634), the video encoding device encodes the third first transform coefficient (S1640). The video encoding device can quantize and entropy - encode the third first transform coefficient to generate a bitstream of the third first transform coefficient.

[0237] Figure 17 FIG. is a flowchart showing a method for inverse - transforming a transform block by a video decoding device according to an embodiment of the present disclosure.

[0238] The video decoding device obtains the inverse - quantized transform coefficients of the transform block for the current block (S1700).

[0239] The video decoding device decodes the NSPT flag from the bitstream (S1702). Here, the NSPT flag indicates whether to apply the non - separable first transform.

[0240] The video decoding device checks the NSPT flag (S1704).

[0241] If the NSPT flag is true ("Yes" in S1704), the video decoding device performs the following steps.

[0242] The video decoding device determines a non-separable inverse transform kernel based on the size of the transform block, the intra prediction mode of the current block, and the characteristics of the quantized transform coefficients (S1706).

[0243] The video decoding device selects a mode set including the intra prediction mode of the current block from a predefined set of modes. The video decoding device selects a set of kernels based on the size of the transform block and the selected mode set.

[0244] The video decoding device adaptively determines the number of inverse transform kernel candidates for the set of kernels based on the sum of the absolute values of the inverse quantized transform coefficients or the position of the first non-zero transform coefficient.

[0245] If the set of kernels includes a single inverse transform kernel candidate, the video decoding device determines the single kernel candidate as the non-separable inverse transform kernel.

[0246] Alternatively, if the set of kernels includes multiple inverse transform kernel candidates, the video decoding device decodes the NSPT index from the bitstream. The video decoding device may determine the candidate indicated by the NSPT index among the multiple inverse transform kernel candidates as the non-separable inverse transform kernel.

[0247] The video decoding device generates a residual signal by applying the non-separable inverse transform kernel to the transform coefficients to perform an inverse transform (S1708).

[0248] The video decoding device packs all or part of the transform coefficients into a one-dimensional vector according to a predefined scan order based on the size and type of the non-separable inverse transform kernel to generate a vector of the first-level transform coefficients. The video decoding device performs matrix multiplication between the vector of the first-level transform coefficients and the non-separable inverse transform kernel to generate a vector of the residual signal. The video decoding device distributes the vector of the residual signal to the transform block according to the predefined scan order to generate a residual signal.

[0249] On the other hand, if the NSPT flag is false ("No" in S1704), the video decoding device performs the following steps.

[0250] The video decoding device decodes an explicit transform flag from the bitstream (S1720). Here, the explicit transform flag indicates whether a pair of inverse transform kernels in the horizontal and vertical directions is explicitly used.

[0251] The video decoding device checks the explicit transform flag (S1722).

[0252] If the explicit transform flag is true ("Yes" in S1722), the video decoding device performs the following steps.

[0253] The video decoding device decodes an index indicating a pair of inverse transform kernels in the horizontal direction and the vertical direction from the bitstream (S1724). Here, the index indicates one of a plurality of pairs of inverse transform kernels in the horizontal direction and the vertical direction.

[0254] The video decoding device applies the pair of inverse transform kernels in the horizontal direction and the vertical direction indicated by the index to the transform coefficients to generate a residual signal (S1726).

[0255] On the other hand, if the explicit transform flag is false ("No" in step S1722), the video decoding device applies predefined one-time inverse transform kernels in the horizontal direction and the vertical direction to the transform coefficients to generate a residual signal (S1730).

[0256] After that, the video decoding device may add the residual signal to the predicted block of the current block to generate a reconstructed block of the current block.

[0257] Although the steps in each flowchart are described as being executed sequentially, these steps merely illustrate the technical concept of some embodiments of the present disclosure. Therefore, those of ordinary skill in the art to which the present disclosure pertains can execute the steps by changing the order described in the corresponding drawings or by executing two or more steps in parallel. Therefore, the steps in each flowchart are not limited to the shown time sequence.

[0258] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in the present disclosure are marked with "…… unit" to strongly emphasize the possibility of their independent implementation.

[0259] On the other hand, various methods or functions described in some embodiments can be implemented as instructions stored in a non-transitory recording medium, which can be read and executed by one or more processors. The non-transitory recording medium may include various types of recording devices in which data is stored in a form readable by a computer system. For example, the non-transitory recording medium may include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.

[0260] Although the embodiments of the present disclosure have been described for illustrative purposes, it will be understood by those skilled in the art that various modifications, additions and substitutions are possible without departing from the concept and scope of the present disclosure. Therefore, for the sake of brevity and clarity, the embodiments of the present disclosure have been described. The scope of the technical concept of the embodiments of the present disclosure is not limited by the examples. Therefore, it will be understood by those skilled in the art that the scope of the present disclosure should not be limited by the embodiments explicitly described above, but by the claims and their equivalents.

[0261] (reference numerals)

[0262] 140: Transformer

[0263] 165: Inverse Converter

[0264] 530: Inverter

[0265] 606: Sub-block divider

[0266] 610: Inverse Transform Kernel Determiner

[0267] 612: Transformation unit determiner

[0268] 614: Inverse Transformation Actuator

[0269] CROSS-REFERENCE TO RELATED APPLICATIONS

[0270] This application claims priority to and the benefit of Korean Patent Application No. 10-2022-0124290 filed on September 29, 2022 and Korean Patent Application No. 10-2023-0105462 filed on August 11, 2023, which are hereby incorporated by reference in their entirety.

Claims

1. A method for reconstructing a current block performed by a video decoding apparatus, the method comprising the following steps: Obtain inverse quantization transform coefficients of a transform block for the current block; Decode a non-separable primary transform (NSPT) flag from a bitstream, wherein the NSPT flag indicates whether to apply a non-separable primary transform; Check the NSPT flag; When the NSPT flag is true, determine a non-separable primary inverse transform kernel based on the size of the transform block, the intra prediction mode of the current block, and the characteristics of the inverse quantization transform coefficients; And Perform a primary inverse transform by generating a residual signal by applying the non-separable primary inverse transform kernel to the transform coefficients.

2. The method according to claim 1, further comprising the following steps: When the NSPT flag is false, explicitly or implicitly obtain a pair of primary inverse transform kernels for the vertical and horizontal directions of the transform block; And Generate the residual signal by applying the pair of primary inverse transform kernels for the vertical and horizontal directions to the transform coefficients.

3. The method according to claim 1, wherein The step of determining the non-separable primary inverse transform kernel includes: Select a mode set including the intra prediction mode of the current block from a predefined set of modes; Select a set of kernels based on the size of the transform block and the selected mode set.

4. The method according to claim 3, wherein, The step of determining the non-separable primary inverse transform kernel includes: Adaptively determine the number of inverse transform kernel candidates for the set of kernels based on the sum of the absolute values of the inverse quantization transform coefficients or the position of the first non-zero transform coefficient.

5. The method according to claim 3, further comprising the following steps: When the set of kernels includes multiple inverse transform kernel candidates, decode an NSPT index, wherein the step of determining the non-separable primary inverse transform kernel includes: Determine the candidate indicated by the NSPT index among the multiple inverse transform kernel candidates as the non-separable primary inverse transform kernel.

6. The method according to claim 1, wherein, The non-separable primary inverse transform kernel is a matrix of size "P × the size of the transform block", where the size of the transform block is defined as the product of the width and height of the transform block, and the size of P corresponds to the size of a one-dimensional vector constructed by packing the transform coefficients.

7. The method according to claim 6, wherein, The size of P is less than or equal to the size of the transform block and is determined based on the size of the transform block.

8. The method according to claim 1, wherein When the width and height of the transform block are greater than or equal to a predefined minimum size applicable to NSPT and less than or equal to a predefined maximum size of NSPT, perform decoding of the NSPT flag, wherein the predefined minimum size applicable to NSPT represents the minimum size of a transform block to which a non-separable primary transform is applied, and the predefined maximum size of NSPT represents the maximum size of a transform block for decoding a non-separable primary transform flag.

9. The method according to claim 8, wherein When the width and height of the transform block are greater than or equal to a predefined minimum size applicable to NSPT and less than or equal to a predefined maximum size applicable to NSPT, perform the step of determining the non-separable primary inverse transform kernel, Among them, the predefined maximum size applicable to NSPT is less than or equal to the predefined maximum size of NSPT, and represents the maximum size of the transform block to which the non-separable one-time transform is applied.

10. The method according to claim 9, wherein When the width and height of the transform block are greater than the predefined maximum size applicable to NSPT, the steps of determining the non-separable one-time inverse transform kernel include: Implicitly and recursively splitting the transform block to generate sub-blocks until the width and height of the sub-blocks become less than the predefined maximum size applicable to NSPT.

11. The method according to claim 10, wherein, Performing inverse transforms on the sub-blocks in sequence according to the z-scan order, and the z-scan order is determined based on the intra-frame prediction mode of the current block.

12. The method according to claim 1, wherein The steps of performing a one-time inverse transform include: Based on the size and type of the non-separable one-time inverse transform kernel, packing all or part of the transform coefficients into a one-dimensional vector according to a predefined first scan order to generate a one-time transform coefficient vector; Performing matrix multiplication between the one-time transform coefficient vector and the non-separable one-time inverse transform kernel to generate a vector of the residual signal; and According to a predefined second scan order, assigning the vector of the residual signal to the transform block to generate the residual signal.

13. The method according to claim 1, wherein, When the NSPT flag is true, implicitly setting the quadratic transform flag or the quadratic transform index to 0, wherein the quadratic transform flag or the quadratic transform index indicates whether to perform a quadratic transform.

14. A method for encoding a current block executed by a video encoding device, the method comprising the following steps: Obtaining a residual signal for a transform block of the current block; Determining a non-separable one-time transform kernel based on the size of the transform block, the intra-frame prediction mode of the current block, and the characteristics of the quantized transform coefficients; Generating first one-time transform coefficients by applying the non-separable one-time transform kernel to the residual signal; Explicitly or implicitly determining a pair of one-time transform kernels for the vertical direction and the horizontal direction of the transform block; and Generating second one-time transform coefficients by applying the pair of one-time transform kernels to the residual signal.

15. The method according to claim 14, further comprising the following steps: Determining a non-separable one-time transform NSPT flag based on the first one-time transform coefficients and the second one-time transform coefficients, wherein the NSPT flag indicates whether to apply the non-separable one-time transform; and Encoding the NSPT flag.

16. The method according to claim 15, further comprising the following steps: Encoding the first one-time transform coefficients or the second one-time transform coefficients based on the NSPT flag.

17. A computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising the following steps: Obtaining a residual signal for a transform block of a current block; Determining a non-separable one-time transform kernel based on the size of the transform block, the intra-frame prediction mode of the current block, and the characteristics of the quantized transform coefficients; Generating first one-time transform coefficients by applying the non-separable one-time transform kernel to the residual signal; Explicitly or implicitly determine a first transform check for a vertical direction and a horizontal direction of the transform block; and Generate second first transform coefficients by applying the first transform check to the residual signal.

Citation Information

Patent Citations

  • Video remote communication system

    KR100268614B1

  • System and method for debugging the delivery of content items

    KR1020220124290A

  • Monitoring Device with Trip-Link of Protection Distribution Panel

    KR1020230105462A