Image encoding / decoding method and apparatus, and recording medium for storing bit stream

By adopting inseparable main transformation and dimensionality reduction technologies in image compression technology, the efficient compression problem of high-resolution images is solved, and efficient decoding of high-definition and ultra-high-definition images is achieved.

CN120345255APending Publication Date: 2025-07-18LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086975.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-12-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing image compression techniques are insufficient in high resolution and high-quality image processing, especially when applying inseparable main transformations, it is difficult to effectively compress and decode high-definition and ultra-high-definition images.

Method used

Inseparable main transform (NSPT) is used for transformation, and image blocks are reconstructed by obtaining residual information and applying backward NSPT, combined with dimensionality reduction NSPT cores to improve coding efficiency.

Benefits of technology

By using inseparable main transformation and dimensionality reduction technology, the performance and encoding efficiency of image compression are improved, and are suitable for efficient decoding of high-definition and ultra-high-definition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345255A_ABST
    Figure CN120345255A_ABST
Patent Text Reader

Abstract

The image decoding method and apparatus disclosed herein may: acquire residual information from a bitstream; deriving a transform coefficient of the current block based on the residual information; applying an inverse inseparable main transform to at least one of the transform coefficients of the current block in order to derive a residual sample of the current block; and reconstructing the current block based on the residual sample of the current block. Here, an inverse non-separable main transform may be applied based on the size of the current block belonging to a group of one or more block sizes to which the non-separable main transform can be applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. Background Art

[0002] Recently, the demand for high-resolution and high-quality images such as HD (high definition) images and UHD (ultra high definition) images has been increasing in various application fields, and thus, efficient image compression techniques are being discussed.

[0003] There are various techniques, such as an inter prediction technique that predicts a pixel value included in a current picture from a picture before or after the current picture by using a video compression technique, an intra prediction technique that predicts a pixel value included in the current picture by using pixel information in the current picture, an entropy coding technique that assigns a short symbol to a value with a high occurrence frequency and a long symbol to a value with a low occurrence frequency, etc., and these image compression techniques can be used to effectively compress image data and transmit or store it. Summary of the Invention

[0004] Technical Problem

[0005] The present disclosure aims to provide a method and apparatus for performing a transform by using an inseparable principal transform.

[0006] The present disclosure aims to provide a method and apparatus for performing a transform by using a reduced-dimensional inseparable principal transform kernel.

[0007] The present disclosure aims to provide a method and apparatus for determining an inseparable principal transform kernel based on encoding parameters.

[0008] Technical Solution

[0009] An image decoding method and apparatus according to the present disclosure may: obtain residual information from a bitstream, derive transform coefficients of a current block based on the residual information, apply a backward non-separable principal transform (NSPT) to at least one of the transform coefficients of the current block to derive residual samples of the current block, and reconstruct the current block based on the residual samples of the current block.

[0010] In the image decoding method and apparatus according to the present disclosure, the backward NSPT may be applied based on the size of the current block belonging to a group of one or more block sizes to which the NSPT is applicable.

[0011] In the image decoding method and apparatus according to the present disclosure, the NSPT set for the backward NSPT may be determined as any one of 35 predefined NSPT sets.

[0012] In the image decoding method and apparatus according to the present disclosure, each of the 35 NSPT sets may include three NSPT kernel candidates.

[0013] In the image decoding method and apparatus according to the present disclosure, the group may include at least one of 4x4, 4x8, 8x4, 8x8, 16x8, or 8x16.

[0014] In the image decoding method and apparatus according to the present disclosure, based on the size of the current block being 8x16, the number of transform coefficients to which the backward NSPT is applied may be 40.

[0015] The image encoding method and apparatus according to the present disclosure may: derive residual samples of a current block, apply a non-separable primary transform (NSPT) to the residual samples of the current block to derive transform coefficients of the current block, generate residual information regarding the transform coefficients of the current block, and encode the residual information to generate a bitstream. Here, the NSPT may be applied based on the size of the current block belonging to a group of one or more block sizes to which the NSPT is applicable.

[0016] There is provided a computer-readable digital storage medium storing encoded video / image information, which causes an image decoding method to be performed by a decoding apparatus according to the present disclosure.

[0017] There is provided a computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the present disclosure.

[0018] There is provided a method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure.

[0019] Advantageous Effects

[0020] The present disclosure may improve the performance of the transform by using a non-separable primary transform as the primary transform.

[0021] The present disclosure may improve the performance of the transform by performing the transform using a dimension-reduced non-separable primary transform kernel.

[0022] The present disclosure may improve the coding efficiency by effectively determining or signaling a non-separable primary transform kernel based on coding parameters such as an intra prediction mode or block size / shape. Description of the Drawings

[0023] Figure 1 Shows a video / image compilation system according to the present disclosure.

[0024] Figure 2 Shows a schematic block diagram of an encoding apparatus to which embodiments of the present disclosure are applicable and which performs encoding of video / image signals.

[0025] Figure 3Schematic block diagram of a decoding device to which embodiments of the present disclosure are applicable and which decodes video / image signals.

[0026] Figure 4 Illustration of an image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure.

[0027] Figure 5 Exemplary illustration of an intra prediction mode and its prediction direction according to the present disclosure.

[0028] Figure 6 Illustration of a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure.

[0029] Figure 7 Illustration of an image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure.

[0030] Figure 8 Illustration of a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure.

[0031] Figure 9 Illustration of an example of a content streaming system to which embodiments of the present disclosure can be applied. Detailed Description

[0032] Since the present disclosure can be made in various changes and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all changes, equivalents, and alternatives included in the spirit and technical scope of the present disclosure. When describing each drawing, like reference numerals are used for like components.

[0033] Terms such as first, second, etc. may be used to describe various components, but the components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the rights of the present disclosure, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term and / or includes any one or a combination of more than one of the related recited items.

[0034] When a component is referred to as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but there may also be another component in the middle. On the other hand, when a component is referred to as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in the middle.

[0035] The terms used in this application are only for describing specific embodiments and are not intended to limit the present disclosure. Unless otherwise clearly indicated in the context, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the existence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not preclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof in advance.

[0036] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein can be applied to the methods disclosed in the Versatile Video Coding (VVC) standard. Additionally, the methods / embodiments disclosed herein can be applied to the methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Second Generation Audio Video Coding Standard (AVS2), or the next-generation video / image coding standards (such as H.267 or H.268, etc.).

[0037] This specification presents various embodiments of video / image coding, and unless otherwise stated, these embodiments can be combined with each other for implementation.

[0038] Here, video can refer to a collection of a series of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a picture in coding. A slice / tile can include at least one Coding Tree Unit (CTU). A picture can be composed of at least one slice / tile. A tile is a rectangular region composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular region of CTUs having the same height as the picture and a width assigned by the syntax requirements of the Picture Parameter Set. A tile row is a rectangular region of CTUs having a height assigned by the Picture Parameter Set and the same width as the picture. The CTUs within a tile can be arranged continuously according to the CTU raster scan, while the tiles within a picture can be arranged continuously according to the tile raster scan. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that can be exclusively included in a single NAL unit. At the same time, a picture can be divided into at least two sub-pictures. A sub-picture can be a rectangular region of at least one slice within a picture.

[0039] A pixel, pel, or picture element can refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0040] A unit may represent a basic unit for image processing. The unit may include at least one of a specific area of a picture and information related to the corresponding area. A unit may include one luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, the unit may be used interchangeably with terms such as block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or an array of samples) consisting of M columns and N rows.

[0041] Here, "A or B" may refer to "only A", "only B", or "both A and B". In other words, here, "A or B" may be interpreted as "A and / or B". For example, here, "A, B, or C" may refer to "only A", "only B", "only C", or any combination of A, B, and C.

[0042] The slashes ( / ) or commas used herein may refer to "and / or". For example, "A / B" may refer to "A and / or B". Thus, "A / B" may refer to "only A", "only B", or "both A and B". For example, "A, B, C" may refer to "A, B, or C".

[0043] Here, "at least one of A and B" may refer to "only A", "only B", or "both A and B". Additionally, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0044] Furthermore, here, "at least one of A, B, and C" may refer to "only A", "only B", "only C", or any combination of A, B, and C. Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may refer to "at least one of A, B, and C".

[0045] Moreover, the parentheses used in this document may refer to "for example". Specifically, when it is indicated as "prediction (intra prediction)", "intra prediction" may be presented as an example of "prediction". In other words, "prediction" here is not limited to "intra prediction", and "intra prediction" may be presented as an example of "prediction". Additionally, even when it is indicated as "prediction (i.e., intra prediction)", "intra prediction" may be presented as an example of "prediction".

[0046] Here, the technical features described separately in one drawing may be implemented separately or simultaneously.

[0047] Figure 1 A video / image compilation system according to the present disclosure is shown.

[0048] Reference Figure 1, a video / image compilation system may include a first device (source device) and a second device (receiving device).

[0049] The source device may send the encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or a network. The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of a separate device or an external component.

[0050] The video source may obtain video / images through the process of capturing, synthesizing, or generating video / images. The video source may include devices for capturing video / images and devices for generating video / images. Devices for capturing video / images may include at least one camera, a video / image archive including previously captured video / images, etc. Devices for generating video / images may include computers, tablets, smartphones, etc., and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer, etc., and in this case, the process of capturing video / images may be replaced by the process of generating relevant data.

[0051] The encoding device may encode the input video / images. The encoding device may perform a series of processes such as prediction, transformation, quantization, etc. for compression and compilation efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0052] The sending unit may send the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The sending unit may include elements for generating a media file in a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit may receive / extract the bitstream and send it to the decoding device.

[0053] The decoding device may decode the video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.

[0054] The renderer may render the decoded video / images. The rendered video / images may be displayed through the display unit.

[0055] Figure 2A rough block diagram of an encoding apparatus that shows embodiments to which the present disclosure can be applied and that performs encoding of video / image signals is shown.

[0056] Referring Figure 2 , the encoding apparatus 200 may be composed of an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the above-described image splitter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0057] The image splitter 210 may partition an input image (or picture, frame) input to the encoding apparatus 200 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure.

[0058] For example, one coding unit may be partitioned into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding process according to this specification may be performed based on the final coding unit that is no longer partitioned. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit may be directly used as the final coding unit, or if necessary, the coding unit may be recursively partitioned into coding units with a deeper depth, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later.

[0059] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or split respectively from the above-mentioned final compilation unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0060] In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an MxN block may represent a set of transform coefficients or samples composed of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample may be used as a term corresponding to a pixel or a cell in a picture (or an image).

[0061] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.

[0062] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a prediction block including predictions of the prediction samples for the current block. The predictor 220 may determine whether to apply intra-frame prediction or inter-frame prediction in units of the current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, etc., and send it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0063] The intra-frame predictor 222 may predict the current block by referring to samples within the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be located at a certain distance from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and a plurality of directional modes. The non-directional mode may include at least one of the DC mode or the planar mode. Depending on the level of detail of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is only an example, and more or fewer directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0064] The inter - frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes, and for example, for the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. For the skip mode, different from the merge mode, the residual signal may not be sent. For the motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0065] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra - frame prediction or inter - frame prediction to predict a block, but also apply intra - frame prediction and inter - frame prediction simultaneously. It can be referred to as the combined inter - frame and intra - frame prediction (CIIP) mode. Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video compilation such as games, such as screen content compilation (SCC), etc. IBC basically performs prediction within the current picture, but it can be performed similar to inter - frame prediction because it derives a reference block within the current picture. In other words, IBC can use at least one of the inter - frame prediction techniques described herein. The palette mode can be considered an example of intra - frame compilation or intra - frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and the palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstructed signal or a residual signal.

[0066] The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transform process may be applied to square pixel blocks of the same size or may be applied to non-square blocks of variable size.

[0067] The quantizer 233 may quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a 1D vector form based on the coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in 1D vector form.

[0068] The entropy encoder 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (e.g., values of syntax elements, etc.) in addition to the transform coefficients quantized together or separately.

[0069] Encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of Network Abstraction Layer (NAL) units. The video / image information may further include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. Here, the information and / or syntax elements sent / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be sent over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (not shown) for sending and / or the storage unit (not shown) for storing the signal output from the entropy encoder 240 may be configured as internal / external elements of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.

[0070] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the dequantization and inverse transformation can be applied to the quantized transform coefficients by the dequantizer 234 and the inverse transformer 235 to reconstruct the residual signal (residual block or residual samples). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture and can also be used for inter-frame prediction of the next picture through filtering described later. Meanwhile, the Luminance Mapping with Chroma Scaling (LMCS) can be applied during the picture encoding and / or reconstruction process.

[0071] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270, specifically in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about the filtering and send it to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0072] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied thereto, the encoding device can avoid prediction mismatches in the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0073] The DPB of the memory 270 can store the modified reconstructed picture to use it as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.

[0074] Figure 3 A rough block diagram of a decoding device that can apply embodiments of the present disclosure and perform decoding of video / image signals is shown.

[0075] Reference Figure 3 , the decoding device 300 can be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321.

[0076] According to an embodiment, the above entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 can be configured by one hardware component (e.g., a decoder chipset or a processor). Additionally, the memory 360 can include a decoded picture buffer (DPB) and can be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0077] When a bitstream including video / image information is input, the decoding device 300 can reconstruct an image in response to the process of processing the video / image information in the Figure 2 encoding device. For example, the decoding device 300 can derive units / blocks based on the relevant information of block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using the processing units applied in the encoding device. Therefore, the decoded processing unit can be a compilation unit, and the compilation unit can be divided from a compilation tree unit or a maximum compilation unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. At least one transform unit can be derived from the compilation unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.

[0078] The decoding device 300 can receive a signal output from the Figure 2 encoding device in the form of a bitstream, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information can further include information about various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information can further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information sent / received by signal and / or the syntax elements described later in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements necessary for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element from the bitstream, determine the context model by using the information of the syntax element to be decoded, the neighboring blocks, and the decoding information of the block to be decoded, or the information of the symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins about the context model for the next symbol / bin. Among the information decoded in the entropy decoder 310, the information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values for which entropy decoding is performed in the entropy decoder 310, i.e., the quantized transform coefficients and the related parameter information, can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.

[0079] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0080] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 may perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.

[0081] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0082] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the information about prediction output from the entropy decoder 310, and determine a specific intra-frame / inter-frame prediction mode.

[0083] The predictor 320 may generate a prediction signal based on various prediction methods described later. For example, the predictor 320 may not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. It may be referred to as a combined inter-frame and intra-frame prediction (CIIP) mode. Additionally, the predictor may be based on the intra-block copy (IBC) prediction mode or may be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode may be used for content image / video compilation such as games, such as screen content compilation (SCC), etc. IBC basically performs prediction within the current picture, but it may be performed similar to inter-frame prediction because it derives a reference block within the current picture. In other words, IBC may use at least one of the inter-frame prediction techniques described herein. The palette mode may be considered an example of intra-frame compilation or intra-frame prediction. When the palette mode is applied, information about the palette table and palette index may be included in the video / image information and signaled.

[0084] The intra predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the samples referred to can be located near the current block or can be located at a certain distance from the current block. In intra prediction, the prediction mode can include at least one non - directional mode and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0085] The inter predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter - prediction mode for the current block.

[0086] The adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when the skip mode is applied, the prediction block can be used as the reconstructed block.

[0087] The adder 340 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for intra prediction of the next block to be processed in the current picture, can be output through filtering described later, or can be used for inter prediction of the next picture. Meanwhile, a luminance mapping with chroma scaling (LMCS) can be applied during the picture decoding process.

[0088] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and send the modified reconstructed picture to the memory 360, specifically to the DPB of the memory 360. Various filtering methods can include de - blocking filtering, sample - adaptive offset, adaptive loop filter, bilateral filter, etc.

[0089] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-predictor 332. The memory 360 can store the motion information of the blocks from which the motion information in its current picture is derived (or decoded) and / or the motion information of the blocks in the pre-reconstructed picture. The stored motion information can be sent to the inter-predictor 260 to be used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-predictor 331.

[0090] Here, the embodiments described in the filter 260, the inter-predictor 221, and the intra-predictor 222 of the encoding device 200 can also be equally or correspondingly applied to the filter 350, the inter-predictor 332, and the intra-predictor 331 of the decoding device 300, respectively.

[0091] Figure 4 The figure illustrates an image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure.

[0092] Reference Figure 4 , the transform coefficients of the current block can be derived from the bitstream (S400). That is, the bitstream can include the residual information of the current block, and the transform coefficients of the current block can be derived by decoding the residual information.

[0093] Reference Figure 4 , the residual samples of the current block can be derived by performing at least one of dequantization and inverse transformation on the transform coefficients of the current block (S410).

[0094] When adaptive multiple transform selection (MTS) is applied, the inverse transformation can be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc. can be referred to as transform types, transform kernels, or transform cores.

[0095] In the present disclosure, the inverse transformation can mean a separable transformation. However, it is not limited thereto, the inverse transformation can mean a non-separable transformation, or can be a concept including separable transformation and non-separable transformation. In addition, the inverse transformation in the present disclosure means the main transformation, but is not limited thereto, and can be applied to the secondary transformation by being modified to the same / similar form.

[0096] For example, as a method for inverse transformation, only DCT-2 and non-separable transformation can be used, or non-separable transformation can be used in addition to at least one of DCT-2, DST-7, or DCT-8, or non-separable transformation can replace the transform kernel of one or more of DCT-2, DST-7, or DCT-8.

[0097] As a more specific embodiment, when there are (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), (DCT-8, DCT-8) as candidates for the transform kernel for separable transform, the non-separable transform can replace or be added to one or more of the five transform kernel candidates. Here, the notation (transform 1, transform 2) indicates that transform 1 is applied in the horizontal direction and transform 2 is applied in the vertical direction. When the non-separable transform replaces some of the transform kernel candidates, the remaining transform kernel candidates other than (DCT-2, DCT-2) and (DST-7, DST-7) can be replaced by the non-separable transform. However, the above transform kernel candidates are only examples, and may include other types of DCT and / or DST, and may include transform skip as a transform kernel candidate.

[0098] The non-separable transform can mean a transform or inverse transform based on a non-separable transform matrix. That is, different from the separable transform that independently performs horizontal and vertical transforms by separating the vertical and horizontal transforms, the non-separable transform can perform horizontal and vertical transforms simultaneously.

[0099] For example, when performing a non-separable transform on a 4x4 block, the input data X to the non-separable transform is as shown in Equation 1 below.

[0100] [Equation 1]

[0101]

[0102] When the input data X is expressed in vector form, the vector X' can be expressed as follows.

[0103] [Equation 2]

[0104]

[0105] In this case, the non-separable transform can be performed as in Equation 3 below.

[0106] [Equation 3]

[0107]

[0108] In Equation 3, F represents the transform coefficient vector, T represents the 16x16 non-separable transform matrix, and • represents the product of the matrix and the vector.

[0109] The 16x1 transform coefficient vector F can be derived from Equation 3, and F can be reconfigured into a 4x4 block according to a predetermined scanning order. The scanning order can be horizontal scanning, vertical scanning, diagonal scanning, z-scanning, raster scanning, or predefined scanning.

[0110] The inseparable transform set and / or the transform kernel for the inseparable transform can be configured differently based on at least one of a prediction mode (e.g., an intra mode, an inter mode, etc.), the width, height, or number of pixels of the current block, the position of a sub-block within the current block, a syntax element signaled explicitly, statistical characteristics of neighboring samples, whether a secondary transform is used, or a quantization parameter (QP).

[0111] Specifically, for the intra mode, predefined intra prediction modes can be grouped to correspond to n inseparable transform sets, and each inseparable transform set can include k transform kernel candidates. Here, n and k can be arbitrary constants according to rules (conditions) defined the same for an encoding device and a decoding device.

[0112] The number of inseparable transform sets and / or the number of transform kernel candidates included in the inseparable transform set can be configured differently depending on the width and / or height of the current block. For example, for a 4x4 block, n1 inseparable transform sets and k1 transform kernel candidates can be configured. For a 4x8 block, n2 inseparable transform sets and k2 transform kernel candidates can be configured. Additionally, the number of inseparable transform sets and the number of transform kernel candidates included in each inseparable transform set can be configured differently depending on the product of the width and height of the current block. For example, when the product of the width and height of the current block is equal to or greater than 256, n3 inseparable transform sets and k3 transform kernel candidates can be configured, and otherwise, n4 inseparable transform sets and k4 transform kernel candidates can be configured. That is, since the degree of change in the statistical characteristics of the residual signal varies depending on the block size, the number of inseparable transform sets and transform kernel candidates can be configured differently to reflect this.

[0113] When the current block is divided into multiple sub-blocks, the statistical characteristics of the residual signal may be different for each sub-block, and thus the number of inseparable transform sets and transform kernel candidates can be configured differently. For example, when a 4x8 or 8x4 block is divided into two 4x4 sub-blocks and an inseparable transform is applied to each sub-block, n5 inseparable transform sets and k5 transform kernel candidates can be configured for the upper left 4x4 sub-block, and n6 inseparable transform sets and k6 transform kernel candidates can be configured for the other 4x4 sub-block.

[0114] Based on the syntax elements signaled explicitly, the number of non-separable transform sets and transform kernel candidates can be configured differently. As the syntax elements, information indicating one of multiple non-separable transform configurations can be used. For example, when three non-separable transform configurations are supported (i.e., n7 non-separable transform sets and k7 transform kernel candidates, n8 non-separable transform sets and k8 transform kernel candidates, n9 non-separable transform sets and k9 transform kernel candidates), the syntax element can have values of 0, 1, and 2, and the non-separable transform configuration applied to the current block can be determined based on the value of the syntax element signaled.

[0115] Based on whether a secondary transform is applied and / or which secondary transform is applied, the number of non-separable transform sets and transform kernel candidates can be configured differently. For example, when no secondary transform is applied, a non-separable transform configuration including n 10 non-separable transform sets and k 10 transform kernel candidates can be applied. When a secondary transform is applied, a non-separable transform configuration including n 11 non-separable transform sets and k 11 transform kernel candidates can be applied.

[0116] Based on the quantization parameter (QP) and / or the range to which the QP value belongs, different non-separable transform configurations can be applied. For example, when the QP value has a small value, a non-separable transform configuration including n 12 non-separable transform sets and k 12 transform kernel candidates can be applied. On the other hand, when the QP value has a large value, a non-separable transform configuration including n 13 non-separable transform sets and k 13 transform kernel candidates can be applied. When the QP value is less than or equal to a threshold (e.g., 32), the case is classified as having a small QP value, and otherwise, the case is classified as having a large QP value. Alternatively, the range of the QP value can be divided into three or more, and different non-separable transform configurations can be applied to each range.

[0117] For a large block, instead of using the non-separable transform corresponding to the width and height of the block, the block can be divided into multiple sub-blocks and the non-separable transform corresponding to the width and height of the sub-blocks can be used. For example, when performing a non-separable transform on a 4x8 block, the 4x8 block can be divided into two 4x4 sub-blocks, and the non-separable transform based on the 4x4 block can be used for each of the 4x4 sub-blocks. Alternatively, an 8x16 block can be divided into two 8x8 sub-blocks, and the non-separable transform based on the 8x8 block can be used.

[0118] The non-separable transform set can be determined based on the intra prediction mode of the current block and a mapping table. The mapping table can define a mapping relationship between predefined intra prediction modes and non-separable transform sets. The predefined intra prediction modes can include two non-directional modes and 65 directional modes. Generally, non-separable transforms have a larger transform kernel size than separable transforms. This means that the computational complexity required for the transform process is high and the memory required to store the transform kernel is large. At the same time, while separable transforms may only consider the statistical characteristics present in the horizontal and / or vertical directions, non-separable transforms can consider the statistical characteristics in a two-dimensional space including both horizontal and vertical directions, thereby providing better compression efficiency. Since the statistical characteristics and diversity of the residual depend on the orientation of the intra prediction mode, there may be cases where non-separable transforms are absolutely necessary, and there may be intra prediction modes where the residual characteristics can be fully identified only by separable transforms. Therefore, by predefining which transform to use based on the intra prediction mode in the encoding device and the decoding device, a transform process with optimized complexity and memory requirements can be designed. The non-directional modes can include the planar mode numbered 0 and the DC mode numbered 1, and the directional modes can include the intra prediction modes numbered 2 to 66. However, this is only an example, and the present disclosure can also be applied to cases where the numbers of the predefined intra prediction modes are different.

[0119] Due to the application of wide-angle intra prediction (WAIP), the predefined intra prediction modes can further include intra prediction modes from -14 to -1 and intra prediction modes from 67 to 80.

[0120] Figure 5 Exemplarily shown are the intra prediction modes according to the present disclosure and their prediction directions. Refer to Figure 5 , modes -14 to -1 and 2 to 33 and modes 35 to 80 are symmetric with respect to mode 34 in terms of the prediction direction. For example, mode 10 and mode 58 are symmetric with respect to the direction corresponding to mode 34, and mode -1 is symmetric with mode 67. Therefore, for vertical directional modes that are symmetric with respect to the horizontal directional mode with respect to mode 34, the input data can be transposed and used. Transposing the input data means that the rows and columns in the input data MxN of the two-dimensional block become columns and rows, respectively, to form NxM data.

[0121] For example, when using a 4x4 block, the 16 data forming the 4x4 block can be appropriately arranged to form a 16x1 1-dimensional vector for non-separable transform. In this case, the 1-dimensional vector can be formed in row-major order or column-major order. The residual samples caused by the non-separable transform can be arranged in the above order to form a two-dimensional block.

[0122] For patterns - 14 to - 1 and 2 to 33, when the data arrangement order used to form a 16x1 input vector is in row - major order, for patterns 35 to 80, the input vector can be formed in column - major order.

[0123] Pattern 34 cannot be regarded as either a horizontally - oriented pattern or a vertically - oriented pattern, but in the present disclosure, it is classified as belonging to the horizontally - oriented pattern. That is, for patterns - 14 to - 1 and 2 to 33, the input data arrangement method for the horizontally - oriented pattern is used, that is, row - major order, and for the vertically - oriented pattern symmetric to pattern 34, the input data can be transposed and used.

[0124] For non - square blocks, the symmetry in square blocks (i.e., the symmetry between pattern P and pattern (68 - P) in an NxN block (2 <= P <= 33) or the symmetry between pattern Q and pattern (66 - Q) (- 14 <= Q <= - 1)) cannot be utilized. Therefore, in addition to the symmetry based only on the intra - frame prediction pattern, the symmetry between block shapes in a transposed relationship with each other can be utilized, that is, the symmetry between a KxL block and an LxK block. Specifically, there is a symmetric relationship between the KxL block predicted by pattern P and the LxK block predicted by pattern (68 - P). Alternatively, there is a symmetric relationship between the KxL block predicted by pattern Q and the LxK block predicted by pattern (66 - Q).

[0125] Since the KxL block with pattern 2 and the LxK block with pattern 66 can be regarded as symmetric to each other, the same transform kernel can be applied to the KxL block and the LxK block. If the non - separable transform set for the intra - frame prediction pattern of the KxL block is mapped, then in order to apply the non - separable transform to the LxK block, the non - separable transform set can be derived through the mapping table corresponding to the KxL block based on pattern (68 - P) instead of pattern P applied to the LxK block. Alternatively, the non - separable transform set can be derived through the mapping table corresponding to the KxL block based on pattern (66 - Q) instead of pattern Q applied to the LxK block.

[0126] For example, in order to apply the non - separable transform to the LxK block, the non - separable transform set can be selected based on pattern 2 instead of pattern 66. Additionally, for the KxL block, the input data can be read in a predetermined order (e.g., row - major order or column - major order) to form a 1D vector, and then the corresponding non - separable transform can be applied. For the LxK block, the input data can be read in a transposed order to form a 1D vector and then the corresponding non - separable transform can be applied. That is, when the KxL block is read in row - major order, the LxK block can be read in column - major order. Conversely, when the KxL block is read in column - major order, the LxK block can be read in row - major order.

[0127] In addition, when Pattern 34 is applied to a KxL block, an inseparable transform set can be determined based on Pattern 34, and the input data can be read in a predetermined order to form a 1D vector and perform the corresponding inseparable transform. When Pattern 34 is applied to an LxK block, an inseparable transform set can be determined based on Pattern 34, but the input data can be read in a transposed order to form a 1D vector and perform the corresponding inseparable transform.

[0128] In the present disclosure, a method for determining an inseparable transform set and a method for forming input data are described based on a KxL block. However, the inseparable transform can be performed based on an LxK block by utilizing the symmetry of the KxL block described above. Alternatively, a block having a width greater than its height can be restricted to be used as a reference block. Alternatively, the symmetry can be restricted not to be utilized in the case of a non-square block. In this case, a non-square block can use a different number of inseparable transform sets and / or transform kernel candidates from those of a square block, and can use a different mapping table from that of a square block to select an inseparable transform set.

[0129] An example of a mapping table for selecting an inseparable transform set is as follows:

[0130] [Table 1]

[0131]

[0132] Table 1 shows an example of assigning an inseparable transform set to each intra-frame prediction mode when there are five inseparable transform sets. The value of predModeIntra means the value of the intra-frame prediction mode considering WAIP, and TrSetIdx is an index indicating a specific inseparable transform set. In Table 1, it can be confirmed that the same inseparable transform set is applied to the modes located in the symmetric direction according to the intra-frame prediction mode. Table 1 is only an example using five inseparable transform sets and does not limit the total number of inseparable transform sets for the inseparable transform.

[0133] Alternatively, as shown in Table 2, for compression performance, the inseparable transform may not be applied to WAIP.

[0134] [Table 2]

[0135]

[0136] Alternatively, as shown in Table 3, instead of configuring a separate inseparable transform set for WAIP, an inseparable transform set corresponding to an adjacent intra-frame prediction mode can be shared.

[0137] [Table 3]

[0138]

[0139] The non-separable transform set may include a plurality of transform kernel candidates, and one of the plurality of transform kernel candidates may be selectively used. For this purpose, an index signaled via a bitstream may be used. Alternatively, one of the plurality of transform kernel candidates may be implicitly determined based on context information of the current block. Here, the context information may mean the size of the current block or whether a non-separable transform is applied to neighboring blocks. Here, the size of the current block may be defined as the width, the height, the maximum / minimum of the width and the height, the sum of the width and the height, or the product of the width and the height.

[0140] Hereinafter, a method for determining a transform kernel for inverse transform of a current block will be described in detail.

[0141] Example 1

[0142] As described above, the inverse transform may be divided into a separable transform and a non-separable transform. The separable transform means performing a transform in the horizontal direction and the vertical direction on a two-dimensional block respectively, and the non-separable transform may mean performing a single auxiliary transform on samples constituting the whole or part of the two-dimensional block. When expressing the separable transform, it may be expressed as a pair of a horizontal transform and a vertical transform, and in the present disclosure, it is expressed as (horizontal transform, vertical transform).

[0143] A plurality of transform sets may be defined for the inverse transform of the current block. Each transform set may include one or more transform kernel candidates.

[0144] For example, one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) may be applied as the separable transform, and the above four transform kernel candidates may be regarded as one transform set. In addition, (DCT-2, DCT-2) may be regarded as one transform set. The transform skip that does not apply a transform may also be regarded as one transform set, and (DCT-2, DCT-2) and the transform skip may be regarded as one transform set. In the present disclosure, the transform kernel may refer to a single transform (e.g., DCT-2, DST-7) or may refer to a pair of two transforms (e.g., (DCT-2, DCT-2)).

[0145] As another example of the transform set, there may be the above-mentioned non-separable transform set. In the present disclosure, the non-separable transform applied as the main transform may be represented as a non-separable main transform (NSPT). In the NSPT, multiple non-separable transform sets may be configured, and each non-separable transform set may include one or more transform kernels as transform kernel candidates. In the case of NSPT, one of the multiple non-separable transform sets is selected based on the intra prediction mode, and the multiple non-separable transform sets for NSPT may be represented as an NSPT set list. This is as described above, and its detailed description will be omitted here.

[0146] A group of one or more transform sets available for the current block may be configured from multiple predefined transform sets. The group of one or more transform sets may be configured in a predetermined region unit to which the current block belongs, and is hereinafter referred to as a set. Here, the predetermined region unit may be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU).

[0147] For example, the transform set composed of (DCT-2, DCT-2) is called S1, and the transform sets composed of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) are called S2. Additionally, the above NSPT set list may include N non-separable transform sets, and the N non-separable transform sets are respectively called S 3,1 、S 3,2 、...、S 3,N . Here, N may be 35, but is not limited thereto.

[0148] When S3,13 is selected as the non-separable transform set of NSPT based on the intra prediction mode of the current block, the transform kernel applicable to the current block may belong to one of S1, S2, or S 3,13 . In this case, the set available for the current block may be represented as {S1, S2, S 3,13}.

[0149] As described above, since the set according to the present disclosure is a group of one or more transform sets available for the current block, the set may be configured differently based on the context of the current block. Here, the context may include at least one of shape, size, or intra prediction mode. If a total of K contexts are defined, K sets may be generated, and each set may be represented as Ci (i = 1, 2,..., N). For example, when the sizes of the blocks to which NSPT is applicable are 4x4, 8x8, 16x16, and 32x32 and one of a total of 35 non-separable transform sets is selected based on the intra prediction mode, if different transform kernels are applied to each block size, a total of 4x35 = 140 contexts may be defined.

[0150] Based on the context configuration set of the current block, and in this case, a process of selecting one of multiple transform sets belonging to the set and selecting one of multiple transform kernel candidates belonging to the selected transform set can be performed. Here, the selection of the transform set and the transform kernel candidate can be implicitly performed based on the context of the current block, or can be performed based on an index explicitly signaled. Alternatively, the process of selecting one of multiple transform sets belonging to the set and the process of selecting one of multiple transform kernel candidates belonging to the selected transform set can be performed separately. For example, an index for selecting a transform set can be signaled first, and one of multiple transform sets belonging to the set can be selected based on the index. Then, an index indicating one of multiple transform kernel candidates belonging to the transform set can be signaled, and one transform kernel candidate can be selected from the transform set based on the signaled index. The transform kernel of the current block can be determined based on the selected transform kernel candidate. Alternatively, selecting one transform set from the set can be implicitly performed based on the context of the current block, and selecting one transform kernel candidate from the selected transform set can be performed based on the signaled index. Alternatively, selecting one transform set from the set can be performed based on the signaled index, and selecting one transform kernel candidate from the selected transform set can be implicitly performed based on the context of the current block. Alternatively, selecting one transform set from the set can be implicitly performed based on the context of the current block, and selecting one transform kernel candidate from the selected transform set can be implicitly performed based on the context of the current block. Of course, when the number of transform sets belonging to the set is 1, the index for selecting the transform set does not need to be signaled. Similarly, when the number of transform kernel candidates belonging to the selected transform set is 1, the index for indicating the transform kernel candidate does not need to be signaled. Alternatively, an index indicating one of all transform kernel candidates belonging to the current set can be signaled. In this case, the process of selecting one transform set from the set can be omitted. In this case, all transform sets belonging to the set can be shuffled considering the priority. For example, in the case of assigning a small-length binary code to a small-value index such as a truncated unary code, it may be advantageous to assign the small-value index to a transform kernel candidate that is more favorable for improving the compilation performance. When shuffling all transform kernel candidates belonging to the set according to the priority, different shuffles can be applied to each set. Additionally, instead of shuffling all transform kernel candidates belonging to the set, some of them can be selectively shuffled only.

[0151] Example 2

[0152] The transform kernel for the inverse transform of the current block can be determined based on MTS (Multiple Transform Selection).

[0153] The MTS according to the present disclosure may use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (identity transform) as a transform kernel. Additionally, the MTS according to the present disclosure may further include a transform kernel of DCT-2.

[0154] In the present disclosure, multiple MTS sets for the MTS may be defined. Based on the size of the current block and / or the intra prediction mode, one of the multiple MTS sets may be determined. For example, when determining an MTS set, 16 transform block sizes may be considered, and for the directional mode, the symmetry between the shape of the transform block and the intra prediction mode may be considered. For the WAIP (wide-angle intra prediction) mode (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), the MTS set corresponding to mode 2 may be applied to modes -1 to -14 (or -15), and the MTS set corresponding to mode 66 may be applied to modes 67 to 80 (or 81). A separate MTS set may be assigned for the MIP (matrix-based intra prediction) mode.

[0155] For example, MTS sets according to the transform block size and the intra prediction mode may be assigned / defined as shown in Table 4 below.

[0156] [Table 4]

[0157]

[0158] Table 4 shows the assignment of MTS sets according to 16 transform block sizes and the intra prediction mode. The number of predefined MTS sets is 80, and the index indicating one of the 80 MTS sets may have a value from 0 to 79, as shown in Table 4.

[0159] [Table 5]

[0160]

[0161]

[0162]

[0163] Table 5 shows the transform kernel candidates included in each MTS set described in Table 4. Each MTS set may consist of six transform kernel candidates. The transform kernel candidate index has a value from 0 to 5 and may indicate one of the six transform kernel candidates. Here, each transform kernel candidate may be a combination of a horizontal transform kernel and a vertical transform kernel for a separable transform, and 25 transform kernel candidates with indices from 0 to 24 may be defined.

[0164] [Table 6]

[0165]

[0166] Table 6 is an example of 25 transform kernel candidates described in Table 5. Specifically, the horizontal transform and the vertical transform of the transform kernel candidate are expressed as (horizontal transform, vertical transform). For each transform kernel candidate index, the horizontal / vertical transform when the intra prediction mode is less than 35 may be opposite to the horizontal / vertical transform when the intra prediction mode is greater than or equal to 35. When the value of the intra prediction mode is greater than or equal to 35, a mode symmetric to mode 34 may be derived, and the MTS set may be selected from Table 4 based on this mode. In addition, the symmetry of the block shape may be additionally considered. When the original transform block has a size of WxH, it may be regarded as having a size of HxW by symmetrizing the original transform block, and the MTS set may be selected from Table 4. Here, the value of the intra prediction mode may be the value of the modified intra prediction mode. That is, as the mode value of WAIP, for values from -14 (or -15) to -1, it is modified to mode 2, for values from 67 to 80 (or 81), it is modified to mode 66, and for the remaining modes, the value of the original intra prediction mode may be set as the value of the modified intra prediction mode. In this case, since the extended mode of WAIP is also configured to be symmetric with respect to mode 34, the symmetry with respect to mode 34 may be used for all directional modes except the planar mode and the DC mode.

[0167] For example, when predicting a 16x32 block based on mode 54, mode 14 (=68 - 54) may be derived as the mode symmetric to mode 54, and the block size may be regarded as 32x16. In this case, the MTS set with index 72 may be selected as defined in Table 4.

[0168] When applying the MIP mode, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, when applying the MIP mode, the MTS set assigned to the MIP mode may be selected based on the symmetric block size considering the symmetry of the block shape. For example, when applying the MIP mode to an 8x16 block, the 8x16 block may be regarded as a 16x8 block symmetric to it, and the MTS set with index 49 may be selected as defined in Table 4. Alternatively, when applying the MIP mode, the intra prediction mode may be regarded as the planar mode. In this case, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, the MTS set assigned to the MIP mode may be selected based on the symmetric block size considering the symmetry of the block shape.

[0169] For the MIP mode, a flag can be used to indicate whether the MIP mode is applied in the transposed mode. When the MIP mode is applied to the current MxN block and the flag indicates the application of the transposed mode, the intra prediction mode can be regarded as the planar mode, and the current MxN block can be regarded as an NxM block. That is to say, from Table 4, the MTS set corresponding to the block size of NxM and the planar mode can be selected. As described in Table 6, when the value of the intra prediction mode is greater than or equal to 35, the horizontal transform and the vertical transform are exchanged, but since the intra prediction mode of the current block is regarded as the planar mode, the horizontal transform and the vertical transform of the transform kernel candidate may not be exchanged. Alternatively, when the MIP mode is applied to the current MxN block and the flag indicates the application of the transposed mode, the intra prediction mode may not be regarded as the planar mode, and the current MxN block can be regarded as an NxM block. That is to say, from Table 4, the MTS set corresponding to the block size of NxM and the MIP mode can be selected.

[0170] In Table 5, the transform kernel candidate selected by the transform kernel candidate index can be set as the transform kernel of the current block. Alternatively, based on the size of the current block, at least one of the horizontal transform or the vertical transform of the selected transform kernel candidate can be changed to another transform kernel. For example, when the transform kernel candidate index is 3 and both the width and height of the current block are less than or equal to 16, at least one of the horizontal transform or the vertical transform of the transform kernel candidate corresponding to the transform kernel candidate index 3 can be changed to another transform kernel. In this case, the horizontal transform and the vertical transform can be changed independently of each other. When the difference (or the absolute value of the difference) between the value of the intra prediction mode of the current block and the value of the horizontal mode is less than or equal to a predetermined threshold, the vertical transform of the selected transform kernel candidate can be changed to the IDT (identity transform). When the difference (or the absolute value of the difference) between the value of the intra prediction mode of the current block and the value of the vertical mode is less than or equal to a predetermined threshold, the horizontal transform of the selected transform kernel candidate can be changed to the IDT (identity transform). Here, the threshold can be determined based on the width and height of the current block, as shown in Table 7 below.

[0171] [Table 7]

[0172]

[0173] Table 7 is used to change the horizontal transform and / or the vertical transform of the transform kernel candidate selected by the transform kernel candidate index to another transform kernel, and defines the threshold according to the size of the transform block.

[0174] The six transform kernel candidates that make up an MTS set can be distinguished by transform kernel candidate indices from 0 to 5, as defined in Table 5. The transform kernel candidate indices can be signaled via the bitstream. A flag (MTS enable flag or MTS flag) indicating whether the MTS set is available / applied can be signaled, and when the flag indicates that the MTS set is available / applied, the transform kernel candidate indices can be signaled. The MTS flag can consist of one bin, and one or more contexts (hereinafter referred to as CABAC contexts) for CABAC-based entropy coding can be assigned to the bin. For example, different CABAC contexts can be assigned to the non-MIP mode and the MIP mode, respectively.

[0175] Based on the context of the current block described above, the number of transform kernel candidates available for the current block can be set differently. For example, as the context of the current block, the sum of the absolute values of all or part of the transform coefficients in the current block can be considered. The sum of the absolute values of the transform coefficients is called AbsSum. When AbsSum is less than or equal to T1, only one transform kernel candidate corresponding to the transform kernel candidate index 0 is available. When AbsSum is greater than T1 and less than or equal to T2, the transform kernel candidates corresponding to the transform kernel candidate indices 0 to 3 may be available. When AbsSum is greater than T2, six transform kernel candidates corresponding to the transform kernel candidate indices 0 to 5 may be available. Here, T1 can be 6 and T2 can be 32, but this is only an example.

[0176] When AbsSum is less than or equal to T1, since the number of transform kernel candidates available for the current block is 1, the transform kernel candidate corresponding to the transform kernel candidate index 0 can be set as the transform kernel of the current block without signaling the transform kernel candidate index. When AbsSum is greater than T1 and less than or equal to T2, since four transform kernel candidates are available, one of the four transform kernel candidates can be selected based on the transform kernel candidate index having two bins. That is, the transform kernel candidate indices 0 to 3 can be signaled as 00, 01, 10, and 11, respectively. For these two bins, the MSB (Most Significant Bit) can be signaled first and the LSB (Least Significant Bit) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, a CABAC context other than the CABAC context assigned for the MTS flag can be assigned to each of the two bins. Alternatively, bypass coding can be applied without assigning a CABAC context to the two bins. When AbsSum is greater than T2, the transform kernel candidate index has values from 0 to 5, so the transform kernel candidate index cannot be expressed using only two bins. In this case, the transform kernel candidate index can be expressed by assigning two or more bins, such as truncated binary coding. For each bin assigned by the truncated binary coding method, a CABAC context can be assigned, or bypass coding can be applied without assigning a CABAC context. Alternatively, a CABAC context can be assigned to some of the multiple bins (e.g., the first bin, or the first bin and the second bin), and bypass coding can be applied to the remaining bins.

[0177] Example 3

[0178] The transform kernel of the current block can be determined based on a transform set including one or more transform kernel candidates. The transform kernel of the current block can be derived as one of the one or more transform kernel candidates belonging to the transform set.

[0179] The process of determining the transform kernel of the current block can include at least one of the following: 1) the process of determining the transform set of the current block, or 2) the process of selecting one transform kernel candidate from the transform set of the current block. The process of determining the transform set can be the process of selecting one of a plurality of predefined identical transform sets in the encoding device and the decoding device. Alternatively, the process of determining the transform set can be the process of configuring one or more transform sets available for the current block from among a plurality of predefined identical transform sets in the encoding device and the decoding device and selecting one of the configured transform sets. Alternatively, the process of determining the transform set can be the process of configuring a transform set based on the transform kernel candidates available for the current block from among a plurality of predefined identical transform kernel candidates in the encoding device and the decoding device.

[0180] When the transform set of the current block includes multiple transform kernel candidates, a process of selecting one of the multiple transform kernel candidates for the current block may be performed. However, when the transform set of the current block includes one transform kernel candidate (i.e., when the number of transform kernel candidates available for the current block is 1), the transform kernel of the current block may be set to the corresponding transform kernel candidate.

[0181] The transform set according to the present disclosure may refer to the (non-separable) transform set in Embodiment 1 above, or may refer to the MTS set in Embodiment 2. Alternatively, the transform set may be defined separately from the (non-separable) transform set in Embodiment 1 or the MTS set in Embodiment 2. In this case, the transform set may include one or more specific transform kernels as transform kernel candidates. A specific transform kernel may be defined as a pair of a transform kernel for horizontal transform and a transform kernel for vertical transform, or may be defined as one transform kernel that is equally applicable to both horizontal and vertical transforms.

[0182] In an embodiment of the present disclosure, a process of applying NSPT (non-separable transform) applicable to a primary transform is described in detail. NSPT may be applied to the entire transform block or a partial transform block. Based on the forward NSPT, residual samples existing in the region where NSPT is applied may be input as a 1D vector of NSPT. In other words, residual samples existing in all or part of a single transform block (referred to as a region of interest (ROI) in the present disclosure) may be collected as a 1D vector and configured as an input. Then, when the forward NSPT is applied, primary transform coefficients may be obtained. Conversely, when the backward NSPT is applied to the primary transform coefficients, a 1D vector output may be obtained. Residual samples for the ROI may be obtained by arranging each element value configuring the corresponding output vector at a position determined within the 2D transform block.

[0183] For the non-separable transform kernel for NSPT, the matrix dimension may be determined according to the size of the ROI. In the present disclosure, the transform kernel may be referred to as a transform type or a transform matrix, and the non-separable transform kernel for NSPT may be referred to as an NSPT kernel. For example, when the current block is an MxN transform block, the ROI is a region for the entire MxN transform block, and square NSPT is applied, the dimension of the corresponding transform matrix may be MN x MN. For example, when the ROI is a region for the entire 8x8 transform block, the dimension of the NSPT kernel may be 64 x 64.

[0184] According to an embodiment of the present disclosure, when NSPT is applied to residuals generated by intra prediction, the NSPT kernel may be adaptively determined according to the intra prediction mode. Since the statistical characteristics of the residual block may vary depending on the intra prediction mode, the compression efficiency may be improved by adaptively determining the NSPT kernel according to the intra prediction mode.

[0185] It can be configured to share the NSPT kernel applied to at least one intra prediction mode. As described above, the non-separable transform set can be determined based on the intra prediction mode of the current block and the mapping table. The mapping table can define the mapping relationship between the predefined intra prediction mode and the non-separable transform set. The predefined intra prediction mode can include two non-directional modes and 65 directional modes.

[0186] As an example, the intra prediction modes can be grouped into intra prediction mode groups. One NSPT kernel can be assigned to an intra prediction mode group, or multiple NSPT kernels can be assigned to an intra prediction mode group. In other words, the non-separable transform set (NSPT set) including at least one NSPT kernel can be assigned to the intra prediction mode group. The non-separable transform set can be mapped to the intra prediction mode, and one of the N NSPT kernels included in the non-separable transform set can be selected.

[0187] As an example, the intra prediction group can include adjacent prediction modes (e.g., modes 17, 18, and 19). Additionally, the intra prediction group can include modes with symmetry. For example, in the above Figure 5 , the directional modes can be symmetric about the diagonal mode (i.e., intra prediction mode 34). In this case, two symmetric modes can form a group (or a pair). For example, mode 18 and mode 50 can be included in the same group because they are symmetric about mode 34. However, for modes with symmetry, the process of adding a transposed 2D input block and then configuring a 1D input vector can be performed before applying the forward NSPT kernel. For example, when the intra prediction mode is less than or equal to 34, a one-dimensional input vector can be derived from the corresponding input block in row-major order without transposing the 2D input block. When the intra prediction mode is greater than 34, a 1D input vector can be configured by first transposing the 2D input block and then reading the corresponding input block in row-major order, or by keeping the 2D input block as it is and reading the corresponding input block in column-major order.

[0188] Table 8 below illustrates the mapping table for assigning NSPT sets according to the intra prediction mode. Referring to Table 8, a total of 35 NSPT sets can be defined from 0 to 34. The NSPT set assigned to the nearest general directional mode can be assigned to the extended WAIP mode (i.e., Figure 5 modes -14 to -1 and modes 67 to 80 in ). In other words, NSPT set 2 can be assigned to the extended WAIP mode.

[0189] [Table 8]

[0190]

[0191] The NSPT set may include at least one NSPT core (or core candidate). In other words, the NSPT set may include N NSPT core candidates. For example, N may be set to a value greater than or equal to 1, such as 1, 2, 3, 4, etc. Among the at least one NSPT core included in the NSPT set, the core applied to the current block may be signaled using an index. In the present disclosure, the corresponding index may be referred to as the NSPT index. For example, the NSPT index may have values of 0, 1, 2, …, N−1.

[0192] Additionally, as an example, when the number of NSPT core candidates is 1, the NSPT index value may be fixed to 0. In this case, the NSPT index may be inferred without being signaled separately. Additionally, a flag indicating whether NSPT is applied may be signaled separately from the NSPT index. In the present disclosure, the corresponding flag may be referred to as the NSPT flag.

[0193] When the NSPT flag value is 1, NSPT may be applied. When the NSPT flag value is 0, NSPT may not be applied. When the NSPT flag is not signaled, the NSPT flag value may be inferred as 0. For example, when the NSPT flag value is 1, the NSPT index may be applied. One of the N core candidates included in the NSPT set selected by the intra prediction mode may be specified based on the signaled NSPT index.

[0194] In an embodiment, an entropy coding method of the NSPT index may be defined in various ways by considering the number (N) of NSPT cores included in the NSPT set. For example, as a method for mapping values from 0 to N−1 to a bin string (i.e., a binarization method), truncated unary binarization, truncated binarization, and fixed-length binarization methods may be used.

[0195] For example, when the value of the number of core candidates N configuring the NSPT set is 2, one of the two candidates may be specified with one bin. For example, 0 may indicate the first candidate and 1 may indicate the second candidate. Additionally, when N has a value of 3 and truncated unary binarization is applied, the candidates may be specified with two bins. For example, the first, second, and third candidates may be binarized to 0, 10, and 11, respectively, and signaled. As an example, the binarized bins may be coded by using context coding or bypass coding.

[0196] In the present disclosure, a reduced principal transform (RPT) method using a dimension-reduced transform kernel as the principal transform is described. As described above, when applying the forward NSPT, samples belonging to a 2D residual block can be arranged (or rearranged) as a 1D vector according to row-major order (or column-major order). Thereafter, the transform matrix for NSPT can be multiplied by the arranged vector. When the corresponding 2D residual block is an M x N block (M is the horizontal length and N is the vertical length), the length of the rearranged 1D vector can be M*N. In other words, the corresponding 2D residual block can also be represented as a column vector with dimensions of M*N x 1. In the present disclosure, for convenience, M*N can be expressed as MN. In this case, the dimension of the corresponding transform matrix can be MN x MN. In summary, the forward NSPT transform can be operated in such a way that an MN x 1 transform coefficient vector is obtained by multiplying the left side of the MN x 1 vector by the corresponding MN x MN transform matrix.

[0197] When applying the RPT transform, r transform coefficients can be obtained by multiplying by an r x MN matrix, rather than multiplying by an MN x MN matrix as in the above-mentioned forward NSPT transform matrix. Here, r represents the number of rows of the transform matrix, and MN represents the number of columns of the transform matrix. According to an embodiment of the present disclosure, the value of r can be set to be less than or equal to MN. In other words, the existing forward NSPT transform matrix includes MN rows, and each row is a 1 x MN row vector and the transform basis vector of the corresponding NSPT transform matrix. The corresponding transform coefficients can be obtained by multiplying each transform basis vector by the MN x 1 sample column vector.

[0198] Since the existing forward NSPT transform matrix is composed of MN row vectors, MN transform coefficients (i.e., an MN x 1 transform coefficient column vector) can be obtained by applying the forward NSPT transform. At the same time, for the forward RPT, the transform matrix can be composed of r transform basis vectors, rather than MN transform basis vectors. Therefore, when applying the forward RPT transform, r, rather than MN, transform coefficients (i.e., an r x 1 transform coefficient column vector) can be obtained.

[0199] The RPT kernel can be configured by selecting r transform basis vectors, which are part of the transform basis vectors that configure the MN x MN forward NSPT kernel. In the present disclosure, a transform kernel may be referred to as a transform type or a transform matrix, and a non-separable transform kernel for NSPT may be referred to as an RPT kernel. In other words, when r 1 x MN row vectors are selected from the MN x MN forward NSPT kernel, it may be advantageous to select the most important transform basis vectors from a compilation performance perspective. Specifically, in terms of energy compaction through transformation, by multiplying the forward NSPT transform matrix, more energy can be concentrated on the transform coefficients that appear first. In other words, the transform basis vectors located at the top of the forward NSPT transform matrix can generate transform coefficients with greater energy. Considering this, an r x MN forward RPT kernel can be configured (or derived) by taking r from the top of the forward NSPT kernel.

[0200] The RPT according to the present disclosure only takes a part (i.e., r) of the transform coefficients obtained by applying the existing NSPT, and thus the energy of the original signal may be partially lost. In other words, distortion between the original signal and the natural signal may occur through the corresponding process. However, since only r transform coefficients are generated by applying the RPT instead of MN, the amount of bits required to compile the corresponding transform coefficients can be reduced. Therefore, for signals in which a large amount of energy is concentrated on a small number of transform coefficients (e.g., image residual signals), the gain obtained by reducing the signal bits may be significantly large, thus improving the compilation performance.

[0201] The backward NSPT is a transform matrix and can be the transposed matrix of the above-mentioned forward NSPT kernel. In this case, the input data can be a transform coefficient signal instead of a sample signal such as a residual signal. Specifically, when the forward NSPT transform matrix is G and the sample signal rearranged as a 1D vector is x, the transform coefficient vector obtained by multiplying the corresponding transform matrix on the left side can be expressed as in Equation 4 below.

[0202] [Equation 4]

[0203]

[0204] Referring to Equation 4, x and y can be MN x 1 column vectors. G can have the form of an MN x MN matrix. The backward NSPT process can be expressed using the same variables as in Equation 5 below.

[0205] [Equation 5]

[0206]

[0207] In Equation 5, G TMeans the transpose matrix of G. The forward RPT operation and the backward RPT operation according to the present disclosure can also be expressed by these two equations. However, when applying RPT, y is an r x 1 column vector, rather than an MN x 1 column vector, and G is an r x MN matrix, rather than an MM x MN matrix. In other words, even when applying RPT instead of NSPT, the dimension of the sample signal (e.g., the image residual signal) does not change, which may mean that the original number of sample signals (i.e., MN sample signals) can be reconstructed through the backward RPT using only r transform coefficients. In other words, by only compiling r transform coefficients less than MN, the original MN sample signals can be reconstructed, which can improve the compilation performance.

[0208] In an embodiment of the present disclosure, an RPT structure is proposed, which defines the value of r by considering the statistical characteristics of the residual block, and derives the residual block of the existing transform block size from the residual block with a reduced size determined according to the defined value of r. If another additional transform (i.e., the auxiliary transform) is applied to predict the statistical distribution of the main transform coefficients, a quantization process is applied to the main transform coefficients, so the quantized non-zero coefficients may be concentrated in the relatively low frequency domain. Therefore, the reduced auxiliary transform for the statistical distribution of the main transform coefficients can relatively simply define the statistical characteristics of the main transform coefficients in the form of setting the value of r for a given low frequency domain. However, as a technique for defining the value of r by considering the statistical characteristics of the samples within the residual block having characteristics significantly different from the distribution of the main transform coefficients, the RPT according to the present disclosure is fundamentally different from the reduced auxiliary transform. Hereinafter, various embodiments for determining the RPT kernel as the transform matrix with reduced dimensions are described. In other words, a method for determining or defining the value of r in RPT is described below.

[0209] In an embodiment of the present disclosure, the value of r in RPT can be determined by considering the worst-case complexity allowed by the transform system. As an example, the worst-case complexity can be calculated based on the number of multiplications per sample. Based on the M x N block, applying RPT in both the forward and backward directions requires MN * r multiplications. Since the 2D block consists of MN samples in total, the number of multiplications per sample can be calculated as (MN * r) / MN = r. Therefore, the value of r can be configured to be less than or equal to the maximum number of multiplications allowed per sample. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16x16 block, the value of r can be determined to be less than or equal to 16. In other words, the forward RPT kernel can be set to 16 x 256.

[0210] In another embodiment, the memory usage can be regarded as a measure of the worst-case complexity. For example, the memory size allowed for each core can be set. For example, when each core coefficient (in the present disclosure, each element configuring the transform core is referred to as a core coefficient) requires p bytes, and the memory usage is set to be less than or equal to q bytes per core, the value of r can be set to be less than or equal to q / (MN * p). For example, when p is 1 byte for a 16x16 block forward RPT core and the memory usage is set to be less than or equal to 8 KB (q = 8 KB = 2 13 to the power of 13 bytes) per core, the value of r can be set to be less than or equal to 32.

[0211] In addition, as another example, the memory usage and / or the number of multiplications per sample can be regarded as a measure of the worst-case complexity. For example, when the maximum possible number of multiplications per sample for a 16x16 block is set to 16 and the memory usage is set to be less than or equal to 8 KB per core (the core coefficients are expressed as 1 byte), the value of r can be set to be less than or equal to 16.

[0212] In addition, in an embodiment, the r value configuring the RPT core can be determined by specific information. In other words, the r value configuring the RPT core can be determined based on predefined coding parameters. For example, the r value can be determined based on the size of the block. In other words, the RPT core can be variably determined based on the size of the block. Here, the block can be at least one of a compilation block, a transform block, and a prediction block. In addition, for example, the r value can be determined based on prediction information. Here, the prediction information can include information about inter / intra prediction, intra prediction mode information, etc. In addition, for example, the r value can be determined based on the information signaled (the value of the syntax element). For example, the r value can be variably determined according to the quantization parameter value. In addition, in terms of complexity improvement, a predefined fixed value can be used as the r value, and the predefined fixed value can be determined based on the information signaled.

[0213] When multiplying the sample signal by the RPT core r x MN, r transform coefficients can be obtained. The obtained r transform coefficients can be arranged according to a predefined scanning order of the transform coefficients (for example, forward / backward zigzag scanning order, forward / backward horizontal scanning order, forward / backward vertical scanning order, forward / backward diagonal scanning order, scanning order specified based on the intra prediction mode, etc.). When the transform coefficients obtained by applying the forward RPT are arranged according to this scanning order (for example, a scanning order in units of coefficient groups (CG) can also be applied), if the value of r is less than MN, the inside of the M x N block may not be completely filled with r transform coefficients, and thus blank spaces may occur. As an embodiment of the present disclosure, the above blank spaces can be predicted in the following manner by considering the characteristics of the residual signal.

[0214] - The value of the blank space can be filled by using the values of available neighboring pixels.

[0215] - The value of the blank space can be filled based on the values of available neighboring pixels and the intra prediction mode. For example, the value of the blank space can be predicted by performing intra prediction based on the values of available neighboring pixels and the intra prediction mode.

[0216] - The value of the blank space can be filled by using a predefined fixed value (e.g., 0).

[0217] - The value of the blank space can be filled from available neighboring pixels by using a predefined intra prediction mode (e.g., planar mode).

[0218] In the present disclosure, filling the blank space with 0 among the above examples may be referred to as a zero-out process. When filling the blank space with 0, the following embodiments can be applied. When a non-zero transform coefficient is detected (or parsed) in the corresponding blank space portion during the parsing of transform coefficients at the decoding device side, it can be considered (or inferred) that RPT is not applied. In other words, when a non-zero transform coefficient exists in a predefined region representing the corresponding blank space, it can be considered that RPT is not applied. In this case, signaling (or parsing) for a flag indicating whether RPT is applied and / or an index specifying one of multiple RPT kernel candidates may not be performed. As an example, when a non-zero transform coefficient exists in a predefined region representing the corresponding blank space, the value of a predefined variable can be updated, and it can be inferred that RPT is not applied based on the updated variable value.

[0219] In an embodiment of the present disclosure, whether to apply RPT can be determined based on the size and / or form of a block. Additionally, the RPT kernel can be variably determined according to the size and / or form of the block. Since the value of r may be different according to the size and / or form of the block (i.e., for each M x N block), the blank space may be different according to the size and / or form of the block. Therefore, the region for checking whether a non-zero transform coefficient is detected can be defined differently for each size and / or form of the block. In other words, the zero-out region can be variably determined.

[0220] For example, when a 16x64 matrix is applied as the forward RPT matrix for an 8x8 block, the value of r can be 16. In this case, when the CG is a 4x4 sub-block, only the upper left 4x4 block can be filled with non-zero RPT transform coefficients, and the remaining three 4x4 sub-blocks that are blank spaces (i.e., the upper right, lower left, and lower right sub-blocks) can be filled with 0 values. In this case, when non-zero transform coefficients are detected in the corresponding remaining three 4x4 sub-block regions during the decoding process, it can be considered that the RPT is not applied. Also, as described above, the flag indicating whether the RPT is applied or the index specifying one of multiple RPT kernel candidates may not be signaled.

[0221] Additionally, for example, when a 32x128 matrix is applied as the forward RPT matrix for a 16x8 block (i.e., the value of r is 32) and the CG is a 4x4 sub-block, only two CGs can be filled with non-zero RPT transform coefficients in the scan order. For example, the upper left 4x4 sub-block and the 4x4 sub-block adjacent to the bottom of the upper left sub-block can be filled with the corresponding RPT transform coefficients. The regions that are blank spaces filled with 0 can be determined as the remaining regions excluding the corresponding two 4x4 sub-blocks. The RPT kernel can be variably determined according to the size and / or form of the block, and as described above, the blank spaces can be determined differently for 8x8 blocks and 16x8 blocks.

[0222] As an example, when the value of r is a multiple of the CG size and the transform coefficients are scanned in units of CG, if non-zero transform coefficients are detected in the CGs belonging to the blank space, the flag and / or index related to the RPT may not be signaled. In other words, the transform coefficients inside the CG can be scanned in the order specified for each CG, and after moving to the next CG in the scan order for the CG unit, the transform coefficients inside the CG can be scanned in the same way. In the existing image compression technology, since the flag indicating whether non-zero transform coefficients exist in the corresponding CG is signaled for each CG first, it can be determined whether the RPT is applied only using the corresponding information, which can reduce the signaling overhead and related implementation complexity.

[0223] As described above, when the RPT is applied, if non-zero transform coefficients are detected in the blank space region filled with 0, the RPT may not be applied. In this case, the signaling for the information related to the RPT can be omitted. However, since it cannot be determined whether the RPT is applied when non-zero transform coefficients are not detected in the corresponding blank space, the flag indicating whether the RPT is applied can be parsed (or signaled) after parsing (or signaling) the relevant transform coefficients to finally determine whether the RPT is applied.

[0224] As an example, a forward auxiliary transform may be additionally applied to the transform coefficients generated by applying RPT. Alternatively, the forward auxiliary transform may be additionally applied to the region where the corresponding generated transform coefficients in the M x N block are located. In the present disclosure, from the perspective of the forward auxiliary transform, the corresponding region or a part of the corresponding region may be referred to as an ROI. For the backward direction, the backward auxiliary transform may be applied first and then the backward RPT. Specifically, the region where r transform coefficients arranged by applying the forward RPT or a part of the corresponding region may be set as the ROI to apply the forward auxiliary transform. In this case, when a 16x64 forward RPT transform matrix is applied to an 8x8 region, the 16 generated transform coefficients may be located in the upper left 4x4 sub-block, and the corresponding sub-block region may be set as the ROI to apply the forward auxiliary transform to the corresponding ROI.

[0225] In addition, the RPT kernel may adjust the coefficient values by considering operations such as integer operations or fixed-point operations. In other words, the RPT kernel may be configured to perform a transform by integer operations (or fixed-point operations) in an actual encoding / decoding system by appropriately scaling the kernel coefficients belonging to the corresponding kernel, rather than a theoretical orthogonal transform or non-orthogonal transform (here, the orthogonal transform and non-orthogonal transform represent a transform in which the norm of each transform basis vector is 1). Even when applying the RPT as much as the scaling factor multiplied when applying a separable transform in existing image compression techniques, this can be similarly reflected. In this case, a separable transform or a non-separable transform (including RPT) may be performed while retaining other processes except for the transform (e.g., quantization and dequantization processes).

[0226] By multiplying the transform basis vector by the above scaling value, the integerized coefficients of the RPT kernel can be obtained. As an example, multiplying the scaling value may include applying operations such as rounding, floor, or ceiling to each kernel coefficient. In other words, the integerized RPT kernel obtained by the above method may be defined and used for the transform / inverse transform process. As described above, when the kernel coefficients of the scaling integer value are obtained by operations such as rounding, floor, or ceiling, the maximum value and the minimum value of all kernel coefficients can be obtained, and thus the number of bits sufficient to represent all kernel coefficients can be obtained from the maximum value and the minimum value. For example, when the maximum value is less than or equal to 127 and the minimum value is greater than or equal to -128, all integer kernel coefficients can be represented by 8 bits (specifically, represented by two's complement, etc.).

[0227] Generally, when the maximum value is less than or equal to (2 (N-1) - 1) and the minimum value is greater than or equal to -2 (N-1) , all integer kernel coefficients can be represented by N bits. When the maximum value is greater than (2 (N-1) - 1) or the minimum value is less than -2(N-1) When, not all integer kernel coefficients can be expressed with N bits. In such a case, 1) all kernel coefficients can be additionally multiplied by a scaling value to adjust them to fall within the range of N bits; or 2) the number of bits required to express the kernel coefficients can be increased (i.e., N + 1 bits or more). When all kernel coefficients need to be multiplied by 2 -p (p >= 1) to be expressed with N bits, they can be compensated by subsequently multiplying by 2 -p so that they can be incorporated into the existing encoding / decoding process. As an example, multiplying by 2 p can be achieved by additionally performing an operation of shifting left by p bits or reducing the amount of right shift applied in the quantization or dequantization process by p.

[0228] All kernel coefficients can be expressed with 8 bits, 9 bits, 10 bits, etc. by using the above method. And of course, the scaling value of the kernel coefficients can be set differently for each block size or kernel, and the number of bits used to express the kernel coefficients can also be set differently.

[0229] The above NSPT can be applied based on at least one of the size of the current block, the tree type, or the component type. For example, it can be determined whether to apply NSPT based on at least one of the size of the current block, the tree type, or the component type. The NSPT index can be signaled based on at least one of the size of the current block, the tree type, or the component type. The NSPT set or NSPT kernel can be derived based on at least one of the size of the current block, the tree type, or the component type.

[0230] The pre - defined allowed transform block sizes in the decoding device can be roughly divided into two groups. Any one of the two groups (hereinafter referred to as the first group) can mean the set of block sizes to which NSPT is applicable. The first group can consist of any one of the allowed transform block sizes or can consist of two or more of the allowed block sizes. The block sizes to which NSPT is applicable can be defined as the block sizes in which at least one of the width and height is less than or equal to a predetermined threshold. Alternatively, the block sizes to which NSPT is applicable can be defined as the block sizes in which the product of the width and height is less than or equal to a predetermined threshold. Alternatively, the block sizes to which NSPT is applicable can be defined as the block sizes in which the maximum value of the width and height is less than or equal to a predetermined threshold. The threshold can be an integer such as 4, 8, 16, 32, 64, 128 or higher.

[0231] The other of the two groups (hereinafter referred to as the second group) can mean the set of block sizes to which NSPT is not applied. The above - mentioned separable main transform can be applied to the block sizes belonging to the second group. Additionally, the non - separable secondary transform can be applied to all or part of the block sizes belonging to the second group.

[0232] For example, when the size of the current block belongs to the first group, a backward NSPT can be applied to the (dequantized) transform coefficients of the current block. When the size of the current block belongs to the second group, a backward separable main transform can be applied to the (dequantized) transform coefficients of the current block. Alternatively, when the size of the current block belongs to the second group, a backward non-separable secondary transform (e.g., low-frequency non-separable transform, LFNST) can be first applied to the (dequantized) transform coefficients of the current block, and a backward separable main transform (e.g., DCT-2) can be applied to the transform coefficients thus obtained.

[0233] As an example, the first set, as a set of block sizes applicable to NSPT, can be defined as the set of 4x4, 4x8, 8x4, and 8x8. Alternatively, the first set can be defined as the set of 4x8, 8x4, and 8x8. Alternatively, the first set can be defined as the set of 4x8 and 8x4. Alternatively, the first set can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x8, and 16x4. Alternatively, the first set can be defined as the set of 4x8, 4x16, 8x4, 8x8, and 16x4. Alternatively, the first set can be defined as the set of 4x8, 4x16, 8x4, and 16x4. Alternatively, the first set can be defined as the set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, and 16x16. Alternatively, the first set can be defined as the set of 4x4, 4x8, 8x4, 8x8, 8x16, and 16x8. Alternatively, the first set can be defined as the set of 4x8, 8x4, 8x8, 8x16, and 16x8. Alternatively, the first set can be defined as the set of 4x8, 8x4, 8x16, and 16x8. Alternatively, the first set can be defined as the set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, 32x16, and 32x32. Alternatively, the first set can be defined as the set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first set can be defined as the set of 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first set can be defined as the set of 4x8, 8x4, 8x16, 16x8, 16x16, 16x32, and 32x16. Alternatively, the first set can be defined as the set of 4x8, 8x4, 8x16, 16x8, 16x32, and 32x16. Alternatively, the first set can be defined as the set of 4x4, 4x8, 4x16, 8x4, and 16x4. Alternatively, the first set can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, and 16x8. Alternatively, the first set can be defined as the set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, and 16x8. Alternatively, the first set can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, and 16x8. Alternatively, the first set can be defined as the set of 4x8, 4x16, 8x4, 8x16, 16x4, and 16x8. Alternatively, the first set can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first set can be defined as the set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, and 16x16.Alternatively, the first group can be defined as a set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, and 16x16. Alternatively, the first group can be defined as a set of 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, and 16x16.

[0234] As in the example, NSPT can be applied to MxN blocks and NxM blocks that are non-square blocks. For example, NSPT can be applied to 4x8 blocks and 8x4 blocks. Alternatively, NSPT can be applied to 4x16 blocks and 16x4 blocks, or NSPT can be applied to 8x16 blocks and 16x8 blocks, or NSPT can be applied to 16x32 blocks and 32x16 blocks.

[0235] By applying NSPT to a specific block size belonging to the first group, the transformation can be performed more precisely and the compilation performance can be improved. When applying the forward LFNST, the main transformation coefficients in the remaining regions except for the region where the LFNST is applied (i.e., the region of interest, ROI) can be cleared. Additionally, the LFNST can be composed of a small number of transformation basis vectors. In this case, when applying a separable main transformation such as DCT-2 and a non-separable secondary transformation such as LFNST to the corresponding block size instead of NSPT, performance degradation may occur. When applying NSPT instead of LFNST to the corresponding case, the clearing process can be omitted, and the compilation performance can be improved compared to the case of applying LFNST. Additionally, performance improvement can be expected through the method used to apply NSPT. NSPT or LFNST can be applied by using the symmetry described below. Here, for LFNST, a transpose operation is performed on the corresponding input block by using symmetry only for the ROI region. On the other hand, for NSPT, a transpose operation is performed on the entire block by using symmetry. Therefore, for NSPT, symmetry in a more precise manner can be used to train and apply the corresponding NSPT kernel, and thus performance improvement can be expected.

[0236] Additionally, when applying LFNST instead of NSPT to an 8x8 block, from the perspective of the forward transformation, a 32x64 transformation matrix can be applied instead of a 16x64 transformation matrix. Here, the 16x64 transformation matrix can be configured by sampling the first 16 rows of the 32x64 transformation matrix. When applying the LFNST based on the 16x64 transformation matrix to an 8x8 block, each sample requires 16 multiplications to apply the LFNST, but when using the 32x64 transformation matrix, each sample requires 32 multiplications to apply the LFNST. However, when using the 32x64 transformation matrix like this, an improvement in compilation performance can be expected.

[0237] When the tree type of the current block is a single tree, NSPT can be applied to the luminance component of the current block, and NSPT cannot be applied to the chrominance component of the current block. When the tree type of the current block is a double tree, NSPT can be applied to both the luminance component and the chrominance component of the current block.

[0238] Alternatively, regardless of whether the tree type of the current block is a single tree, NSPT can be applied to the luminance component of the current block, and NSPT may not be applied to the chrominance component of the current block. Alternatively, regardless of whether the tree type of the current block is a single tree, NSPT can be applied to both the luminance component and the chrominance component of the current block.

[0239] As an example, when the tree type of the current block is a single tree and NSPT is allowed for both the luminance component and the chrominance component and the size of the current block belongs to the first group, an NSPT index can be signaled and the luminance component and the chrominance component of the current block can share the corresponding NSPT index. Here, the NSPT index can be an index for selecting any one of the transform kernel candidates for NSPT. When the sizes of the luminance block and the chrominance block of the current block belong to the first group, the transform kernel candidate selected by the same NSPT index can be applied to both the luminance component and the chrominance component. When the tree type of the current block is a single tree and NSPT is only applied to the luminance component, LFNST may not be applied and a separate transform can be applied to the chrominance component of the current block. Alternatively, when the tree type of the current block is a single tree and NSPT is only applied to the luminance component, LFNST can be applied to the chrominance component of the current block.

[0240] A single tree may have a high correlation between the luminance component and the chrominance component. In this case, by applying NSPT only to the luminance component or by jointly applying the transform kernel candidate selected by one NSPT index to both the luminance component and the chrominance component, unnecessary signaling can be reduced and the compression efficiency can be improved. On the other hand, for a non-single tree, the luminance component and the chrominance component have independent partitioning and coding structures respectively. In this case, by signaling the NSPT index for each component, the characteristics of each component can be reflected and the compression efficiency can be improved.

[0241] The NSPT kernel for NSPT can be derived based on at least one of the symmetry between intra prediction modes or the symmetry between block shapes. As an example, the NSPT kernel can be derived as an NSPT kernel corresponding to at least one of the following: a mode symmetric to the intra prediction mode of the current block or a block shape symmetric to the block shape of the current block. Alternatively, the NSPT kernel can be derived based on an NSPT set including one or more NSPT kernel candidates, where the NSPT set can be derived as an NSPT set corresponding to at least one of the following: a mode symmetric to the intra prediction mode of the current block or a block shape symmetric to the block shape of the current block. Any one of the one or more NSPT kernel candidates belonging to the NSPT set can be set as the NSPT kernel of the current block. For this purpose, an NSPT index specifying any one of the one or more NSPT kernel candidates belonging to the NSPT set can be used. The NSPT index can be signaled by a bitstream or can be derived based on the above symmetry.

[0242] There can be symmetry between at least two of the intra prediction modes predefined in the decoding device. Hereinafter, for convenience of description, the symmetry around the upper left diagonal mode (i.e., mode 34) is described. Refer to Figure 5 , there is symmetry between the directional modes. Excluding the planar mode numbered 0 and the DC mode numbered 1, all modes have prediction directions. Modes 2 to 66 can be named normal directional modes (which can be expressed as [2, 66]), and modes -14 to -1 (which can be expressed as [-14, -1]) and modes 67 to 80 (which can be expressed as [67, 80]) can be named wide directional modes. The wide directional modes can include at least one of a mode having a value less than -14 or a mode having a value greater than 80. Refer to Figure 5 , all modes excluding mode 0 and mode 1 are symmetric around mode 34. Specifically, mode x and mode (68 - x) are symmetric for mode [2, 66], and mode x and mode (66 - x) are symmetric between mode [-14, -1] and mode [67, 80]. The same symmetric relationship can be established between mode [N, -1] and mode [67, 66 - N]. Here, N can be an integer less than or equal to -14.

[0243] Meanwhile, regarding the symmetry between block shapes, an MxN block and an NxM block can be defined as blocks that are symmetric to each other. Here, M and N can be the same or different. Alternatively, when the aspect ratio (M1 / N1) of an M1xN1 block is the same as the aspect ratio (N2 / M2) of an M2xN2 block, the M1xN1 block and the M2xN2 block can be defined as blocks that are symmetric to each other. Alternatively, when the aspect ratio (M1 / N1) of an M1xN1 block is the same as the aspect ratio (M2 / N2) of an M2xN2 block, the M1xN1 block and the M2xN2 block can be defined as blocks that are symmetric to each other.

[0244] In a square block, symmetric patterns can share at least one of the NSPT set, the NSPT index, or the NSPT core. In other words, at least one of the NSPT set, the NSPT index, or the NSPT core for any one symmetric pattern can be equally applied to another symmetric pattern.

[0245] For example, symmetric patterns can share an NSPT core. However, for any one symmetric pattern, the corresponding NSPT core can be applied to the input data, and for the other symmetric pattern, the corresponding NSPT core can be applied after applying a transpose operation to the input data. Specifically, when pattern x belongs to pattern [2, 33], for pattern x, a 1D vector can be configured for the MxM block as the input data in column-major order, and the NSPT core can be applied to the corresponding 1D vector. Here, configuring the 1D vector in column-major order can read the input data column by column from the MxM block as the input data to obtain M columns and arrange them in order to configure the 1D vector. On the other hand, for the pattern (68 - x) symmetric to pattern x, a 1D vector can be configured in row-major order, and the same corresponding NSPT core can be applied to the corresponding 1D vector. Here, configuring the 1D vector in row-major order can read the input data row by row from the MxM block as the input data to obtain M rows and arrange them in order to configure the 1D vector. When pattern x belongs to pattern [N, -1] (N ≤ -14), for the pattern (66–x) symmetric to pattern x, a 1D vector can be configured in row-major order, and the same NSPT core as pattern x can be applied to the corresponding 1D vector. Column-major order or row-major order can be applied to pattern 0 and pattern 1, and column-major order or row-major order can also be applied to pattern 34. Additionally, row-major order can be applied to the intra prediction pattern belonging to pattern [2, 33], and column-major order can be applied to the pattern symmetric to the corresponding intra prediction pattern. Row-major order can be applied to the intra prediction pattern belonging to pattern [N, -1], and column-major order can be applied to the pattern symmetric to it.

[0246] For non-square blocks, in addition to the symmetry between intra prediction modes, symmetry between block shapes can be further considered. A non-square block with width and height of M and N respectively can be considered to have a symmetric relationship with a non-square block with width and height of N and M respectively. For example, in mode [2, 66], there may be symmetry between mode x of the MxN block and mode (68 - x) of the NxM block. Similarly, when mode x of the MxN block belongs to mode [N, -1] (N ≤ -14), there may be symmetry between mode x of the MxN block and mode (66 - x) of the NxM block.

[0247] The method for configuring a 1D vector from an input data block is as described above. In other words, when column-major order is applied to mode x, row-major order can be applied to the symmetric mode. Alternatively, when row-major order is applied to mode x, column-major order can be applied to the symmetric mode. Specifically, when column-major order is applied to mode x, M columns can be obtained by reading the input data column by column from the MxN block as the input data, and they can be arranged in order to configure a 1D vector. Here, the length of each column can be N. For the mode symmetric to mode x, N rows can be obtained by reading the input data row by row from the MxN block as the input data, and they can be arranged in order to configure a 1D vector. Here, the length of each row can be M. Alternatively, when row-major order is applied to mode x, N rows can be obtained by reading the input data row by row from the MxN block as the input data, and they can be arranged in order to configure a 1D vector. Here, the length of each row can be M. For the mode symmetric to mode x, M columns can be obtained by reading the input data column by column from the MxN block as the input data, and they can be arranged in order to configure a 1D vector. Here, the length of each column can be N.

[0248] When the current block is an MxN block with mode x and the above symmetry is used for the current block, the NSPT set and / or NSPT kernel of the current block can be determined based on at least one of the intra prediction mode symmetric to mode x or the block size of NxM symmetric to the block size of MxN. Here, the NSPT kernel can be set to the NSPT kernel for the NxM block instead of the NSPT kernel for the MxN block. In other words, when symmetry is used for the current block, the NSPT set and / or NSPT kernel of the block having symmetry with the current block can be used in the same way. As described above, a 1D vector can be configured from the input data block according to a predetermined priority, which can correspond to the input of the NSPT kernel.

[0249] In addition, there may be a restriction that symmetry is used only when the value of the intra prediction mode of the current block is greater than 34. In other words, when the value of the intra prediction mode of the current block is greater than 34, a transpose operation can be applied when configuring a 1D vector from the input data block, and an NSPT set or NSPT kernel corresponding to the block shape and / or mode symmetric to the current block can be used. Specifically, when the intra prediction mode of the current block belongs to mode [N, -1] and mode [2, 34], symmetry may not be used for the current block. On the other hand, when the intra prediction mode of the current block belongs to mode [35, 66] and mode [67, 66 - N], symmetry can be used for the current block. Here, N can be an integer less than or equal to -14.

[0250] Derivation of the symmetry-based NSPT set or NSPT kernel can be adaptively performed based on the size of the current block. As an example, for 4x4 blocks and 8x8 blocks, the NSPT set or NSPT kernel can be derived based on symmetry; and for 4x8 blocks and 8x4 blocks, the NSPT set or NSPT kernel can be derived without using symmetry.

[0251] Depending on whether symmetry is used, the number of available NSPT sets may be different. As an example, when symmetry is used, the number of available NSPT sets can be 35; and when symmetry is not used, the number of available NSPT sets can be 67.

[0252] Table 9 below relates to an example of determining the NSPT set by using symmetry, and shows the mapping relationship between the NSPT set and the intra prediction mode when the number of available NSPT sets is 35.

[0253] [Table 9]

[0254]

[0255] Referring to Table 9, when the value of the intra prediction mode of the current block (X) is less than 0, the NSPT set of the current block can be determined as the NSPT set among the 35 NSPT sets with an NSPT set index of 2. When the value of the intra prediction mode of the current block (X) is greater than or equal to 0 and less than or equal to 34, the NSPT set of the current block can be determined as the NSPT set among the 35 NSPT sets with an NSPT set index of X. When the value of the intra prediction mode of the current block (X) is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set among the 35 NSPT sets with an NSPT set index of (68 - X). When the value of the intra prediction mode of the current block (X) is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block can be the same as the NSPT set corresponding to the value of the symmetric mode of the intra prediction mode of the current block (68 - X). Similarly, when the value of the intra prediction mode of the current block (X) is greater than 66, the NSPT set of the current block can be determined as the NSPT set among the 35 NSPT sets with an NSPT set index of 2. When the value of the intra prediction mode of the current block (X) is greater than 66, the NSPT set of the current block can be the same as the NSPT set corresponding to the symmetric mode of the intra prediction mode of the current block.

[0256] Table 10 below relates to an example of determining the NSPT set without using symmetry and shows the mapping relationship between the NSPT set and the intra prediction mode when the number of available NSPT sets is 67.

[0257] [Table 10]

[0258]

[0259] Referring to Table 10, when the value of the intra prediction mode of the current block (X) is less than 0, the NSPT set of the current block can be determined as the NSPT set among the 67 NSPT sets with an NSPT set index of 2. When the value of the intra prediction mode of the current block (X) is greater than or equal to 0 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set among the 67 NSPT sets with an NSPT set index of X. Similarly, when the value of the intra prediction mode of the current block (X) is greater than 66, the NSPT set of the current block can be determined as the NSPT set among the 67 NSPT sets with an NSPT set index of 66.

[0260] Symmetry can be utilized to save the memory size required to store the transform kernel while maintaining performance according to the application of the transform. For example, when using 35 NSPT sets instead of 67 NSPT sets by utilizing symmetry, the memory size required to store the NSPT kernel can be significantly reduced.

[0261] The number of available NSPT sets and / or the number of NSPT core candidates belonging to an NSPT set can vary depending on the block size. For example, the number of available NSPT sets for a 4x4 block can be 35, the number of available NSPT sets for 4x8 and 8x4 blocks can be 19, and the number of available NSPT sets for an 8x8 block can be 10. The NSPT set for a 4x4 block can consist of three NSPT core candidates, the NSPT set for 4x8 and 8x4 blocks can consist of three or two NSPT core candidates, and the NSPT set for an 8x8 block can consist of one NSPT core candidate.

[0262] The size of the transform kernel may increase as the block size increases. Therefore, the number of available NSPT sets and / or the number of NSPT core candidates belonging to an NSPT set can be reduced to save the memory size required to store the transform kernel. Additionally, as the block size increases, the characteristics of the residual signal within the corresponding block tend to become more generalized. Therefore, reducing the number of available NSPT sets and / or the number of NSPT core candidates belonging to an NSPT set can help maintain compression efficiency while reducing implementation complexity by reflecting these statistical characteristics.

[0263] The NSPT core can be configured with 8-bit precision. The range of coefficients within the NSPT core can be greater than or equal to -128 and less than or equal to 127. When the corresponding precision is increased to more than 8 bits, the resulting value obtained through matrix multiplication can be right-shifted by the increased precision. For example, when the value obtained after matrix multiplication based on an NSPT core with 8-bit precision is right-shifted by S bits and stored in a buffer, if the core coefficients are configured with N-bit precision, it can be right-shifted by (S + (N - 8)) bits and stored in the buffer.

[0264] When the NSPT core is configured with 8-bit precision, it is possible to prevent an excessive increase in internal precision in the encoder / decoder that performs the transform, thereby reducing implementation complexity in terms of memory requirements and computational load, and at the same time minimizing the reduction in compression efficiency.

[0265] When applying the backward NSPT to a current block of size NxN, the size of the NSPT core (or NSPT matrix) can be expressed as MN x r. Here, MN can mean the product of the width and height of the current block. This can mean the output length of the NSPT or the number of residual samples generated by the NSPT. Additionally, r can mean the input length of the NSPT or the number of (dequantized) transform coefficients to which the NSPT is applied. r can be an integer greater than or equal to 0 and less than or equal to MN. The following are examples of NSPT matrices of MN x r according to the block size.

[0266] The NSPT matrix for a 4x4 block can be composed of a 16x16 matrix. The NSPT matrix for a 4x8 block and an 8x4 block can be composed of a 32x20 matrix, a 32x16 matrix, a 32x24 matrix, a 32x28 matrix, or a 32x32 matrix. The NSPT matrix for an 8x8 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. The NSPT matrix for a 4x16 block and a 16x4 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. The NSPT matrix for an 8x16 block and a 16x8 block can be composed of a 128x96 matrix, a 128x64 matrix, a 128x48 matrix, or a 128x32 matrix. The NSPT matrix for a 16x16 block can be composed of a 256x128 matrix, a 256x96 matrix, or a 256x64 matrix. The NSPT matrix for a 16x32 block and a 32x16 block can be composed of a 512x256 matrix or a 512x128 matrix. The NSPT matrix for a 32x32 block can be composed of a 1024x512 matrix, a 1024x256 matrix, or a 1024x128 matrix. Alternatively, a 16x16 matrix can be applied to 4xN blocks and Nx4 blocks. Here, N can be an integer greater than or equal to 4. A 64x16 matrix can be applied to an 8x8 block. A 64x32 matrix can be applied to 8xN blocks and Nx8 blocks. Here, N can be an integer greater than or equal to 16. A 96x32 matrix can be applied to 16xN blocks and Nx16 blocks. Here, N can be an integer greater than or equal to 16.

[0267] Table 11 is an example of an NSPT kernel that can be applied to an 8x16 block. The coefficients configuring the NSPT kernel are represented with 8-bit precision. The coefficients of the NSPT kernel can have values ranging from -128 to 127. In this example, the total number of NSPT sets predefined equally in the encoding device and the decoding device is 35, and each NSPT set can be composed of three NSPT kernel candidates.

[0268] In this example, g_nspt8x16

[35] [3]

[40]

[128] can represent the NSPT kernel that can be applied to an 8x16 block. However, the NSPT sets and / or NSPT kernels can be determined considering the above symmetries. When the current block is an 8x16 block and has a specific intra prediction mode (e.g., a mode with vertical directivity), the NSPT sets and / or NSPT kernels for the block that has symmetry with the current block (i.e., the 16x8 block) can be applied to the current block in the same way. In this case, the input data from the current block is transposed and then multiplied by the corresponding NSPT kernel.

[0269] In the following g_nspt8x16

[35] [3]

[40]

[128] ,

[35] can indicate that it consists of 35 NSPT sets, and [3] can indicate that each NSPT set is composed of 3 NSPT core candidates. For example, in Table 11, / k can refer to the NSPT set with the set index k, where the range of k can be from 0 to 34. The three matrices belonging to / k can respectively represent the NSPT core candidates with the transform indices 0, 1, and 2. Additionally,

[40]

[128] can represent a 40x128 matrix for the forward NSPT matrix. Specifically, the first

[40] can represent 40 transform basis vectors, and the second

[128] can represent that each transform basis vector is a one-dimensional vector composed of 128 coefficients. From the perspective of the forward NSPT, this can indicate that the input to the forward NSPT is a one-dimensional vector with a length of 128. The backward NSPT matrix can be derived by transposing the NSPT cores selected from the following g_nspt8x16 array. In this case, the backward NSPT matrix becomes a 128x40 matrix.

[0270] [Table 11]

[0271]

[0272]

[0273]

[0274]

[0275]

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289]

[0290] The transform coefficients can be derived by applying a forward NSPT transform to the residual samples of an MxN block. In this case, due to zeroing, the number of the derived transform coefficients may be less than or equal to the value of (M*N). In other words, the forward NSPT matrix can be defined as an r x (M*N) matrix, where r can mean the output length of the NSPT or the number of transform coefficients derived by the NSPT, and (M*N) can mean the input length of the NSPT or the number of residual samples to which the NSPT is applied.

[0291] The derived transform coefficients can be arranged in the MxN block according to a predetermined scan order, and the regions where the transform coefficients are not filled can be filled with 0 (i.e., zeroed). Therefore, during the process of the decoding device scanning the transform coefficients, if a non-zero transform coefficient is found in the region that has been filled with 0 after the NSPT has been applied (or, when the scan position of the last valid coefficient in the MxN block is greater than or equal to r), it is considered that the NSPT has not been applied to the corresponding MxN block, and the NSPT index may not be signaled.

[0292] The transform kernel of the current block can be determined based on any one of the above-mentioned Embodiments 1 to 3. Alternatively, within the scope where the inventions according to the above-mentioned Embodiments 1 to 3 do not conflict with each other, the transform kernel of the current block can be determined based on a combination of at least two of Embodiments 1 to 3.

[0293] Reference Figure 4 , the current block can be reconstructed based on the residual samples of the current block (S420).

[0294] The prediction samples of the current block can be derived based on the intra prediction mode of the current block. The reconstructed samples of the current block can be generated based on the prediction samples and the residual samples of the current block.

[0295] Figure 6 FIG. illustrates a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure.

[0296] Reference Figure 6 , the decoding device (300) according to the present disclosure can include a transform coefficient exporter (600), a residual sample exporter (610), and a reconstructed block generator (620). The transform coefficient exporter (600) can be configured to be inFigure 3 In the entropy decoder (310), the residual sample exporter (610) may be configured in Figure 3 the residual processor (320), and the reconstruction block generator (620) may be configured in Figure 3 the adder (340).

[0297] The transform coefficient exporter (600) may obtain the residual information of the current block from the bitstream and decode it to export the transform coefficients of the current block.

[0298] The residual sample exporter (610) may export the residual samples of the current block by performing at least one of dequantization or inverse transformation on the transform coefficients of the current block.

[0299] The residual sample exporter (610) may determine the transform kernel for the inverse transformation of the current block by a predetermined transform kernel determination method and export the residual samples of the current block based on this. It is the same as that described by referring to Figure 4 and its detailed description will be omitted here.

[0300] The reconstruction block generator (620) may reconstruct the current block based on the residual samples of the current block.

[0301] Figure 7 The figure illustrates an image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure.

[0302] Referring to Figure 7 , the residual samples of the current block (S700) may be exported.

[0303] The residual samples of the current block may be exported by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples may be exported based on a predetermined intra prediction mode.

[0304] Referring to Figure 7 , the transform coefficients of the current block (S710) may be exported by performing at least one of transformation or quantization on the residual samples of the current block.

[0305] The transform method according to the present disclosure may be understood as the inverse process of the inverse transformation described by referring to Figure 4 . The method for determining the transform kernel for transformation is the same as that described by referring to Figure 4 and its detailed description will be omitted here.

[0306] For example, one or more transform sets for the transform of the current block can be defined / configured, and each transform set can include one or more transform kernel candidates. In this case, one of the multiple transform sets can be selected as the transform set for the current block. One of the multiple transform kernel candidates belonging to the transform set of the current block can be selected. The selection can be implicitly performed based on the context of the current block. Alternatively, the best transform set and / or transform kernel candidate for the current block can be selected, and its index can be signaled.

[0307] Alternatively, the transform kernel of the current block can be determined based on the MTS set. One of the multiple MTS sets can be selected based on at least one of the size of the current block or the intra prediction mode. The selected MTS set can include one or more transform kernel candidates. One of the one or more transform kernel candidates can be selected, and the transform kernel of the current block can be determined based on the selected transform kernel candidate. The selection of the transform kernel candidate can be performed using the transform kernel candidate index derived from the context of the current block. Alternatively, the best transform kernel candidate for the current block can be selected, and the transform kernel candidate index indicating the selected transform kernel candidate can be signaled.

[0308] Alternatively, the transform kernel of the current block can be determined based on the non-separable primary transform (NSPT) kernel. When the size of the current block belongs to the first group of block sizes applicable to NSPT, forward NSPT can be applied to the current block; and when the size of the current block belongs to the second group, forward NSPT is not applied to the current block. When the size of the current block belongs to the second group, forward separable primary transform (e.g., DCT-2) can be applied to the residual samples of the current block to derive transform coefficients. Forward LFNST can be additionally applied to all or part of the transform coefficients derived by the separable primary transform.

[0309] In addition, NSPT can be applied based on at least one of the tree type or component type of the current block. The NSPT kernel (or NSPT matrix) for NSPT can be determined by using the symmetry between intra prediction modes or the symmetry between block shapes. When forward NSPT is applied to the current block of MxN, the NSPT kernel can be expressed as r x MN. Here, r means the output length of NSPT or the number of transform coefficients generated by NSPT, and MN is the product of the width and height of the current block, which can mean the input length of NSPT or the number of residual samples to which NSPT is applied. The method for determining the size of the NSPT kernel is the same as that described by referring to Figure 4 the same as described.

[0310] Referring to Figure 7 , a bitstream (S720) can be generated by encoding the transform coefficients of the current block.

[0311] Residual information regarding transform coefficients may be generated based on the transform coefficients of a current block, and a bitstream may be generated by encoding the residual information.

[0312] Figure 8 FIG. shows a schematic configuration of an encoding apparatus (200) that executes an image encoding method according to the present disclosure.

[0313] Reference Figure 8 , according to the present disclosure, the encoding apparatus (200) may include a residual sample extractor (800), a transform coefficient extractor (810), and a transform coefficient encoder (820). The residual sample extractor (800) and the transform coefficient extractor (810) may be configured in Figure 2 the residual processor (230), and the transform coefficient encoder (820) may be configured in Figure 2 the entropy encoder (240).

[0314] The residual sample extractor (800) may extract the residual samples of a current block by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples may be extracted based on a predetermined intra prediction mode.

[0315] The transform coefficient extractor (810) may extract the transform coefficients of a current block by performing at least one of transformation or quantization on the residual samples of the current block. The transform coefficient extractor 810 may determine the transform kernel of the current block based on at least one of the above-described Embodiments 1 to 3, and apply the transform kernel to the residual samples of the current block to extract the transform coefficients.

[0316] The transform coefficient encoder (820) may encode the transform coefficients of a current block to generate a bitstream.

[0317] In the above embodiments, the method is described as a series of steps or blocks based on a flowchart, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in an order different from other steps as described above. In addition, those skilled in the art may understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.

[0318] The above method according to an embodiment of the present disclosure can be implemented in the form of software, and the encoding apparatus and / or decoding apparatus according to the present disclosure may be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0319] In the present disclosure, when an embodiment is implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuitry, and / or a data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments described herein can be executed by being implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be executed by being implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0320] In addition, a decoding device and an encoding device applying the embodiments of the present disclosure can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a device for providing a video-on-demand (VoD) service, an over-the-top (OTT) device, a device for providing an Internet streaming service, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a videophone video device, a transportation tool terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process video signals or data signals. For example, an over-the-top (OTT) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet, a digital video recorder (DVR), and so on.

[0321] In addition, the processing method according to an embodiment of the present disclosure can be generated in the form of a program executable by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium may include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical media storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmitted via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network.

[0322] In addition, an embodiment of the present disclosure can be implemented by a computer program product through program code, and the program code can be executed by a computer according to an embodiment of the present disclosure. The program code can be stored on a computer-readable carrier.

[0323] Figure 9 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown.

[0324] Reference Figure 9 , a content streaming system to which an embodiment of the present disclosure is applied may mainly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0325] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a video camera, etc. into digital data and sends it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0326] A bitstream can be generated by applying an encoding method or a bitstream generation method according to an embodiment of the present disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0327] The streaming server sends multimedia data to the user device via the web server based on the user's request, and the web server serves as a medium for notifying the user of what services are available. When the user requests a required service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the commands / responses between each device in the content streaming system.

[0328] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming media service, the streaming server may store the bitstream for a specific period of time.

[0329] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0330] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received from each server may be distributed and processed.

[0331] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. Additionally, the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a method.

Claims

1. An image decoding method, comprising: Obtaining residual information from a bitstream; Deriving transform coefficients of a current block based on the residual information; Applying a backward non-separable principal transform (NSPT) to at least one of the transform coefficients of the current block to derive residual samples of the current block; And Reconstructing the current block based on the residual samples of the current block, wherein the backward NSPT is applied based on the size of the current block belonging to a group of one or more block sizes to which the NSPT is applicable.

2. The image decoding method according to claim 1, wherein, The NSPT set for the NSPT is determined to be any one of 35 predefined NSPT sets.

3. The image decoding method according to claim 2, wherein, Each of the 35 NSPT sets includes three NSPT kernel candidates.

4. The image decoding method according to claim 1, wherein, The group includes at least one of 4x4, 4x8, 8x4, 8x8, 16x8, or 8x16.

5. The image decoding method according to claim 1, wherein, Based on the size of the current block being 8x16, the number of transform coefficients to which the backward NSPT is applied is 40.

6. An image encoding method, comprising: Deriving residual samples of a current block; Applying a non-separable principal transform (NSPT) to the residual samples of the current block to derive transform coefficients of the current block; Generating residual information about the transform coefficients of the current block; And Encoding the residual information to generate a bitstream, wherein the NSPT is applied based on the size of the current block belonging to a group of one or more block sizes to which the NSPT is applicable.

7. The image encoding method according to claim 6, wherein, The NSPT set for the NSPT is determined to be any one of 35 predefined NSPT sets.

8. The image encoding method according to claim 7, wherein, Each of the 35 NSPT sets includes three NSPT kernel candidates.

9. The image encoding method according to claim 6, wherein, The group includes at least one of 4x4, 4x8, 8x4, 8x8, 16x8, or 8x16.

10. The image encoding method according to claim 6, wherein, Based on the size of the current block being 8x16, the number of transform coefficients derived by the NSPT is 40.

11. A computer-readable storage medium storing a bitstream generated by the image encoding method according to claim 6.

12. A method for transmitting data, comprising: Obtaining a bitstream for image information, wherein the bitstream is generated by: deriving residual samples of a current block, applying a non-separable principal transform (NSPT) to the residual samples of the current block to derive transform coefficients of the current block, generating residual information related to the transform coefficients of the current block, and encoding the residual information; and Transmitting the data including the bitstream, wherein the NSPT is applied based on the size of the current block belonging to a group of one or more block sizes to which the NSPT is applicable.