Image encoding / decoding method and device, and recording medium storing bitstream
By employing non-separable transformations and adaptively selecting reduced-dimensional transform kernels based on encoding parameters, the method addresses the challenges of compressing high-resolution images, achieving improved encoding efficiency and quality.
Patent Information
- Application Number
- PCT/KR2024/019492
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-05
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently compressing and transmitting high-resolution, high-quality images such as HD and UHD images, due to limitations in transformation methods and kernel selection.
The proposed solution involves using a non-separable first-order transformation and a reduced-dimensional non-separable transform kernel, which is determined and signaled based on encoding parameters. This method allows for the selection of appropriate transform sets for inter and intra blocks, improving the efficiency of transformation and encoding.
The use of non-separable transformations and adaptive kernel selection enhances the performance of video encoding and decoding by improving compression efficiency and quality, particularly for high-resolution images.
Smart Images

Figure KR2024019492_05062025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream. Recently, the demand for high-resolution, high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed. There are various technologies such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information in the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, etc., and using these image compression technologies, image data can be effectively compressed and transmitted or stored. The present disclosure seeks to provide a method and apparatus for performing a transformation using a non-separable first-order transformation. The present disclosure seeks to provide a method and apparatus for performing a transform using a reduced-dimensional non-separable first-order transform kernel. The present disclosure provides a method and device for determining / signaling a non-separable transform kernel based on encoding parameters. The video decoding method and device according to the present disclosure can derive transform coefficients of a current block from a bitstream, derive residual samples of the current block based on an inverse transform for the transform coefficients of the current block, and reconstruct the current block based on the residual samples of the current block. Here, the inverse transform can be performed based on a non-separable transform. A transform set for the non-separable transform can be selected based on a predetermined intra prediction mode for the current block. In the image decoding method and device according to the present disclosure, the number of available transform kernel candidates in the transform set can be determined based on whether the current block is an inter block. In the video decoding method and device according to the present disclosure, when the current block is an inter block, the number of available transform kernel candidates in the transform set can be determined based on information signaled through a bitstream. In the image decoding method and device according to the present disclosure, the number of available transform kernel candidates in the transform set can be determined based on information about the transform coefficients. In the video decoding method and device according to the present disclosure, the number of available transform kernel candidates in the transform set can be determined based on at least one of a flag for at least one of a luma component or a chroma component of the current block or a position of a last valid transform coefficient in the current block. The flag can indicate whether at least one non-zero transform coefficient exists. In the image decoding method and device according to the present disclosure, the number of transform coefficients input to the non-separable transform can be determined based on whether the current block is an inter block. In the image decoding method and device according to the present disclosure, the transformation set can be selected as any one of pre-defined transformation sets. In the video decoding method and device according to the present disclosure, transform sets for inter blocks and transform sets for intra blocks can be defined respectively. In the video decoding method and device according to the present disclosure, the pre-defined transformation sets can be equally applied to inter blocks and intra blocks. In the video decoding method and device according to the present disclosure, the intra prediction mode for the current block can be converted to an extended intra prediction mode based on whether the current block is a non-square block. The video encoding method and device according to the present disclosure can derive residual samples of a current block, derive transform coefficients of the current block based on a transform for the residual samples of the current block, and encode the transform coefficients of the current block. Here, the transform can be performed based on a non-separable transform. A transform set for the non-separable transform can be selected based on a predetermined intra prediction mode for the current block. A computer-readable digital storage medium is provided having encoded video / image information stored thereon, which causes a decoding device according to the present disclosure to perform a video decoding method. A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided. A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided. The present disclosure can improve the performance of transformation by using a non-separable primary transformation as a primary transformation. The present disclosure can improve the performance of transformation by performing the transformation using a non-separable first-order transformation kernel of reduced dimension. The present disclosure can improve encoding efficiency by effectively determining and / or signaling a non-separable transform kernel based on encoding parameters. FIG. 1 illustrates a video / image coding system according to the present disclosure. FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed. FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed. FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure. FIG. 5 exemplarily illustrates an intra prediction mode and its prediction direction according to the present disclosure. FIG. 6 is a flowchart illustrating a method for deriving a DIMD mode according to the present disclosure. Fig. 7 illustrates a filter for inducing a DIMD mode according to the present disclosure. FIG. 8 illustrates a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure. FIG. 9 illustrates an image encoding method performed by an encoding device (200) as an embodiment according to the present disclosure. FIG. 10 illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure. FIG. 11 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied. The present disclosure may have various modifications and various embodiments, and thus specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to specific embodiments, but should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present disclosure. In describing each drawing, similar reference numerals are used for similar components. The terms first, second, etc. may be used to describe various components, but the components should not be limited by the terms. The terms are only used to distinguish one component from another. For example, without departing from the scope of the present disclosure, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term and / or includes any combination of a plurality of related described items or any item among a plurality of related described items. When it is said that a component is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. On the other hand, when it is said that a component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprises" or "has" and the like are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof. The present disclosure relates to video / image coding. For example, the method / embodiment disclosed in this specification can be applied to a method disclosed in the versatile video coding (VVC) standard. In addition, the method / embodiment disclosed in this specification can be applied to a method disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation of audio video coding standard (AVS2) or the next generation video / image coding standard (e.g., H.267 or H.268, etc.). This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other. In this specification, video may mean a collection of images over time. A picture generally means a unit representing one image at a specific time, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of a picture parameter set. A tile row is a rectangular area of CTUs having a height equal to the width of the picture and a width equal to the width of the picture specified by a picture parameter set. While CTUs within a tile may be arranged sequentially according to a CTU raster scan, tiles within a picture may be arranged sequentially according to a tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture that may be exclusively contained in a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture. A pixel, or pel, can mean the smallest unit that constitutes a picture (or image). Also, a 'sample' can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (ex. cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows. As used herein, “A or B” can mean “only A”, “only B”, or “both A and B”. In other words, as used herein, “A or B” can be interpreted as “A and / or B”. For example, as used herein, “A, B or C” can mean “only A”, “only B”, “only C”, or “any combination of A, B and C”. As used herein, a slash ( / ) or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Accordingly, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C". As used herein, “at least one of A and B” can mean “only A”, “only B” or “both A and B”. Additionally, as used herein, the expressions “at least one of A or B” or “at least one of A and / or B” can be interpreted identically to “at least one of A and B”. Additionally, in this specification, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Additionally, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C". In addition, parentheses used in this specification may mean "for example". Specifically, when it is indicated as "prediction (intra prediction)", "intra prediction" may be suggested as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be suggested as an example of "prediction". In addition, even when it is indicated as "prediction (i.e., intra prediction)", "intra prediction" may be suggested as an example of "prediction". Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously. FIG. 1 illustrates a video / image coding system according to the present disclosure. Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component. The video source can obtain the video / image through a process of capturing, compositing, or generating the video / image. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device can include a computer, a tablet, a smart phone, etc., and can (electronically) generate the video / image. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated. The encoding device can encode input video / image. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream. The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the reception unit of the receiving device through a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcasting / communication network. The reception unit can receive / extract the bitstream and transmit it to a decoding device. The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device. The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit. FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed. Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The adding unit (250) may be called a reconstructor or a reconstructed block generator. The image segmenting unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the adding unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component. The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a coding unit may be split into a plurality of coding units having deeper depths based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be used as the final coding unit, or, if necessary, the coding unit may be split recursively into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient. The term unit may be used interchangeably with terms such as block or area, depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image). The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input image signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied to the current block or CU unit. The prediction unit (220) can generate various information about prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information about prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream. The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located a certain distance away from the current block depending on the prediction mode. In the intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of the DC mode or the planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail of the prediction direction. However, this is only an example, and a number of directional modes greater or less than that may be used depending on the setting. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block. The inter prediction unit (221) can derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the above temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of the skip mode and the merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of one block, and can also apply intra prediction and inter prediction at the same time. This can be called a combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT means a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size. The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately in addition to quantized transform coefficients. Encoded information (e.g. encoded video / image information) can be transmitted or stored in the form of a bitstream as a unit of a network abstraction layer (NAL). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by the transmitting unit (not shown) and / or the storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240). The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be restored. The adding unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the target block to be processed, such as when the skip mode is applied, the predicted block can be used as a reconstructed block. The adding unit (250) can be called a reconstructed unit or a reconstructed block generating unit. The generated restoration signal can be used for intra prediction of the next processing target block in the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during the picture encoding and / or restoration process. The filtering unit (260) can apply filtering to the restoration signal to improve subjective / objective picture quality. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, a sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream. The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency. The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit (222). FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed. Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter prediction unit (331) and an intra prediction unit (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), adding unit (340), and filtering unit (350) described above may be configured by one hardware component (e.g., a decoding device chipset or processor) according to an embodiment. In addition, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component. When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device. The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model by using information on syntax elements to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in a previous step, and predicts an occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after the context model is determined. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), that is, quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310). Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the adding unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331). The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients. In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array). The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra prediction or inter prediction is applied to the current block based on the information about the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode. The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of one block, and can also apply intra prediction and inter prediction at the same time. This can be called a combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be signaled and included in the video / image information. The intra prediction unit (331) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In the intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks. The inter prediction unit (332) can derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit (332) can configure a motion information candidate list based on neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block. The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). When there is no residual for the target block to be processed, such as when the skip mode is applied, the prediction block can be used as the restoration block. The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block in the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process. The filtering unit (350) can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit (350) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and transmit the modified restoration picture to the memory (360), specifically, the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331). In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) may be applied identically or correspondingly to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300), respectively. FIG. 4 illustrates an image decoding method performed by a decoding device according to an embodiment of the present disclosure. Referring to FIG. 4, transform coefficients of the current block can be derived from the bitstream (S400). That is, the bitstream can include residual information of the current block, and the transform coefficients of the current block can be derived by decoding the residual information. Referring to FIG. 4, residual samples of the current block can be derived by performing at least one of dequantization and inverse transform on the transform coefficients of the current block (S410). When Adaptive Multiple Transform Selection (MTS) is applied, the inverse transform can be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc. can be called a transform type, a transform kernel, or a transform core. In the present disclosure, the inverse transform may mean a separable transform. However, it is not limited thereto, and the inverse transform may mean a non-separable transform, or may be a concept including a separable transform and a non-separable transform. In addition, the inverse transform in the present disclosure means a primary transform, but is not limited thereto, and may be applied to a secondary transform by being transformed into an identical / similar form. For example, as a method for inverse transformation, only DCT-2 and a non-separable transform may be used, or a non-separable transform may be used in addition to at least one of DCT-2, DST-7, or DCT-8, or a non-separable transform may replace the transform kernel of one or more of DCT-2, DST-7, or DCT-8. As a more specific example, if there are (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), (DCT-8, DCT-8) as transform kernel candidates for a separable transform, a non-separable transform can replace or be added to one or more of the five transform kernel candidates. Here, the notation (transform1, transform2) indicates that transform1 is applied in the horizontal direction and transform2 is applied in the vertical direction. If the non-separable transform replaces some of the transform kernel candidates, the remaining transform kernel candidates except (DCT-2, DCT-2) and (DST-7, DST-7) can be replaced with the non-separable transform. However, the transform kernel candidates are only examples, and other types of DCT and / or DST may be included, and a transform skip may be included as the transform kernel candidate. A non-separable transform can mean a transform or inverse transform based on a non-separable transform matrix. That is, unlike a separable transform that performs horizontal and vertical transforms independently by separating vertical and horizontal transforms, a non-separable transform can perform horizontal and vertical transforms at once. For example, when a non-separable transformation is performed on a 4x4 block, the input data X to the non-separable transformation is as shown in the following mathematical expression 1. When the above input data X is expressed in vector form, vector X' can be expressed as follows. In this case, the non-separable transformation can be performed as in the following mathematical expression 3. In mathematical expression 3, F represents a transformation coefficient vector, T represents a 16x16 non-separable transformation matrix, and ㆍ represents the multiplication of a matrix and a vector. A 16x1 transform coefficient vector F can be derived through the above mathematical expression 3, and the F can be reconstructed into 4x4 blocks according to a predetermined scan order. The scan order can be a horizontal scan, a vertical scan, a diagonal scan, a z-scan, a raster scan, or a pre-defined scan. The non-separable transform set and / or transform kernel for the above non-separable transform can be variously configured based on at least one of a prediction mode (e.g., intra mode, inter mode, etc.), the width, height, or number of pixels of the current block, the position of a sub-block within the current block, explicitly signaled syntax elements, statistical characteristics of surrounding samples, whether a second transform is used, or a quantization parameter (QP). Specifically, for the intra mode, the pre-defined intra prediction modes are grouped to correspond to n non-separable transform sets, and each non-separable transform set may include k transform kernel candidates, where n and k may be arbitrary constants according to rules (conditions) defined identically for the encoding device and the decoding device. The number of non-separable transformation sets and / or the number of transformation kernel candidates included in the non-separable transformation sets can be configured differently depending on the width and / or height of the current block. For example, for a 4x4 block, n 1 A set of non-separable transformations and k 1 A number of transformation kernel candidates can be constructed. For a 4x8 block, n 2 A set of non-separable transformations and k 2 A number of transformation kernel candidates can be configured. In addition, the number of non-separable transformation sets and the number of transformation kernel candidates included in each non-separable transformation set can be configured differently depending on the product of the width and height of the current block. For example, if the product of the width and height of the current block is 256 or more (or exceeds), n 3 A set of non-separable transformations and k 3 A candidate transformation kernel can be constructed, otherwise, n 4 A set of non-separable transformations and k 4 A number of transform kernel candidates can be configured. That is, since the degree of change in the statistical characteristics of the residual signal is different depending on the block size, the number of non-separable transform sets and transform kernel candidates can be configured differently to reflect this. If the current block is divided into multiple sub-blocks, the statistical characteristics of the residual signal may be different for each sub-block, so the number of non-separable transform sets and transform kernel candidates can be configured differently. For example, if a 4x8 or 8x4 block is divided into two 4x4 sub-blocks and a non-separable transform is applied to each sub-block, n for the upper left 4x4 sub-block5 A set of non-separable transformations and k 5 A number of transform kernel candidates can be constructed, n for different 4x4 sub-blocks. 6 A set of non-separable transformations and k 6 A set of transformation kernel candidates can be constructed. Based on the explicitly signaled syntax elements, the number of non-separable transformation sets and transformation kernel candidates can be differently configured. As the syntax elements, information indicating one of a plurality of non-separable transformation configurations can be used. For example, if three types of non-separable transformation configurations are supported (i.e., n 7 A set of non-separable transformations and k 7 Transform kernel candidates of dogs, n 8 A set of non-separable transformations and k 8 Transform kernel candidates of dogs, n 9 A set of non-separable transformations and k 9 (a candidate for a dog transformation kernel) where the corresponding syntax element can have values of 0, 1, and 2, and the non-separable transformation configuration to be applied to the current block can be determined based on the value of the signaled syntax element. Depending on whether and / or which secondary transformation is applied, the number of non-separable transformation sets and transformation kernel candidates can be configured differently. For example, if no secondary transformation is applied, n 10 A set of non-separable transformations and k 10 A non-separable transformation configuration including a transformation kernel candidate can be applied. When a second transformation is applied, n 11 A set of non-separable transformations and k 11 A non-separable transformation configuration including a transformation kernel candidate can be applied. Depending on the range of quantization parameters (QP) and / or QP values, different non-separable transform configurations can be applied. For example, when the QP value has a small value, n 12 A set of non-separable transformations and k 12A non-separable transformation configuration including a transformation kernel candidate can be applied. On the other hand, if the QP value has a large value, n 13 A set of non-separable transformations and k 13 A non-separable transform configuration including a candidate transform kernel can be applied. If the QP value is less than or equal to a threshold (e.g., 32), the case can be classified as having a small QP value, otherwise, the case can be classified as having a large QP value. Alternatively, the range of QP values can be divided into three or more ranges, and a different non-separable transform configuration can be applied to each range. For relatively large blocks, instead of using a non-separable transform corresponding to the width and height of the block, the block can be divided into multiple sub-blocks and a non-separable transform corresponding to the width and height of the sub-blocks can be used. For example, when performing a non-separable transform for a 4x8 block, the 4x8 block can be divided into two 4x4 sub-blocks and a 4x4 block-based non-separable transform can be used for each of the 4x4 sub-blocks. Alternatively, for an 8x16 block, the block can be divided into two 8x8 sub-blocks and an 8x8 block-based non-separable transform can be used. The above non-separable transform set can be determined based on the intra prediction mode of the current block and the mapping table. The mapping table can define the mapping relationship between the pre-defined intra prediction modes and the non-separable transform sets. The pre-defined intra prediction modes can include two non-directional modes and 65 directional modes. In general, the non-separable transform has a larger transform kernel size than the separable transform. This means that the computational complexity required for the transform process is high and the memory required for storing the transform kernel is large. Meanwhile, the separable transform can consider only statistical characteristics existing in the horizontal and / or vertical directions, but the non-separable transform can simultaneously consider statistical characteristics in the two-dimensional space including the horizontal and vertical directions, thereby providing better compression efficiency. Since the statistical characteristics and diversity of the residual are different depending on the directionality of the intra prediction mode, there may be cases where the non-separable transform is absolutely necessary, and there may exist an intra prediction mode where the characteristics of the residual can be sufficiently identified by the separable transform alone. Therefore, by predefining which transform to use according to the intra prediction mode in the encoding device and the decoding device, the transform process can be designed with optimized complexity and memory requirements. The non-directional mode can include the planar mode of number 0 and the DC mode of number 1, and the directional mode can include the intra prediction modes of numbers 2 to 66. However, this is merely an example, and the present disclosure can also be applied to cases where the number of pre-defined intra prediction modes is different. Due to the application of wide angle intra prediction (WAIP), the pre-defined intra prediction modes can further include intra prediction modes from -14 to -1 and intra prediction modes from 67 to 80. FIG. 5 exemplarily shows intra prediction modes and prediction directions thereof according to the present disclosure. Referring to FIG. 5 , modes -14 to -1 and 2 to 33 and modes 35 to 80 are symmetrical with respect to the prediction direction with respect to mode 34. For example, modes 10 and 58 are symmetrical with respect to the direction corresponding to mode 34, and mode -1 is symmetrical with mode 67. Accordingly, for a vertical mode that is symmetrical with respect to a horizontal mode with respect to mode 34, the input data can be transposed and used. Transposing the input data means that rows in the input data MxN of a two-dimensional block become columns and columns become rows to form NxM data. For example, when a 4x4 block is used, 16 data forming a 4x4 block can be appropriately arranged to form a 16x1 1-dimensional vector for non-separable transformation. At this time, the 1-dimensional vector can be formed in row-major order or in column-major order. The residual samples resulting from the non-separable transformation can be arranged in the above order to form a 2-dimensional block. For modes -14 to -1 and 2 to 33, if the data arrangement order for constructing a 16x1 input vector is row-major order, for modes 35 to 80, the input vector can be constructed according to column-major order. Mode 34 can be considered as neither a horizontal mode nor a vertical mode, but in this disclosure, it is classified as belonging to a horizontal mode. That is, for modes -14 to -1 and 2 to 33, the input data alignment method for the horizontal mode, i.e., the row-major order, is used, and for the vertical mode that is symmetrical about mode 34, the input data can be transposed and used. For non-square blocks, the symmetry in square blocks (i.e., the symmetry between the P mode and the (68-P) mode in an NxN block (2<=P<=33) or the symmetry between the Q mode and the (66-Q) mode (-14<=Q<=-1)) cannot be utilized. Therefore, in addition to the symmetry based only on the intra prediction mode, the symmetry between block shapes that are in a transpose relationship with each other, i.e., the symmetry between a KxL block and an LxK block, can also be utilized. Specifically, a symmetry relationship exists between a KxL block predicted by the P mode and an LxK block predicted by the (68-P) mode. Or, a symmetry relationship exists between a KxL block predicted by the Q mode and an LxK block predicted by the (66-Q) mode. Since a KxL block having mode 2 and an LxK block having mode 66 can be viewed as symmetrical to each other, the same transformation kernel can be applied to the KxL block and the LxK block. If a non-separable transformation set for the intra prediction mode of the KxL block is mapped, in order to apply a non-separable transformation to the LxK block, the non-separable transformation set can be derived through a mapping table corresponding to the KxL block based on the (68-P) mode instead of the P mode applied to the LxK block. Alternatively, the non-separable transformation set can be derived through a mapping table corresponding to the KxL block based on the (66-Q) mode instead of the Q mode applied to the LxK block. For example, in order to apply a non-separable transformation to an LxK block, the non-separable transformation set can be selected based on mode 2 instead of mode 66. In addition, for a KxL block, the input data can be read in a pre-determined order (e.g., row-major order or column-major order) to form a one-dimensional vector and then the corresponding non-separable transformation can be applied. For an LxK block, the input data can be read in the transposed order to form a one-dimensional vector and then the corresponding non-separable transformation can be applied. That is, if the KxL block is read in row-major order, the LxK block can be read in column-major order. Conversely, if the KxL block is read in column-major order, the LxK block can be read in row-major order. In addition, when mode 34 is applied to the KxL block, a non-separable transformation set can be determined based on mode 34, and the input data can be read in a pre-determined order to form a one-dimensional vector and perform the corresponding non-separable transformation. When mode 34 is applied to the LxK block, a non-separable transformation set can be determined based on mode 34, but the input data can be read in a transposed order to form a one-dimensional vector and perform the corresponding non-separable transformation. In the present disclosure, a method for determining a non-separable transformation set and a method for organizing input data are described based on a KxL block. However, the non-separable transformation may be performed based on an LxK block by utilizing the symmetry described above for a KxL block. Alternatively, a block having a width greater than its height may be restricted to be used as a reference block. Alternatively, the symmetry may be restricted not to be utilized in the case of non-square blocks. In this case, a non-square block may use a different number of non-separable transformation sets and / or transformation kernel candidates than a square block, and may select a non-separable transformation set using a different mapping table than a square block. An example of a mapping table for selecting a set of non-separable transformations is as follows: predModeIntraTrSetIdxpredModeIntra < 040 <= predModeIntra <= 102 <= predModeIntra <= 12113 <= predModeIntra <= 23224 <= predModeIntra <= 44345 <= predModeIntra <= 55256 <= predModeIntra <= 66167 <= predModeIntra <= 804 Table 1 shows an example of allocating non-separable transform sets by intra prediction mode when there are five non-separable transform sets. The value of predModeIntra means the value of the intra prediction mode considering WAIP, and TrSetIdx is an index indicating a specific non-separable transform set. In Table 1, it can be confirmed that the same non-separable transform set is applied to modes located in symmetrical directions according to the intra prediction mode. Table 1 is only an example of using five non-separable transform sets, and does not limit the total number of non-separable transform sets for non-separable transforms. Alternatively, as shown in Table 2, the non-separable transform may not be applied to WAIP for compression performance. predModeIntraTrSetIdx0 <= predModeIntra <= 102 <= predModeIntra <= 12113 <= predModeIntra <= 23224 <= predModeIntra <= 44345 <= predModeIntra <= 55256 <= predModeIntra <= 661 Alternatively, as shown in Table 3, instead of constructing a separate non-separable transform set for WAIP, a non-separable transform set corresponding to adjacent intra prediction modes may be shared. predModeIntraTrSetIdxpredModeIntra < 010 <= predModeIntra <= 102 <= predModeIntra <= 12113 <= predModeIntra <= 23224 <= predModeIntra <= 44345 <= predModeIntra <= 55256 <= predModeIntra <= 801 The above non-separable transform set may include a plurality of transform kernel candidates, and any one of the plurality of transform kernel candidates may be selectively used. For this purpose, an index signaled through a bitstream may be used. Alternatively, any one of the plurality of transform kernel candidates may be implicitly determined based on context information of a current block. Here, the context information may mean a size of a current block or whether a non-separable transform is applied to a neighboring block. Here, the size of the current block may be defined by a width, a height, a maximum / minimum value of the width and the height, a sum of the width and the height, or a product of the width and the height. Below, we will take a closer look at how to determine the transformation kernel for the inverse transformation of the current block. Example 1 As mentioned above, the inverse transformation can be divided into a separable transformation and a non-separable transformation. A separable transformation means performing transformations in the horizontal direction and the vertical direction respectively for a two-dimensional block, and a non-separable transformation can mean performing a single transformation for samples constituting the entire or a part of a two-dimensional block. When expressing a separable transformation, it can be expressed as a pair of horizontal transformation and vertical transformation, and in this disclosure, it will be expressed as (horizontal transformation, vertical transformation). Multiple transformation sets can be defined for the inverse transformation of the current block. Each transformation set can contain one or more transformation kernel candidates. For example, any one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) can be applied as a separate transform, and the four transform kernel candidates can be regarded as one transform set. In addition, (DCT-2, DCT-2) can be regarded as one transform set. A transform skip that does not apply a transform can also be regarded as one transform set, and (DCT-2, DCT-2) and the transform skip can be regarded as one transform set. In the present disclosure, a transform kernel may refer to one transform (eg, DCT-2, DST-7) or may refer to two transform pairs (eg, (DCT-2, DCT-2)). As another example of a transform set, there may be the aforementioned non-separable transform set. In the present disclosure, a non-separable transform applied as a primary transform may be denoted as a Non-Separable Primary Transform (NSPT). In NSPT, a plurality of non-separable transform sets may be configured, and each non-separable transform set may include one or more transform kernels as transform kernel candidates. In the case of NSPT, one of the plurality of non-separable transform sets is selected according to the intra prediction mode, and the plurality of non-separable transform sets for NSPT may be denoted as an NSPT set list. This is as discussed above, and a detailed description thereof will be omitted here. A group of one or more transform sets available to a current block can be formed from a plurality of pre-defined transform sets. The group of one or more transform sets can be formed by a predetermined area unit to which the current block belongs, and is hereinafter referred to as a collection. Here, the predetermined area unit can be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU). For example, let S be a set of transforms consisting of (DCT-2, DCT-2). 1 , (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) are the set of transforms S 2 Let us call them respectively. In addition, the aforementioned NSPT set list can include N non-separable transformation sets, and the N non-separable transformation sets are called S 3,1 , S 3,2 , ..., S 3,N Let us call them respectively. Here, N can be 35, but is not limited thereto. S by intra prediction mode of the current block 3,13 When selected as a non-separable transformation set for this NSPT, the transformation kernel applicable to the current block is S 1 , S 2 , or S 3,13 It may belong to one of the following collections. In this case, the current block is the collection available to {S 1 , S 2 , S 3,13} can be expressed as . As described above, since a collection according to the present disclosure is a group of one or more transformation sets available to a current block, the collection may be differently configured depending on the context of the current block. Here, the context may include at least one of a shape, a size, or an intra prediction mode. If a total of K contexts are defined, K collections may be generated, each collection including C i can be denoted as (i=1, 2, ..., N). For example, if the sizes of blocks to which NSPT can be applied are 4x4, 8x8, 16x16, and 32x32, and one of a total of 35 non-separable transformation sets is selected by the intra prediction mode, a total of 4 x 35 = 140 contexts can be defined if different transformation kernels are applied to each block size.A collection may be formed based on the context of the current block, and at this time, a process of selecting one of a plurality of transformation sets belonging to the collection and selecting one of a plurality of transformation kernel candidates belonging to the selected transformation set may be performed. Here, the selection of the transformation set and the transformation kernel candidate may be performed implicitly based on the context of the current block, or may be performed based on an index that is explicitly signaled. Alternatively, the process of selecting one of a plurality of transformation sets belonging to the collection and the process of selecting one of a plurality of transformation kernel candidates belonging to the selected transformation set may be performed separately. For example, an index for selecting a transformation set may be first signaled, and based on this, one of a plurality of transformation sets belonging to the collection may be selected. Then, an index indicating one of a plurality of transformation kernel candidates belonging to the transformation set may be signaled, and based on the signaled index, one of the transformation kernel candidates may be selected from the transformation set. A transformation kernel of the current block may be determined based on the selected transformation kernel candidate. Alternatively, selection of any one of the transformation sets from the collection may be implicitly performed based on the context of the current block, and selection of any one of the transformation kernel candidates from the selected transformation set may be performed based on a signaled index. Alternatively, selection of any one of the transformation sets from the collection may be implicitly performed based on the signaled index, and selection of any one of the transformation kernel candidates from the selected transformation set may be implicitly performed based on the context of the current block. Alternatively, selection of any one of the transformation sets from the collection may be implicitly performed based on the context of the current block, and selection of any one of the transformation kernel candidates from the selected transformation set may also be implicitly performed based on the context of the current block. Of course, if the number of transformation sets belonging to the collection is 1, an index for selecting a transformation set may not be signaled.Similarly, if the number of transformation kernel candidates belonging to the selected transformation set is 1, the index indicating the transformation kernel candidate may not be signaled. Alternatively, an index indicating any one of all transformation kernel candidates belonging to the current collection may be signaled. In this case, the process of selecting any one transformation set from the collection may be omitted. At this time, all transformation sets belonging to the collection may be rearranged in consideration of priorities. For example, in the case of assigning a small-length binary code to a small-value index, such as a truncated unary code, it may be advantageous to assign a small-value index to a transformation kernel candidate that is more advantageous for improving coding performance. When rearranging all transformation kernel candidates belonging to the collection in accordance with priorities (shuffling), different shuffling may be applied to each collection. In addition, instead of rearranging all transformation kernel candidates belonging to the collection, only some of them may be selectively rearranged. Example 2 The transformation kernel for the inverse transformation of the current block can be determined based on MTS (Multiple Transform Selection). The MTS according to the present disclosure may use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (identity transform) as a transform kernel. In addition, the MTS according to the present disclosure may further include a transform kernel of DCT-2. In the present disclosure, a plurality of MTS sets for MTS can be defined. Based on the size of a current block and / or an intra prediction mode, one of the plurality of MTS sets can be determined. For example, in determining one MTS set, 16 transform block sizes can be considered, and for a directional mode, the shape of the transform block and the symmetry between the intra prediction modes can be considered. For the WAIP (Wide Angle Intra Prediction) mode (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), an MTS set corresponding to mode 2 can be applied for modes -1 to -14 (or -15), and an MTS set corresponding to mode 66 can be applied for modes 67 to 80 (or 81). A separate MTS set can be allocated for the MIP (Matrix-based Intra Prediction) mode. For example, MTS sets according to transform block size and intra prediction mode can be allocated / defined as shown in Table 4 below. Block size Intra prediction mode Width Height [0, 1] [2, 12] [13, 23] [24, 34] MIP440 12 3 4 4 8 5 6 7 8 9 4 16 10 1 1 2 1 3 1 4 4 3 2 1 5 16 17 18 19 8 4 2 0 2 1 2 2 2 3 2 4 8 8 2 5 26 27 28 29 8 16 30 31 32 33 34 8 32 35 36 37 38 39 16 4 4 0 4 1 4 2 4 3 4 1 6 8 4 5 4 6 4 7 4 8 4 9 16 16 5 0 5 1 5 2 5 3 5 4 1 6 32 5 5 5 6 5 7 5 8 5 9 32 4 6 0 6 16 2 6 3 6 4 3 2 8 6 5 6 6 6 7 6 8 6 9 32 16 7 0 7 1 7 2 7 3 7 4 3 2 3 2 7 5 7 6 7 7 7 8 7 9 Table 4 shows the allocation of MTS sets according to 16 transform block sizes and intra prediction modes. The number of pre-defined MTS sets is 80, and the index indicating one of the 80 MTS sets can have a value from 0 to 79, as shown in Table 4. MTS set index transformation kernel candidate index Table 5 shows transformation kernel candidates included in each MTS set examined in Table 4. Each MTS set can be composed of six transformation kernel candidates. The transformation kernel candidate index has a value of any one of 0 to 5 and can indicate any one of the six transformation kernel candidates. Here, each transformation kernel candidate can be a combination of a horizontal transformation kernel and a vertical transformation kernel for a separate transformation, and 25 transformation kernel candidates having indices of 0 to 24 can be defined. Kernel Combination IndexIf the value of the intra prediction mode is less than 35If the value of the intra prediction mode is greater than or equal to 350(DCT-8, DCT-8)(DCT-8, DCT-8)1(DST-7, DCT-8)(DCT-8, DST-7)2(DCT-5, DCT-8)(DCT-8, DCT-5)3(DST-4, DCT-8)(DCT-8, DST-4)4(DST-1, DCT-8)(DCT-8, DST-1)5(DCT-8, DST-7)(DST-7, DCT-8)6(DST-7, DST-7)(DST-7, DST-7)7(DCT-5, DST-7)(DST-7, DCT-5)8(DST-4, DST-7)(DST-7, DST-4)9(DST-1, DST-7)(DST-7, DST-1)10(DCT-8, DCT-5)(DCT-5, DCT-8)11(DST-7, DCT-5)(DCT-5, DST-7)12(DCT-5, DCT-5)(DCT-5, DCT-5)13(DST-4, DCT-5)(DCT-5, DST-4)14(DST-1, DCT-5)(DCT-5, DST-1)15(DCT-8, DST-4)(DST-4, DCT-8)16(DST-7, DST-4)(DST-4, DST-7)17(DCT-5, DST-4)(DST-4, DCT-5)18(DST-4, DST-4)(DST-4, DST-4)19(DST-1, DST-4)(DST-4, DST-1)20(DCT-8, DST-1)(DST-1, DCT-8)21(DST-7, DST-1)(DST-1, DST-7)22(DCT-5, DST-1)(DST-1, DCT-5)23(DST-4, DST-1)(DST-1, DST-4)24(DST-1, DST-1)(DST-1, DST-1) Table 6 is an example of the 25 transform kernel candidates examined in Table 5. Specifically, the horizontal transformation and vertical transformation of the transform kernel candidate are indicated as (horizontal transformation, vertical transformation). For each transform kernel candidate index, the horizontal / vertical transformation when the intra prediction mode is less than 35 may be the opposite of the horizontal / vertical transformation when the intra prediction mode is 35 or more. When the value of the intra prediction mode is 35 or more, a mode symmetrical with respect to mode 34 may be derived, and an MTS set may be selected from Table 4 based on the mode. In addition, the symmetry of the block shape may be additionally considered. When the original transform block has a WxH size, the original transform block may be considered to have a HxW size by symmetrizing it, and an MTS set may be selected from Table 4. Here, the value of the intra prediction mode may be the value of the modified intra prediction mode. That is, as mode values for WAIP, -14 (or -15) to -1 are modified to mode 2, 67 to 80 (or 81) are modified to mode 66, and the remaining modes can be set to the values of the modified intra prediction modes as the values of the original intra prediction modes. In this case, since the extended modes for WAIP are also configured symmetrically around mode 34, the symmetry around mode 34 can be utilized for all directional modes except for the Planar mode and the DC mode. For example, if a 16x32 block is predicted to be mode 54, mode 14 (=68-54) can be derived as a mode symmetric to mode 54, and the block size can be considered as 32x16. In this case, an MTS set with an index of 72 can be selected, as defined in Table 4. When the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, when the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the symmetric block size considering the symmetry of the block shape. For example, when the MIP mode is applied for an 8x16 block, the 8x16 block may be regarded as a symmetrical 16x8 block thereof, and an MTS set having an index of 49 may be selected as defined in Table 4. Alternatively, when the MIP mode is applied, the intra prediction mode may be regarded as the Planar mode. In this case, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, the MTS set assigned to the MIP mode may be selected based on the symmetrical block size considering the symmetry of the block shape. In the case of the MIP mode, a flag may be used to indicate whether the MIP mode is applied in the transpose mode. If the MIP mode is applied to the current block of MxN and the flag indicates application of the transpose mode, the intra prediction mode is regarded as the Planar mode, and the current block of MxN may be regarded as an NxM block. That is, from Table 4, an MTS set corresponding to the block size of NxM and the Planar mode may be selected. As seen in Table 6, if the value of the intra prediction mode is 35 or more, the horizontal transformation and the vertical transformation are swapped, but since the intra prediction mode of the current block is regarded as the Planar mode, the horizontal transformation and the vertical transformation of the transformation kernel candidate may not be swapped. Alternatively, if the MIP mode is applied to the current block of MxN and the flag indicates application of the transpose mode, the intra prediction mode is not regarded as the Planar mode, and the current block of MxN may be regarded as an NxM block. That is, from Table 4, an MTS set corresponding to the block size of NxM and the MIP mode may be selected. In Table 5, a transformation kernel candidate selected by a transformation kernel candidate index may be set as a transformation kernel of the current block. Alternatively, depending on the size of the current block, at least one of the horizontal transformation or the vertical transformation of the selected transformation kernel candidate may be changed to another transformation kernel. For example, if the transformation kernel candidate index is 3 and both the width and the height of the current block are 16 or less, at least one of the horizontal transformation or the vertical transformation of the transformation kernel candidate corresponding to the transformation kernel candidate index of 3 may be changed to another transformation kernel. At this time, the horizontal transformation and the vertical transformation may be changed independently of each other. If the difference (or the absolute value of the difference) between the value of the intra prediction mode of the current block and the value of the horizontal mode is less than or equal to a predetermined threshold, the vertical transformation of the selected transformation kernel candidate may be changed to an IDT (identity transformation). If the difference (or the absolute value of the difference) between the value of the intra prediction mode of the current block and the value of the vertical mode is less than or equal to a predetermined threshold, the horizontal transformation of the selected transformation kernel candidate may be changed to an IDT (identity transformation). Here, the threshold can be determined based on the width and height of the current block, as shown in Table 7 below. Block Size Threshold Width Height 44848641648488888166164416821616-1 Table 7 defines thresholds according to the size of a transform block for changing the horizontal transformation and / or vertical transformation of a transform kernel candidate selected by a transform kernel candidate index to another transform kernel. Six transform kernel candidates composing one MTS set can be distinguished by transform kernel candidate indices from 0 to 5 as defined in Table 5. The transform kernel candidate indices can be signaled via a bitstream. A flag indicating whether the MTS set is available / applied (MTS enabled flag or MTS flag) can be signaled, and a transform kernel candidate index can be signaled when the flag indicates the availability / applicability of the MTS set. The MTS flag can be composed of one bin, and one or more contexts for CABAC-based entropy coding (hereinafter, referred to as CABAC contexts) can be allocated to the bin. For example, different CABAC contexts can be allocated for non-MIP mode and MIP mode, respectively. Depending on the context of the current block described above, the number of transform kernel candidates available to the current block may be set differently. For example, as the context of the current block, the sum of the absolute values of all or part of the transform coefficients in the current block may be considered. The sum of the absolute values of the transform coefficients is referred to as AbsSum. If AbsSum is less than or equal to T1, only one transform kernel candidate corresponding to the transform kernel candidate index of 0 may be available. If AbsSum is greater than T1 and less than or equal to T2, four transform kernel candidates corresponding to the transform kernel candidate indices of 0 to 3 may be available. If AbsSum is greater than T2, six transform kernel candidates corresponding to the transform kernel candidate indices of 0 to 5 may be available. Here, T1 may be 6 and T2 may be 32, but this is only an example. When AbsSum is less than or equal to T1, since the number of transformation kernel candidates available for the current block is 1, the transformation kernel candidate corresponding to the transformation kernel candidate index of 0 can be set as the transformation kernel of the current block without signaling the transformation kernel candidate index. When AbsSum is greater than T1 and less than or equal to T2, since four transformation kernel candidates are available, any one of the four transformation kernel candidates can be selected based on the transformation kernel candidate index having two bins. That is, the transformation kernel candidate indices of 0 to 3 can be signaled as 00, 01, 10, and 11, respectively. For the two bins, the Most Significant Bit (MSB) can be signaled first, and the Least Significant Bit (LSB) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, for two bins, a CABAC context other than the CABAC context allocated for the MTS flag may be allocated to each bin. Alternatively, a CABAC context may not be allocated to the two bins and bypass coding may be applied. When AbsSum is greater than T2, the transform kernel candidate index has a value from 0 to 5, so the transform kernel candidate index cannot be expressed with only two bins. In this case, the transform kernel candidate index may be expressed by allocating two or more bins, such as truncated binary coding. For each bin allocated by the truncated binary coding method, a CABAC context may be allocated, or bypass coding may be applied without allocating a CABAC context. Alternatively, a CABAC context may be allocated to some of a plurality of bins (e.g., the first bin, or the first and second bins), and bypass coding may be applied to the remaining bins. Example 3 The transformation kernel of the current block can be determined based on a transformation set including one or more transformation kernel candidates. The transformation kernel of the current block can be derived from any one or more transformation kernel candidates belonging to the transformation set. The process of determining a transformation kernel of a current block may include at least one of 1) a process of determining a transformation set of the current block or 2) a process of selecting one transformation kernel candidate from the transformation set of the current block. The process of determining the transformation set may be a process of selecting one of a plurality of transformation sets that are identically pre-defined for the encoding device and the decoding device. Alternatively, the process of determining the transformation set may be a process of configuring one or more transformation sets available to the current block from among a plurality of transformation sets that are identically pre-defined for the encoding device and the decoding device, and selecting one of the configured transformation sets. Alternatively, the process of determining the transformation set may be a process of configuring one transformation set based on a transformation kernel candidate available to the current block from among a plurality of transformation kernel candidates that are identically pre-defined for the encoding device and the decoding device. If the transformation set of the current block includes multiple transformation kernel candidates, a process of selecting one of the multiple transformation kernel candidates for the current block may be performed. However, if the transformation set of the current block includes one transformation kernel candidate (i.e., the current block has one transformation kernel candidate available), the transformation kernel of the current block may be set to the corresponding transformation kernel candidate. The transform set according to the present disclosure may mean the (non-separable) transform set in the aforementioned embodiment 1, or may mean the MTS set in the embodiment 2. Alternatively, the transform set may be defined separately from the (non-separable) transform set in the embodiment 1 or the MTS set in the embodiment 2. In this case, the transform set may include one or more specific transform kernels as transform kernel candidates. One specific transform kernel may be defined as a pair of a transform kernel for horizontal transform and a transform kernel for vertical transform, or may be defined as one transform kernel that is equally applied to horizontal and vertical transforms. In the embodiment of the present disclosure, the process of applying NSPT, which is a non-separable transform applied as a primary transform, is described in detail. NSPT can be applied to all or part of a transform block. Based on the forward NSPT, residual samples existing in the region where NSPT is applied can be input as a one-dimensional vector of NSPT. That is, residual samples existing in all or part of a transform block (referred to as Region Of Interest, ROI, in the present disclosure) can be collected as a one-dimensional vector and configured as an input. Thereafter, by applying the forward NSPT, a primary transform coefficient can be obtained. Conversely, by applying the backward NSPT to the primary transform coefficient, a one-dimensional vector output can be obtained. By arranging each element value constituting the corresponding output vector at a predetermined position within the 2D transform block, a residual sample for the ROI can be obtained. The non-separable transform kernel for NSPT may have a matrix dimension determined according to the size of the ROI. In the present disclosure, the transform kernel may be referred to as a transform type and a transform matrix, and the non-separable transform kernel for NSPT may be referred to as an NSPT kernel. For example, if the current block is an MxN transform block, the ROI is an area of the entire MxN transform block, and a square NSPT is applied, the dimension of the corresponding transform matrix may be MN x MN. For example, if the ROI is an area of the entire 8x8 transform block, the dimension of the NSPT kernel may be 64 x 64. According to one embodiment of the present disclosure, when NSPT is applied to a residual generated by intra prediction, the NSPT kernel can be adaptively determined according to the intra prediction mode. Since the statistical characteristics of the residual block can vary depending on the intra prediction mode, the compression efficiency can be improved by adaptively determining the NSPT kernel according to the intra prediction mode. It can be configured to share an NSPT kernel that is applied to one or more intra prediction modes. As described above, the non-separable transform set can be determined based on the intra prediction mode of the current block and the mapping table. The mapping table can define a mapping relationship between the pre-defined intra prediction modes and the non-separable transform sets. The pre-defined intra prediction modes can include two non-directional modes and 65 directional modes. As an example, intra prediction modes can be grouped into intra prediction mode groups. One NSPT kernel or multiple NSPT kernels can be assigned to an intra prediction mode group. In other words, a non-separable transform set (NSPT set) including one or more NSPT kernels can be assigned to an intra prediction mode group. The non-separable transform set is mapped to an intra prediction mode, and one of N NSPT kernels included in the non-separable transform set can be selected. For example, an intra prediction group may include adjacent prediction modes (eg, modes 17, 18, and 19). In addition, an intra prediction group may include modes having symmetry. For example, directional modes may be symmetrical with respect to the diagonal mode (i.e., intra prediction mode 34) of FIG. 5 described above. In this case, two modes having symmetry may form one group (or pair). For example, modes 18 and 50 may be included in the same group because they are symmetrical with respect to mode 34. However, for modes having symmetry, a process of transposing a 2D input block and then forming a one-dimensional input vector may be added before applying the forward NSPT kernel. For example, when the intra prediction mode is 34 or less, a one-dimensional input vector may be derived from the corresponding input block in the row first order without transposing the 2D input block. If the intra prediction mode is greater than 34, the one-dimensional input vector can be constructed by first transposing the 2D input block and then reading the input block in the row-first order, or by leaving the 2D input block as it is and reading the input block in the column-first order. The following Table 8 illustrates an example of an allocation mapping table for NSPT sets according to intra prediction modes. Referring to Table 8, a total of 35 NSPT sets from 0 to 34 can be defined. The extended WAIP mode (i.e., modes from -14 to -1 and modes from 67 to 80 in Fig. 5) can be allocated the NSPT set allocated to the nearest normal directional mode. That is, the extended WAIP mode can be allocated 2 NSPT sets. An NSPT set may include one or more NSPT kernels (or kernel candidates). That is, an NSPT set may include N NSPT kernel candidates. For example, N may be set to a value greater than or equal to 1, such as 1, 2, 3, 4, etc. A kernel applied to a current block among one or more NSPT kernels included in the NSPT set may be signaled using an index. In the present disclosure, the index may be referred to as an NSPT index. For example, the NSPT index may have values of 0, 1, 2, ..., N - 1. In addition, as an example, if the number of NSPT kernel candidates is 1, the NSPT index value may be fixed to 0. In this case, the NSPT index may be inferred without being separately signaled. In addition, a flag indicating whether NSPT is applied may be signaled separately from the NSPT index. In the present disclosure, the flag may be referred to as an NSPT flag. When the NSPT flag value is 1, NSPT may be applied. When the NSPT flag value is 0, NSPT may not be applied. When the NSPT flag is not signaled, the NSPT flag value may be inferred as 0. As an example, when the NSPT flag value is 1, the NSPT index may be signaled. Based on the signaled NSPT index, one of the N kernel candidates included in the NSPT set selected by the intra prediction mode may be specified. In one embodiment, the entropy coding method of the NSPT index can be defined in various ways considering the number (N) of NSPT kernels included in the NSPT set. For example, as a method of mapping values from 0 to N-1 to empty strings (i.e., binarization method), truncated unary binarization, truncated binarization, and fixed-length binarization methods can be used. For example, when the number N of kernel candidates constituting the NSPT set is 2, one bin can be used to specify one of the two candidates. For example, 0 can indicate the first candidate, and 1 can indicate the second candidate. In addition, when the N value is 3 and truncated unary binarization is applied, the candidates can be specified by two bins. For example, the first, second, and third candidates can be binarized to 0, 10, and 11, respectively, and signaled. As an example, the binarized bins can be coded using context coding or bypass coding. In this disclosure, a reduced primary transform (RPT) method using a transform kernel of a reduced dimension by a primary transform is described. As described above, when a forward NSPT is applied, samples belonging to a 2D residual block can be arranged (or rearranged) into a 1D vector according to the row priority (or column priority). Then, a transformation matrix for NSPT can be multiplied by the arranged vector. When the corresponding 2D residual block is an M x N block (M is the width length, N is the height length), the length of the rearranged 1D vector can be M*N. That is, the corresponding 2D residual block can also be represented as a column vector having a dimension of M*N x 1. In this disclosure, M*N can be conveniently denoted as MN. In this case, the dimension of the corresponding transformation matrix can be MN x MN. In summary, forward NSPT can work by multiplying the left side of an MN x 1 vector by the corresponding MN x MN transformation matrix to obtain an MN x 1 transformation coefficient vector. When RPT is applied, instead of multiplying the MN x MN matrix as the forward NSPT transform matrix described above, r transform coefficients can be obtained by multiplying an r x MN matrix. Here, r represents the number of rows of the transform matrix, and MN represents the number of columns of the transform matrix. According to an embodiment of the present disclosure, the value of r can be set to be less than or equal to MN. That is, the existing forward NSPT transform matrix includes MN rows, and each row is a 1 x MN row vector, which is a transform basis vector of the corresponding NSPT transform matrix. The corresponding transform coefficients can be obtained by multiplying each transform basis vector by MN x 1 sample column vector. Since the conventional forward NSPT transformation matrix is composed of MN row vectors, MN transformation coefficients (i.e., MN x 1 transformation coefficient column vectors) can be obtained by applying the forward NSPT. On the other hand, in the case of the forward RPT, the transformation matrix can be composed of r transformation basis vectors instead of MN transformation basis vectors. Accordingly, when the forward RPT is applied, r transformation coefficients (i.e., r x 1 transformation coefficient column vectors) can be obtained instead of MN. The RPT kernel can be configured by selecting r transformation basis vectors, which are some of the transformation basis vectors constituting the MN x MN forward NSPT kernel. In the present disclosure, the transformation kernel can be referred to as a transformation type and a transformation matrix, and the non-separable transformation kernel for NSPT can be referred to as an RPT kernel. That is, when selecting r 1 x MN row vectors from the MN x MN forward NSPT kernel, it can be advantageous to select the transformation basis vectors that are most important from the viewpoint of coding performance. Specifically, in terms of energy compaction through transformation, more energy can be concentrated on the transformation coefficients that appear first by multiplying the forward NSPT transformation matrix. In other words, the transformation basis vectors located on the upper side of the forward NSPT transformation matrix can generate transformation coefficients having larger energy. Considering this point, the rx MN forward RPT kernel can be configured (or derived) by taking r from the upper side of the forward NSPT kernel. The RPT according to the present disclosure takes only a part (i.e., r) of the transform coefficients obtained by applying the existing NSPT, which may result in a loss of a part of the energy of the original signal. That is, distortion may occur between the original signal and the original signal through the process. Nevertheless, by applying the RPT, only r transform coefficients are generated instead of MN, so that the amount of bits required to code the transform coefficients can be reduced. Therefore, in the case of a signal in which a large amount of energy is concentrated in a small number of transform coefficients (e.g., an image residual signal), the gain obtained by reducing the signaling bits can be significantly large, thereby improving the coding performance. The reverse NSPT may be the transpose matrix of the forward NSPT kernel described above as a transformation matrix. At this time, the input data may be a transformation coefficient signal instead of a sample signal such as a residual signal. Specifically, if the forward NSPT transformation matrix is G and the sample signal rearranged into a 1D vector is x, the transformation coefficient vector obtained by multiplying the transformation matrix on the left side can be expressed as in the following mathematical expression 4. Referring to Equation 4, x and y can be MN x 1 column vectors. G can have the form of MN x MN matrix. The backward NSPT process can be expressed as Equation 5 below using the same variables. In mathematical expression 5, G T denotes the transpose matrix of G. The forward RPT operation and the backward RPT operation according to the present disclosure can also be expressed by the two mathematical equations above. However, when the RPT is applied, y is an rx 1 column vector instead of an MN x 1 column vector, and G is an rx MN matrix instead of an MM x MN matrix. That is, even if the RPT is applied instead of the NSPT, the dimension of the sample signal (e.g., the image residual signal) does not change, which may mean that the original number of sample signals (i.e., MN sample signals) can be restored with only r transform coefficients through the backward RPT. That is, the original MN sample signals can be restored by coding only r transform coefficients that are less than MN, which may lead to an improvement in coding performance. In one embodiment of the present disclosure, an RPT structure is proposed that defines an r value considering the statistical characteristics of a residual block, and derives a residual block of an existing transform block size from a residual block of a reduced size determined according to the defined r value. If an additional transformation (i.e., a secondary transform) is applied to predict a statistical distribution of primary transform coefficients, quantized non-zero coefficients may be concentrated in a relatively low frequency region since a quantization process is applied to the primary transform coefficients. Accordingly, a reduced secondary transform for statistical distribution of primary transform coefficients can define the statistical characteristics of primary transform coefficients relatively simply in the form of setting an r value for a given low frequency region. However, in the case of the RPT according to the present disclosure, it is a technique for defining an r value considering the statistical characteristics of samples in a residual block that have very different characteristics from the distribution of primary transform coefficients, and thus has a fundamental difference from the reduced secondary transform. Below, we describe various embodiments for determining the RPT kernel, which is a reduced-dimensional transformation matrix. In other words, we describe below a method for determining or defining the r value in the RPT. In one embodiment of the present disclosure, the r value in the RPT can be determined by considering the worst case complexity allowed by the conversion system. As one embodiment, the worst case complexity can be calculated based on the number of multiplications per sample. MN * r multiplications are required to apply the RPT in both the forward and backward directions based on an M x N block. Since a 2D block is composed of a total of MN samples, the number of multiplications per sample can be calculated as (MN * r) / MN = r. Therefore, the r value can be configured to be maintained less than or equal to the maximum number of multiplications per sample allowed. For example, when the maximum possible number of multiplications per sample is set to 16 for a 16x16 block, the r value can be determined to be less than or equal to 16. That is, the forward RPT kernel can be set to 16 x 256. In another embodiment, memory usage can be considered as a measure of worst-case complexity. As an example, the allowed memory size per kernel can be set. For example, if p bytes are required for each kernel coefficient (each element constituting a transform kernel is referred to as a kernel coefficient in this disclosure) and the memory usage is set to q bytes or less per kernel, the value of r can be set to q / (MN * p) or less. For example, if p is 1 byte for a forward RPT kernel for a 16x16 block and the memory usage is set to 8 KB or less per kernel (q = 8 KB = 2 13 bytes), the r value can be set to 32 or less. Also, as another example, memory usage and / or number of multiplications per sample can be considered as a measure of worst-case complexity. For example, if the maximum possible number of multiplications per sample is set to 16 for a 16x16 block, and the memory usage is set to 8KB or less per kernel (kernel coefficients are expressed in 1 byte), then the value of r can be set to 16 or less. In addition, in one embodiment, the r value constituting the RPT kernel may be determined by specific information. In other words, the r value constituting the RPT kernel may be determined based on a predefined encoding parameter. For example, the r value may be determined according to the size of the block. In other words, the RPT kernel may be variably determined according to the size of the block. Here, the block may be at least one of a coding block, a transform block, and a prediction block. In addition, for example, the r value may be determined based on prediction information. Here, the prediction information may include information about inter / intra prediction, intra prediction mode information, etc. In addition, for example, the r value may be determined based on signaled information (value of a syntax element). For example, the r value may be variably determined according to a quantization parameter value. In addition, in terms of complexity improvement, a predefined fixed value may be used as the r value, and the predefined fixed value may be determined based on the signaled information. When the RPT kernel rx MN is multiplied by the sample signal, r transform coefficients can be obtained. The obtained r transform coefficients can be arranged according to a scan order of the predefined transform coefficients (e.g., forward / backward zig-zag scan order, forward / backward horizontal scan order, forward / backward vertical scan order, forward / backward diagonal scan order, a scan order specified based on an intra prediction mode, etc.). When the transform coefficients obtained from the forward RPT application are arranged according to such a scan order (e.g., a scan order per coefficient group (CG) unit can also be applied), if the value of r is smaller than MN, the inside of the M x N block cannot be completely filled with the r transform coefficients, and thus a blank space may be generated. As an embodiment of the present disclosure, the above-described blank space can be predicted by the following method in consideration of the characteristics of the residual signal. - The values of empty spaces can be filled using the values of available surrounding pixels. - The values of empty spaces can be filled based on the values of available surrounding pixels and the intra prediction mode. For example, the values of empty spaces can be predicted by performing intra prediction based on the values of available surrounding pixels and the intra prediction mode. - You can fill in the empty space values using predefined fixed values (e.g. 0). - The empty space values can be filled in from available surrounding pixels using a predefined intra prediction mode (e.g., planar mode). In the present disclosure, filling in the blank space with 0 among the examples described above may be referred to as a zero-out process. In the case of filling in the blank space with 0, the following embodiment may be applied. If a non-zero transform coefficient is detected (or parsed) in the blank space portion during parsing of the transform coefficients on the decoding device side, it may be considered (or inferred) that the RPT is not applied. In other words, if a non-zero transform coefficient exists in the predefined area representing the corresponding blank space, it may be considered that the RPT is not applied. In this case, signaling (or parsing) for a flag indicating whether to apply the RPT and / or an index designating one of a plurality of RPT kernel candidates may not be performed. As an example, if a non-zero transform coefficient exists in the predefined area representing the corresponding blank space, a predefined variable value may be updated, and it may be inferred that the RPT is not applied based on the updated variable value. In one embodiment of the present disclosure, whether to apply RPT can be determined depending on the size and / or shape of a block. In addition, the RPT kernel can be variably determined depending on the size and / or shape of the block. Since the r value can be different depending on the size and / or shape of the block (i.e., for each M x N block), the empty space can be different depending on the size and / or shape of the block. Accordingly, the area for checking whether a non-zero transform coefficient is detected depending on the size and / or shape of the block can be defined differently. In other words, the zero-out area can be variably determined. For example, when a 16x64 matrix is applied as a forward RPT matrix for an 8x8 block, the r value may be 16. In this case, when the CG is a 4x4 sub-block, only the upper left 4x4 block may be filled with a non-zero RPT transform coefficient, and the remaining three 4x4 sub-blocks (i.e., the upper right, lower left, and lower right sub-blocks), which are empty spaces, may be filled with 0 values. In this case, if a non-zero transform coefficient is detected in the remaining three 4x4 sub-block areas during the decoding process, it may be considered that the RPT is not applied. In addition, as described above, a flag indicating whether to apply the RPT or an index specifying one of a plurality of RPT kernel candidates may not be signaled. Also, as an example, when a 32x128 matrix is applied as a forward RPT matrix for a 16x8 block (i.e., the r value is 32), and a CG is a 4x4 sub-block, non-zero RPT transform coefficients may be filled only for two CGs in the scan order. For example, the RPT transform coefficients may be filled in the upper-left 4x4 sub-block and the 4x4 sub-block adjacent to the lower side of the upper-left sub-block. The area to be filled with zero as an empty space may be determined as the remaining area excluding the two 4x4 sub-blocks. The RPT kernel may be variably determined depending on the size and / or shape of the block, and as discussed above, the empty space may be determined differently for an 8x8 block and a 16x8 block. As an example, if the value of r is a multiple of the CG size and the transform coefficients are scanned in CG units, the flag and / or index related to the RPT may not be signaled if it is detected that non-zero transform coefficients exist in CGs belonging to the empty space. That is, the transform coefficients within the CG may be scanned in a specified order for each CG, and the transform coefficients within the CG may be scanned in the same manner for the next CG according to the scan order for the CG unit. In the conventional image compression technology, since a flag for whether a non-zero transform coefficient exists within each CG is signaled first, whether to apply the RPT can be determined based on the information alone, which can reduce the signaling overhead and the related implementation complexity. As described above, if a non-zero transform coefficient is detected in an empty space area that is filled with 0 when RPT is applied, RPT may not be applied. In this case, signaling for information related to RPT may be omitted. However, since it is not possible to determine whether RPT is applied if a non-zero transform coefficient is not detected in the empty space area, the flag indicating whether RPT is applied may be parsed after parsing (or signaling) the related transform coefficients to finally determine whether RPT is applied. In one embodiment, a forward secondary transform may be additionally applied to the transform coefficients generated through the application of RPT. Alternatively, a forward secondary transform may be additionally applied to a region where the generated transform coefficients are located in an M x N block. In the present disclosure, the region or a part of the region may be referred to as an ROI in terms of the forward secondary transform. For the reverse direction, the backward secondary transform may be applied first and then the backward RPT may be applied. Specifically, a region or a part of the region where r transform coefficients generated by the application of the forward RPT are located may be set as an ROI and the forward secondary transform may be applied. In this case, when a 16x64 forward RPT transform matrix is applied to an 8x8 region, the generated 16 transform coefficients may be located in the upper left 4x4 sub-block, and the sub-block region may be set as an ROI and the forward secondary transform may be applied to the ROI. In addition, the RPT kernel can adjust the coefficient values considering operations such as integer operations or fixed-point operations. That is, the RPT kernel can be configured to perform a transformation through integer operations (or fixed-point operations) in a practical codec system by appropriately scaling the kernel coefficients belonging to the kernel, rather than a theoretical orthogonal transformation or non-orthogonal transformation (wherein, the orthogonal transformation and the non-orthogonal transformation represent transformations in which the norm of each transformation basis vector is 1). When applying a separable transformation in a conventional image compression technology, the scaling factor multiplied can be equally reflected when applying the RPT. In this case, a separable transformation or a non-separable transformation (including RPT) can be performed while maintaining other processes (e.g., quantization and dequantization processes) other than the transformation. Integer coefficients of the RPT kernel can be obtained by multiplying the transformation basis vector by the scaling value described above. As an example, multiplying the scaling value may include applying operations such as rounding, flooring, and ceilinging to each kernel coefficient. That is, the integerized RPT kernel obtained through the above-described method is defined and can be used in the transformation / inverse transformation process. As described above, if kernel coefficients of scaled integer values are obtained through operations such as rounding, flooring, and ceilinging, the maximum and minimum values can be obtained for all kernel coefficients, so that the number of bits that can express all kernel coefficients can be obtained from the maximum and minimum values. For example, if the maximum value is less than or equal to 127 and the minimum value is greater than or equal to -128, all integer kernel coefficients can be expressed with 8 bits (in particular, through 2's complement representation, etc.). In general, the maximum value is (2 (N-1) - 1) and the minimum value is -2 (N-1) If this is the case, all integer kernel coefficients can be expressed with N bits. If the maximum value is (2 (N-1) - 1) Greater than or equal to the minimum value of -2 (N-1) For smaller cases, it may not be possible to express all the integer kernel coefficients with N bits. In this case, 1) all the kernel coefficients can be additionally multiplied by a scaling value to fit within the N-bit range, or 2) the number of bits required to express the kernel coefficients can be increased (i.e., more than N+1 bits). If all the kernel coefficients are multiplied by 2 to represent them with N bits, -p If we need to multiply by that amount (p >= 1), then we can merge it into the existing encoding / decoding process, and then multiply by 2. p It can be compensated by multiplying the amount. As an example, 2 pMultiplying by p can be implemented by performing an additional shift operation to the left by p bits, or by reducing the amount of right shift applied during quantization or dequantization by p. All kernel coefficients can be expressed as 8-bit, 9-bit, 10-bit, etc. using the method described above. Of course, the scaling value of the kernel coefficients can be set differently for each block size or kernel, and the number of bits for expressing the kernel coefficients can be set differently. The above-described NSPT can be applied based on at least one of the size, tree type, or component type of the current block. For example, whether to apply the NSPT can be determined based on at least one of the size, tree type, or component type of the current block. An NSPT index can be signaled based on at least one of the size, tree type, or component type of the current block. An NSPT set or an NSPT kernel can be derived based on at least one of the size, tree type, or component type of the current block. Allowed transform block sizes defined in the decoding device can be broadly divided into two groups. One of the two groups (hereinafter referred to as the first group) may mean a set of block sizes to which NSPT can be applied. The first group may be composed of one of the allowed transform block sizes, or may be composed of two or more block sizes among the allowed block sizes. The block size to which NSPT can be applied may be defined as a block size in which at least one of the width and the height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT can be applied may be defined as a block size in which the product of the width and the height is less than or equal to a predetermined threshold. Alternatively, the block size to which NSPT can be applied may be defined as a block size in which the maximum value of the width and the height is less than or equal to a predetermined threshold. The threshold may be an integer of 4, 8, 16, 32, 64, 128, or higher. The other of the above two groups (hereinafter referred to as the second group) may mean a set of block sizes to which NSPT is not applied. The above-described separable primary transformation may be applied to the block sizes belonging to the second group. In addition, a non-separable secondary transformation may be applied to all or part of the block sizes belonging to the second group. For example, if the size of the current block belongs to the first group, the reverse NSPT can be applied to the (dequantized) transform coefficients of the current block. If the size of the current block belongs to the second group, the reverse separable primary transform can be applied to the (dequantized) transform coefficients of the current block. Alternatively, if the size of the current block belongs to the second group, the reverse non-separable secondary transform (eg, low frequency non-separable transform, LFNST) can be first applied to the (dequantized) transform coefficients of the current block, and then the reverse separable primary transform (eg, DCT-2) can be applied to the transform coefficients obtained therefrom. For example, the first group, which is a set of block sizes to which NSPT can be applied, can be defined as a set of 4x4, 4x8, 8x4, 8x8. Alternatively, the first group can be defined as a set of 4x8, 8x4, 8x8. Alternatively, the first group can be defined as a set of 4x8, 8x4. Alternatively, the first group can be defined as a set of 4x4, 4x8, 4x16, 8x4, 8x8, 16x4. Alternatively, the first group can be defined as a set of 4x8, 4x16, 8x4, 8x8, 16x4. Alternatively, the first group can be defined as a set of 4x8, 4x16, 8x4, 16x4. Alternatively, the first group can be defined as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16. Alternatively, the first group can be defined as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8. Alternatively, the first group can be defined as a set of 4x8, 8x4, 8x8, 8x16, 16x8. Alternatively, the first group can be defined as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8. Alternatively, the first group can be defined as a set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, 32x16, 32x32. Alternatively, the first group can be defined as the set of 4x4, 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, 32x16. Alternatively, the first group can be defined as the set of 4x8, 8x4, 8x8, 8x16, 16x8, 16x16, 16x32, 32x16. Alternatively, the first group can be defined as the set of 4x8, 8x4, 8x16, 16x8, 16x16, 16x32, 32x16. Alternatively, the first group can be defined as the set of 4x8, 8x4, 8x16, 16x8, 16x32, 32x16. Alternatively, the first group can be defined as a set of 4x4, 4x8, 4x16, 8x4, 16x4.Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8. Alternatively, the first group can be defined as the set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, 16x8. Alternatively, the first group can be defined as the set of 4x8, 4x16, 8x4, 8x16, 16x4, 16x8. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, 16x16. Alternatively, the first group can be defined as the set of 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, 16x16. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, 16x16. Alternatively, the first group can be defined as the set of 4x8, 4x16, 8x4, 8x16, 16x4, 16x8, 16x16. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 32x4, 32x8. Alternatively, the first group can be defined as the set of 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 32x4, 32x8. The first group can be defined as the set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 32x4, 32x8. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x32, 32x4, 32x8, 32x16.Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 16x32, 32x4, 32x8, 32x16. Alternatively, the first group can be defined as the set of 4x8, 4x16, 4x32, 8x4, 8x16, 8x32, 16x4, 16x8, 16x32, 32x4, 32x8, 32x16. Alternatively, the first group can be defined as the set of 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 16x16, 16x32, 32x4, 32x8, 32x16, 32x32. For block sizes belonging to the first group, an NSPT matrix (or NSPT kernel) having a predetermined dimension can be applied. Here, the NSPT matrix can be expressed as a matrix having a dimension of PxQ as a reverse transformation matrix, and the PxQ matrix represents a matrix in which the number of rows and the number of columns are P and Q, respectively. For example, for a 4x4 block, an NSPT matrix of 16x16 may be applied. For at least one of the 4x8 block or the 8x4 block, an NSPT matrix of 32x20 may be applied. For at least one of the 4x16 block or the 16x4 block, an NSPT matrix of 64x24 may be applied. For an 8x8 block, an NSPT matrix of 64x32 may be applied. For at least one of the 8x16 block or the 16x8 block, an NSPT matrix of 128x40 may be applied. For a 16x16 block, an NSPT matrix of 256x44 may be applied. For a 4x32 block or a 32x4 block, an NSPT matrix of 128x36, 128x38, or 128x40 may be applied. For 8x32 blocks or 32x8 blocks, an NSPT matrix of 256x48 can be applied. For 16x32 blocks or 32x16 blocks, an NSPT matrix of 512x52 or 512x54 can be applied. For at least one of the 4x32 blocks or the 32x4 blocks, an NSPT matrix of 128x36, 128x38, 128x40, or 128x(36-(4*n)) can be applied. Alternatively, for at least one of the 4x32 blocks or the 32x4 blocks, an NSPT matrix of 128x(36-(4*n)) can be applied instead of the NSPT matrix of 128x36, 128x38, or 128x40. Here, * denotes multiplication, and n can be an integer greater than or equal to 0. For example, as the NSPT matrix for at least one of a 4x32 block or a 32x4 block, at least one of a 128x36 matrix, a 128x32 matrix, a 128x28 matrix, a 128x24 matrix, a 128x20 matrix, a 128x16 matrix, a 128x12 matrix, a 128x8 matrix, or a 128x4 matrix can be used. Also, for at least one of the 8x32 blocks or the 32x8 blocks, an NSPT matrix of 256x48 or 256x(48-(4*m)) can be applied. Or, for at least one of the 8x32 blocks or the 32x8 blocks, instead of the NSPT matrix of 256x48, an NSPT matrix of 256x(48-(4*m)) can be applied. Here, * denotes multiplication, and m can mean an integer greater than or equal to 0. For example, as an NSPT matrix for at least one of an 8x32 block or a 32x8 block, at least one of a 256x48 matrix, a 256x44 matrix, a 256x40 matrix, a 256x36 matrix, a 256x32 matrix, a 256x28 matrix, a 256x24 matrix, a 256x20 matrix, a 256x16 matrix, a 256x12 matrix, a 256x8 matrix, or a 256x4 matrix can be used. The combination of the structures of NSPT matrices for 4x32 blocks and 32x4 blocks and the structures of NSPT matrices for 8x32 blocks and 32x8 blocks can be formed by combinations of the above-described matrices. That is, the combination of the structures of NSPT matrices for 4x32 blocks and 32x4 blocks and the structures of NSPT matrices for 8x32 blocks and 32x8 blocks can be defined by a combination of one or more 128x(36-(4*n)) NSPT matrices and one or more 256x(48-(4*m)) NSPT matrices. Here, n can be one or more integers in the range of 0 to 8, and m can be one or more integers in the range of 0 to 11. For example, the NSPT matrix for at least one of a 4x32 block or a 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of an 8x32 block or a 32x8 block can be a 256x36 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x32 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x28 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x24 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x20 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x24 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x16 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x36 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x32 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x28 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x24 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x20 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x20 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x16 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x36 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x32 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x28 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x24 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x20 matrix. Alternatively, the NSPT matrix for at least one of the 4x32 block or the 32x4 block can be a 128x16 matrix, and the NSPT matrix for at least one of the 8x32 block or the 32x8 block can be a 256x16 matrix. Since the above PxQ matrix is a reverse NSPT matrix, a Px1 output vector can be obtained by applying the PxQ matrix to the Qx1 input vector (i.e., (PxQ matrix) x (Qx1 input vector)). Here, the Qx1 input vector may correspond to (inverse quantized) transform coefficients in the current block to which the NSPT is applied. At this time, the value of Q may mean the number of transform coefficients to which the NSPT is applied, and may be less than or equal to the product of the width and the height of the current block. The value of Q may be variably determined based on the size of the current block among the block sizes belonging to the first group described above. Alternatively, the value of Q may be set to be the same for the block sizes belonging to the first group. The Px1 output vector may correspond to a residual signal (or, decoded residual samples). The value of P may be equal to the product of the width and the height of the current block. In contrast, the forward NSPT matrix can be expressed as a QxP matrix, which is a transpose matrix of the PxQ matrix. A QxP output vector can be obtained by applying the QxP matrix to the Px1 input vector (i.e., (QxP matrix)x(Px1 input vector)). Here, the Px1 input vector can correspond to residual samples in the current block to which the NSPT is applied. The value of P can be equal to the product of the width and the height of the current block. The Qx1 output vector can correspond to transform coefficients in the current block induced through the NSPT. At this time, the value of Q can mean the number of transform coefficients output through the NSPT, and can be less than or equal to the product of the width and the height of the current block. Similarly, the value of Q can be variably determined based on the size of the current block among the block sizes belonging to the first group described above. Alternatively, the value of Q can be set identically for the block sizes belonging to the first group. As in the above examples, NSPT can be applied to non-square blocks, MxN blocks and NxM blocks. For example, NSPT can be applied to 4x8 blocks and 8x4 blocks. Alternatively, NSPT can be applied to 4x16 blocks and 16x4 blocks, or NSPT can be applied to 8x16 blocks and 16x8 blocks, or NSPT can be applied to 16x32 blocks and 32x16 blocks. By applying NSPT to specific block sizes belonging to the first group, more precise transformation can be performed and coding performance can be improved. When forward LFNST is applied, primary transformed transform coefficients of the remaining regions except for the region to which LFNST is applied (i.e. Region-Of-Interest, ROI) can be zeroed out. In addition, LFNST can be composed of a small number of transform basis vectors. In this case, if a separable primary transform such as DCT-2 and a non-separable secondary transform such as LFNST are applied instead of NSPT to the corresponding block sizes, performance degradation may occur. In this case, if NSPT is applied instead of LFNST, the zero-out process is omitted, and coding performance can be improved compared to the case of applying LFNST. In addition, performance improvement can also be expected by the method of applying NSPT. NSPT or LFNST can be applied using the symmetry described below. Here, in the case of LFNST, the transpose operation is performed on the corresponding input block only for the ROI region by utilizing symmetry. On the other hand, in the case of NSPT, the transpose operation is performed on the entire block by utilizing symmetry. Therefore, in the case of NSPT, the corresponding NSPT kernel can be trained and applied by utilizing a more sophisticated symmetry, so performance improvement can be expected. Also, when applying LFNST instead of NSPT to an 8x8 block, a 32x64 transformation matrix can be applied instead of a 16x64 transformation matrix from the perspective of forward transformation. Here, the 16x64 transformation matrix can be constructed by sampling the upper 16 rows of the 32x64 transformation matrix. When applying LFNST based on the 16x64 transformation matrix to an 8x8 block, 16 multiplications are required per sample to apply LFNST, but when using a 32x64 transformation matrix, 32 multiplications are required per sample to apply LFNST. However, when using a 32x64 transformation matrix in this way, improvement in coding performance can be expected. If the tree type of the current block is a single tree, NSPT can be applied to the luma component of the current block, and NSPT can not be applied to the chroma component of the current block. If the tree type of the current block is a dual tree, NSPT can be applied to the luma component and chroma component of the current block. Alternatively, NSPT may be applied to the luma component of the current block and not applied to the chroma component of the current block, regardless of whether the tree type of the current block is a single tree. Alternatively, NSPT may be applied to the luma component and chroma component of the current block, regardless of whether the tree type of the current block is a single tree. For example, if the tree type of the current block is single tree, NSPT is allowed for luma and chroma components, and the size of the current block belongs to the first group, one NSPT index can be signaled, and the luma and chroma components of the current block can share the NSPT index. Here, the NSPT index can be an index for selecting any one of the transformation kernel candidates for the NSPT. If the sizes of the luma block and the chroma block of the current block belong to the first group, the transformation kernel candidate selected by the same NSPT index can be applied to the luma and chroma components. If the tree type of the current block is single tree and NSPT is applied only to the luma component, LFNST may not be applied to the chroma component of the current block, and a separate transformation may be applied. Alternatively, if the tree type of the current block is single tree and NSPT is applied only to the luma component, LFNST may be applied to the chroma component of the current block. In the case of a single tree, there may be a high correlation between the luma component and the chroma component. In this case, by applying NSPT only to the luma component or by applying the transformation kernel candidate selected by one NSPT index to the luma component and the chroma component in common, unnecessary signaling can be reduced and compression efficiency can be improved. On the other hand, in the case of a non-single tree, the luma component and the chroma component have independent partitioning and encoding structures. In this case, by signaling the NSPT index for each component, the characteristics of each component can be reflected and compression efficiency can be improved. The NSPT kernel for the above NSPT can be derived based on at least one of symmetry between intra prediction modes or symmetry between block shapes. For example, the NSPT kernel can be derived as an NSPT kernel corresponding to at least one of a mode symmetric to the intra prediction mode of the current block or a block shape symmetric to the block shape of the current block. Alternatively, the NSPT kernel can be derived based on an NSPT set including one or more NSPT kernel candidates, wherein the NSPT set can be derived as an NSPT set corresponding to at least one of a mode symmetric to the intra prediction mode of the current block or a block shape symmetric to the block shape of the current block. Any one of the one or more NSPT kernel candidates belonging to the NSPT set can be set as the NSPT kernel of the current block. For this purpose, an NSPT index specifying any one of the one or more NSPT kernel candidates belonging to the NSPT set can be used. The NSPT index can be signaled through a bitstream or can be derived based on the aforementioned symmetry. There may be symmetry between at least two intra prediction modes among intra prediction modes pre-defined in a decoding device. Hereinafter, for the convenience of explanation, the symmetry will be described centered on the upper left diagonal mode (i.e., mode 34). Referring to Fig. 5, there is symmetry between directional modes. All modes except for the Planar mode (0) and the DC mode (1) have a prediction direction. Modes 2 to 66 may be named normal directional modes (which may be expressed as [2, 66]), and modes -14 to -1 (which may be expressed as [-14, -1]) and modes 67 to 80 (which may be expressed as [67, 80]) may be named wide directional modes. The wide directional mode may include at least one of a mode having a value smaller than -14 or a mode having a value larger than 80. Referring to Fig. 5, all modes except mode 0 and mode 1 are symmetrical with respect to mode 34. Specifically, for mode [2, 66], mode x and mode (68 - x) are symmetrical, and mode x and mode (66 - x) are symmetrical between mode [-14, -1] and mode [67, 80]. The same symmetry relationship can be established between mode [N, -1] and mode [67, 66 - N]. Here, N can be an integer less than or equal to -14. Meanwhile, with respect to the symmetry between block shapes, an MxN block and an NxM block can be defined as blocks that are symmetric to each other. Here, M and N can be the same or different. Or, if the ratio of the width and height of an M1xN1 block (M1 / N1) and the ratio of the height and width of an M2xN2 block (N2 / M2) are the same, an M1xN1 block and an M2xN2 block can be defined as blocks that are symmetric to each other. Or, if the ratio of the width and height of an M1xN1 block (M1 / N1) and the ratio of the width and height of an M2xN2 block (M2 / N2) are the same, an M1xN1 block and an M2xN2 block can be defined as blocks that are symmetric to each other. In a square block, modes that are symmetric to each other can share at least one of the NSPT set, the NSPT index, or the NSPT kernel. That is, at least one of the NSPT set, the NSPT index, or the NSPT kernel for one of the symmetric modes can be applied equally to another of the symmetric modes. For example, modes that are symmetric to each other can share a single NSPT kernel. However, for one of the symmetric modes, the corresponding NSPT kernel can be applied to the input data, and for the other mode, a transpose operation can be applied to the input data and then the corresponding NSPT kernel can be applied. Specifically, if the xth mode belongs to the [2, 33] mode, for the xth mode, a 1D vector can be constructed in the column-first order for the MxM block, which is the input data, and the NSPT kernel can be applied to the 1D vector. Here, the construction of the 1D vector in the column-first order can be done by reading the input data from the MxM block, which is the input data, in units of columns, obtaining M columns, and arranging them in order to construct the 1D vector. On the other hand, for the (68 - x) mode that is symmetric to the xth mode, a 1D vector can be constructed in the row-first order and the same NSPT kernel can be applied to the 1D vector. Here, the configuration of a 1D vector according to the row-major order can be to read the input data in units of rows from the MxM block of input data, obtain M rows, and sequentially arrange them to configure a 1D vector. If the x-th mode belongs to the [N, -1] mode (N ≤ -14), a 1D vector can be configured in the column-first order for the (66 - x) mode which is symmetric to the x-th mode, and the same NSPT kernel as that of the x-th mode can be applied to the 1D vector. The column-major order or the row-major order can be applied to the modes 0 and 1, and the column-major order or the row-major order can be applied to the mode 34.In addition, the row-major order may be applied to the intra prediction mode belonging to the [2, 33] mode, and the column-major order may be applied to the mode symmetrical to the intra prediction mode. The row-major order may be applied to the intra prediction mode belonging to the [N, -1] mode, and the column-major order may be applied to the mode symmetrical to it. For non-square blocks, in addition to the symmetry between intra prediction modes, symmetry between block shapes may be further considered. A non-square block whose width and height are M and N, respectively, can be viewed as having a symmetry relationship with a non-square block whose width and height are N and M, respectively. For example, in the [2, 66] mode, there may be symmetry between the xth mode of an MxN block and the (68 - x)th mode of an NxM block. Similarly, when the xth mode of an MxN block belongs to the [N, -1] mode (N ≤ -14), there may be symmetry between the xth mode of the MxN block and the (66 - x)th mode of the NxM block. The method of constructing a 1D vector from an input data block is as discussed above. That is, if the column-major order is applied to the x-th mode, the row-major order can be applied to the mode that is symmetric to it. Alternatively, if the row-major order is applied to the x-th mode, the column-major order can be applied to the mode that is symmetric to it. Specifically, if the column-major order is applied to the x-th mode, the input data can be read from the MxN block of input data in units of columns to obtain M columns, which can be arranged in order to construct a 1D vector. Here, each column can have a length of N. For the mode that is symmetric to the x-th mode, the input data can be read from the MxN block of input data in units of rows to obtain N rows, which can be arranged in order to construct a 1D vector. Here, each row can have a length of M. Alternatively, if the row-major order is applied to the x-th mode, the input data can be read from the MxN block of input data in units of rows to obtain N rows, which can be arranged in order to construct a 1D vector. Here, each row can have a length of M. For a mode symmetrical to mode x, input data can be read column by column from the MxN block of input data to obtain M columns, and these can be arranged in order to form a 1D vector. Here, each column can have a length of N. If the current block is an MxN block with x mode and the symmetry described above is utilized for the current block, the NSPT set and / or the NSPT kernel of the current block can be determined based on at least one of an intra prediction mode symmetric to the x mode or a block size of NxM symmetric to the block size of MxN. Here, the NSPT kernel can be set as the NSPT kernel for the NxM block, not the NSPT kernel for the MxN block. That is, if the symmetry is utilized for the current block, the NSPT set and / or the NSPT kernel for the block having symmetry with the current block can be utilized in the same manner. As described above, a 1D vector can be constructed from an input data block according to a predetermined priority order, and this can correspond to the input of the NSPT kernel. In addition, the symmetry may be restricted to be utilized only when the value of the intra prediction mode of the current block is greater than 34. That is, when the value of the intra prediction mode of the current block is greater than 34, a transpose operation may be applied when constructing a 1D vector from an input data block, and an NSPT set or NSPT kernel corresponding to a mode and / or block shape having symmetry with the current block may be utilized. Specifically, when the intra prediction mode of the current block belongs to the [N, -1] mode and the [2, 34] mode, the symmetry may not be utilized for the current block. On the other hand, when the intra prediction mode of the current block belongs to the [35, 66] mode and the [67, 66 - N] mode, the symmetry may be utilized for the current block. Here, N may be an integer less than or equal to -14. The derivation of the NSPT set or NSPT kernel based on the above symmetry can be adaptively performed based on the size of the current block. For example, for 4x4 blocks and 8x8 blocks, the NSPT set or NSPT kernel may be derived based on symmetry, and for 4x8 blocks and 8x4 blocks, the NSPT set or NSPT kernel may not be derived based on symmetry. Depending on whether the above symmetry is utilized, the number of available NSPT sets may be different. For example, if symmetry is utilized, the number of available NSPT sets may be 35, and if symmetry is not utilized, the number of available NSPT sets may be 67. Table 9 below shows an example of how the NSPT set is determined using symmetry, and shows the mapping relationship between intra prediction modes and NSPT sets when the number of available NSPT sets is 35. Intra prediction modeNSPT set indexX < 020 ≤ X ≤ 34X35 ≤ X ≤ 6668 - XX > 662 Referring to Table 9, if the value (X) of the intra prediction mode of the current block is less than 0, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of 2 among the 35 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than or equal to 0 and less than or equal to 34, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of X among the 35 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of (68-X) among the 35 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than or equal to 35 and less than or equal to 66, the NSPT set of the current block may be identical to the NSPT set corresponding to the value (68-X) of the mode symmetrical to the intra prediction mode of the current block. Similarly, if the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block may be determined as the NSPT set having an NSPT set index of 2 among the 35 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block may be identical to the NSPT set corresponding to the mode symmetrical to the intra prediction mode of the current block. The following Table 10 shows an example in which the NSPT set is determined without using symmetry, and shows the mapping relationship between intra prediction modes and NSPT sets when the number of available NSPT sets is 67. Intra prediction modeNSPT set indexX < 020 ≤ X ≤ 66XX > 6666 Referring to Table 10, if the value (X) of the intra prediction mode of the current block is less than 0, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of 2 among the 67 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than or equal to 0 and less than or equal to 66, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of X among the 67 NSPT sets. If the value (X) of the intra prediction mode of the current block is greater than 66, the NSPT set of the current block can be determined as the NSPT set having the NSPT set index of 66 among the 67 NSPT sets. By utilizing the above symmetry, the memory size required for storing the transformation kernel can be saved while maintaining the performance according to the transformation application. For example, if 35 NSPT sets are used instead of 67 NSPT sets by utilizing the symmetry, the memory size required for storing the NSPT kernel can be significantly reduced. The number of available NSPT sets and / or the number of NSPT kernel candidates in an NSPT set may vary depending on the block size. For example, the number of available NSPT sets for a 4x4 block may be 35, the number of available NSPT sets for 4x8 blocks and 8x4 blocks may be 19, and the number of available NSPT sets for 8x8 blocks may be 10. The NSPT set for a 4x4 block may consist of three NSPT kernel candidates, the NSPT sets for 4x8 blocks and 8x4 blocks may consist of three or two NSPT kernel candidates, and the NSPT set for an 8x8 block may consist of one NSPT kernel candidate. As the block size increases, the size of the transform kernel may increase. Accordingly, by reducing the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to the NSPT set, the memory size required for storing the transform kernel can be saved. In addition, as the block size increases, the residual signal characteristics within the corresponding block tend to become more generalized. Therefore, reducing the number of available NSPT sets and / or the number of NSPT kernel candidates belonging to the NSPT set can help maintain compression efficiency while reducing the implementation complexity by reflecting these statistical characteristics. Example 4 A non-separable transform may be applied to the current block. Here, the current block may be a block encoded with inter prediction. However, the present invention is not limited thereto, and the method according to the present disclosure may be applied in the same / similar manner even when the current block is a block encoded with intra prediction. The non-separable transform may include at least one of the above-described LFNST or NSPT. For a block encoded with intra prediction, a transform set for non-separable transformation can be selected based on an intra prediction mode used for intra prediction of the block. For a block encoded with inter prediction, an intra prediction mode can be derived based on a decoder side intra mode derivation (DIMD) method, and a transform set for non-separable transformation can be selected based on the derived intra prediction mode. According to the DIMD method, a predetermined filter is applied to at least one of a prediction block or a surrounding area of a current block to derive a gradient value in a horizontal and / or vertical direction, and a specific intra prediction mode is derived based on the derived gradient value. A transform set applied to the current block can be selected based on the specific intra prediction mode. The intra prediction mode derived based on the DIMD method may be called a virtual intra prediction mode (VIPM) or a DIMD mode. Hereinafter, a method of deriving an intra prediction mode based on the DIMD method will be described in detail with reference to FIG. 6. Referring to FIG. 6, an initial sample position may be set, and accumulated intensity values for all intra prediction modes may be initialized (S600). An intensity value for a current sample position may be calculated (S610). An intra prediction mode for accumulating the calculated intensity value may be selected (S620). The intensity value calculated in S610 may be added to the accumulated intensity value for the selected intra prediction mode (S630). The processes of S610 to S630 described above may be performed for each of all or some sample positions belonging to a current block until a next sample position does not exist. If a next sample position does not exist, an intra prediction mode having the largest accumulated intensity value may be selected (S640). The process of applying the DIMD method to the surrounding area of the current block is specifically described as follows. For convenience of explanation, it is assumed that the surrounding area is a block whose width and height are W and H. However, as described above, the DIMD method may also be applied based on the prediction block of the current block, and the "surrounding area" hereinafter may be understood as being replaced with "prediction block". A filter with a size of PxQ can be applied inside a WxH block. The PxQ filter can be a 2D filter or a 1D filter. However, there may be cases where the filter goes beyond the block boundary for samples adjacent to the block boundary. The filter can be applied only to the internal area of the WxH block excluding samples adjacent to the block boundary. Specifically, the filter can be applied only to sample positions belonging to the (W-P+1)x(H-Q+1) area, which is the internal area of the WxH block. For example, when the size of the filter is 3x3, a 3x3 2D filter can be applied only to sample positions belonging to the (W-2)x(H-2) block excluding an edge of length 1 in the WxH block. In this case, the 2D filter can be applied to a 3x3 area having a sample at the sample position (hereinafter referred to as a reference sample) as a central sample. The above reference sample and at least one surrounding sample adjacent to the reference sample can be input to the 2D filter. Here, the surrounding sample can include a sample adjacent to at least one of the top, left, bottom, right, top left, top right, bottom left, or bottom right of the reference sample. A filter can be applied to each sample position belonging to the internal region to derive an intensity value of a specific intra prediction mode. The derived intensity value is added to a previously derived intensity value for the specific intra prediction mode, and through this process, an accumulated intensity value can be derived for the specific intra prediction mode. When a filter is applied to all sample positions belonging to the internal region, accumulated intensity values can be derived for all intra prediction modes (or directional modes excluding the non-directional mode). An intra prediction mode having the largest accumulated intensity value can be selected, and this can be set as an intra prediction mode derived based on the DIMD method (hereinafter, referred to as a DIMD mode). Two types of 3x3 filters, as shown in Fig. 7, can be used as filters for inducing the DIMD mode, and these are called filterY and filterX, respectively. If the position where filterY and filterX are applied (i.e., the position of the reference sample) is E, the positions of the samples related to the application of the two filters can be represented as A, B, C, D, E, F, G, H, and I. If the values obtained by applying filterY and filterX are represented as iDy and iDx, respectively, iDy and iDx can be calculated as follows. [Mathematical Formula 6] iDy = A + 2*D + G - C - 2*F - I iDx = G + 2*H + I - A - 2*B - C If iDy and iDx are 0, the subsequent steps (i.e., obtaining the intensity value, selecting a specific intra prediction mode, and adding the intensity value to the accumulated intensity value for the intra prediction mode) can be skipped and the filter can be moved to the next sample position to be applied. The function abs can be a function that calculates and returns the absolute value of the input value. The value of iAmp can be calculated as follows. This can correspond to step S610 of Fig. 6. [Mathematical formula 7] iAmp = abs(iDx) + abs(iDy) If either iDx or iDy is 0, the value of iAngUneven, which is a value of a specific intra prediction mode to be selected, can be determined as in the following mathematical expression 8. This may correspond to step S620 of FIG. 6 for the case where either iDx or iDy is 0. [Mathematical formula 8] iAngUneven = ( iDx == 0 ) ? VER_IDX : HOR_IDX In mathematical expression 8, VER_IDX and HOR_IDX can mean vertical mode and horizontal mode, respectively. For example, VER_IDX and HOR_IDX can correspond to mode 50 and mode 18 in FIG. 5, respectively. A value of iDx of 0 can mean that the amount of variation in the vertical direction is 0 or almost none. In other words, it can mean that it is considered to have the same sample values in the vertical direction, and thus prediction is good in the vertical direction. Conversely, a value of iDy of 0 (i.e., when the value of iDx is not 0) can mean that the amount of variation in the horizontal direction is 0 or almost none. In other words, it can mean that it is considered to have the same sample values in the horizontal direction, and thus prediction is good in the horizontal direction. If both iDx and iDy are not 0, the value of iAngUneven, which is the value of the particular intra prediction mode to be selected, can be determined as follows. This may correspond to step S620 of FIG. 6 for the case where both iDx and iDy are not 0. First, the values of intra prediction modes (especially, the values of directional modes) can be divided into the following four groups. The first group (region 0) can be composed of modes having a specific number of horizontal directionality. For example, the first group can be composed of modes having a value less than or equal to the horizontal mode. In Fig. 5, the modes having a horizontal directionality are modes 2 to 34 (however, mode 34 may not be included in the modes having a horizontal directionality), and the value of the horizontal mode is 18 (i.e., the horizontal mode is mode 18). When the specific number is N, the first group can be composed of modes {18, 18-1, 18-2, ... , 18-(N-1)}. When N is 17, the first group can be composed of modes {18, 17, 16, ... , 2}. The second group (region 1) can be composed of modes having a specific number of horizontal orientations. For example, the second group can be composed of modes greater than or equal to the value of the horizontal mode. If the specific number is N, the second group can be composed of modes {18, 18+1, 18+2, ... , 18+(N-1)}. If N is 17, the second group can be composed of modes {18, 19, 20, ... , 34}. The third group (region 2) can be composed of modes having a specific number of vertical directionality. For example, the third group can be composed of modes having a value less than or equal to the vertical mode. In Fig. 5, the modes having a vertical directionality are modes 34 to 66 (however, mode 34 may not be included in the modes having a vertical directionality), and the value of the vertical mode is 50 (i.e., the vertical mode is mode 50). When a specific number is N, the third group can be composed of modes {50, 50-1, 50-2, ..., 50-(N-1)}. When N is 17, the third group can be composed of modes {50, 49, 48, ..., 34}. The fourth group (region 3) can be composed of modes having a specific number of vertical orientations. For example, the fourth group can be composed of modes having a value greater than or equal to the vertical mode. If the specific number is N, the fourth group can be composed of modes {50, 50+1, 50+2, ..., 50+(N-1)}. If N is 17, the fourth group can be composed of modes {50, 51, 52, ..., 66}. When defining groups for intra prediction modes as above, an identifier indicating a specific group can be calculated as shown in Table 11 below. The identifiers corresponding to the first to fourth groups are 0, 1, 2, and 3, respectively. signx = iDx < 0 ? 1 : 0signy = iDy < 0 ? 1 : 0absx = iDx < 0 ? -iDx : iDxabsy = iDy < 0 ? -iDy : iDygtY = absx > absy ? 1: 0mapXgrY1[0][0] = 1, mapXgrY1[0][1] = 0, mapXgrY1[1][0] = 0, mapXgrY1[1][1] = 1mapXgrY0[0][0] = 2, mapXgrY0[0][1] = 3, mapXgrY0[1][0] = 3, mapXgrY0[1][1] = 2region = gtY ? mapXgrY1[signy][signx] : mapXgrY0[signy][signx] In Table 11, gtY can indicate whether the amount of variation in the vertical direction is greater than the amount of variation in the horizontal direction. That is, if absx is greater than absy, gtY can be derived as 1, and otherwise, gtY can be derived as 0. Here, if gtY is 1, this indicates that the amount of variation in the vertical direction is greater than the amount of variation in the horizontal direction, and if gtY is 0, this indicates that the amount of variation in the vertical direction is less than or equal to the amount of variation in the horizontal direction. If the amount of variation in the vertical direction is large, it may mean that the prediction in the vertical direction is less likely to be good. In this case, region 0 or region 1 may be selected. That is, mapXgrY1[signy][signx] may be selected as an identifier for the region. If the amount of variation in the horizontal direction is large, it may mean that the prediction in the horizontal direction is less likely to be good. In this case, region 2 or region 3 may be selected. That is, mapXgrY0[signy][signx] can be selected as an identifier for the region. Also, for the horizontal direction, the direction pointed by the positive change amount can be right, and the direction pointed by the negative change amount can be left. For the vertical direction, the direction pointed by the positive change amount can be downward, and the direction pointed by the negative change amount can be upward. When the sign for iDx and the sign for iDy are the same, region 1 or region 2 can be selected, and when the sign for iDx and the sign for iDy are different, region 0 or region 3 can be selected. Next, the ratio value scaled to an integer value can be obtained as shown in Table 12 below. fRatio = gtY ? ( absy / absy ) : ( absx / absy )fRatioScaled = fRatio * (1 << 16)ratio = intFunc(fRatioScaled) In Table 12, (1 << 16) means shifting 1 to the left by 16, which is 2 16 represents 65536. The intFunc function can be a function that converts fRatioScaled expressed as a decimal value to an integer value. Operations such as round, ceiling, and floor can be applied to convert to an integer value. It can also be converted to an integer value using a casting function such as int provided by the C / C++ library. In Table 12, a division operation ( / ) may be required to obtain the ratio of absy and absx. However, when implementing a codec (especially when implementing with hardware), there is a problem that the implementation cost for the division operation is expensive or the implementation is not easy. Therefore, it is often advantageous in terms of implementation to approximate the division operation as a combination of several integer operations. Therefore, the equations for obtaining the above ratio can be approximated by the equations in Table 13 below. int g_gradDivTable
[0016] = { 0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0};s0 = gtY ? absy : absxs1 = gtY ? absx : absyx = floorLog2(s1)norm = ((s1 << 4) >> x) & 15int v = g_gradDivTable[norm] | 8x = x + (norm != 0)shift = 13 - xif (shift < 0){shift = -shiftadd = (1 << (shift - 1))ratio = (s0 * v + add) >> shift}else{ratio = (s0 * v) << shift} The integerized ratio value can be determined through the process shown in Table 13. The position to which the ratio value is closest in angTable can be determined as shown in Table 14 below. int angTable
[0017] = { 0, 2048, 4096, 6144, 8192, 12288, 16384, 20480, 24576, 28672, 32768, 36864, 40960, 47104, 53248, 59392, 65536}idx = 16;for( int i = 1; i < 17; i++ ){if( ratio <= angTable[i] ){idx = ratio - angTable[i - 1] < angTable[i] - ratio ? i - 1 : i;break;}} According to Table 14, angTable can be composed of 17 entries. If each group described above is composed of 17 modes, each intra prediction mode constituting each group can correspond to an entry. In Table 14, the entry of angTable closest to the ratio value can be found, and the value of idx can be derived based on the index of angleTable corresponding to the entry. int offsets[4] = { HOR_IDX, HOR_IDX, VER_IDX, VER_IDX}int dirs[4] = { -1, 1, -1, 1}iAngUneven = offsets[region] + dirs[region] * idx Table 15 is described according to the C / C++ grammar. iAngUneven can be assigned a value of an intra prediction mode for accumulating an intensity value for a current sample position in a current block to which a filter is applied. The region derived above can be an identifier indicating one of four groups. offsets[region] can indicate a start value of an intra prediction mode for a group indicated by the region, and dirs[region] can indicate a direction in which the corresponding intra prediction mode increases or decreases. idx can indicate a position in the angleTable. Therefore, the value of iAngUneven can be determined through a formula such as Table 15. The process of adding the accumulated intensity values for a particular selected intra prediction mode can be performed as follows. [Mathematical formula 9] piHistogram[iAngUneven] += iAmp In the mathematical expression 9, the piHistogram array is an array that stores the accumulated intensity values for all intra prediction modes, and after being initialized to 0, the intensity values can be calculated by looping over the sample positions of the internal area within the current block to which the filter is applied. At this time, the intensity value calculated for the current sample position is added to the accumulated intensity value for the selected specific intra prediction mode. As described above, iAngUneven may store the value of the intra prediction mode selected for the current sample position to which the filter is applied, and iAmp may store the intensity value calculated for the corresponding sample position. piHistogram[iAngUneven] may store the accumulated intensity value for the intra prediction mode indicated by iAngUneven. The mathematical expression 9 may correspond to step S630 of FIG. 6. After looping over all the sample locations in the internal region within the block to which the filter is applied, the piHistogram array stores the cumulative intensity values for all intra prediction modes. One or more modes with the largest cumulative intensity value can be selected. The selected mode can be called a DIMD mode. This can correspond to step S640 in Fig. 6. firstAmp = 0, curAmp = 0;firstMode = 0, curMode = 0;for (int i = 0; i < NUM_LUMA_MODE; i++){curAmp = piHistogram[i];curMode = i;if (curAmp > firstAmp){firstAmp = curAmp;firstMode = curMode;}} Table 16 is described according to the C / C++ grammar. NUM_LUMA_MODE can represent the number of all intra prediction modes available as DIMD mode. For example, NUM_LUMA_MODE can be 67. According to Table 16, the value of the intra prediction mode with the largest accumulated intensity value can be assigned to the firstMode variable. Accordingly, the intra prediction mode corresponding to the value of the finally determined firstMode variable can be set as the DIMD mode for the WxH block. An encoding device and a decoding device may define a mapping table that specifies a mapping relationship between intra prediction modes and transform sets. A transform set corresponding to an intra prediction mode derived in advance from the mapping table may be selected without signaling information for selecting a transform set. The above selected transformation set may include a plurality of transformation kernel candidates. An index indicating one of the plurality of transformation kernel candidates may be signaled. The index here may be expressed as an NSPT index or an LFNST index depending on the type of the non-separable transformation. Alternatively, the index may be expressed as a transformation index. For example, the value of the index may fall in the range of 0 to N. An index of 0 may indicate that a non-separable transformation is not applied, and an index of k may indicate a k-th transformation kernel candidate ((1 ≤ k ≤ N). The number of transformation kernel candidates available to a block encoded with inter prediction (hereinafter referred to as an inter block) within a transformation set may be the same as the number of transformation kernel candidates available to a block encoded with intra prediction (hereinafter referred to as an intra block). Alternatively, the number of transformation kernel candidates available to an inter-block within the transformation set may be less than the number of transformation kernel candidates available to an intra-block. For example, the number of transformation kernel candidates available to an intra-block may be three, and the number of transformation kernel candidates available to an inter-block may be one or two. Since the inter mode has higher compression efficiency than the intra mode, the number of transformation kernel candidates for an inter-block may be relatively smaller than that for an intra-block in terms of rate-distortion efficiency. If the current block is an inter-block, the number of transformation kernel candidates available to the current block within the transformation set may be less than the number of transformation kernel candidates belonging to the transformation set. The transformation kernel candidates for the inter-block may be all or some of the transformation kernel candidates belonging to the selected transformation set. Among the N transformation kernel candidates belonging to the transformation set, M transformation kernel candidates may be available for the inter-block. Here, M may be less than or equal to N and greater than or equal to 1. The M transformation kernel candidates may be specified differently for each transformation set. For example, N transformation kernel candidates belonging to a transformation set may be respectively assigned numbers from 0 to (N-1). The M transformation kernel candidates may be transformation kernel candidates having numbers from 0 to (M-1) among the N transformation candidates. Since a transformation kernel candidate that decorrelates a statistically more frequently occurring data pattern may be assigned a smaller number, it may be advantageous in terms of coding performance to preferentially select a transformation kernel candidate having a relatively smaller number to configure the M transformation kernel candidates. Alternatively, if a transform set consists of N transform kernel candidates, the number of transform kernel candidates for the inter block may be M1, and the number of transform kernel candidates for the intra block may be M2. Here, M2 may be greater than or equal to M1 and less than or equal to N. Alternatively, M1 may be greater than or equal to M2 and less than or equal to N. Since the amount of data of the inter block is smaller than the amount of data of the intra block, setting M1 to be less than or equal to M2 may be advantageous in terms of coding performance. The smaller the number of available transform kernel candidates, the lower the cost of signaling an index. In certain cases (e.g., when applying a non-separable transform to an inter-block), only some transform kernel candidates belonging to a transform set may be mainly selected, and the selection frequency of the remaining transform kernel candidates may be relatively low. In such cases, reducing the number of available transform kernel candidates can reduce the bit rate required to signal an index while minimizing the degradation of image quality. Alternatively, the number of transform kernel candidates available to an inter-block may be greater than the number of transform kernel candidates available to an intra-block. In certain cases (e.g., when the QP value is low), the pattern of the residual data may be more complex and diverse. If the reduction in image quality degradation due to the increase in the number of transform kernel candidates is greater than the increase in the cost due to the signaling of the index, it may be more advantageous to increase the number of available transform kernel candidates. The cost of signaling an index may vary depending on the number of transformation kernel candidates. For example, it is assumed that truncated unary coding is applied as a binarization method of the index. When the number of transformation kernel candidates is L (L≥1), it can be configured to code L bins. At this time, it can be coded including the case where a non-separable transformation is not applied. For the case where a non-separable transformation is not applied, an index of 0 can be coded, and the binary code of the index of 0 can be 0. When L is 1, an index of 1 can be coded for the case where a non-separable transformation is applied, and 1 can be assigned as the binary code of the index of 1. If the number of available transformation kernel candidates decreases, the cost of signaling an index may also decrease. On the other hand, if the number of available transformation kernel candidates increases, the cost of signaling an index may also increase. The number of transform kernel candidates can be set differently based on the properties of the current block. Here, the properties can include at least one of a size, a shape, whether the block is encoded with inter prediction, or an inter prediction mode. The size can be defined as a width, a height, a product of the width and the height, a ratio of the width and the height, or a sum of the width and the height. The inter prediction mode is an inter mode pre-defined in the encoding device and the decoding device, and can be, for example, CIIP, Affine, BDOF, etc. Information about the number of transform kernel candidates for an inter block may be encoded and signaled via a bitstream. The information may be signaled at least at a high level or a low level. Here, the high level may include at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). The low level may include at least one of a slice, a tile, a coding tree unit, a coding unit, or a transform unit. The inter block may be a block encoded based on a predetermined inter prediction mode. Here, the inter prediction mode may be a mode pre-defined to an encoding device and a decoding device, and may refer to a mode for encoding based on inter prediction. Alternatively, the inter prediction mode may refer to an inter prediction mode (e.g., GPM, Affine, CIIP) based on a specific coding tool. Alternatively, if the index is signaled after a residual coding process is performed, the number of transform kernel candidates for the inter block can be determined based on information about transform coefficients obtained from the residual coding process. The residual coding process can be a process of decoding residual information of a bitstream to derive quantized transform coefficients. The information about the transform coefficients can include at least one of a sum of absolute values of transform coefficients, position information about a non-zero transform coefficient, or whether a non-zero transform coefficient exists at a sample position other than the DC position in at least one of the color component blocks (or transform blocks) constituting one coding unit. For example, the interval to which S belongs can be determined based on a comparison between the sum of the absolute values of the transform coefficients (S) and a predetermined threshold value. Based on the number of transform kernel candidates corresponding to the determined interval, the number of transform kernel candidates for the inter block can be determined. Specifically, three sections can be defined based on two thresholds. A first section smaller than or equal to a first threshold, a second section larger than the first threshold and smaller than or equal to a second threshold, and a third section larger than the second threshold can be defined, respectively. When S belongs to the first section, the number of transform kernel candidates for the inter block can be determined as 1. When S belongs to the second section, the number of transform kernel candidates for the inter block can be determined as 2. When S belongs to the third section, the number of transform kernel candidates for the inter block can be determined as 3. However, this is only an example, and two intervals may be defined based on one threshold value. In addition, the number of transformation kernel candidates per interval is intended to show that different numbers of transformation kernel candidates may be defined per interval, and a number different from the above-described example may be defined. As described above, if a certain number of transformation kernel candidates are determined to be available among transformation kernel candidates belonging to the transformation set, the certain number of candidates may be selected in ascending order of the numbers assigned to the transformation kernel candidates. For example, if M transformation kernel candidates are determined to be available, the top M transformation kernel candidates may be selected in ascending order of the numbers among the transformation kernel candidates belonging to the transformation set. Alternatively, the index may be signaled before the residual coding process is performed. In this case, the index may be signaled before the transform coefficients are acquired. For example, after the coded block flag (CBF), the transform skip flag, the position of the last valid transform coefficient in the transform block, etc. are signaled for each component constituting the coding unit, the index may be signaled, and then the absolute value and sign information of the transform coefficients may be signaled. Here, the CBF may indicate whether there is at least one non-zero transform coefficient in the transform block of the corresponding component. If the index is signaled before the residual coding process is performed, the absolute sum of the transform coefficients is not known at the time of signaling or parsing the index. In this case, the number of transform kernel candidates for the inter-block can be determined based on at least one of the aforementioned CBF, the transform skip flag, or the position of the last valid transform coefficient within the transform block. For example, the interval to which X belongs can be determined based on a comparison between a number (X) regarding the last valid transform coefficient position and a predetermined threshold value. Based on the number of transform kernel candidates corresponding to the determined interval, the number of transform kernel candidates for the inter block can be determined. Specifically, three intervals can be defined based on two thresholds. A first interval less than or equal to a first threshold, a second interval greater than the first threshold and less than or equal to a second threshold, and a third interval greater than the second threshold can be defined. When X belongs to the first interval, the number of transform kernel candidates for the inter block can be determined as 1. When X belongs to the second interval, the number of transform kernel candidates for the inter block can be determined as 2. When X belongs to the third interval, the number of transform kernel candidates for the inter block can be determined as 3. If the current block is encoded based on a non-separable transform, the positions where the last valid transform coefficient can exist in the current block can be from the first position to the Kth position in a given scan order. If a forward non-separable transform is applied to an MxN block, only R transform coefficients can be generated. Here, R can be less than (M*N). If the number for the first position starts from 0, the value of K can be (R-1). Or, if the number for the first position starts from 1, the value of K can be R. If the position of the last valid transform coefficient exists after the Kth position in the scan order, this means that the non-separable transform is not applied to the current block, and therefore, an index may not be signaled for the current block. The first threshold value and the second threshold value can be less than K. Alternatively, two intervals may be defined based on one threshold. A first interval smaller than or equal to a first threshold and a second interval larger than the first threshold may be defined. If X belongs to the first interval, the number of transform kernel candidates for the inter block may be determined as 1. If X belongs to the second interval, the number of transform kernel candidates for the inter block may be determined as 3. However, this is only an example. The number of transform kernel candidates per interval is intended to show that different numbers of transform kernel candidates may be defined per interval, and a number different from the above-described example may be defined. For example, if X belongs to the first interval, the number of transform kernel candidates for the inter block may be determined as 2. If X belongs to the second interval, the number of transform kernel candidates for the inter block may be determined as 3. Alternatively, if X belongs to the first section, the number of transformation kernel candidates for the inter block may be determined as 1. If X belongs to the second section, the number of transformation kernel candidates for the inter block may be determined as 2. As described above, if it is determined that a certain number of candidates among the transformation kernel candidates belonging to the transformation set are available, the certain number of candidates may be selected in ascending order of the numbers assigned to the transformation kernel candidates. If the current block is encoded with a single tree structure, the number of transform kernel candidates can be determined based on the CBF for at least one of the luma component or chroma component of the current block. The CBF for the luma component is referred to as CBF_Y, and the CBFs for the two chroma components are referred to as CBF_Cb and CBF_Cr, respectively. Non-separable transforms can be applied to up to seven cases, as follows. (CBF_Y, CBF_Cb, CBF_Cr) = (0, 0, 1), (0, 1, 0), (0, 1, 1), (1, 0, 0), (1, 0, 1), (1, 1, 0), (1, 1, 1) If the current block is encoded with a single tree structure, it can be configured so that a non-separable transform is applied only to the luma component of the current block. In this case, if the value of CBF_Y is not 0 (e.g., (1, 0, 0), (1, 0, 1), (1, 1, 0), (1, 1, 1)), the process of determining the number of the aforementioned transform kernel candidates can be performed. If the value of CBF_Y is 0, the index may not be signaled for the current block. Alternatively, the number of transform kernel candidates can be determined based on the values of (CBF_Cb, CBF_Cr). For example, if (CBF_Cb, CBF_Cr) is (0, 0), the number of transform kernel candidates can be set to N1. When (CBF_Cb, CBF_Cr) are (1, 0) or (0, 1), the number of transformation kernel candidates can be set to N2. When (CBF_Cb, CBF_Cr) are (1, 1), the number of transformation kernel candidates can be set to N3. The values of (N1, N2, N3) can be (1, 2, 3), (1, 1, 2), (1, 2, 2), (2, 2, 3), (2, 3, 3), (1, 1, 3), or (1, 3, 3). If the current block is encoded with a single tree structure, it can be configured so that a non-separable transform is applied to the luma component and the chroma component of the current block. In this case, the number of transform kernel candidates can be determined based on the values of (CBF_Y, CBF_Cb, CBF_Cr). For example, if only one of CBF_Y, CBF_Cb, or CBF_Cr is 1, the number of transform kernel candidates can be set to M1. If any two of CBF_Y, CBF_Cb, or CBF_Cr are 1, the number of transform kernel candidates can be set to M2. If all of CBF_Y, CBF_Cb, and CBF_Cr are 1, the number of transform kernel candidates can be set to M3. The above (M1, M2, M3) values can be (1, 2, 3), (1, 1, 2), (1, 2, 2), (2, 2, 3), (2, 3, 3), (1, 1, 3), or (1, 3, 3). Compared to an intra block encoded based on a non-separable transform, the number of transform coefficients obtained for an inter block encoded based on a non-separable transform may be smaller. Specifically, when an RxS matrix is applied as a forward non-separable transform for an intra block, R transform coefficients may be generated. Here, the RxS matrix may mean a transform matrix having an input length S and an output length R. The input length may represent the number of transform coefficients (or residual samples) input to the non-separable transform, and the output length may represent the number of transform coefficients output from the non-separable transform. On the other hand, a TxS matrix may be applied as a forward non-separable transform for an inter block, so that T transform coefficients may be generated. Here, T may be less than or equal to R. From the decoder perspective, for an intra block, the number of transform coefficients input to the reverse non-separable transform can be R. On the other hand, for an inter block, the number of transform coefficients input to the reverse non-separable transform can be T. Compared to an intra block, the number of transform coefficients input to the non-separable transform of an inter block can be smaller. The TxS matrix for the above inter block may be obtained by sampling T rows out of R rows of the RxS matrix for the intra block. For example, the TxS matrix may be constructed by sampling or extracting T rows from the top row of the RxS matrix. If a smaller number of transform coefficients are generated for the inter block based on the forward non-separable transform, the bit amount of the transform coefficients to be encoded can be reduced. If the pattern of the residual data found in the inter block can be sufficiently covered even with a smaller number of transform basis vectors, the coding performance can be improved by reducing the required bit amount while minimizing image quality degradation. In general, since the residual data of the inter block may have less information than the residual data of the intra block, it may be appropriate to apply a transform matrix composed of a smaller number of transform basis vectors to the inter block. The number of transformation sets available to an inter block may be different from the number of transformation sets available to an intra block. For example, the number of transformation sets for an inter block may be less than the number of transformation sets for an intra block. Assuming that the number of transformation sets for an intra block is 35, the number of transformation sets for an inter block may be less than 35. In this case, depending on whether the current block is an inter block, transformation sets corresponding to the same intra prediction mode may be different from each other. Alternatively, 35 transformation sets may be available for an intra block, and different transformation sets may be mapped to mode 18 and mode 19. On the other hand, if 18 transformation sets are available for an inter block, the same transformation set may be mapped to mode 18 and mode 19. However, the present invention is not limited thereto, and the number of transformation sets for an inter block may be more than the number of transformation sets for an intra block. In order to apply a non-separable transform to an inter-block, an intra-prediction mode can be derived. The derived intra-prediction mode can belong to general intra-prediction modes (e.g., modes 0 to 66 illustrated in FIG. 5). If the inter-block is a non-square block, the derived intra-prediction mode can be converted into one of the extended intra-prediction modes (e.g., modes -14 to 80 illustrated in FIG. 5). A prediction method based on such a transformation may be called Wide Angle Intra Prediction (WAIP). A transformation set can be selected based on the converted mode. Alternatively, even if the inter-block is a non-square block, a transformation set can be selected based on a previously derived intra-prediction mode (e.g., VIPM) rather than the converted mode. According to WAIP, in the case of a non-square block whose width is longer than its height, some modes having a prediction direction from the lower left to the upper right (for example, some of the modes greater than or equal to mode 2 and less than or equal to the horizontal mode) can be converted into modes having a prediction direction from the upper right to the lower left. Hereinafter, the converted modes can be named wide-angle modes. Conversely, in the case of a non-square block whose height is longer than its width, some modes having a prediction direction from the upper right to the lower left (for example, some of the modes less than or equal to mode 66 and greater than or equal to the vertical mode) can be converted into modes having a prediction direction from the lower left to the upper right. Among the wide-angle modes, the modes having a prediction direction from the upper right to the lower left can have a mode value of 67 or greater. Among the wide-angle modes, the modes having a prediction direction from the lower left to the upper right can have a mode value of -1 or less. When the above-described transformed mode is utilized to select a transform set of an inter block, the transform set can be selected by utilizing the symmetry with respect to the values of the previously derived intra prediction mode (e.g., VIPM) and the symmetry between MxN blocks and NxM blocks. Specifically, a wide-angle mode (A) having a prediction direction from the upper right to the lower left in an MxN block may be symmetrical to a wide-angle mode (B) having a prediction direction from the lower left to the upper right in an NxM block. In this case, a transformation set corresponding to the wide-angle mode (A) may be set to a transformation set corresponding to the wide-angle code (B). At this time, since blocks and modes that are symmetrical to each other are used, a transpose operation may be performed on an input data block to configure input data of a transformation matrix for a non-separable transformation. Here, the transpose operation may mean configuring a 1-D input data vector by reading data in a column-first order instead of reading data in a row-first order for one 2-D block. Alternatively, the transpose operation may mean configuring a 1-D input data vector by reading data in a row-first order instead of reading data in a column-first order for one 2-D block. Reading in row-first order can mean reading data sequentially starting from the top row. Reading in column-first order can mean reading data sequentially starting from the leftmost column. Conversely, for a wide-angle mode (A) having a prediction direction from the lower left to the upper right in an MxN block, it can be symmetrical to a wide-angle mode (B) having a prediction direction from the upper right to the lower left in an NxM block. In this case, the transformation set corresponding to the wide-angle mode (A) can be set as a transformation set corresponding to the wide-angle code (B). At this time, since blocks and modes that are symmetrical to each other are used, a transpose operation can be performed on the input data block to configure the input data of the transformation matrix for the non-separable transformation. Intra blocks and inter blocks can share the same transform sets. For example, if the intra prediction mode of an intra block and the intra prediction mode derived for an inter block are the same, the same transform set can be selected for the intra block and the inter block. Alternatively, for the inter-block, different transformation sets may be defined from the transformation sets for the intra-block. The mapping table for the inter-block may be defined separately from the mapping table for the intra-block. For example, the number of available transformation sets may be different. Or, the number of available transformation sets may be the same, but the transformation kernel candidates belonging to each transformation set may be different. Or, the number of available transformation sets and the transformation kernel candidates belonging to the transformation sets may be different. In this case, even if the intra prediction mode of the intra-block and the intra prediction mode derived for the inter-block are the same, different transformation sets (or transformation kernel candidates) may be selected for the intra-block and the inter-block. Even if the intra prediction mode of the intra-block and the intra prediction mode derived for the inter-block are the same and the transformation set of the same index is selected from the mapping table, the transformation set for the intra-block may include different transformation kernel candidates from the transformation set for the inter-block. Alternatively, for some intra prediction modes, the same set of transforms (or transform kernel candidates) may be selected for the intra block and the inter block. On the other hand, for other intra prediction modes, different sets of transforms (or transform kernel candidates) may be selected for the intra block and the inter block. For example, the same set of transforms (or transform kernel candidates) may be selected for the directional mode. On the other hand, different sets of transforms (or transform kernel candidates) may be selected for the non-directional mode. In the mapping table for the inter block, the set of transforms (or transform kernel candidates) corresponding to the directional mode may be identical to the set of transforms (or transform kernel candidates) corresponding to the directional mode in the mapping table for the intra block. On the other hand, the set of transforms (or transform kernel candidates) corresponding to the non-directional mode in the mapping table for the inter block may be different from the set of transforms (or transform kernel candidates) corresponding to the non-directional mode in the mapping table for the intra block. In the present disclosure, if a current block is an inter block, a non-separable transform may be applied to the current block, and the non-separable transform may include at least one of LFNST or NSPT. For example, both LFNST and NSPT may be configured to be applied to the inter block. Alternatively, only LFNST may be configured to be applied to the inter block. Alternatively, only NSPT may be configured to be applied to the inter block. Alternatively, whether to apply LFNST and / or NSPT to the inter block may be adaptively determined based on the size of the current block. The NSPT kernel can be configured with 8-bit precision. The range of the coefficients in the NSPT kernel can be greater than or equal to -128 and less than or equal to 127. If the precision is increased to more than 8 bits, the result obtained through matrix multiplication can be shifted to the right by the increased precision. For example, if the value obtained after matrix multiplication based on the NSPT kernel with 8-bit precision is shifted to the right by S bits and stored in the buffer, if the coefficients of the kernel are configured with N-bit precision, they can be shifted to the right by (S+(N-8)) bits and stored in the buffer. When configuring the NSPT kernel with 8-bit precision, excessive increase in internal precision within the encoder / decoder performing the conversion can be prevented, thereby minimizing the decrease in compression efficiency while reducing implementation complexity in terms of memory requirements and computational amount. When applying the reverse NSPT to the current block of the size of MxN, the size of the NSPT kernel (or NSPT matrix) can be expressed as MN xr. Here, MN can mean the product of the width and the height of the current block. This can mean the output length of the NSPT or the number of residual samples generated by the NSPT. In addition, r can mean the input length of the NSPT or the number of (inverse quantized) transform coefficients to which the NSPT is applied. r can be an integer greater than or equal to 0 and less than or equal to MN. The following is an example of the NSPT matrix of MN xr according to the block size. The NSPT matrix for a 4x4 block can be composed of a 16x16 matrix. The NSPT matrix for a 4x8 block and an 8x4 block can be composed of a 32x20 matrix, a 32x16 matrix, a 32x24 matrix, a 32x28 matrix, or a 32x32 matrix. The NSPT matrix for an 8x8 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. The NSPT matrix for a 4x16 block and a 16x4 block can be composed of a 64x16 matrix, a 64x24 matrix, a 64x32 matrix, a 64x40 matrix, a 64x48 matrix, a 64x56 matrix, or a 64x64 matrix. The NSPT matrix for 8x16 blocks and 16x8 blocks can be a 128x96 matrix, a 128x64 matrix, a 128x48 matrix, or a 128x32 matrix. The NSPT matrix for 16x16 blocks can be a 256x128 matrix, a 256x96 matrix, or a 256x64 matrix. The NSPT matrix for 16x32 blocks and 32x16 blocks can be a 512x256 matrix or a 512x128 matrix. The NSPT matrix for 32x32 blocks can be a 1024x512 matrix, a 1024x256 matrix, or a 1024x128 matrix. Alternatively, a 16x16 matrix can be applied for 4xN blocks and Nx4 blocks, where N can be an integer greater than or equal to 4. A 64x16 matrix can be applied for 8x8 blocks. A 64x32 matrix can be applied for 8xN blocks and Nx8 blocks, where N can be an integer greater than or equal to 16. A 96x32 matrix can be applied for 16xN blocks and Nx16 blocks, where N can be an integer greater than or equal to 16. Alternatively, the value of r in the NSPT matrix of the MN xr may be determined according to a predetermined criterion. Here, the criterion may be (1) ensuring that the sum of the amount of computation for the first transformation and the amount of computation for the second transformation is below a certain level, and (2) ensuring that the number of multiplications per sample required for the NSPT operation is below a certain number. Based on the reverse transformation, if a separate linear transformation is performed on an MxN block by matrix multiplication, (M+N) multiplications are required per sample to perform the linear transformation. In addition, if LFNST is applied to a certain ROI (Region-Of-Interest) area, assuming that the LFNST matrix of the reverse transformation is a PxQ matrix, (P*Q) / (M*N) multiplications are required per sample. Here, a PxQ matrix can mean a matrix in which the number of rows is P and the number of columns is Q. When applying NSPT instead of DCT-2 transform (or separable transform such as KLT) and LFNST for MxN blocks, the value of r that ensures that the number of multiplications per sample for applying NSPT is less than or equal to the number of multiplications per sample for applying DCT-2 transform and LFNST can be determined as follows. If the value of r is set to the maximum value while satisfying the above mathematical expression 10 (i.e., r = M + N + (P*Q) / (M*N)), the value of r in the NSPT matrix by block size can be set as follows. In mathematical expression 10, if the value of (P*Q) / (M*N) is not an integer, an integer close to the value of (P*Q) / (M*N) can be used. For example, a floor operation can be applied to the value of (P*Q) / (M*N). In this case, r can be set to (M + N + floor((P*Q) / (M*N))). Here, floor(x) can mean the largest integer that does not exceed x. Alternatively, a round operation can be applied to the value of (P*Q) / (M*N). In this case, r can be set to (M + N + round((P*Q) / (M*N))). Here, round(x) may mean a rounded value of x. Alternatively, a ceil operation may be applied to the value of (P*Q) / (M*N). In this case, r may be set to (M + N + ceil((P*Q) / (M*N))). Here, ceil(x) may mean the smallest integer greater than or equal to x. When a floor operation is applied, the inequality of Mathematical Expression 10 may be satisfied. However, when a round operation or a ceil operation is applied, the inequality of Mathematical Expression 10 may not be satisfied. For NSPT for 4x4 blocks, the maximum value of r is 24. However, the value of r must be less than or equal to 16, so the value of r can be set to 16. For NSPT for 4x8 blocks and 8x4 blocks, the maximum value of r is 20. The value of r can be set to 20. For NSPT for 8x8 blocks, the maximum value of r is 32. The value of r can be set to 32. For NSPT for 4x16 blocks and 16x4 blocks, the maximum value of r is 24. The value of r can be set to 24. For NSPT for 8x16 blocks and 16x8 blocks, the maximum value of r is 40. The value of r can be set to 40. For NSPT for 16x16 blocks, the maximum value of r is 44. The value of r can be set to 44. For NSPT for 16x32 blocks and 32x16 blocks, the maximum value of r is 54. The value of r can be set to 54. For NSPT on 32x32 blocks, the maximum value of r is 67. The value of r can be set to 67. For NSPT for 4x32 blocks and 32x4 blocks, the maximum value of r is 38. The value of r can be set to 38. Alternatively, the value of r can be set to 20. For NSPT for 8x32 blocks and 32x8 blocks, the maximum value of r is 48. The value of r can be set to 48. Alternatively, the value of r can be set to 24. In the NSPT matrix for each block size described above, there may be cases where the value of r is not a multiple of 4. For ease of implementation, it may be advantageous for the value of r to be a multiple of 4. For example, in implementing parallel processing through SIMD (Single Instruction Multiple Data) instructions, etc., when processing the inner product for four transformation basis vectors simultaneously (i.e., generating four transformation coefficients simultaneously) in the process of applying forward NSPT, it may be advantageous to set the value of r to be a multiple of 4. For example, for NSPT for 16x32 blocks and 32x16 blocks, the value of r can be set to 52 or 56 instead of 54. For NSPT for 32x32 blocks, the value of r can be set to 64 or 68 instead of 67. For NSPT for 4x32 blocks and 32x4 blocks, the value of r can be set to 36 or 40 instead of 38. More generally, the value of r can be set to be a multiple of K, where K can be an integer greater than or equal to 1. For example, the value of r can be set to be a multiple of K that satisfies the inequality in the following mathematical expression (11). In the above mathematical expression 11, func() can be floor, round, or ceil as described above. The value of r according to the above-mentioned predetermined criterion is for the case where zero-out is not considered. That is, when forward LFNST is applied, the primary transformed transform coefficients of the remaining areas except for the area where LFNST is applied are zeroed out, so the actual amount of calculation required to apply DCT-2 and LFNST may be less than the above-mentioned amount of calculation. Therefore, when the above-mentioned zero-out is considered, the value of r may be set to a value smaller than the value of r according to the above-mentioned predetermined criterion. Since zero-out is not performed for 4x4 blocks, the value of r can be set to a value less than or equal to 16. For a 4x8 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 8 (((4x4x8)+(4x8x4)) / (4x8)=8), and the number of sample-wise multiplications required in LFNST is 8 ((16x16) / 32=8). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 16, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 16. For an 8x4 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 6 ((4x8x4)+(4x4x4) / (8x4)=6), and the number of sample-wise multiplications required in LFNST is 8 ((16x16) / 32=8). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 14, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 14. For 8x8 blocks, zero-out may not be performed for the separable primary transform, in which case the value of r may be set to a value less than or equal to 32. For an 8x16 block, zero-out can be performed on the remaining areas except for the upper left 8x8 block based on the forward transform, and a 64x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 8x8 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 16 ((8x8x16)+(8x16x8) / (8x16)=16), and the number of sample-wise multiplications required in LFNST is 16 ((64x32) / 128=16). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 32, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 32. For a 16x8 block, zero-out can be performed on the remaining areas except for the upper left 8x8 block based on the forward transform, and a 64x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 8x8 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 12 ((8x16x8)+(8x8x8) / (16x8)=12), and the number of sample-wise multiplications required in LFNST is 16 ((64x32) / 128=16). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 28, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 28. For a 16x16 block, zero-out can be performed on the remaining areas except for the upper left 12x12 block based on the forward transform, and a 96x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 12x12 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 21 ((12x16x16)+(12x16x12) / (16x16)=21), and the number of sample-wise multiplications required in LFNST is 12 ((96x32) / 256=12). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 33, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 33. For a 4x16 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 8 (((4x4x16)+(4x16x4)) / (4x16)=8), and the number of sample-wise multiplications required in LFNST is 4 ((16x16) / 64=4). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 12, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable primary transform and LFNST, the value of r can be set to a value less than or equal to 12. For a 16x4 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 5 (((4x16x4)+(4x4x4)) / (16x4)=5), and the number of sample-wise multiplications required in LFNST is 4 ((16x16) / 64=4). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 9, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable first-order transform and LFNST, the value of r can be set to a value less than or equal to 9. For a 4x32 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 8 (((4x4x32)+(4x32x4)) / (4x32)=8), and the number of sample-wise multiplications required in LFNST is 2 ((16x16) / 128=2). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 10, which is the sum of the number of sample-wise multiplications in the separable linear transform and the number of sample-wise multiplications in the LFNST. Since the same amount of computation is required when applying the reverse separable first-order transform and LFNST, the value of r can be set to a value less than or equal to 10. For a 32x4 block, zero-out can be performed on the remaining areas except for the upper left 4x4 block based on the forward transform, and a 16x16 matrix, which is the forward LFNST matrix, can be applied to the upper left 4x4 block. When such zero-out is performed, the number of multiplications per sample required in the forward separable linear transform is 4.5 (((4x32x4)+(4x4x4)) / (32x4)=4.5), and the number of multiplications per sample required in LFNST is 2 ((16x16) / 128=2). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 6.5. Since the same amount of computation is required when applying the backward separable linear transform and LFNST, the value of r can be set to a value less than or equal to 6.5. Here, the value of r is a value that constitutes the matrix dimension, so it can be set to an integer of 6 or 7 instead of 6.5. For an 8x32 block, zero-out can be performed on the remaining areas except for the upper left 8x8 block based on the forward transform, and a 64x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 8x8 block. When such zero-out is performed, the number of multiplications per sample required in the forward separable linear transform is 16 (((8x8x32)+(8x32x8)) / (8x32)=16), and the number of multiplications per sample required in LFNST is 8 ((64x32) / 256=8). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 24. Since the same amount of computation is required when applying the backward separable linear transform and LFNST, the value of r can be set to a value less than or equal to 24. For a 32x8 block, zero-out can be performed on the remaining areas except for the upper left 8x8 block based on the forward transform, and a 64x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 8x8 block. When such zero-out is performed, the number of multiplications per sample required in the forward separable linear transform is 10 (((8x32x8)+(8x8x8)) / (32x8)=10), and the number of multiplications per sample required in LFNST is 8 ((64x32) / 256=8). Therefore, when replacing the separable linear transform and LFNST with NSPT, the value of r can be set to a value less than or equal to 18. Since the same amount of computation is required when applying the backward separable linear transform and LFNST, the value of r can be set to a value less than or equal to 18. For a 16x32 block, zero-out can be performed on the remaining areas except for the upper left 12x12 block based on the forward transform, and a 96x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 12x12 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable linear transform is 21 (((12x16x32)+(12x32x12)) / (16x32)=21), and the number of sample-wise multiplications required in LFNST is 6 ((96x32) / 512=6). Therefore, if the separable linear transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 27. Since the same amount of computation is required when applying the backward separable linear transform and LFNST, the value of r can be set to a value less than or equal to 27. For a 32x16 block, zero-out can be performed on the remaining areas except for the upper left 12x12 block based on the forward transform, and a 96x32 matrix, which is the forward LFNST matrix, can be applied to the upper left 12x12 block. When such zero-out is performed, the number of sample-wise multiplications required in the forward separable first-order transform is 13.5 (((12x32x12)+(12x16x12)) / (32x16)=13.5), and the number of sample-wise multiplications required in LFNST is 6 ((96x32) / 512=6). Therefore, if the separable first-order transform and LFNST are replaced with NSPT, the value of r can be set to a value less than or equal to 19.5. Since the same amount of computation is required when applying the reverse separable first-order transform and LFNST, the value of r can be set to a value less than or equal to 19.5. Here, the value of r is a value that constitutes the matrix dimension, so it can be set to an integer of 19 or 20 instead of 19.5. In the NSPT matrix for block sizes described above, there may be cases where the value of r is not a multiple of K. Here, K can be 4. In this case, the value of r may be set to be a multiple of K for ease of implementation. The value of r set through the method described above may be set to r prev In this case, the value of r, which is a multiple of K, can be set as follows. In the above mathematical expression 12, func() can be floor, round, or ceil as described above. As described above, the value of r in the NSPT matrix for an MxN block may be different from the value of r in the NSPT matrix for an NxM block. For example, the reverse NSPT matrix for a 4x8 block may be a 32x16 matrix, and the reverse NSPT matrix for an 8x4 block may be a 32x14 matrix. In this case, the NSPT matrix may be determined by utilizing the symmetry between the MxN block and the NxM block. Assume that the current block is an MxN block with x mode. In the case of utilizing the symmetry described above for the current block, instead of applying the NSPT matrix corresponding to the x mode or the block size of MxN, an NSPT matrix corresponding to at least one of a mode symmetric to the x mode or a block size of NxM symmetric to the block size of MxN may be applied. In this case, the NSPT matrix corresponding to the block size of NxM for the current block may be applied as is. Alternatively, the NSPT matrix corresponding to the block size of NxM may be applied, but the value of r may be the value of r in the NSPT matrix corresponding to the block size of MxN. For example, the reverse NSPT matrix for a 4x8 block may be a 32x16 matrix (i.e., the value of r in the NSPT matrix is 16), and the reverse NSPT matrix for an 8x4 block may be a 32x14 matrix (i.e., the value of r in the NSPT matrix is 14). If the current block is an 8x4 block with mode x, the NSPT matrix for at least one of a mode symmetric to mode x for the current block or a 4x8 block symmetric to the 8x4 block may be applied. In this case, the 32x16 matrix, which is the reverse NSPT matrix for the 4x8 block, may be used as is, or a 32x14 matrix having the value of r in the reverse NSPT matrix for the 8x4 block may be used. Here, the 32x14 matrix may be derived by sampling 14 rows from the left in the 32x16 matrix. In this way, when a 32x14 matrix is applied to the current block with a block size of 8x4, the aforementioned criteria are satisfied. Conversely, if the current block is a 4x8 block with mode x, then an NSPT matrix for at least one of a mode symmetric to mode x for the current block or an 8x4 block symmetric to the 4x8 block can be applied. In this case, for the current block, a 32x14 matrix for an 8x4 block can be applied instead of a 32x16 matrix for a 4x8 block. This allows NSPT to be performed using fewer multiplications than are allowed for a 4x8 block. For NSPT for MxN blocks and NxM blocks, if the values of r satisfying the aforementioned conditions are r1 and r2, respectively, the reverse NSPT matrix for MxN blocks and NxM blocks can be set to MN x max(r1, r2). Here, max(r1, r2) can mean selecting a value that is greater than or equal to r1 and r2. For example, the backward NSPT matrix for a 4x8 block can be a 32x16 matrix (i.e., the value of r in the NSPT matrix is 16), and the backward NSPT matrix for an 8x4 block can be a 32x14 matrix (i.e., the value of r in the NSPT matrix is 14). If the current block is a 4x8 block with mode x, the NSPT matrix for at least one of a mode symmetric to the x mode for the current block or an 8x4 block symmetric to the 4x8 block can be applied. In this case, the 32x16 matrix, which is the backward NSPT matrix for the 8x4 block, can be used. If the backward NSPT matrix is not constructed as MN x max(r1, r2), the 32x14 matrix will be used as the NSPT matrix for the 8x4 block. However, if the NSPT matrix for 4x8 blocks and the NSPT matrix for 8x4 blocks are configured as 32 x max(16, 14) matrices, the 32x16 matrix can be fully applied. Conversely, if the current block is an 8x4 block with mode x, the NSPT matrix for at least one of a mode symmetric to mode x for the current block or a 4x8 block symmetric to the 8x4 block may be applied. In this case, a 32x16 matrix, which is the reverse NSPT matrix for the 4x8 block, or a 32x14 matrix may be used. Here, the 32x14 matrix can be derived by sampling 14 rows from the left in the 32x16 matrix. As described above, when constructing an NSPT matrix, a transformation consisting of the maximum number of transformation basis vectors can be applied while satisfying the above-mentioned conditions, thereby maximizing coding performance. In the above-described embodiment, the value of r can be set to be a multiple of 16. For example, in the case of reverse NSPT for 4x8 blocks and 8x4 blocks, a 32x16 matrix instead of a 32x20 matrix can be applied. The transform coefficients of the transform block can be encoded in units of a predetermined coefficient group (CG). Here, a CG can be defined as a group of 16 transform coefficients, and for example, a CG can be a sub-block of a size such as 4x4, 2x8, or 8x2. There may be a case where there is no non-zero transform coefficient in a CG, and in this case, the encoding process of the transform coefficient can be skipped for the corresponding CG. Therefore, when the value of r is set to a multiple of 16, there is an advantage of reducing the complexity of the implementation. Transform coefficients can be derived by applying forward NSPT to residual samples of an MxN block. At this time, the number of derived transform coefficients can be less than or equal to the value of (M*N) due to zero-out. That is, the forward NSPT matrix can be defined as an r x (M*N) matrix, where r means the output length of NSPT or the number of transform coefficients induced through NSPT, and (M*N) means the input length of NSPT or the number of residual samples to which NSPT is applied. The above-described derived transform coefficients can be arranged in an MxN block according to a predetermined scan order, and an area where the transform coefficients are not filled can be filled with 0 (i.e., zero-out). Therefore, during the process of scanning the transform coefficients in a decoding device, if a non-zero transform coefficient is found in an area that would have been filled with 0 if NSPT had been applied (or, if the scan position of the last valid coefficient in the MxN block is greater than or equal to r), it is considered that NSPT is not applied to the corresponding MxN block, and the NSPT index may not be signaled. An index indicating a scan position of 0 may be assigned to the upper-left coefficient (i.e., the DC component coefficient) in the MxN block, and indices increased by 1 may be assigned to the remaining coefficients in the MxN block according to a predetermined order. One or more values of r can be defined for the block sizes to which NSPT can be applied. For example, one or more values of r can be defined for each of the block sizes to which NSPT can be applied. Alternatively, one value of r can be defined for each block size to which NSPT can be applied, and the value of r for one of the block sizes to which NSPT can be applied can be different from the value of r for another one. Alternatively, one value of r can be defined for some of the block sizes to which NSPT can be applied, and two or more values of r can be defined for the rest. When multiple values of r are available, an index specifying one of the multiple values of r or r itself can be signaled. The index can be signaled in a high level syntax (HLS) such as VPS, SPS, PPS, PH, SH, or can be signaled at a block level such as CTU, CU, TU. When the value of r is within a specific range, it can be signaled by allocating enough bits to encompass the range. For example, when the value of r is within the range of 1 to 256, it can be signaled by designating 8 bits as a fixed-length. The transformation kernel of the current block may be determined based on any one of the embodiments 1 to 4 described above. Alternatively, the transformation kernel of the current block may be determined based on a combination of at least two of the embodiments 1 to 4, within a range where the inventions according to the embodiments 1 to 4 described above do not conflict with each other. A transform index for the inverse transform of the current block can be signaled. Here, the transform index can specify one or more transform kernels (or transform matrices) belonging to a transform set. Here, the transform index can mean an NSPT index specifying one or more NSPT kernels belonging to an NSPT set. Alternatively, the transform index can mean an LFNST index specifying one or more LFNST kernels belonging to an LFNST set. Whether the above transformation index corresponds to the NSPT index can be determined based on whether the size of the current block is one of the block sizes belonging to the first group described above. This is based on the assumption that the block sizes to which NSPT is applicable and the block sizes to which LFNST is applicable are distinguished from each other. In this case, if the size of the current block belongs to the first group, the transformation index signaled for the current block corresponds to the NSPT index, and the NSPT kernel can be determined from the NSPT set based on the transformation index. On the other hand, if the size of the current block does not belong to the first group, the transformation index signaled for the current block corresponds to the LFNST index, and the LFNST kernel can be determined from the LFNST set based on the transformation index. If the size of the current block does not belong to the first group, this may mean that the size of the current block belongs to the second group described above. Alternatively, if the size of the current block does not belong to the first group, this may mean that the size of the current block corresponds to a block size to which LFNST is applicable among the block sizes belonging to the second group. In this way, the NSPT index and the LFNST index can be configured as a single integrated syntax rather than as separate syntaxes. For example, assume that the block sizes to which NSPT belonging to the first group is applicable are 4x4, 4x8, 8x4, and 8x8. For the block sizes belonging to the first group, NSPT can be applied instead of LFNST. Specifically, NSPT can be applied instead of the combination of separable primary transform (e.g., DCT-2, separable KLT) and LFNST. For the four block sizes belonging to the first group, the NSPT index can be signaled, and for the remaining block sizes (where LFNST is allowed), the LFNST index can be signaled. In this way, when signaling the NSPT index and the LFNST index with a single unified syntax, the amount of information to be encoded can be reduced. In addition, by applying at least one of the binarization, CABAC context, or initial value for entropy coding equally to the NSPT / LFNST index, the complexity of the implementation can also be reduced. Alternatively, the NSPT index and LFNST index can be signaled separately as separate syntaxes. In this case, the implementation complexity may increase somewhat, but compression performance can be improved by performing optimized entropy coding for each index. If the number of LFNST kernel candidates in the LFNST set is the same as the number of NSPT kernel candidates in the NSPT set, the same binarization can be applied to the LFNST index and the NSPT index. The same CABAC context (or CABAC context increment) can be assigned to the bins of the LFNST index and the NSPT index. Different binarization and / or CABAC contexts may be used for the LFNST index and the NSPT index. Different CABAC initial values may be assigned for the LFNST index and the NSPT index. For example, one of the LFNST index and the NSPT index may be binarized based on fixed-length binarization, and the other may be binarized based on truncated unary binarization. Even if the binarization of the LFNST index and the NSPT index are the same, different CABAC contexts and / or CABAC initial values may be assigned. When the number of LFNST kernel candidates belonging to the LFNST set and the number of NSPT kernel candidates belonging to the NSPT set are different, different binarization and / or CABAC contexts may be used for the LFNST index and the NSPT index. The number of NSPT kernel candidates belonging to the NSPT set may be set differently for each block size. Alternatively, the block sizes belonging to the first group may be divided into a plurality of subgroups. In this case, the number of NSPT kernel candidates belonging to the NSPT set may be set differently for each of the plurality of subgroups. Here, at least one of the plurality of subgroups may include a plurality of different block sizes. Depending on the number of NSPT kernel candidates in the NSPT set, the binarization applied to the NSPT index may be different. For example, if the number of NSPT kernel candidates in the NSPT set for a specific block size is three, the NSPT index can have any one of 0 to 3. If the value of the NSPT index is 0, this may indicate that NSPT is not applied to the current block. If the value of the NSPT index is not 0, this may indicate an NSPT kernel candidate corresponding to the NSPT index among the three NSPT kernel candidates. A bin can be assigned to distinguish between the cases where NSPT is applied and the cases where it is not. If the value of the bin is 0, it may correspond to the case where the value of the NSPT index is 0. On the other hand, if the value of the bin is 1, it may correspond to the case where the value of the NSPT index is 1, 2, or 3. In this case, truncated unary binarization can be applied to distinguish between the three NSPT kernel candidates. That is, by assigning two bins, the three NSPT kernel candidates can be distinguished as 0, 10, and 11. When the number of NSPT kernel candidates in the NSPT set for a specific block size is two, the NSPT index can have any one of 0 to 2. If the value of the NSPT index is 0, it may indicate that NSPT is not applied to the current block. If the value of the NSPT index is not 0, it may indicate an NSPT kernel candidate corresponding to the NSPT index among the two NSPT kernel candidates. One bin can be assigned to distinguish between the cases where NSPT is applied and those where it is not. The two NSPT kernel candidates can be distinguished by assigning one bin representing one of the two NSPT kernel candidates. If the number of NSPT kernel candidates in the NSPT set for a specific block size is 1, the NSPT index can have either 0 or 1. If the value of the NSPT index is 0, it can indicate that NSPT is not applied to the current block. If the value of the NSPT index is 1, it can indicate 1 NSPT kernel candidate. In this case, whether NSPT is applied and the NSPT kernel candidate can be specified with only one bin. The inverse transform of the current block may be a separable linear transform and / or LFNST-based inverse transform. That is, a reverse LFNST may be applied to all or part of the (inverse quantized) transform coefficients of the current block, and then a reverse separable linear transform may be applied to the transform coefficients derived via the LFNST to derive residual samples. For example, a reverse LFNST may be applied to (inverse quantized) transform coefficients belonging to a part of the current block. Herein, the part of the region means a part to which the forward LFNST is applied, and is referred to as a region of interest (ROI) hereinafter. The transform coefficients derived via the LFNST may be arranged in the ROI region according to a predetermined scan order. The predetermined scan order may be a row-first order or a column-first order. A reverse separable primary transform may be applied to the transform coefficients derived through the LFNST and the transform coefficients belonging to the remaining area except the ROI area within the current block. Alternatively, in the forward transform process, if zero-out is performed on the remaining area except the ROI area within the current block (i.e., if the transform coefficients within the remaining area are set to 0), a reverse separable primary transform may be applied to the transform coefficients derived through the LFNST. The size of the ROI area may be determined based on at least one of the width or the height of the current block. Here, the size of the ROI area may mean at least one of the width or the height of the ROI area, or may mean the number of sample locations belonging to the ROI area. Referring to FIG. 4, the current block can be restored based on the residual sample of the current block (S420). The current block can be restored based on the predicted block and residual block of the current block. A prediction block of the current block can be derived based on at least one of inter prediction or intra prediction. For example, the current block can be divided into multiple partitions, and a prediction block of the current block can be generated based on a prediction for each partition. The current block can be divided into a plurality of partitions based on one or more dividing lines. The dividing lines can include at least one of a vertical line and a horizontal line. Alternatively, when geometric partitioning is applied to the current block, the current block can be divided into two partitions by a predetermined dividing line. The dividing line for the geometric partitioning can be defined based on a predetermined dividing direction (or, dividing angle) and a distance from the center of the current block. The current block can be a coding block that is no longer divided through tree-based block partitioning. Assume that the current block is divided into two partitions, i.e., a first partition and a second partition. In this case, the prediction block of the current block can be generated as a weighted sum of the first prediction block for the first partition and the second prediction block for the second partition. Here, the first and second prediction blocks can be generated based on intra prediction. Alternatively, the first and second prediction blocks can be generated based on inter prediction. Alternatively, either the first prediction block or the second prediction block can be generated based on intra prediction, and the other can be generated based on inter prediction. For example, if the current block is divided into two partitions based on geometric partitioning, a prediction block for each partition can be generated based on inter prediction, and a prediction block of the current block can be generated based on a weighted sum of the generated prediction blocks. Alternatively, if the current block is divided into two partitions based on geometric partitioning, a prediction block for each partition can be generated based on intra prediction, and a prediction block of the current block can be generated based on a weighted sum of the generated prediction blocks. Hereinafter, this will be called Spatial Geometric Partitioning Mode (SGPM). When SGPM is applied to the current block, the partition type of the current block can be determined based on a partition type index that specifies a partitioning direction and a position of a partitioning line. The partition type index can indicate any one of pre-defined partition type candidates. The intra prediction mode of each partition in the current block can be derived based on a mode index that indicates any one of a plurality of intra prediction mode candidates. The mode index can be defined for each partition in the current block. The above partition type index and the mode index for each partition can be signaled via the bitstream. In this case, the partition type index can be represented as partition_mode_idx, and the mode index for each partition can be represented as intra_pred_mode0_idx and intra_pred_mode1_idx, respectively. Alternatively, the partition type index and the mode index for the first partition can be signaled via the bitstream, and the mode index for the second partition can be derived based on the mode index for the first partition. At least one of the above-described partition type indexes or mode indexes can be signaled based on a flag (cu_sgpm_flag) indicating whether SGPM is applied to the current block. For example, when cu_sgpm_flag is 1, the partition type index and the mode index can be signaled, and when cu_sgpm_flag is 0, the partition type index and the mode index can not be signaled. Alternatively, a candidate list for SGPM may be constructed. The candidate list may include a plurality of candidates, and each candidate may include one partition type index and two mode indices. For example, the plurality of candidates in the candidate list may be derived from combinations of pre-defined partition type candidates (e.g., 26 partition type candidates) and predetermined intra prediction mode candidates (e.g., 3 intra prediction mode candidates). The maximum number of candidates that can be included in the candidate list may be 16. Based on any one of the plurality of candidates, a partition type index for a current block and a mode index for each partition may be derived. A candidate index indicating any one of the plurality of candidates may be signaled through a bitstream. The above candidate list can be rearranged based on a predetermined template region. For example, a SAD between a predicted sample and a restored sample of the template region can be calculated. The SAD can be calculated for each of a plurality of candidates belonging to the candidate list. The plurality of candidates in the candidate list can be rearranged in ascending order of the SAD. The template region can include at least one of an upper peripheral region or a left peripheral region adjacent to a current block. The height of the upper peripheral region and the width of the left peripheral region can be fixed to a predetermined length (e.g., 1). The above intra prediction mode candidates can be configured in an IPM list. The IPM list can be configured for each partition in the current block. At least one of the intra prediction mode candidates belonging to the IPM list of the first partition can be different from the intra prediction mode candidates belonging to the IPM list of the second partition. Alternatively, one IPM list can be configured for the current block, and the partitions belonging to the current block can share the one IPM list. The IPM list can include three or more intra prediction mode candidates. SGPM can be applied when the size of the current block satisfies a given condition. The condition can include at least one of the following conditions 1 to 5. Here, width and height can mean the width and height of the current block, respectively. (Condition 1) 4<=width<=64 (Condition 2) 4<=height<=64 (Condition 3) width <height*8 (condition 4) height <width*8 (Condition 5) width*height>=32 A flag may be defined indicating whether blending between the prediction block of the first partition and the prediction block of the second partition for the current block is allowed. If the above flag indicates that blending between prediction blocks is allowed (e.g., if the flag is false), adaptive blending can be utilized in SGPM. The blending depth for the adaptive blending can be derived based on the size of the current block. For example, if the minimum of the width and the height of the current block is 4, the blending depth can be derived as 1 / 2τ. If the minimum of the width and the height of the current block is 8, the blending depth can be derived as τ. If the minimum of the width and the height of the current block is 16, the blending depth can be derived as 2τ. If the minimum of the width and the height of the current block is 32, the blending depth can be derived as 4τ. If the minimum of the width and the height of the current block is greater than 32, the blending depth can be derived as 8τ. Here, τ can be any integer greater than 0. On the other hand, if the flag indicates that blending between prediction blocks is not allowed (e.g., if the flag is true), the blending depth can be derived as a default value (e.g., 1 / 4τ). This indicates that blending is not used when the dividing lines of the geometric partitioning correspond to vertical or horizontal lines, and the width of the region to which blending is applied becomes relatively narrower when the dividing lines of the geometric partitioning do not correspond to vertical and horizontal lines (i.e., when the dividing lines have different dividing directions). FIG. 8 illustrates a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure. Referring to FIG. 8, a decoding device (300) according to the present disclosure may include a transform coefficient derivation unit (800), a residual sample derivation unit (810), and a restoration block generation unit (820). The transform coefficient derivation unit (800) may be configured in the entropy decoding unit (310) of FIG. 3, the residual sample derivation unit (810) may be configured in the residual processing unit (320) of FIG. 3, and the restoration block generation unit (820) may be configured in the adding unit (340) of FIG. 3. The transform coefficient derivation unit (800) can obtain residual information of the current block from the bitstream and decode it to derive the transform coefficient of the current block. The residual sample derivation unit (810) can derive a residual sample of the current block by performing at least one of inverse quantization or inverse transformation on the transform coefficient of the current block. The residual sample derivation unit (810) can determine a transformation kernel for the inverse transformation of the current block through a predetermined transformation kernel determination method, and derive the residual sample of the current block based on this. This has been described with reference to Fig. 4, and a detailed description thereof will be omitted here. The restoration block generation unit (820) can restore the current block based on the residual sample of the current block. FIG. 9 illustrates an image encoding method performed by an encoding device (200) as an embodiment according to the present disclosure. Referring to FIG. 9, residual samples of the current block can be derived (S900). The residual sample of the current block can be derived by differentiating prediction samples from original samples of the current block. Here, the prediction samples can be derived based on inter prediction or intra prediction. The current block can be divided into multiple partitions, and a prediction block of the current block can be generated based on the prediction for each partition. As discussed with reference to Fig. 4, when a current block is divided into two partitions (i.e., a first partition and a second partition), a prediction block of the current block can be generated as a weighted sum of a first prediction block for the first partition and a second prediction block for the second partition. Here, each of the first and second prediction blocks can be generated based on intra prediction or inter prediction. When the aforementioned SGPM is applied to a current block, a partition type of the current block and an intra prediction mode of each partition can be determined. A partition type index for indicating the determined partition type can be encoded in a bitstream. A mode index indicating the determined intra prediction mode among a plurality of intra prediction mode candidates can be encoded in the bitstream. The mode index can be encoded for each partition. The mode index for the second partition can be encoded based on the mode index for the first partition. At least one of the partition type index or the mode index can be encoded based on a flag (cu_sgpm_flag) indicating whether SGPM is applied to the current block. As described with reference to FIG. 4, a candidate list for SGPM of a current block may be constructed. Each candidate in the candidate list may include one partition type index and two mode indices. Based on any one of the plurality of candidates in the candidate list, a partition type index for the current block and a mode index for each partition may be derived. A candidate index indicating any one of the plurality of candidates may be encoded in a bitstream. In addition, the candidate list may be reordered based on a predetermined template region. As discussed with reference to Fig. 4, the intra prediction mode candidates can be organized into an intra prediction mode (IPM) list. SGPM can also be applied when the size of the current block satisfies a certain condition. It can be determined whether blending between the prediction block of the first partition and the prediction block of the second partition for the current block is allowed. If it is determined that blending between the prediction blocks is allowed, adaptive blending can be used in the SGPM, and a blending depth for the adaptive blending can be derived based on the size of the current block. On the other hand, if it is determined that blending between the prediction blocks is not allowed, the blending depth can be derived as a default value (e.g., 1 / 4τ). Based on the determination, a flag indicating whether blending between the prediction blocks of the first and second partitions is allowed can be encoded in the bitstream. Referring to FIG. 9, transform coefficients of the current block can be derived by performing at least one of transformation or quantization on the residual sample of the current block (S910). The transformation method according to the present disclosure can be understood as the reverse process of the inverse transformation examined with reference to Fig. 4. The method of determining the transformation kernel for the above transformation is as examined with reference to Fig. 4. A detailed description thereof will be omitted here. For example, one or more transformation sets for transformation of the current block can be defined / configured, and each transformation set can include one or more transformation kernel candidates. At this time, one from the plurality of transformation sets can be selected as the transformation set of the current block. One of the plurality of transformation kernel candidates belonging to the transformation set of the current block can be selected. The selection can be performed implicitly based on the context of the current block. Alternatively, an optimal transformation set and / or transformation kernel candidate for the current block can be selected, and an index indicating the selection can be signaled. Alternatively, the transform kernel of the current block can be determined based on the MTS set. One of the plurality of MTS sets can be selected based on at least one of the size of the current block or the intra prediction mode. The selected MTS set can include one or more transform kernel candidates. One of the one or more transform kernel candidates can be selected, and the transform kernel of the current block can be determined based on the selected transform kernel candidate. The selection of the transform kernel candidate can be performed using a transform kernel candidate index derived based on the context of the current block. Alternatively, an optimal transform kernel candidate for the current block can be selected, and a transform kernel candidate index indicating the selected transform kernel candidate can be signaled. Alternatively, the transform kernel of the current block may be determined based on a non-separable primary transform (NSPT) kernel. If the size of the current block belongs to the first group, which is a set of block sizes to which the NSPT can be applied, the forward NSPT may be applied to the current block, and if the size of the current block belongs to the second group, the forward NSPT may not be applied to the current block. If the size of the current block belongs to the second group, the forward separable primary transform (e.g., DCT-2) may be applied to the residual sample of the current block to derive transform coefficients. The forward LFNST may be additionally applied to all or part of the transform coefficients derived through the separable primary transform. In addition, the NSPT can be applied based on at least one of the tree type or component type of the current block. The NSPT kernel (or NSPT matrix) for the NSPT can be determined by utilizing the symmetry between intra prediction modes or the symmetry between block shapes. When the forward NSPT is applied to the current block of MxN, the NSPT kernel can be expressed as rx MN. Here, r means the output length of the NSPT or the number of transform coefficients generated by the NSPT, and MN can mean the input length of the NSPT or the number of residual samples to which the NSPT is applied as the product of the width and the height of the current block. The method for determining the size of the NSPT kernel is as described with reference to FIG. 4. The LFNST index and / or NSPT index for transformation may be encoded as a single integrated syntax, or the LFNST index and NSPT index may be encoded separately and inserted into the bitstream. Binarization for the LFNST index and NSPT index, and assignment of CABAC context and initial value are as described with reference to FIG. 4. A non-separable transform can be applied to the current block encoded with inter prediction (or intra prediction), as discussed with reference to Fig. 4. Additionally, we have looked at a method of signaling a transformation index with reference to FIG. 4, which can be equally applied to a method of encoding a transformation index. Referring to Fig. 9, a bitstream can be generated by encoding the transform coefficients of the current block (S920). Based on the transform coefficients of the current block, residual information about the transform coefficients can be generated, and a bitstream can be generated by encoding the residual information. FIG. 10 illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure. Referring to FIG. 10, an encoding device (200) according to the present disclosure may include a residual sample derivation unit (1000), a transform coefficient derivation unit (1010), and a transform coefficient encoding unit (1020). The residual sample derivation unit (1000) and the transform coefficient derivation unit (1010) may be configured in the residual processing unit (230) of FIG. 2, and the transform coefficient encoding unit (1020) may be configured in the entropy encoding unit (240) of FIG. 2. The residual sample derivation unit (1000) can derive a residual sample of the current block by differentiating a prediction sample from an original sample of the current block. Here, the prediction sample may be derived based on a predetermined intra prediction mode. The transform coefficient derivation unit (1010) can derive the transform coefficient of the current block by performing at least one of transform and quantization on the residual sample of the current block. The transform coefficient derivation unit (1010) can determine the transform kernel of the current block based on at least one of the above-described embodiments 1 to 4, and derive the transform coefficient by applying the transform kernel to the residual sample of the current block. The transform coefficient encoding unit (1020) can generate a bitstream by encoding the transform coefficient of the current block. In the above-described embodiments, the methods are described based on a flow chart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps depicted in the flow chart are not exclusive, and other steps may be included or one or more steps of the flow chart may be deleted without affecting the scope of the embodiments of the present document. The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or the decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc. When the embodiments in this document are implemented as software, the above-described method may be implemented as a module (process, function, etc.) that performs the above-described function. The module may be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, a logic circuit, and / or a data processing device. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing may be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm may be stored on a digital storage medium. In addition, the decoding device and the encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), a medical video device, and may be used to process a video signal or a data signal. For example, the OTT video (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), and the like. In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a media implemented in the form of a carrier wave (for example, transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. In addition, the embodiment(s) of the present specification can be implemented as a computer program product by program code, and the program code can be executed on a computer by the embodiment(s) of the present specification. The program code can be stored on a carrier readable by a computer. FIG. 11 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied. Referring to FIG. 11, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device. The encoding server compresses content input from multimedia input devices such as smartphones, cameras, camcorders, etc. into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices such as smartphones, cameras, camcorders, etc. directly generate a bitstream, the encoding server may be omitted. The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream. The above streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary that informs the user of any available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system. The above streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time. Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head mounted displays)), digital TVs, desktop computers, digital signage, etc. Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be distributedly processed. The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. In addition, the technical features of the method claims of this specification and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims of this specification and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
1. A step of deriving transform coefficients of the current block from the bitstream; A step of deriving residual samples of the current block based on inverse transformation of the transform coefficients of the current block; and A step of restoring the current block based on residual samples of the current block, The above inverse transformation is performed based on a non-separable transformation, A method wherein a set of transformations for the above non-separable transformation is selected based on a given intra prediction mode for the current block.
2. In paragraph 1, A method wherein the number of available transformation kernel candidates in the above transformation set is determined based on whether the current block is an inter-block.
3. In paragraph 1, A method in which the number of available transformation kernel candidates within the transformation set is determined based on information signaled via a bitstream when the current block is an inter-block.
4. In paragraph 1, A method wherein the number of available transformation kernel candidates in the above transformation set is determined based on information about the transformation coefficients.
5. In paragraph 1, The number of available transformation kernel candidates in the transformation set is determined based on at least one of a flag for at least one of a luma component or a chroma component of the current block or the position of the last valid transformation coefficient in the current block, A method wherein the above flag indicates whether at least one non-zero transform coefficient exists.
6. In paragraph 1, A method in which the number of transform coefficients input to the above non-separable transform is determined based on whether the current block is an inter block.
7. In paragraph 1, A method wherein the above transformation set is selected from among pre-defined transformation sets.
8. In paragraph 7, A method in which transformation sets for inter blocks and transformation sets for intra blocks are respectively defined.
9. In paragraph 7, The above-described pre-defined transformation sets are applied equally to inter-blocks and intra-blocks.
10. In paragraph 1, A method wherein the intra prediction mode for the current block is converted to an extended intra prediction mode based on whether the current block is a non-square block.
11. Step of deriving residual samples of the current block; A step of deriving transform coefficients of the current block based on transforms of residual samples of the current block; and Including a step of encoding the transform coefficients of the current block, The above transformation is performed based on non-separable transformation, A method wherein a set of transformations for the above non-separable transformation is selected based on a given intra prediction mode for the current block.
12. A computer-readable storage medium storing a bitstream generated by the method according to Article 11.
13. A step of obtaining a bitstream for image information; wherein the bitstream is obtained based on a step of deriving residual samples of a current block, a step of deriving transform coefficients of the current block based on a transform of the residual samples of the current block, and a step of encoding transform coefficients of the current block, and Including a step of transmitting data including the above bitstream, The above transformation is performed based on non-separable transformation, A method wherein a set of transformations for the above non-separable transformation is selected based on a given intra prediction mode for the current block.
Citation Information
Patent Citations
Card wallet for easy withdrawal of cards
KR1020230052132A
Method and apparatus for control of vehicle
KR1020230110041A
Mobile phone elevator control system for the convenience of the disabled
KR1020250001107A
Military Important Facilities Vigilance Operation System
KR102633616B1
Brazier for charcoal fire roasting
KR102684525B1