Transform-based image coding method, and device therefor

The transform-based image/video coding method addresses the challenge of high data size in high-resolution media by using Multiple Transform Selection to enhance compression efficiency, reducing transmission and storage costs.

WO2025150797A1PCT designated stage expired Publication Date: 2025-07-17LX SEMICON CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000183
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-03
Filing Date
2025-01-03
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images and immersive media has led to a rise in data size, resulting in higher transmission and storage costs, necessitating a more efficient image/video compression technology.

Method used

A transform-based image/video coding method and device that employs Multiple Transform Selection (MTS) to determine the most efficient transform pair for each block, improving compression efficiency by selecting from a set of transform candidates based on intra prediction modes and signaling transformation-related information effectively.

Benefits of technology

This approach enhances video/image compression efficiency by optimizing transformation performance and reducing data size, thereby lowering transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000183_17072025_PF_FP_ABST
    Figure KR2025000183_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method performed by a decoding device, according to the present disclosure, comprises the steps of: receiving image information including residual information for the current block; deriving transform coefficients for the current block on the basis of the residual information; deriving residual samples for the current block by performing inverse transform on the basis of the transform coefficients; and generating reconstructed samples for the current block on the basis of the residual samples, wherein the image information includes multiple transform selection (MTS)-related information, one transform pair among transform pair candidates of a transform set for the current block is determined on the basis of the MTS-related information, and the number of transform pair candidates of the transform set is one or more.
Need to check novelty before this filing date? Find Prior Art

Description

Transform-based image coding method and device thereof The present disclosure relates to an image / video coding method and a device therefor. Image / video coding is used in various applications such as digital storage media, television broadcasting, video streaming services, and real-time communications, and recently, the demand for high-resolution, high-quality images / videos is increasing in various fields. As the image / video becomes higher resolution and higher quality, the data size of the image / video increases, and the amount of information or bits transmitted relatively increases. Therefore, when transmitting image data using media such as existing wired / wireless broadband lines or storing image / video data using existing storage media, the transmission cost and storage cost increase. In addition, interest in and demand for immersive media such as VR (virtual reality), AR (artificial reality), MR (mixed reality) content and holograms have been increasing recently, and attempts to provide immersive experiences using immersive media in games, education, medicine, real estate, marketing, etc. are increasing. Accordingly, a highly efficient image / video compression technology is required to effectively compress, transmit, store, and play back information of high-resolution, high-quality images / videos with various characteristics as described above. According to one embodiment of the present disclosure, a method and device for improving video / image coding efficiency are provided. According to one embodiment of the present disclosure, a transform-based video / image coding method and device are provided. According to one embodiment of the present disclosure, a video decoding method performed by a decoding device is provided. The method includes the steps of receiving video information including residual information for a current block, deriving transform coefficients for the current block based on the residual information, deriving residual samples for the current block by performing inverse transform based on the transform coefficients, and generating restoration samples for the current block based on the residual samples, wherein the video information includes MTS (Multiple Transform Selection) related information, and one transform pair among transform pair candidates of a transform set for the current block is determined based on the MTS related information, and the number of candidates of the transform pair of the transform set is at least one. According to one embodiment of the present disclosure, a video encoding method performed by an encoding device is provided. The method includes a step of deriving prediction samples for a current block, a step of deriving residual samples for the current block based on the prediction samples, a step of deriving transform coefficients for the current block by performing a transform based on the residual samples, a step of generating residual information for the current block based on the transform coefficients, and a step of encoding image information including the residual information, wherein the image information includes MTS (Multiple Transform Selection) related information, and one transform pair among transform pair candidates of a transform set for the current block is determined based on the MTS related information, and the number of candidates of the transform pair of the transform set is at least one. According to one embodiment of the present disclosure, a decoding device for image decoding is provided. The decoding device includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform the steps of: receiving image information including residual information for a current block, deriving transform coefficients for the current block based on the residual information, deriving residual samples for the current block by performing inverse transform based on the transform coefficients, and generating restoration samples for the current block based on the residual samples, wherein the image information includes MTS (Multiple Transform Selection) related information, and one transform pair among transform pair candidates of a transform set for the current block is determined based on the MTS related information, and the number of candidates of the transform pair of the transform set is at least one. According to one embodiment of the present disclosure, an encoding device for video encoding is provided. The encoding device includes a memory and at least one processor connected to the memory, and the at least one processor is configured to perform a step of deriving prediction samples for a current block, a step of deriving residual samples for the current block based on the prediction samples, a step of deriving transform coefficients for the current block by performing a transform based on the residual samples, a step of generating residual information for the current block based on the transform coefficients, and a step of encoding image information including the residual information, wherein the image information includes MTS (Multiple Transform Selection) related information, and one transform pair among transform pair candidates of a transform set for the current block is determined based on the MTS related information, and the number of candidates of the transform pair of the transform set is at least one. According to one embodiment of the present disclosure, a method of transmitting video / video data including a bitstream generated by a video / video encoding method according to at least one of the embodiments of the present disclosure is provided. According to one embodiment of the present disclosure, a device is provided for transmitting video / video data including a bitstream generated by a video / video encoding method according to at least one of the embodiments of the present disclosure. According to one embodiment of the present disclosure, a computer-readable storage medium having stored thereon a program for performing a method according to at least one of the embodiments of the present disclosure may be provided. According to one embodiment of the present disclosure, a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method according to at least one of the embodiments of the present disclosure is provided. According to one embodiment of the present disclosure, a computer-readable digital storage medium is provided having encoded information or encoded video / image information stored thereon, which causes a decoding device to perform a video / image decoding method according to at least one of the embodiments of the present disclosure. According to one embodiment of the present disclosure, the overall video / image compression efficiency can be improved. According to one embodiment of the present disclosure, the transformation performance for a current block can be improved. According to one embodiment of the present disclosure, transformation related information can be efficiently signaled. According to one embodiment of the present disclosure, the efficiency of transformation can be improved by determining an MTS transformation set based on various intra prediction modes. According to one embodiment of the present disclosure, the efficiency of transformation can be improved by determining a transformation set of a current block based on an intra prediction mode derived based on a template. According to one embodiment of the present disclosure, a transformation pair can be determined by efficiently signaling transformation-related information based on the number of transformation sets, thereby increasing the efficiency of transformation. Figure 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure can be applied. FIG. 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied. FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied. Figure 4 illustrates an intra prediction procedure as an example. Figure 5 shows examples of intra prediction based video / image encoding methods. Figure 6 shows examples of intra prediction based video / image decoding methods. Figure 7 shows examples of video / image encoding methods based on residual processing. Figure 8 shows examples of residual processing-based video / image decoding methods. FIG. 9 illustrates an example of a multiple transformation technique according to the present disclosure. Figure 10 is a diagram illustrating an example of LFNST. Figure 11 shows an example of NSPT kernels according to block size. Figure 12 shows an example of ROI for LFNST 16. Figure 13 shows an example of ROI for LFNST 8. Figure 14 shows an example of a MIP prediction sample for constructing HoG. Figure 15 illustrates an example of context modeling for LFNST and NSPT transform coefficients. Figure 16 shows an example of deriving an MTS set. Figure 17 shows an example of deriving DIMD-based intra modes for MTS and LFNST. FIG. 18 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure. FIG. 19 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure. This disclosure is susceptible to various modifications and embodiments, and thus specific embodiments are illustrated and described in detail by way of illustration in the drawings. However, this is not intended to limit the embodiments of the disclosure to the specific embodiments. The terminology used herein is only used to describe specific embodiments, and is not intended to be limiting of the technical idea of the disclosure. The singular forms "a" and "an" as used herein are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term "and / or" as used herein includes any one or a combination of two or more of the associated listed items. The terms "comprises," "consists," and "contains" as used herein specify the presence of stated features, numbers, operations, elements, components, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, elements, components, and / or combinations thereof. The use of the term "can" or "may" in connection with an example or embodiment (e.g., what the example or embodiment can include or implement) in this disclosure means that there is at least one example or embodiment that includes or implements such feature, but not all examples are limited thereto and such feature or configuration may be omitted. Meanwhile, each component in the drawings described in the present disclosure is independently depicted for the convenience of explaining different characteristic functions, and does not mean that each component is implemented with separate hardware or separate software. For example, two or more components among each component may be combined to form one component, or one component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included in the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure. In this disclosure, "A or B" can mean "only A", "only B", or "both A and B". In other words, "A or B" in this disclosure can be interpreted as "A and / or B". For example, "A, B or C" in this disclosure can mean "only A", "only B", "only C", or "any combination of A, B and C". The slash ( / ) or comma used in this disclosure can mean "and / or". For example, "A / B" can mean "A and / or B". Accordingly, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C". In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Additionally, in this disclosure, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted identically to "at least one of A and B". Additionally, in the present disclosure, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.” In addition, parentheses used in the present disclosure may mean "for example". Specifically, when it is indicated as "prediction (intra prediction)", "intra prediction" may be suggested as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" may be suggested as an example of "prediction". In addition, even when it is indicated as "prediction (i.e., intra prediction)", "intra prediction" may be suggested as an example of "prediction". Technical features individually described in a single drawing in this disclosure may be implemented individually or simultaneously. The present disclosure relates to video / image coding. For example, the method / embodiment described in the present disclosure can be applied to a method disclosed in the enhanced compression model (ECM) or H.267 standard. In addition, the method / embodiment disclosed in the present disclosure can be applied to a method disclosed in the AV2 (AOMedia Video 2) standard, or a next-generation video / image coding standard (e.g. H.268, H.269, etc.). In the present disclosure, coding may include encoding and / or decoding. In the present disclosure, image coding may be interchangeably used with video coding. In the present disclosure, a video may mean a collection of a series of images over time. A picture generally means a unit representing one image of a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A tile may represent a rectangular area of CTUs within a specific tile row and a specific tile column within a picture. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture. A pixel or pel can mean the smallest unit that constitutes a picture (or image). Also, a 'sample' can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (ex. cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows. Hereinafter, with reference to the attached drawings, embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals may be used for the same components in the drawings, and redundant descriptions of the same components may be omitted. Figure 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure can be applied. Referring to FIG. 1, a video / image coding system may include a first device (an encoding device) and a second device (a decoding device). The first device may transmit encoded video / image information or data to the second device through a digital storage medium or a network in the form of a file or streaming. The video / image coding system may further include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or may be configured as a separate device or an external component. The video / image renderer may be included in the decoding device, or may be configured as a separate device or an external component. The first device may include the transmission unit as an internal component, or as a separate device or external component. The second device may include the receiver as an internal component, or may include it as a separate device or external component. The encoder may be referred to as an encoding device, and the decoder may be referred to as a decoding device. The transmitting unit may be included in the encoding device. The receiving unit may be included in the decoding device. The renderer may include a display unit, and the display unit may be comprised of a separate device or an external component. The decoding device and the encoding device to which the embodiment(s) of the present disclosure are applied may be included in a multimedia broadcasting transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), a medical video device, and may be used to process a video signal or a data signal. For example, the OTT video (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), and the like. A video / image capture device can capture a video / image source. The video / image capture device can capture the video / image through a process of capturing, synthesizing, or generating the video / image. The video / image capture device can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device can include, for example, a camcorder, a computer, a tablet, a smart phone, etc., and can (electronically) generate the video / image. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data. The video / image source can also perform a video / image preprocessing process in order to input an optimized video / image to an encoder. The encoding device can encode input video / image. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream. The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the reception unit of the receiving device through the network in the form of a file or streaming. The encoded video / image information or data output in the form of a bitstream can also be transmitted to the reception unit through a streaming server. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file through a predetermined file format and can include an element for transmission through a broadcasting / communication network. The reception unit can receive / extract the bitstream and transmit it to a decoding device. The streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium that informs the user of any available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system. The above streaming server can receive content from a media storage and / or an encoding device. For example, when receiving content from the encoding device, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time. The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device. The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit. FIG. 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied. The encoding device hereinafter may include an image encoding device and / or a video encoding device. Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit and an intra prediction unit. The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processor (230) may further include a subtractor (subtractor) 231. The addition unit (250) may be called a reconstructor or a reconstructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. In addition, the memory (270) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component. The image segmentation unit (210) can segment an input image (or, picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit can further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can be divided or partitioned from the final coding unit described above, respectively.The above prediction unit may be a unit for sample prediction, and the above transformation unit may be a unit for deriving a transformation coefficient and / or a unit for deriving a residual signal from the transformation coefficient. The term unit may be used interchangeably with terms such as block or area, depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image). The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from a prediction unit from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input image signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied per current block or CU unit. The prediction unit can generate various information about prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information about prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream. The intra prediction unit can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In the intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail of the prediction direction. However, this is only an example, and a number of directional prediction modes more or less than that may be used depending on the setting. The intra prediction unit may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block. An inter prediction unit can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in an inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. A reference picture including the reference block and a reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit may configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of the skip mode and the merge mode, the inter prediction unit may use the motion information of the neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of one block, and can also apply intra prediction and inter prediction at the same time. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for SCC (screen content coding), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loe've Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit (240) may encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The above bitstream can be transmitted through a network or stored in a digital storage medium. Here, the network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (240) can be configured as an internal / external element of the encoding device (200) such as a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit can be included in the entropy encoding unit (240). The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit. When there is no residual for the target block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (250) can be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next target block in the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process. The filtering unit (260) can apply filtering to the restoration signal to improve subjective / objective picture quality. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate information about the filtering and transmit it to the entropy encoding unit (240). The information about the filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream. The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit. Through this, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device when inter prediction is applied, and can also improve encoding efficiency. The memory (270) DPB can store the modified restored picture to be used as a reference picture in the inter prediction unit. The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit. FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied. The decoding device hereinafter may include an image decoding device and / or a video decoding device. Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter prediction unit and an intra prediction unit. The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), adding unit (340), and filtering unit (350) described above may be configured by one hardware component (e.g., decoder chipset or processor) according to an embodiment. In addition, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component. When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device. The decoding device (300) can receive a signal output from the encoding device in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model by using information on a syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in a previous step, and predicts an occurrence probability of a bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after the context model is determined. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (330), and residual values on which entropy decoding is performed by the entropy decoding unit (310), that is, quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present disclosure may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the adding unit (340), the filtering unit (350), the memory (360), and the prediction unit (330). The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients. In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array). The prediction unit can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information about the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode. The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of one block, and can also apply intra prediction and inter prediction at the same time. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for SCC (screen content coding), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure. The intra prediction unit can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In the intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block. The inter prediction unit can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can configure a motion information candidate list based on neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating a mode of inter prediction for the current block. The addition unit (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330). When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block. The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process. The filtering unit (350) can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit (350) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and transmit the modified restoration picture to the memory (360), specifically, the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit. The memory (360) can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit. In the present disclosure, the embodiments described in the filtering unit (260) and the prediction unit (220) of the encoding device (200) can be applied identically or correspondingly to the filtering unit (350) and the prediction unit (330) of the decoding device (300), respectively. As described above, prediction is performed in order to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase image coding efficiency by signaling information (residual information) about a residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks. The residual information can be generated through a transformation and quantization procedure. For example, the encoding device can derive a residual block between the original block and the predicted block, perform a transformation procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as value information of the quantized transform coefficients, position information, transformation technique, transformation kernel, and quantization parameter. The decoding device can perform an inverse quantization / inverse transformation procedure based on the residual information to derive residual samples (or residual block). The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inversely quantize / inversely transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a restored picture based on the residual block. In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. If the quantization / dequantization is omitted, the quantized transform coefficient may be called a transform coefficient. If the transform / inverse transform is omitted, the transform coefficient may be called a coefficient or a residual coefficient, or may still be called a transform coefficient for the sake of uniformity of expression. In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or the information about the transform coefficient(s)), and scaled transform coefficients may be derived through inverse transformation (scaling) on the transform coefficients. Residual samples may be derived based on inverse transformation (transformation) on the scaled transform coefficients. This may be similarly applied / expressed in other parts of the present disclosure. Intra prediction may represent a prediction that generates prediction samples for a current block based on reference samples in a picture to which the current block belongs (hereinafter, the current picture). When intra prediction is applied to the current block, peripheral reference samples to be used for intra prediction of the current block may be derived. The peripheral reference samples of the current block may include H+W samples located at the left of a current block of a size WХH, W+H samples located at the top of the current block, and at least one sample neighboring the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples of multiple rows and left peripheral samples of multiple columns. Some of the surrounding reference samples of the current block may not be decoded yet or may not be available. In this case, the decoder can construct surrounding reference samples to be used for prediction by padding or substituting the unavailable samples with the available samples. When the surrounding reference samples are derived, the prediction samples of the current block can be derived based on the surrounding reference samples and the intra prediction mode / type information. Here, the intra prediction mode can represent one of the non-directional prediction modes and the directional prediction modes that indicate spatial correlation for intra prediction. Here, the directional prediction mode can be called an angular prediction mode, and the non-directional prediction mode can be called a non-angular prediction mode. The intra prediction type can represent various prediction types for performing intra prediction. The intra prediction type may include, for example, MRL (multi-reference line), ISP (intra sub-partitions), PDPC (Position dependent intra prediction), MIP (matrix weighted intra prediction or matrix based intra prediction), CCLM (cross-component linear model), MMLM (multi-model linear model), DIMD (Decoder side intra mode derivation), fusion of chroma intra prediction modes, intra template matching, TIMD (fusion for template-based intra mode derivation), intra prediction fusion, CCCM (cross-component convolutional model), CCP (cross-component prediction), SGPM (spatial geometric partitioning mode), etc. In some cases, the intra prediction mode and / or the intra prediction type may be used to perform intra prediction. Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. In addition, a post-filtering step may be performed on the derived prediction sample as needed. Figure 4 illustrates an intra prediction procedure as an example. Referring to FIG. 4, the intra prediction procedure as described above may include an intra prediction mode / type determination step, a reference sample derivation step, and an intra prediction performance (prediction sample generation) step. The intra prediction procedure may be performed in an encoding device and a decoding device as described above. The coding device determines the intra prediction mode / type (S400). The coding device may include an encoding device and / or a decoding device as described above. An encoding device can determine an intra prediction mode / type applied to the current block from among various intra prediction modes / types described in the present disclosure, and can generate prediction-related information. The prediction-related information can include intra prediction mode information indicating an intra prediction mode applied to the current block and / or intra prediction type information indicating an intra prediction type applied to the current block. A decoding device can determine an intra prediction mode / type applied to the current block based on the prediction-related information. For example, when intra prediction is applied, the intra prediction mode to be applied to the current block can be determined using the intra prediction mode of the surrounding block. For example, the coding device can select one of the MPM candidates in the MPM (most probable mode) list derived based on the intra prediction mode of the surrounding block (e.g., the left and / or upper surrounding block) of the current block and / or additional candidate modes based on the received index information, or can select one of the remaining intra prediction modes not included in the MPM candidates based on MPM remind information (remaining intra prediction mode information). The MPM list can be configured to include or not include the planar mode as a candidate. The coding device can construct an MPM (most probable modes) list for the current block. The MPM list may also be expressed as an MPM candidate list. Here, MPM may mean a mode used to improve coding efficiency by considering the similarity between the current block and surrounding blocks during intra prediction mode coding. The encoding device can perform prediction based on various intra prediction modes, and determine an optimal intra prediction mode based on RDO (rate-distortion optimization) based on the prediction. In this case, the encoding device can determine the optimal intra prediction mode using MPM candidates configured in the MPM list, or can determine the optimal intra prediction mode using the remaining intra prediction modes other than the MPM candidates configured in the MPM list. Specifically, for example, if the intra prediction type of the current block is not a normal intra prediction type but a specific type (e.g., DIMD, TIMD, MRL or ISP), the encoding device can determine the optimal intra prediction mode by considering only the MPM candidates as intra prediction mode candidates for the current block. That is, in this case, the intra prediction mode for the current block can be determined only among the MPM candidates, and in this case, the MPM flag may not be encoded / signaled. In this case, the decoding device can estimate that the MPM flag is 1 without being separately signaled the MPM flag. Meanwhile, in general, if the intra prediction mode of the current block is not a planar mode but one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) pointing to one of the MPM candidates. If the intra prediction mode of the current block is not in the MPM list either, the encoding device generates MPM reminder information (remaining intra prediction mode information) pointing to a mode same as the intra prediction mode of the current block among the remaining intra prediction modes not included in the MPM list (and the planar mode). The MPM reminder information may include, for example, an intra_luma_mpm_remainder syntax element. A decoding device obtains intra prediction mode information from a bitstream. The intra prediction mode information may include at least one of an MPM flag, an MPM index, and MPM remind information (remaining intra prediction mode information) as described above. The decoding device may configure an MPM list. The MPM list is configured in the same manner as the MPM list configured in the encoding device. That is, the MPM list may include intra prediction modes of surrounding blocks, and may further include specific intra prediction modes according to a predetermined method. The decoding device can determine the intra prediction mode of the current block based on the MPM list and the intra prediction mode information. For example, when the value of the MPM flag is 1, the decoding device can derive a candidate indicated by the MPM index among the MPM candidates in the MPM list as the intra prediction mode of the current block. As another example, when the value of the MPM flag is 0, the decoding device can derive the intra prediction mode indicated by the remaining intra prediction mode information (which may be called mpm remainder information) from among the remaining intra prediction modes as the intra prediction mode of the current block. The coding device derives reference samples of the current block (S410). The reference samples may include peripheral reference samples of the current block. The peripheral reference samples of the current block may include H+W samples located at the left of a current block of a size WХH, W+H samples located at the top of the current block, and at least one sample neighboring the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples of multiple rows and left peripheral samples of multiple columns. The coding device performs intra prediction on the current block to derive prediction samples (S1220). The coding device can derive the prediction samples based on the intra prediction mode / type and the reference samples. The coding device can derive a reference sample according to the intra prediction mode of the current block among the reference samples of the current block, and derive a prediction sample of the current block based on the reference sample. An encoding procedure based on intra prediction may roughly include, for example: Figure 5 shows examples of intra prediction based video / image encoding methods. Referring to FIG. 5, S500 may be performed by a prediction unit of an encoding device, S505 may be performed by a residual processing unit of the encoding device, and S510 or S515 may be performed by an entropy encoding unit of the encoding device. Specifically, the prediction-related information may be derived by the prediction unit and encoded by the entropy encoding unit. The residual information may be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients through the transform unit of the encoding device, and the transform coefficients may be derived as quantized transform coefficients through the quantization unit. The information about the quantized transform coefficients may be encoded in the entropy encoding unit through a residual coding procedure. The encoding device performs intra prediction on the current block (S500). The encoding device can derive an intra prediction mode / type for the current block, derive reference samples of the current block, and generate prediction samples in the current block based on the intra prediction mode / type and the reference samples. Here, the intra prediction mode / type determination, surrounding reference sample derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedure. The encoding device can determine a mode / type applied to the current block among a plurality of intra prediction modes / types. The encoding device can compare RD costs for the intra prediction modes / types and determine an optimal intra prediction mode / type for the current block. Meanwhile, the encoding device may also perform a prediction sample filtering procedure. The prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted. The encoding device generates residual samples for the current block based on the prediction samples (S505). The encoding device can compare the prediction samples with the original samples of the current block based on phase and derive the residual samples. The encoding device can encode image / video information including information about the intra prediction (prediction-related information) and / or information about the residual samples (residual information) (S510 or S515). The prediction-related information can include the intra prediction mode information and the intra prediction type information. The encoding device can output the encoded image / video information in the form of a bitstream. The output bitstream can be transmitted to a decoding device via a storage medium or a network. The residual information may include a residual coding syntax element. An encoding device may transform / quantize the residual samples to derive quantized transform coefficients. The residual information may include information about the quantized transform coefficients. Meanwhile, as described above, the encoding device can generate a restored picture (including restored samples and restored blocks). To this end, the encoding device can dequantize / inversely transform the quantized transform coefficients again to derive (corrected) residual samples. The reason for performing dequantization / inversely transform on the residual samples after transforming / quantizing them in this way is to derive residual samples that are identical to the residual samples derived from the decoding device as described above. The encoding device can generate a restored block including restored samples for the current block based on the prediction samples and the (corrected) residual samples. A restored picture for the current picture can be generated based on the restored block. As described above, an in-loop filtering procedure, etc. can be further applied to the restored picture. The decoding device can perform operations corresponding to the operations performed in the encoding device. A video / image decoding procedure based on intra prediction may include, for example, the following. Figure 6 shows examples of intra prediction based video / image decoding methods. Referring to FIG. 6, S600 may be performed by an entropy decoding unit of a decoding device, S610 may be performed by a prediction unit of the decoding device, S615 may be performed by a residual processing unit of the decoding device, and S620 may be performed by an adder or restoration unit of the decoding device. Specifically, the decoding device obtains image / video information from the bitstream (S600). The image / video information may include prediction-related information and / or residual information. The decoding device performs intra prediction based on prediction-related information (S610). The decoding device may derive an intra prediction mode / type for a current block based on the prediction-related information, derive reference samples of the current block, and generate prediction samples in the current block based on the intra prediction mode / type and the reference samples. In this case, the decoding device may perform a prediction sample filtering procedure. The prediction sample filtering may be referred to as post filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted. The decoding device performs residual processing based on the residual information (S615). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit of the residual processing unit can perform inverse transformation on the transform coefficients to derive residual samples for the current block. The decoding device generates a restored block / picture (S620). The decoding device can generate restored samples for the current block based on the prediction samples and / or the residual samples, and derive a restored block including the restored samples. A restored picture for the current picture can be generated based on the restored block. As described above, an in-loop filtering procedure, etc. can be further applied to the restored picture. The intra prediction mode information and / or the intra prediction type information can be encoded / decoded through the binarization and coding method described in the present disclosure. For example, the intra prediction mode information and / or the intra prediction type information can be binarized through fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, the intra prediction mode information and / or the intra prediction type information can be encoded / decoded through entropy coding (ex. CABAC, CAVLC) coding. Below, we explain Decoder Side Intra Mode Derivation (DIMD), a method of intra prediction. According to the present disclosure, an intra mode can be derived based on the DIMD technique. In DIMD, a Histogram of Gradient (HoG) can be calculated using a template including surrounding samples of a current block. For example, the HoG can be calculated using a horizontal / vertical Sobel filter based on a three-line template of surrounding samples. In this case, horizontal and vertical Sobel filters can be applied to calculate the HoG. If the template is located in a different CTU, the Sobel filter may not be applied to the upper CTU boundary, or the DIMD technique may not be applied to the current block. For example, for candidate intra modes, n intra modes having the highest histogram value can be selected / extracted through the HoG calculation. In this case, the prediction block for the current block can be derived using the n intra modes. Also, in this case, the prediction block can be derived using the n intra modes and a non-directional mode (e.g., planar mode or DC mode). In this case, a predictor according to each intra mode can be derived, and the prediction block can be derived by weighting the predictors / weighted average. In other words, the predictors of the five selected / extracted intra modes and the predictor of the planar mode can be fused through prediction fusion. Below, we describe Template-based Intra Mode Derivation (TIMD), a method of intra prediction. According to the present disclosure, an intra mode can be derived based on a fusion of template-based intra mode derivation (TIMD) technique. The TIMD technique may be referred to as a TIMD type. In TIMD, a template including surrounding samples of a current block is derived, a predictor of the template according to a candidate prediction mode is derived based on reference samples of the template (i.e., predicted samples of the template are derived), and this is compared with a reconstructed template (i.e., reconstructed samples of the template), so as to derive an optimal intra prediction mode. The template may include an upper template and a left template. Here, the size of the template may be L. However, this is only an example, and the height L2 of the upper template and the width L1 of the left template may vary depending on the size of the current block, etc. For example, if the current block is a non-square block whose width is greater than its height, L2 may be larger than L1. For example, if the current block is a non-square block whose height is greater than its width, L1 may be larger than L2. TIMD candidate intra prediction modes can be determined based on various criteria. For example, TIMD candidate intra prediction modes can be limited to only MPMs. For example, TIMD candidate intra prediction modes can be limited to only candidate modes in the MPM (or PMPM) list. In this case, for example, the timd flag can be signaled after the mpm flag (or pmpm flag) signaling. As another example, predetermined candidate intra prediction modes can be defined. The predetermined candidate intra prediction modes can include some or all of the above-described DIMD derivation modes. The predetermined candidate intra prediction modes can include a default mode. The default mode can include, for example, at least one of a vertical mode, a horizontal mode, a left-down diagonal mode, a left-up diagonal mode, or a right-up diagonal mode. Meanwhile, the TIMD derivation mode can derive only one intra prediction mode when a specific condition is satisfied. For example, if the SATD of the intra prediction mode having the minimum SAD or SATD calculated based on the TIMD is smaller than a threshold value, only one intra prediction mode can be derived. As another example, if the difference between the first SATD of the first intra prediction mode having the minimum SATD calculated based on the TIMD and the second SATD of the second intra prediction mode having the second minimum SATD is larger than a threshold value, only one intra prediction mode can be derived. Below, we describe intra template matching (intraTMP), a method of intra prediction. According to the present disclosure, intraTMP can be applied. IntraTMP can be regarded as one of prediction techniques or prediction types. According to intraTMP, prediction of a corresponding block can be performed by finding a template most similar to a current template (a surrounding reference sample (L-shape) template) within a predefined search range of a reconstructed part within a current picture. In this case, a block vector indicating a position of a matching block (reference block) within the current picture can be derived / stored based on the current block position. In addition, the predefined search range can be located within a left CTU, a current CTU, an upper-left CTU, an upper-right CTU, and an upper-right CTU. For example, intraTMP can be performed in the same way in the encoder / decoder, and the encoder can signal whether intraTMP is used. Meanwhile, in order to perform intraTMP efficiently, a candidate list for intraTMP can be constructed. The candidate list can include up to n (e.g. 19) template matching block vectors, and in this case, they can be sorted in ascending order based on the SAD cost for the template. For example, intra-template matching can be enabled when the CU size height and width are less than or equal to 64. The maximum enabled size can be predefined or signaled at a higher level. In addition, whether intra-template matching (intraTMP) is enabled can be signaled via a CU level flag. For example, it can be signaled when the DIMD mode is not enabled for the current block. That is, intratmp_flag can be signaled when dimd_flag is 0. In addition, a block vector (BV) derived via intraTMP can be stored and used in an IBC candidate list for IBC of subsequent blocks. Below, we describe SGPM (spatial geometric partitioning mode), a method of intra prediction. According to the present disclosure, SGPM can be applied. SGPM can be regarded as one of prediction techniques or prediction types. According to SGPM, a current block can be divided into two partitions (geometric partitions), and different intra predictions can be applied to the two partitions. For example, a first intra prediction mode can be applied to a first partition, and a second intra prediction mode can be applied to a second partition, and the two predictors derived through this can be fused or blended with each other to generate one predicted block. At this time, n partition modes can be used for SGPM, and for example, 26 partition modes can be used. In addition, n partition modes and m intra prediction modes can be combined to form an SGPM candidate list of length k. Also, for example, the length k of the SGPM candidate list can be 16. For example, n can be 26 and m can be 3. The above k can use a predefined value or can be set differently based on the size of the current block. Below, residual processing is described. The residual processing can be performed in an encoding device and a decoding device. The residual processing can include coefficient coding, transform, and / or quantization procedures. The residual processing procedure in the encoding section may include a procedure for generating and / or encoding residual information from residual samples for the derived current block. The residual processing procedure may further include a procedure for deriving the residual samples based on prediction samples. The residual processing procedure in the decoding section may include a procedure for deriving residual samples from residual information of a received bitstream. For example, the residual processing procedure may include a (inverse) transformation and / or a (inverse) quantization procedure. In addition, the residual processing procedure may include an encoding / decoding procedure of the residual information. The residual information may include residual data and / or transformation / quantization related parameters. Specifically, for residual processing, a method is provided for deriving (quantized) transform coefficients within a block and performing generation and (en)coding of residual information based on the same, and for deriving (quantized) residual coefficients within a block to which transform skip is applied and performing generation and (en)coding of residual information for transform skip based on the same. The encoded information can be output in the form of a bitstream as described above. In addition, in the case of the decoding stage, (quantized) transform coefficients or (quantized) residual coefficients within the block can be derived from the residual information included in the bitstream (or residual information for transform skip), and (if necessary) inverse quantization / inverse transformation can be performed to derive residual samples. An encoding procedure based on residual processing may roughly include, for example: Figure 7 shows examples of video / image encoding methods based on residual processing. Referring to FIG. 7, S700 may be performed by a prediction unit of an encoding device, S710 may be performed by a residual processing unit of an encoding device, and S720 may be performed by an entropy encoding unit of an encoding device. Specifically, the prediction-related information may be derived by the prediction unit and encoded by the entropy encoding unit. The residual information may be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients through the transform unit of the encoding device, and the transform coefficients may be derived as quantized transform coefficients through the quantization unit. The information about the quantized transform coefficients may be encoded in the entropy encoding unit through a residual coding procedure. The encoding device derives prediction samples of the current block (S700). The encoding device can derive prediction samples of the current block based on the above-described inter prediction and / or intra prediction. The encoding device can perform residual processing based on the prediction samples (S710). The encoding device can derive residual samples based on the prediction samples. The encoding device can derive the residual samples by comparing the original samples of the current block with the prediction samples. The residual processing includes a transformation and / or quantization process for the residual sample, as described above. The encoding device can generate residual information from the residual samples through the residual processing. The residual information can include information about quantized transform coefficients, as described above. An encoding device encodes image information including prediction-related information and / or residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction-related information can include information related to the prediction procedure. The residual information is information about the residual samples. The residual information can include information about quantized transform coefficients for the residual samples. The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network. Meanwhile, as described above, the encoding device can generate a restored picture (including restored samples and restored blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as performed by the decoding device from the encoding device, and thereby increase coding efficiency. Accordingly, the encoding device can store the restored picture (or restored samples, restored blocks) in memory and utilize it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the restored picture. The decoding device can perform operations corresponding to the operations performed in the encoding device. A video / image decoding procedure based on residual processing may include, for example, the following. Figure 8 shows examples of residual processing-based video / image decoding methods. Referring to FIG. 8, S800 may be performed by an entropy decoding unit of a decoding device, S810 may be performed by a prediction unit of the decoding device, S820 may be performed by a residual processing unit of the decoding device, and S830 may be performed by an adder or restoration unit of the decoding device. Specifically, the decoding device obtains image / video information from the bitstream (S800). The image / video information may include prediction-related information and / or residual information. The decoding device performs prediction (including inter prediction and / or intra prediction) based on prediction-related information (S810). The decoding device may derive a prediction mode / type for a current block based on the prediction-related information, and generate prediction samples in the current block based on the prediction mode / type. In this case, the decoding device may perform a prediction sample filtering procedure. The prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted. The decoding device performs residual processing based on the residual information (S820). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit of the residual processing unit can perform inverse transformation on the transform coefficients to derive residual samples for the current block. The decoding device generates a restored block / picture (S830). The decoding device can generate restored samples for the current block based on the prediction samples and / or the residual samples, and derive a restored block including the restored samples. A restored picture for the current picture can be generated based on the restored block. As described above, an in-loop filtering procedure, etc. can be further applied to the restored picture. The above residual information can be encoded / decoded through the binarization and coding method described in the present disclosure. For example, the above residual information can be binarized through fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, the above residual information can be encoded / decoded through entropy coding (ex. CABAC, CAVLC) or bypass coding. Below we describe the maximum transform size and zeroing out of the transform coefficients. For example, the CTU size and maximum transform size (all MTS kernels) can be extended up to 256. In this case, the maximum size of the intra prediction block can be set to 128X128. In UHD sequences, the maximum CTU size can be set to 256, and otherwise, it can be set to 128. In the first transformation process, zeroing out of the transformation coefficients may not be applied. However, in the second transformation to which LFNST is applied, the first transformation coefficients outside the ROI area to which LFNST is applied may be zeroed out. Meanwhile, zeroing out of the transform coefficients may be applied in the first transformation process. As the size of the transform block increases, zeroing out may be required after the first transformation, and whether or not to zero out may be determined depending on whether MTS is applied. For example, when MTS is applied, zeroing out may be performed to any one of 16, 20, 24, 32, or 40, and when MTS is not applied, zeroing out may be performed to a preset specific value. The specific value may be 32 or 64. When MTS is applied, the number of transform coefficients to be zeroed out may vary depending on the size of the transform block. For example, if the width or height to which MTS is applied in the transform block is 4, 8, or 16 (if it is a 4-point transformation, an 8-point transformation, or a 16-point transformation), zeroing out is not applied, and if the width or height to which MTS is applied in the transform block is 32 (if it is a 32-point transformation), zeroing out may be performed to 16. Or, if MTS is applied to a width or height greater than 32, it may be zeroed out to 32. Additionally, the residual processing may include a transform / inverse transform and / or a quantization / inverse quantization process. According to the present disclosure, a first transform and / or a second transform may be applied to a residual block to derive a transform coefficient block (transform coefficients), and an inverse second transform and / or an inverse first transform may be applied to the transform coefficient block (transform coefficients) to derive a residual block. As described above, the transformation for the residual can be performed via a primary transform and / or a secondary transform that is optionally performed after the primary transform. The primary transform can be called a primary transform and can be a DCT (Discrete Cosine Transform) and a DST (Discrete Sine Transform) that are applied to all rows and all columns of the residual block. After the primary transform, a secondary transform can be additionally applied to a specific transform coefficient at the upper left of the transform block according to the result of the primary transform. The inverse transform performed at the decoding stage can be applied to a specific residual at the upper left of the residual block corresponding to the (inverse quantized) residual information by applying an inverse secondary transform, and an inverse primary transform can be applied to a transform block according to the result of the inverse secondary transform. Meanwhile, as described later, NSPT transformation may be used in some cases. The NSPT transformation may be a transformation that integrates and replaces the first transformation and the second transformation. Below, we explain MTS (Multiple Transform Selection), which is one method of transformation. FIG. 9 illustrates an example of a multiple transformation technique according to the present disclosure. Referring to FIG. 9, the conversion unit may correspond to the conversion unit in the encoding device of FIG. 2 described above, and the inverse conversion unit may correspond to the inverse conversion unit in the encoding device of FIG. 2 described above or the inverse conversion unit in the decoding device of FIG. 3. The transform unit can perform a primary transform based on residual samples (residual sample array) in the residual block to derive (primary) transform coefficients (S900). This primary transform may be referred to as a core transform. Here, the primary transform may be based on multiple transform selection (MTS), and when multiple transforms are applied as the primary transform, it may be referred to as a multiple core transform. The transformation unit can perform a secondary transformation based on the (first) transformation coefficients to derive modified (second) transformation coefficients (S910). The first transformation is a transformation from a spatial domain to a frequency domain, and the second transformation can be expressed as a transformation into a more compact expression by utilizing the correlation existing between the (first) transformation coefficients. For example, the secondary transform may include a non-separable transform. In this case, the secondary transform may be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent a transform that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by second-orderly transforming the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix. Here, the vertical transform and the horizontal transform may not be separately applied to the (primary) transform coefficients based on the non-separable transform matrix (or the horizontal transform and the vertical transform independently), but may be applied at once. That is, the non-separable second-order transform can refer to a transform method that does not separate the vertical and horizontal components of the (first-order) transform coefficients, but rearranges, for example, two-dimensional signals (transform coefficients) into one-dimensional signals along a specific predetermined direction, and then generates modified transform coefficients (or second-order transform coefficients) based on the non-separable transform matrix. In other words, it can refer to a transform method that rearranges into one-dimensional signals along a row-first direction or a column-first direction, and then generates modified transform coefficients (or second-order transform coefficients) based on the non-separable transform matrix. In addition, the inverse transform unit can perform a series of procedures in the reverse order of the procedures performed in the above-described transform unit. The inverse transform unit can receive (inverse quantized) transform coefficients, perform a second (inverse) transform to derive (first) transform coefficients (S920), and perform a first (inverse) transform on the (first) transform coefficients to obtain a residual block (residual samples) (S930). Here, the first transform coefficients can be called modified transform coefficients on the inverse transform unit side. As described above, the encoding device and / or the decoding device can generate a restoration block based on the residual block and the predicted block, and can generate a restoration picture based on the same. In the present disclosure, the primary transform may be referred to as a core transform. Here, the primary transform may be based on multiple transform selection (MTS), and when a transform kernel selected from among a plurality of transform kernel types is applied as the primary transform, it may be referred to as a multi-core transform. The multi-core transform can indicate a method of transforming using Discrete Cosine Transform (DCT) 2 and Discrete Sine Transform (DST) 7, DCT 8, etc. That is, the multi-core transform can indicate a transform method of transforming a residual signal (or residual block) of a spatial domain into transform coefficients (or 1st transform coefficients) of a frequency domain based on a plurality of transform kernels selected from among the DCT 2, DST 7, DCT 8, and DST 1. Here, the 1st transform coefficients can be called temporary transform coefficients from the perspective of a transform unit. In other words, when a conventional transform method is applied, a transformation from a spatial domain to a frequency domain can be applied to a residual signal (or a residual block) based on DCT 2, so that transform coefficients can be generated. In contrast, when a multi-core transform is applied, a transformation from a spatial domain to a frequency domain can be applied to a residual signal (or a residual block) based on DCT 2, DST 7, DCT 8, and / or DST 1, so that transform coefficients (or first-order transform coefficients) can be generated. Here, DCT 2, DST 7, DCT 8, and DST 1, etc. may be called a transform type, a transform kernel, or a transform core. These DCT / DST transform types can be defined based on basis functions. When a multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for a target block (a block to be transformed) may be selected from among transform kernels, and a vertical transform may be performed for the target block based on the vertical transform kernel, and a horizontal transform may be performed for the target block based on the horizontal transform kernel. Here, the horizontal transform may represent a transform for horizontal components of the target block, and the vertical transform may represent a transform for vertical components of the target block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a prediction mode and / or a transform index of a target block (CU or sub-block) including a residual block. In addition, according to an example, when performing the first transformation by applying MTS, specific basis functions can be set to predetermined values, and a mapping relationship for the transformation kernel can be set by combining which basis functions are applied when it is a vertical transformation or a horizontal transformation. For example, when a horizontal transformation kernel is represented as trTypeHor, and a vertical transformation kernel is represented as trTypeVer, a trTypeHor or trTypeVer value of 0 can be set to DCT2, a trTypeHor or trTypeVer value of 1 can be set to DST7, and a trTypeHor or trTypeVer value of 2 can be set to DCT8. In this case, MTS index information may be encoded and signaled to the decoding device to indicate which of a plurality of sets of transform kernels. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that trTypeHor value is 2 and trTypeVer value is 1, an MTS index of 3 may indicate that trTypeHor value is 1 and trTypeVer value is 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2. Such MTS may be explicitly applied by signaling as described above, or implicitly applied according to specific conditions. Implicit MTS may independently derive a transform kernel for each direction according to the width or height of a transform block. For example, when a sub-block transform is applied that applies a transform to only one of the sub-blocks, an implicit MTS may be applied if the width or height satisfies a specific condition, or if the width or height satisfies a condition for a specific intra mode. For example, an implicit MTS may be applied to a block to which SBT is applied and the larger value of the width or height is 32 or 64 or less. Alternatively, an implicit MTS may be applied to a block to which ISP is applied, or to a block to which intra prediction is applied but LFNST and MIP are not applied. In addition, implicit MTS can infer primary transform pairs based on the intra prediction mode of the current block and the size of the TU using the LUT. For example, the same intra prediction mode as explicit MTC can be considered, and the maximum TU size can be 32x32. In this case, the LUT can include transform pairs based on separable primary transforms, and no additional primary transforms can be added. For example, DCT 2, DCT 5, DCT 8, DST 1, DST 4, DST 7 can be used. For example, for an ISP block, the sizeIdx used as input to the LUT can be based on the location of the current ISP subpartition within the ISP block. In addition, as an example, the implicit MTS can be applied only to the ISP block depending on the CTC configuration settings. In another example, the implicit MTS can be applied to the angular intra prediction mode. For example, the implicit MTS for the ISP block can be used by using implicit signaling for all modes (non-CTC). For example, the default implicit MTS (i.e., based on the shape of TU) can be maintained in TIMD and DIMD modes. In addition, DCT 2 can be used for MIP, EIP, SGPM, and IntraTMP modes. In addition, the default implicit MTS can be used for the ISP block. Below, we explain LFNST (Low Frequency Non Separable Transform), which is one of the transformation methods. Meanwhile, the transform unit of the encoding device can perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients. Here, the primary transform is a transform from a spatial domain to a frequency domain, and the secondary transform means transforming into a more compact representation by utilizing the correlation existing between the (primary) transform coefficients. The secondary transform may include a non-separable transform. In this case, the secondary transform may be called a non-separable secondary transform (NSST). The non-separable secondary transform may represent a transform that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix. Here, based on the non-separable transform matrix, the vertical transform and the horizontal transform can be applied at once without separately applying the vertical transform (or the horizontal-vertical transform independently) to the (1st) transform coefficients. In other words, the non-separable second-order transform can represent a transform method of rearranging, for example, two-dimensional signals (transform coefficients) into one-dimensional signals in a specific predetermined direction (e.g., row-first direction or column-first direction) without separating the vertical components and horizontal components of the (1st) transform coefficients, and then generating the modified transform coefficients (or second-order transform coefficients) based on the non-separable transform matrix. For example, the row-first order is to arrange them in a row in the order of the 1st row, the 2nd row, ..., the Nth row for an MxN block, and the column-first order is to arrange them in a row in the order of the 1st column, the 2nd column, ..., the Mth column for an MxN block.The above non-separable second-order transform can be applied to the top-left region of a block composed of (first-order) transform coefficients (hereinafter, referred to as a transform coefficient block). The non-separable second-order transform can be selected based on a mode (mode dependent) in which a transform kernel (or transform core, transform type) is selected. Here, the mode can include an intra-prediction mode and / or an inter-prediction mode. The non-separable second-order transform can be performed based on an 8Х8 transform or a 4Х4 transform determined based on the width (W) and height (H) of a transform coefficient block. The 8x8 transform refers to a transform that can be applied to an 8x8 region included inside a transform coefficient block when both W and H are greater than or equal to 8, and the 8x8 region can be the upper left 8x8 region inside the transform coefficient block. Similarly, the 4x4 transform refers to a transform that can be applied to a 4x4 region included inside a transform coefficient block when both W and H are greater than or equal to 4, and the 4x4 region can be the upper left 4x4 region inside the transform coefficient block. For example, an 8x8 transform kernel matrix can be a 64x64 / 16x64 matrix, and a 4x4 transform kernel matrix can be a 16x16 / 8x16 matrix. Meanwhile, for mode-based transformation kernel selection, a transformation set for non-separable second-order transformation can be set, and k non-separable second-order transformation kernels can be configured per transformation set. For example, the transformation sets can be 4, 35, etc. Selection of a specific set among the transformation sets can be performed based on, for example, an intra prediction mode of a target block (CU or sub-block). For example, if it is determined that a specific set is used for a non-separable transform, one of the k transform kernels in the specific set can be selected via a non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index pointing to the specific transform kernel based on a rate-distortion (RD) check, and signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. The inverse transform unit of the encoding device and the decoding device can perform a series of procedures in the reverse order of the procedures performed in the above-described transform unit. The inverse transform unit can receive (inverse quantized) transform coefficients and perform a second (inverse) transform to derive (first) transform coefficients. The first transform coefficients can be called modified transform coefficients from the standpoint of the inverse transform unit. In the present disclosure, in order to reduce the amount of computation and memory requirement accompanying a non-separable secondary transform, a reduced secondary transform (RST) in which the size of a transform matrix (kernel) is reduced can be applied in the concept of NSST. The RST may be referred to by various terms such as reduced transform, reduced secondary transform, reduction transform, simplified transform, simple transform, etc., and the names by which the RST may be referred to are not limited to the listed examples. Alternatively, since the RST is mainly performed in a low-frequency region including non-zero coefficients in a transform block, it may also be referred to as a low-frequency non-separable transform (LFNST). In LFNST, an N-dimensional vector can be mapped to an R-dimensional vector located in another space to determine a reduced transformation matrix, where R is smaller than N. N can mean the square of the length of one side of a block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value. As an example, the size of the LFNST matrix is RxN, which is smaller than the size NxN of a typical transformation matrix, and can be defined as in the following mathematical expression 1. When the LFNST matrix TRxN is multiplied by the residual samples for the target block of the transformation, the transformation coefficients for the target block can be derived. When the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), LFNST can be expressed by a matrix operation as shown in the mathematical expression 2 below. In mathematical expression 2, r1 to r64 may represent residual samples for the target block, and more specifically, may be transform coefficients generated by applying the primary transform. As a result of the operation of mathematical expression 2, transform coefficients ci for the target block may be derived. ci may be derived from c1 to cR. That is, when R=16, transform coefficients c1 to c16 for the target block may be derived. If a regular transformation, not LFNST, had been applied and a transformation matrix of size 64x64 (NxN) had been multiplied to the residual samples of size 64x1 (Nx1), 64 (N) transformation coefficients for the target block would have been derived. However, since LFNST was applied, only 16 (R) transformation coefficients for the target block are derived. Since the total number of transformation coefficients for the target block is reduced from N to R, the amount of data transmitted from an encoding device to a decoding device is reduced, so that the transmission efficiency between the encoding device and the decoding device can be increased. The size of the inverse LFNST matrix TNxR is NxR, which is smaller than the size NxN of a normal inverse transform matrix, and is in a transpose relationship with the LFNST matrix TRxN shown in mathematical expression 1. Tt may mean the inverse LFNST matrix TRxNT (the superscript T means transpose). When the inverse LFNST matrix TRxNT is multiplied by the transform coefficients for the target block, modified transform coefficients for the target block or residual samples for the target block can be derived. The inverse LFNST matrix TRxNT may also be expressed as (TRxN)TNxR. For example, if the size of the block to which the inverse transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the inverse LFNST can be expressed as a matrix operation as in the following mathematical expression 3. c1 to c16 may represent transform coefficients for the target block. As a result of the operation of mathematical expression 3, rj representing modified transform coefficients for the target block or residual samples for the target block may be derived. rj may be derived from r1 to rN. That is, when N=16, transform coefficients r1 to r64 for the target block may be derived. Figure 10 is a diagram illustrating an example of LFNST. Referring to Fig. 10, 4x4 LFNST can be applied to blocks with min (width, height) < 8, and 8x8 LFNST can be applied to blocks with min (width, height) > 4. For example, 16 (first-order) transform coefficients can be input for 4x4 forward LFNST, and 64 (first-order) transform coefficients can be input for 8x8 forward LFNST. Since 8 or 16 transform coefficients can be derived through these forward LFNSTs, respectively, 8 or 16 transform coefficients can be input for the input of the inverse LFNST, respectively. When 4x4 inverse LFNST is performed, 16 modified transform coefficients can be output from 8 transform coefficients, and when 8x8 inverse LFNST is performed, 64 modified transform coefficients can be output from 16 transform coefficients. Meanwhile, for example, in the transformation of the encoding process, instead of the 16 x 64 transformation kernel matrix for the 64 data constituting the 8 x 8 area, a maximum 16 x 48 transformation kernel matrix can be applied by selecting only 48 data. Here, "maximum" means that the maximum value of m is 16 for the m x 48 transformation kernel matrix that can generate m coefficients. That is, when LFNST is performed by applying the m x 48 transformation kernel matrix (m ≤ 16) to the 8 x 8 area, 48 data can be input and m coefficients can be generated. When m is 16, 48 data are input and 16 coefficients are generated. That is, when 48 data form a 48 x 1 vector, a 16 x 1 vector can be generated by sequentially multiplying the 16 x 48 matrix and the 48 x 1 vector. At this time, 48 data forming an 8 x 8 area can be appropriately arranged to form a 48 x 1 vector. At this time, if a matrix operation is performed by applying a maximum 16 x 48 transformation kernel matrix, 16 modified transformation coefficients are generated. The 16 modified transformation coefficients can be arranged in the upper left 4 x 4 area according to the scanning order, and the upper right 4 x 4 area and the lower left 4 x 4 area can be filled with 0. The transposed matrix of the transformation kernel matrix described above can be used for the inverse transformation of the decoding process. That is, when the inverse LFNST is performed as the inverse transformation process performed in the decoding device, the input coefficient data to which the inverse LFNST is to be applied is configured as a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse LFNST matrix from the left can be arranged in a two-dimensional block according to a predetermined arrangement order. The matrix operation in this case can be expressed as (48 x 16 matrix) * (16x1 transformation coefficient vector) = (48 x 1 modified transformation coefficient vector). Here, since the nx1 vector can be interpreted in the same meaning as an nx1 matrix, it can also be expressed as an nx1 column vector. * indicates a matrix multiplication operation. When these matrix operations are performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left, upper right, and lower left regions excluding the lower right region of the 8x8 region. Meanwhile, LFNST can be applied to sub-blocks. For example, when a coding block is divided into sub-blocks, the size of the sub-blocks can be "width / 2 x height / 2" or "width / 4 x height / 4". In this case, the sub-blocks can include sub-blocks located at corners or third blocks located at the center. For example, LFNST can be applied to sub-blocks located at corners. In this case, the division mode and the location of non-zero sub-blocks can be signaled similarly to the existing SBT. Below, we describe the Non-separable primary transform (NSPT) for transformation. For example, as described above, a separable transform that applies a transform kernel to each of the vertical and horizontal directions as a first transform or an inverse first transform is applied, and a non-separable transform, LFNST, may be applied as a second transform or an inverse second transform, and if MTS is applied as a first transform and an inverse first transform, LFNST may not be applied. Meanwhile, when MTS is not applied and (DCT 2, DCT 2) is applied as the first transform and LFNST is applied, a non-separable primary transform can be applied as the first transform. That is, a non-separable primary transform is performed as the first transform, and no additional secondary transform may be performed. This transform may be applied only to the luma component, or to both the luma and chroma components. Or, considering that the non-separable transform is applied to the intra prediction block, it may be applied only to the intra prediction block. In addition, NSPT may be applied conditionally according to a specific tree structure. For example, in case of dual tree chroma, coded_flag for the chroma component is 0 or transform_skip_flag[x0][y0]

[0001] and transform_skip_flag[x0][y0]

[0002] NSPT can be performed if is 0. In addition, in case of dual tree luma, if coded_flag for luma component is 0 or transform_skip_flag[x0][y0]

[0000] is 0, NSPT can be performed. In addition, in case of single tree structure, if coded_flag for luma component is 0 or transform_skip_flag[x0][y0]

[0000] is 0, and coded_flag for chroma component is 0 or transform_skip_flag[x0][y0]

[0001] and transform_skip_flag[x0][y0]

[0002] are 0, NSPT can be performed. That is, in case of dual tree chroma, chroma block must not be transform skip, and in case of dual tree luma, luma block must not be transform skip, NSPT can be applied. Also, in the case of a single tree, LFNST can be applied only to the luma component, and in this case, NSPT can be applied only if not all blocks are transform skips. Meanwhile, for NSPT, similar to LFNST, 35 or 67 transformation sets can be applied, and 3 transformation candidates can be applied for each transformation set. The transformation kernel can be derived based on the size of the block to which NSPT is applied. Figure 11 shows an example of NSPT kernels according to block size. Referring to FIG. 11, NSPT kernels may include NSPT4x4 applied to 4x4, NSPT4x8 applied to 4x8, NSPT8x4 applied to 8x4, NSPT8x8 applied to 8x8, NSPT4x16 applied to 4x16, NSPT16x4 applied to 16x4, NSPT8x16 applied to 8x16, NSPT16x8 applied to 16x8, NSPT4x32 applied to 4x32, NSPT32x4 applied to 32x4, NSPT8x32 applied to 8x32, and NSPT32x8 applied to 32x8. Specifically, (a) of Fig. 11 may represent NSPT applied to a 4XN / NX4 block, (b) of Fig. 11 may represent NSPT applied to an 8XN / NX8 block, and (c) of Fig. 11 may represent a transform applied to a 16XN / NX16 block. As described above, in the 4XN / NX4 block, the dimensions of the NSPT kernel may be as follows. - NSPT4x4: 16x16 - NSPT4x8 / NSPT8x4: 32x20 - NSPT8x8: 64x32 - NSPT4x16 / NSPT16x4: 64x24 - NSPT8x16 / NSPT16x8: 128x40 - NSPT4x32 / NSPT32x4: 128x20 - NSPT8x32 / NSPT32x8: 256x24 For example, when applying NSPT4x8 / NSPT8x4, 12 transform coefficients may be zeroed out, when applying NSPT8x8, 32 transform coefficients may be zeroed out, when applying NSPT4x16 / NSPT16x4, 40 transform coefficients may be zeroed out, when applying NSPT8x16 / NSPT16x8, 88 transform coefficients may be zeroed out, when applying NSPT4x32 / NSPT32x4, 108 transform coefficients may be zeroed out, and when applying NSPT8x32 / NSPT32x8, 232 transform coefficients may be zeroed out. As shown in Figure 11 and described above, NSPT can be applied only to transform blocks of a specific size, and LFNST may not be applied when NSPT is applied. Meanwhile, the NSPT index (nspt_idx) indicating the NSPT kernel can be signaled only when it satisfies the size condition of the transformation block. At this time, if the NSPT index is 0, it can indicate that NSPT is not applied to the transformation block, and if the NSPT index is not signaled, it can be inferred as 0. For example, since LFNST is not applied when NSPT is applied, the NSPT index may be signaled before lfnst_idx. For example, if the NSPT index is signaled, lfnst_idx may not be signaled. On the other hand, if the NSPT index is not signaled or the NSPT index is 0, lfnst_idx may be signaled so that LFNST is performed as an inverse second transform, and the inverse first transform may be DCT 2. In this case, mts_idx may not be signaled. In another example, lfnst_idx may be signaled first, and NSPT index and mts_idx may be signaled after lfnst_idx. In addition, NSPT index and mts_idx may be signaled when lfnst_idx is not signaled or is 0. In this case, NSPT index and mts_idx may be signaled sequentially or not sequentially. In addition, since NSPT index should be signaled before parsing of transform coefficients, it may be signaled after last_sig_coeff_pos syntax element like lfnst_idx. In addition, NSPT index may be signaled at residual level (residual coding syntax), and mts_idx may be signaled at CU level, TU tree, or TU level. Meanwhile, the NSPT index can be replaced with the signaling of the LFNST index. That is, the NSPT index is not signaled, and if the size of a specific transform block is satisfied, the NSPT kernel can be selected with the value of the LFNST index. For example, as shown in FIG. 8, if the size of a specific block is satisfied, the NSPT, i.e., non-separable transform, is applied for the transform, and in this case, the kernel of the non-separable transform, i.e., NSPT, can be selected according to the value of the signaled LFNST index (or the index information indicating the non-separable transform kernel). If the size of the transform block is a transform block to which the NSPT is not applied, the LFNST can be applied, and in this case, the kernel of the non-separable transform, i.e., LFNST, can be selected according to the value of the signaled LFNST index (or the index information indicating the non-separable transform kernel). In addition, if the LFNST is not applied, the separable transform can be applied, and in this case, the DCT 2 transform or the MTS can be applied. Meanwhile, the NSPT kernel set and the LFNST kernel set can be selected based on the size of the transformed block and the intra prediction mode. For example, NSPT can be used for block shapes of 4x4, 4x8, 4x16, 8x8, 8x16, 8x32 and corresponding transposed blocks (ex. 4x4, 8x4, 16x4, 8x8, 16x8, 32x8 blocks), and LFNST can be used for the remaining block shapes. In addition, NSPT can also be used for most intra prediction tools, such as regular intra prediction, DIMD, TIMD, SGPM, MIP, EIP, IntraTMP, etc. In addition, NSPT can also be applied to inter CUs. In this case, the existing NSPT kernel set can be replaced with three kernel sets by additionally using two additional NSPT kernel sets. For example, the first NSPT kernel set can be applied to a general intra prediction block, the second NSPT kernel set can be applied to a block using TIMD, DIMD, EIP, MIP or SGPM, and the third NSPT kernel set can be applied to a block using IntraTMP and inter CUs. That is, the three kernel sets can be used according to the intra prediction mode, and no additional signal may be required. Additionally, for DIMD, TIMD, SGPM, MIP, EIP, IntraTMP or Inter prediction modes, which are not the case for general intra prediction modes, additional kernel sets may be used and signaling of additional information may not be required. Meanwhile, information indicating three kernel sets may be signaled. Additionally, a flag may be signaled indicating whether an additional set of NSPT transform kernels for DIMD, TIMD, EIP, MIP, SGPM, etc. is used. For example, when the flag value is 1, the additional set of transform kernels may be used, and when the flag value is 0, the additional set of transform kernels may or may not be used. Meanwhile, multiple transform set selection (MTSS) for intra LFNST / NSPT may be used. For example, multiple LFNST / NSPT transform sets may be used for DIMD, TIMD, OBIC, SGPM, MIP, EIP and IntraTMP modes. In this case, the DIMD method for deriving intra prediction modes for transform set selection may operate on subsampled neighboring or prediction block samples. For example, the second intra prediction mode used for selecting LFNST / NSPT transform kernels for MIP, EIP, SGPM and IntraTMP modes may be derived from the second highest HoG of the prediction block or neighboring blocks. In this case, the first transform set and the second transform set may be different transform sets or different transform types. In addition, there may be restrictions on the block size and the number of transformation candidates of the second set. For example, MTSS may be used based on the size of the block. For example, MTSS may be applied to a CU whose width x height is equal to or greater than 128. In addition, for a CU whose width x height is less than 256, only the first two set candidates may be used, and for a larger CU, all three set candidates may be used. In this case, the size of the block to which MTSS may be applied is not limited to 128. In addition, whether MTSS is applied may be determined based on the current state of the block. In addition, the MTSS for LFNST / NSPT may include a first transform set and a second transform set. In this case, the intra prediction mode for the first transform set may vary depending on the prediction mode. For example, in case of DIMD, OBIC, and TIMD, the first intra prediction mode for the first transform set may be a prediction mode used in the current prediction mode, and in case of SGPM, MIP, EIP, and IntraTMP, the first intra prediction mode for the first transform set may be a prediction mode derived based on HoG. In addition, for example, in case of DIMD, OBIC, and TIMD, the second intra prediction mode for the second transform set may be a prediction mode used in the current prediction mode, and in case of SGPM, MIP, EIP, and IntraTMP, the second intra prediction mode for the second transform set may be a prediction mode derived based on HoG. In addition, a flag indicating whether the MTSS is used may be signaled. For example, if the flag value is 1, MTSS may be used, and if the flag value is 0, MTSS may not be used. Additionally, a flag or index indicating the first transformation set or the second transformation set may be signaled. That is, the flag or index value may indicate the first transformation set or the second transformation set. Meanwhile, a transform set for an intra chroma block coded in LFNST / NSPT and CCP modes can be derived. For example, an intra prediction mode of a current block can be derived based on a CCP prediction sample using DIMD. Specifically, DIMD can be applied to CCP prediction samples, and then the 1st DIMD mode can be used to determine the LFNST / NSPT transform set. For example, a horizontal gradient and a vertical gradient are calculated for each prediction sample to derive a HoG, and an LFNST / NSPT transform set can be determined using an intra prediction mode having the largest histogram amplitude value. Meanwhile, in the case of the SGPM mode, the direction of the partitioning may be used to determine the set of transform kernels. As another example, the DIMD process may be applied to the prediction signal of the SPGM to calculate the VIPM (virtual intra prediction mode), and the set of LFNST / NSPT transform kernels may be determined. At this time, the intra prediction modes of the SGPM and the VIPM may be compared using a threshold value. For example, if the value of the VIPM is close to at least one of the intra prediction modes in the SGPM, the VIPM may be used, and the set of LFNST / NSPT transform kernels may be derived using the VIPM. Otherwise, the direction of the partitioning mode may be used to derive the set of LFNST / NSPT transform kernels. Whether the value of the VIPM is close to at least one of the intra prediction modes in the SGPM may be determined based on the threshold value. For example, if the difference between the VIPM value and one of the intra prediction modes in the SGPM is less than N, the set of LFNST / NSPT transform kernels may be derived using the VIPM. At this time, N can be an integer greater than or equal to 0. Below we describe an extension of LFSNT for transformation. For example, the LFNST described above can be extended with transformation sets and transformation kernels. For example, the LFNST transformation sets can be 35, and each transformation set can be composed of three transformation kernels, i.e., transformation candidates. The transformation set (lfnstTrSetIdx) for the intra prediction mode can have a mapping relationship as shown in the table below. Referring to Table 1 above, if the intra prediction mode (predModeIntra) is less than 0, lfnstTrSetIdx is 2, and if the intra prediction mode is 0 to 34, lfnstTrSetIdx is also mapped to 0 to 34, and if the intra prediction mode is 35 to 66, lfnstTrSetIdx can be mapped to (68- predModeIntra). Also, if the intra prediction mode is greater than 66, lfnstTrSetIdx can be mapped to 2. In this case, lfnstTrSetIdx can be called an LFNST set index. Meanwhile, the LFNST kernel can include the LFNST 4 kernel, the LFNST 8 kernel, and the LFNST 16 kernel. For example, the LFNST 16 kernel can be applied in addition to the LFNST 4 kernel and the LFNST 8 kernel. For example, the LFNST 4 kernel can be applied to a block that is 4xN / Nx4 (N≥4), the LFNST 8 kernel can be applied to a block that is 8xN / Nx8 (N≥8), and the LFNST16 kernel can be applied to a block that is 16xN / Nx16 (N≥16). Meanwhile, Region-Of-Interest (ROI) and zero-out can be applied in LFNST. Here, ROI can mean working on a specific part of the image with interest, and can mean a sample in a data set identified for a specific purpose. For example, Forward LFNST is applied to ROI, which is a specific region of interest at the upper left of the target block of transformation. Therefore, when LFMST is applied, the first transformation coefficients existing in areas other than ROI can be zeroed out. Figure 12 shows an example of ROI for LFNST 16. Referring to FIG. 12, the ROI for LFNST 8 can be composed of six 4x4 sub-blocks that are arranged sequentially in the scan direction from the upper left of the target block. Since a total of 96 transform coefficients are input to the Forward LFNST, the dimension of the Forward LFNST matrix can be Rx96. Here, R can be 32, 48, or 64, which is less than 96. For example, R can be 32, and in this case, LFNST 16 is a 32x96 matrix. In addition, for example, when a 32x96 matrix is used for LFNST 16, the first transform coefficients existing in an area other than the ROI can be zeroed out. Figure 13 shows an example of ROI for LFNST 8. Referring to FIG. 13, the ROI for LFNST 8 can be composed of four 4x4 sub-blocks located at the upper left of the target block, i.e., an 8x8 area at the upper left of the target block. Since a total of 64 transform coefficients are input to the Forward LFNST, the dimension of the Forward LFNST matrix can be Rx64. Here, R can be 32 or 48, which is smaller than 64. For example, R can be 32, and in this case, LFNST 8 is a 32x64 matrix. When a 32x64 matrix is used for LFNST 8, the first transform coefficients existing in areas other than the ROI can be zeroed out. In addition, for example, in the case of an 8X8 block, since 64 transform coefficients are input to the Forward LFNST, the entire block is an ROI, so zeroing out may not be applied. Figure 14 shows an example of a MIP prediction sample for constructing HoG. Referring to FIG. 14, a HoG can be constructed using DIMD based on MIP-based prediction samples. For example, for a target block predicted by MIP or IntraTMP, DIMD can be used to derive an intra prediction mode of the target block based on the MIP or IntraTMP-based prediction sample. For example, for a target block to which MIP is applied, DIMD can be applied based on the MIP prediction sample before upsampling. For example, in order to construct a HoG, horizontal slopes and vertical slopes are calculated for each prediction sample. Then, an LFNST transform set and an LFNST transpose flag (LFNST Transpose flag) can be determined using an intra prediction mode having the largest histogram amplitude value. In addition, the LFNST transpose flag can be signaled after the LFNST index signaling as information indicating whether the LFNST kernel is transposed. Alternatively, for example, the MIP transpose flag (mip_transposed_flag) can be used as the LFNST transpose flag. In this case, signaling of the LFNST prefix flag may be omitted. Meanwhile, if the intra prediction mode is IntraTMP, the intra prediction mode derived through DIMD can be applied as the intra prediction mode for determining the LFNST transformation set. At this time, the intra prediction mode with the largest histogram width can be used as the intra prediction mode for determining the LFNST set. Alternatively, as an example, if the intra prediction mode is one of IBC, SGPM (Spatial Geometric partitioning mode), TIMD (template-based intra mode derivation), and Palette mode, DIMD can be applied based on the prediction sample derived through the mode to derive the intra prediction mode for determining the LFNST transform set. Alternatively, as an example, LFNST or NSPT can be applied to inter prediction blocks as well as intra prediction blocks. For example, a set of transformations can be mapped for each inter prediction mode, and LFNST or NSPT can be performed by applying any one of a plurality of transformation kernels to the mapped set of transformations. Alternatively, as an example, DIMD can be applied based on inter prediction samples, and horizontal gradients and vertical gradients can be calculated for each prediction sample to build a HoG. Then, a prediction mode with the largest histogram amplitude value can be used to derive a set of LFNST transformations. Meanwhile, a modified context model that uses the previous five coefficients in the coding order instead of the neighboring 2D coefficients for coding LFNST / NSPT coefficients can be used. That is, when LFNST / NSPT is applied, the five previous transform coefficients in the coding (scan) order can be used instead of the neighboring coefficients at the 2D position for context modeling or deriving context information about the currently parsed transform coefficient (e.g., it can include at least one of sig_coeff_flag, gt1_flag, or gt2_flag). Figure 15 illustrates an example of context modeling for LFNST and NSPT transform coefficients. Referring to (a) of Fig. 15, it can be shown that the surrounding two-dimensional coefficients (6, 7, 9, 10, 13) are used for context modeling of the transformation coefficient (5) to be currently parsed. At this time, the transformation coefficient to be currently parsed may be 5, and the surrounding two-dimensional coefficients may be 6, 7, 9, 10, 13. In addition, referring to (b) of Fig. 15, it can be shown that the previous five coefficients (2, 3, 6, 9, 12) are used for context modeling of the transformation coefficient (5) to be currently parsed. At this time, the transformation coefficient to be currently parsed may be 5, and the surrounding two-dimensional coefficients may be 2, 3, 6, 9, 12. In this case, the values or absolute values of the previous five coefficients may be used. For example, context information (context index or context increment) of information about a transform coefficient can be derived based on whether the sum of the values or absolute values of the five coefficients above is greater than a threshold value. Or, for example, context information (context index or context increment) of information about a transform coefficient can be derived based on whether the average of the values or absolute values of the five coefficients above is greater than a threshold value. Meanwhile, since lfnstIdx (lfnst_idx) is required for parsing transform coefficients, lfnstIdx may be signaled after every last_sig_coeff_pos syntax element of the CU. For example, when LFNST or NSPT is applied, DCT-2 transform coefficients may be placed within a coefficient block using diagonal reordering. Or, when LFNST or NSPT is applied, the transform coefficient scan order may be changed / derived due to zero out. As an example, lfnstIdx and / or NSPT index may also be signaled after last_sig_coeff_pos in the residual coding level. Meanwhile, for non-separable transforms such as LFNST or NSPT, one-dimensional directionality such as horizontal or vertical may not have a significant effect or be important in the coding of transform coefficients. Therefore, when performing a transform with LFNST or NSPT, transform coefficients can be reordered in the diagonal direction and transformed, and context modeling of transform coefficients can be performed by utilizing modeling information of coefficients coded first according to the diagonal scan order reflecting this. Therefore, by signaling the LFNST index first before transform coefficient coding, context modeling of transform coefficients can be performed by dividing the cases where LFNST (or NSPT) is applied and cases where LFNST (or NSPT) is not applied. In this case, mts_idx may be signaled at the same level as lfnstIdx and / or NSPT index, e.g., at residual coding level, or at CU level, TU tree, or TU level. Additionally, mts_idx may be signaled immediately after lfnstIdx signaling or immediately after NSPT index signaling. For example, the conversion-related syntax could be as shown in the table below. Referring to Table 2 above, the lfnst_idx, nspt_idx, mts_flag and / or mts_idx may be signaled in the same syntax, or some of them may be signaled in different syntaxes. For example, after lfnst_idx is signaled, nspt_idx, mts_flag and / or mts_idx may be signaled. Additionally, nspt_idx may be signaled or mts_flag and / or mts_idx may be signaled based on the value of lfnst_idx. For example, if the lfnst apply condition is true, lfnst_idx may be signaled. Additionally, for example, if the value of lfnst_idx is 0 and the nspt apply condition is true, nspt_idx may be signaled, otherwise mts_flag and / or mts_idx may be signaled. As another example, the conversion-related syntax could be as shown in the table below. Referring to Table 3 above, the nspt_idx, lfnst_idx, mts_flag and / or mts_idx may be signaled in the same syntax, or some of them may be signaled in different syntaxes. For example, after nspt_idx is signaled, lfnst_idx, mts_flag and / or mts_idx may be signaled. For example, if the nspt apply condition is true, nspt_idx may be signaled. Additionally, for example, if the value of nspt_idx is 0 and the lfnst apply condition is true, lfnst_idx may be signaled. Additionally, for example, if the value of lfnst_idx is 0, mts_flag and / or mts_idx may be signaled. As another example, the conversion-related syntax could be as shown in the table below. Referring to Table 4 above, the nspt_idx, lfnst_idx, mts_flag and / or mts_idx may be signaled in the same syntax, or some of them may be signaled in different syntaxes. For example, if the nspt apply condition is true, nspt_idx may be signaled. Additionally, if the nspt apply condition is not true and the lfnst apply condition is true, lfnst_idx may be signaled. Additionally, for example, if the value of lfnst_idx is 0, mts_flag and / or mts_idx may be signaled. Meanwhile, lfnst_idx may point to an LFNST transformation kernel, and may point to an NSPT transformation kernel depending on conditions. For example, if the NSPT condition described in the present disclosure (e.g., NSPT application block size condition) is satisfied, the value of lfnst_idx may point to one of the NSPT candidates. As another example, the conversion-related syntax could be as shown in the table below. Referring to Table 5 above, the lfnst_idx, nspt_idx, mts_flag and / or mts_idx may be signaled in the same syntax, or some of them may be signaled in different syntaxes. For example, if the lfnst condition is true, lfnst_idx may be signaled. Otherwise, mts_flag and / or mts_idx may be signaled. As another example, the conversion-related syntax could be as shown in the table below. Referring to Table 6 above, when the value of lfnst_idx is 0 and the value of nspt_idx is 0, mts_idx can be signaled at the coding unit level. Additionally, the signaling of the transform_tree syntax can be as shown in the table below. Referring to Table 7 above, the transform_unit syntax can be signaled at the transform_tree syntax level. Additionally, the signaling of residual_coding syntax can be as shown in the table below. Referring to Table 8 above, the residual_coding syntax can be signaled at the transform_unit syntax level. Also, as another example, the conversion syntax could be as follows: Referring to Table 9 above, nspt_idx can be signaled after lfnst_idx is signaled. For example, LfnstZeroOutSigCoeffFlag may be set to 0 if the last valid coefficient position is located in the lfnst zero out region. Additionally, the LfnstZeroOutSigCoeffFlag may be set to 1 if the last valid coefficient position is located in the lfnst coefficient region (i.e., an area other than the lfnst zero out region). In this case, the last valid coefficient position may be derived based on last valid coefficient position information (e.g., including at least one of last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, or last_sig_coeff_y_suffix). Specifically, last_sig_coeff_x_prefix represents a prefix of the column position of the last significant coefficient in the scanning order within the transform block, last_sig_coeff_y_prefix represents a prefix of the row position of the last significant coefficient in the scanning order within the transform block, last_sig_coeff_x_suffix represents a suffix of the column position of the last significant coefficient in the scanning order within the transform block, and last_sig_coeff_y_suffix represents a suffix of the row position of the last significant coefficient in the scanning order within the transform block. At this time, the effective coefficient can represent the non-zero coefficient. Here, it is explained in units of transformation blocks (TB), but this is for example only and the TB can be used interchangeably with a coding block (CB). For example, if the value of LfnstZeroOutSigCoeffFlag is 1, lfnst_idx can be signaled. Also, if the value of lfnst_idx is 0 and the nspt application condition is true, nspt_idx can be signaled. Additionally, for example, when LFNST and / or NSPT are applied, sb_coded_flag[ xS ][ yS ] for the last subblock containing the last significant coefficient and the subblocks within the DC subblock may be omitted from signaling and their values may be derived as 1. Additionally, some or all of sb_coded_flag[ xS ][ yS ] may be omitted, for example, if the lfnst_idx or nspt_idx value is greater than 0. Meanwhile, transformation-related information can be coded based on context information, and the related context information can be expressed as follows, for example. For example, when the NSPT condition (e.g., NSPT applicable block size condition) is satisfied, even if the value of lfnst_idx can point to one of the NSPT candidates (i.e., when lfnst_idx replaces or is used interchangeably with nspt_idx), different context models (context information) may need to be used depending on whether NSPT or LFNST is applied. For this purpose, the following structure can be used. For example, ApplyNsptFlag can be set to 1 when the NSPT condition (e.g., NSPT applicable block size condition) is satisfied. For example, examples of context models that assign ctxInc to syntax elements using context-coded beans might look like Tables 10 through 13 below. Meanwhile, when using nspt_idx separately from lfnst_idx, examples of context models that assign ctxInc to syntax elements using context-coded beans can be as shown in Tables 14 to 17 below. Below we describe Enhanced MTS for Intra. For example, when MTS is applied to a block predicted intra, the transform kernel may be DCT 5, DST 4 in addition to DCT 2, DST 7, DCT 8. Also, when MTS is applied to a block predicted intra, the transform kernel may be DCT 5, DST 4, DST 1, and identity transform (IDT) in addition to DCT 2, DST 7, DCT 8. In addition, intra MTS candidates, i.e., MTS sets, can be derived based on TU size and intra prediction mode information. For example, TU sizes can be applied in total of 16, and can be classified into 5 groups (classes) according to intra prediction modes for each individual TU size. In this case, if 5 groups are applied to each of the 16 TU sizes, a total of 80 groups can be considered, but since the transformation set can be shared, less than 80 groups, for example, 58 groups, can be considered. In addition, it is not limited to 58 groups, and a positive integer number N of groups can be considered. Also, if the intra prediction mode is a directional mode, the symmetry of TU shape and intra prediction direction can be considered. For example, mode i with AХB and mode j with (i > 34) BХA can be mapped to the same group (j = (68 - i)), and in this case, the transform pairs for vertical and horizontal kernels can be swapped. For example, a 16x4 block whose intra prediction mode number is 18 and a 4x16 block whose intra prediction mode number is 50 can be mapped to one group. If the intra prediction mode is a wide-angle intra prediction mode, the closest existing directional mode can be used for determining the transform set. For example, to derive a group for determining the transform set, a wide-angle intra prediction mode between -2 and -14 can use mode 2, and a wide-angle intra prediction mode between 67 and 80 can use mode 66. In addition, multiple transformation candidates (transformation pairs) can be applied to each group. For example, one, four, or six transformation sets can be applied. The number of transformation sets to be applied can be determined based on the position of the last transformation coefficient or based on the absolute value of the transformation coefficient. For example, the position of the last transformation coefficient can be compared with two thresholds, and based on the comparison result, one of one, four, or six transformation sets to be applied can be selected. Alternatively, the sum of the absolute values of the transformation coefficients can be compared with two thresholds, and one of one, four, or six transformation sets can be selected. In addition, the number of transformation sets can be determined based on the sum of the absolute values of the transformation coefficients. For example, if the sum of the absolute values of the transformation coefficients is less than th0, one transformation set can be applied. If the sum of the absolute values of the transformation coefficients is greater than th0 and less than or equal to th1, four transformation sets can be applied. If the sum of the absolute values of the transformation coefficients is greater than th1, six transformation sets can be applied (1 candidate: sum <= th0, 4 candidates: th0 < sum <= th1, 6 candidates: sum > th1). Here, sum can represent the sum of the absolute values of the transformation coefficients. In addition, th0 can be set to 6 and th1 can be set to 32. Figure 16 shows an example of deriving an MTS set. Referring to Fig. 16, four transform sets can be determined based on the TU size and intra prediction mode information. For example, if the TU sizes are all 16, and the intra prediction modes are 0 to 34, and a MIP mode is added, a total of 36 intra modes can be applied. As described above, the intra prediction mode of the transform block can be mapped to any one of the intra prediction modes 1 to 34 according to the intra prediction mode information. A total of five groups can be mapped for each TU size. For example, intra prediction modes 0 and 1, intra prediction modes 2 to 12, intra prediction modes 13 to 23, intra prediction modes 24 to 34, and the MIP mode can each be mapped to one group. For example, if the TU size is 4x8 and the intra prediction mode is 5, four transform sets can be determined. At this time, if the MTS index information is 1, number 12 is selected, and the transform set of DCT 5, DCT 5 can be selected. Additionally, each transform pair can have five transform kernels applied in pairs: DCT 8, DST 7, DCT 5, DST 4, and DST 1, and each transform pair can be indexed from 0 to 24. As described above, if four transformation sets are applied to one group, a transformation set consisting of four transformation pairs can be applied to each of the 80 groups. At this time, some of the 80 groups can share the transformation sets, and the 80 groups can be reduced to 58 groups. Figure 17 shows an example of deriving DIMD-based intra modes for MTS and LFNST. Referring to Fig. 17, a DIMD-based intra mode can be derived for determining MTS and LFNST sets. Meanwhile, if the intra prediction mode is IntraTMP, the intra prediction mode derived through DIMD can be applied as the intra prediction mode for determining the MTS transform set. For example, the intra prediction mode with the largest histogram width can be used as the intra prediction mode for determining the MTS and LFNST sets. Meanwhile, as another example, the intra mode for IntraTMP can also be derived from a reference block or a neighboring block of the reference block. In addition, if the intra prediction mode is one of IBC, SGPM (Spatial Geometric partitioning mode), TIMD (template-based intra mode derivation), and Palette mode, the intra prediction mode for determining the MTS transformation set can be derived by applying DIMD based on the prediction sample derived through the mode. Meanwhile, for example, MTS can be applied even when the size of the transform block is greater than 32. Or, considering the complexity, the intra prediction modes can be grouped into 3 or 2 groups instead of 5 groups for each size of the transform block. Or, in addition to the transform pairs of Figure 4.3-2, IDT (identity transform) can be applied. For example, one of the vertical or horizontal transforms can be one of DST 7, DCT 8, DCT 5, DST 4, and DST 1, as in Figure 4.3-2, and the other can be the identity transform. In this case, the number of cases of the transform pair can increase. Additionally, regarding the signaling of the MTS index (mts_idx), the following is taken: For example, mts_enabled_flag may be signaled at a higher level, and mts_flag and mts_idx may be signaled at a CU or residual coding level. In this case, if mts_flag is 1, mts_idx may be signaled to determine a transform set, and if mts_flag is 0, DCT 2 may be applied to the first transform. Also, if there are four transformation sets, one transformation pair out of the four can be selected through mts_idx, and if there are six transformation sets, one transformation pair out of the six can be selected through mts_idx. In this case, if there is one transformation set, the first transformation pair out of the four transformation sets or the first transformation pair out of the six transformation sets can be selected without separate signaling. That is, if there is one transformation set, mts_flag can be signaled with a value of 1, and mts_idx can be not signaled. Also, if mts_idx is not signaled, a specific transformation pair can be used, or mts_idx can be derived as 0 or 1. Alternatively, if there is only one transformation set, mts_flag can be signaled with a value of 1, indicating that the first transformation pair of the 4 transformation set is applied if mts_idx is 0, and indicating that the first transformation pair of the 6 transformation set is applied if mts_idx is 1. For example, binarization based on the value of the MTS index (mts_idx) can be as shown in the table below. Also, as another example, without signaling mts_flag, if mts_idx is 0, it can indicate that MTS is not applied, i.e., DCT 2 is applied to the first transform. In the above case, the binarization of the MTS index can be as shown in the table below. The MTS index can be binarized as truncated rice (or truncated unary) as shown in the above tables, or can be binarized based on a fixed-length coding method. Alternatively, as another example, if there is only one transformation set, a specific transformation pair can be used instead of a four-transform set or a six-transform set, and in this case, mts_idx can point to a specific transformation pair. In this case, mts_idx pointing to one transformation set can be signaled as 6 in Table 2 or 7 in Table 3. Alternatively, mts_idx can be signaled as an intermediate value, for example, 3 or 4, in Table 2 or Table 3. Alternatively, if IDT (Identity Transform) is applied and any one of the six transformation kernels is applied, information indicating whether IDT (Identity Transform) is applied in the horizontal direction or the vertical direction may be further signaled after mts_idx. For example, idt_flag indicating whether IDT (Identity Transform) is applied may be signaled, and if idt_flag is 1, the direction in which the identity transform is applied may be further signaled. Alternatively, flag information indicating whether the identity transform is applied in the horizontal direction and the vertical direction may be signaled, respectively. Alternatively, when IDT (Identity Transform) is applied, mts_inter_enabled_flag or mts_intra_enabled_flag can be signaled at the upper level, and kernel indices indicating the six kernels can be individually signaled for each direction at the lower level. The kernel index in the horizontal direction can be signaled as mts_horizontal_idx, for example, and the kernel index in the vertical direction can be signaled as mts_vertical_idx. FIG. 18 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 18 may be performed by the encoding device disclosed in FIG. 2. Specifically, for example, S1800 of FIG. 18 may be performed by the prediction unit (220) of the encoding device (200), S1810 to S1830 may be performed by the residual processing unit (230) of the encoding device (200), and S1840 may be performed by the entropy encoding unit (240) of the encoding device (200). The method disclosed in FIG. 18 may include the embodiments described above in the present disclosure. Referring to FIG. 18, the encoding device derives prediction samples for the current block (S1800). For example, the encoding device can derive the prediction samples for the current block. The encoding device derives residual samples for the current block (S1810). For example, the encoding device can derive residual samples for the current block based on the prediction samples. The encoding device derives transform coefficients for the current block (S1820). The encoding device can derive transform coefficients for the current block by performing a transform based on the residual samples. For example, the transform may be a first-order transform. For example, based on MTS related information, one of the transform pair candidates of the transform pair of the transform set for the current block can be determined. At this time, for example, the number of candidates of the transform pair of the transform set can be at least one. For example, the number of candidates of the transform pair of the transform set can be one of 1, 4, or 6. In addition, the one transform pair can be configured based on DCT 2, DCT 5, DCT 8, DST4, DST 7, or an identity transform related kernel. In addition, the one transform pair can be determined based on the intra prediction mode. For example, based on the prediction mode for the current block being IBC (Intra Block Copy), SGPM (Spatial Geometric Partitioning mode), TIMD (Template-based Intra Mode Derivation) and palette mode, prediction samples can be derived. At this time, the intra prediction mode is derived using the DIMD mode based on the prediction samples, and the one transform pair can be determined based on the intra prediction mode. In addition, MTS can be applied even when the size of the transform block is larger than 32. In this case, the size of the transform block is not limited to being larger than 32. In addition, the intra prediction mode can be grouped into 2 or 3 groups instead of 5 groups for each size of the transform block. In addition to the transform pairs of Table 19 described above, the identity transform can be applied. In this case, one of the vertical or horizontal transforms can be one of DST 7, DCT 8, DCT 5, DST 4, and DST 1 as shown in Table 19, and the other can be the identity transform. In this case, the number of cases of the transform pair can increase. The encoding device generates residual information for the current block (S1830). For example, the encoding device can generate residual information for the current block based on the transform coefficients. The encoding device encodes image information including residual information (S1840). For example, the encoding device can encode the image information including the residual information. For example, the image information may include MTS (Multiple Transform Selection) related information. At this time, the MTS related information may include at least one of MTS available flag information, MTS flag information, or MTS index information. In this case, the MTS flag information and the MTS index information may be signaled at a coding unit level or a residual coding level, and the MTS available flag information may be signaled at a higher level than the MTS flag information and the MTS index information. That is, the MTS index information may be signaled at a coding unit syntax or a residual coding syntax, and the MTS available flag information may be signaled at a higher level syntax than the MTS flag information and the MTS index information. In addition, the MTS available flag information may indicate a syntax element "mts_enabled_flag", the MTS flag information may indicate a syntax element "mts_flag", and the MTS index information may indicate a syntax element "mts_idx". At this time, the MTS index information may be binarized into a truncated rice or a truncated unary. In addition, the MTS index information may also be binarized into a fixed-length coding method. In addition, based on the MTS-related information, one of the transform pair candidates of the transform pair of the transform set for the current block can be determined. At this time, the number of candidates of the transform pair of the transform set can be at least one. For example, the number of candidates of the transform pair of the transform set can be one of 1, 4, or 6. In addition, the one transform pair can be configured based on DCT 2, DCT 5, DCT 8, DST4, DST 7, or an identity transform-related kernel. Additionally, signaling of the MTS-related information may be determined based on the number of candidates in the transformation set. For example, based on the number of candidates in the above transformation set being 1, the value of the MTS flag information may be signaled as 1 and the MTS index information may not be signaled. At this time, based on the fact that the MTS index information is not signaled, a specific transformation pair may be used or the value of the MTS index information may be derived. If the MTS index information is not signaled, the value of the MTS index information may be derived as 0 or 1. Additionally, for example, based on the fact that the number of candidates in the above transformation set is 1, either the first transformation pair among the transformation pair candidates of the four transformation sets or the first transformation pair among the transformation pair candidates of the six transformation sets may be selected. For example, based on the number of candidates of the above transformation set being 1, the MTS flag information can be signaled with a value of 1. In this case, based on the MTS index information being 0, the first transformation pair among the transformation pair candidates of the four transformation sets can be selected. Also, based on the MTS index information being 1, the first transformation pair among the transformation pair candidates of the six transformation sets can be selected. In addition, based on the fact that the number of candidates of the above transformation set is 1, instead of using the above 4 transformation sets or the above 6 transformation sets, a specific transformation pair may be used. At this time, the MTS index information may indicate a specific transformation pair. For example, the MTS index information indicating one transformation set may be signaled as 6 in the above-described Table 18, and may be signaled as 7 in the above-described Table 19. Alternatively, the MTS index information may be signaled as an intermediate value in the above-described Table 18 or Table 19. At this time, the intermediate value may be a value of 3 or 4. Here, one transformation set may be a specific transformation set. Additionally, MTS may not be applied to the transform based on the fact that the MTS flag information is not signaled and the value of the MTS index information is 0. In this case, the transform may be performed based on DCT-2 on the transform coefficients. In addition, when an identity transformation is applied to the above transformation, any one of six transformation kernels may be applied as the identity transformation related kernel. In this case, information indicating whether the identity transformation is applied in the horizontal direction or the vertical direction may be signaled after the MTS index information. In addition, the image information may include flag information indicating whether the identity transformation is applied to the transform coefficients. Based on the value of the flag information being 1, information related to the direction in which the identity transformation is applied may be signaled. In addition, the image information may include information related to whether the identity transformation is applied to the transform coefficients in a horizontal direction and information related to whether the identity transformation is applied to the current block in a vertical direction. Here, the information indicating whether the identity transformation is applied may indicate a syntax element "idt_flag". In addition, based on the identity transformation being applied to the transformation, the image information may include MTS inter available flag information or MTS intra available flag information. At this time, kernel index information representing six transformation kernels may be individually signaled according to the direction in which the identity transformation is applied. For example, the kernel index information may be signaled in a lower level syntax than the MTS inter available flag information or the MTS intra available flag information. At this time, the kernel index information may include horizontal direction kernel index information and / or vertical direction kernel index information. Here, the MTS inter available flag information may indicate a syntax element "mts_inter_enabled_flag", the MTS intra available flag information may indicate a syntax element "mts_intra_enabled_flag", the horizontal direction kernel index information may indicate "mts_horizontal_idx", and the vertical direction kernel index information may indicate "mts_vertical_idx". In addition, MTS can be applied even when the size of the transform block is larger than 32. In this case, the size of the transform block is not limited to being larger than 32. In addition, the intra prediction mode can be grouped into 2 or 3 groups instead of 5 groups for each size of the transform block. In addition to the transform pairs of Table 19 described above, the identity transform can be applied. In this case, one of the vertical or horizontal transforms can be one of DST 7, DCT 8, DCT 5, DST 4, and DST 1 as shown in Table 19, and the other can be the identity transform. In this case, the number of cases of the transform pair can increase. Additionally, the image information may include various information according to an embodiment of the present disclosure. For example, the image information may include information disclosed in at least one of the tables described above. Additionally, encoded image information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or storage medium. In addition, as described above, the encoding device can generate a restored picture (including restored samples and restored blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as performed in the decoding device from the encoding device, and thereby increase coding efficiency. Accordingly, the encoding device can store the restored picture (or restored samples, restored blocks) in memory and utilize it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the restored picture. According to the above-described embodiment(s), in the case of IBC, SGPM, TIMD, palette mode, etc., the MTS transform set can be determined based on the intra prediction mode derived by applying DIMD, so that the MTS transform set can be adaptively determined in various intra prediction modes, thereby improving the efficiency of the transform. In addition, according to the above-described embodiment(s), a transformation pair can be determined by efficiently signaling MTS index information based on the number of transformation sets, signaling of MTS-related information, and / or application of identity transformation, thereby efficiently performing transformation. FIG. 19 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 19 may be performed by the decoding device disclosed in FIG. 3. Specifically, for example, S1900 of FIG. 19 may be performed by the entropy decoding unit (310) of the decoding device (300), S1910 to S1920 may be performed by the residual processing unit (320) of the decoding device (300), and S1930 may be performed by the adding unit (340) of the decoding device (300). The method disclosed in FIG. 19 may include the embodiments described above in the present disclosure. Referring to FIG. 19, the decoding device receives image information including residual information for the current block (S1900). For example, the decoding device may receive image information including the residual information for the current block through a bitstream. The image information may further include prediction-related information as described above. For example, the image information may include MTS (Multiple Transform Selection) related information. At this time, the MTS related information may include at least one of MTS available flag information, MTS flag information, or MTS index information. In this case, the MTS flag information and the MTS index information may be included in a coding unit level or a residual coding level, and the MTS available flag information may be included at a higher level than the MTS flag information and the MTS index information. That is, the MTS index information may be signaled in a coding unit syntax or a residual coding syntax, and the MTS available flag information may be signaled in a higher level syntax than the MTS flag information and the MTS index information. In addition, the MTS available flag information may indicate a syntax element "mts_enabled_flag", the MTS flag information may indicate a syntax element "mts_flag", and the MTS index information may indicate a syntax element "mts_idx". At this time, the MTS index information may be binarized into a truncated rice or a truncated unary. In addition, the MTS index information may also be binarized into a fixed-length coding method. In addition, based on the MTS-related information, one of the transform pair candidates of the transform pair of the transform set for the current block can be determined. At this time, the number of candidates of the transform pair of the transform set can be at least one. For example, the number of candidates of the transform pair of the transform set can be one of 1, 4, or 6. In addition, the one transform pair can be configured based on DCT 2, DCT 5, DCT 8, DST4, DST 7, or an identity transform-related kernel. Additionally, signaling of the MTS-related information may be determined based on the number of candidates in the transformation set. For example, based on the number of candidates of the transformation set being 1, the value of the MTS flag information may be 1 and the MTS index information may not be included in the image information. At this time, based on the fact that the MTS index information is not included in the image information, a specific transformation pair may be used or the value of the MTS index information may be derived. If the MTS index information is not signaled, the value of the MTS index information may be derived as 0 or 1. Additionally, for example, based on the fact that the number of candidates in the above transformation set is 1, either the first transformation pair among the transformation pair candidates of the four transformation sets or the first transformation pair among the transformation pair candidates of the six transformation sets may be selected. For example, based on the number of candidates of the above transformation set being 1, the value of the MTS flag information may be 1. In this case, based on the MTS index information being 0, the first transformation pair among the transformation pair candidates of the four transformation sets may be selected. Additionally, based on the MTS index information being 1, the first transformation pair among the transformation pair candidates of the six transformation sets may be selected. In addition, based on the fact that the number of candidates of the above transformation set is 1, instead of using the above 4 transformation sets or the above 6 transformation sets, a specific transformation pair may be used. At this time, the MTS index information may indicate a specific transformation pair. For example, the MTS index information indicating one transformation set may be signaled as 6 in the above-described Table 18, and may be signaled as 7 in the above-described Table 19. Alternatively, the MTS index information may be signaled as an intermediate value in the above-described Table 18 or Table 19. At this time, the intermediate value may be a value of 3 or 4. Here, one transformation set may be a specific transformation set. The decoding device derives transform coefficients for the current block (S1910). For example, the decoding device can derive transform coefficients for the current block based on the residual information. The decoding device derives residual samples for the current block (S1920). For example, the decoding device may perform an inverse transform based on the transform coefficients to derive residual samples for the current block. For example, the inverse transform may be an inverse first-order transform. For example, the number of candidates for the transform pair of the transform set can be at least one. For example, the number of candidates for the transform pair of the transform set can be one of 1, 4, or 6. In addition, the one transform pair can be configured based on DCT 2, DCT 5, DCT 8, DST4, DST 7, or an identity transform related kernel. In addition, the one transform pair can be determined based on the intra prediction mode. For example, based on the prediction mode for the current block being IBC (Intra Block Copy), SGPM (Spatial Geometric Partitioning mode), TIMD (Template-based Intra Mode Derivation) and palette mode, prediction samples can be derived. At this time, the intra prediction mode is derived using the DIMD mode based on the prediction samples, and the one transform pair can be determined based on the intra prediction mode. In addition, MTS can be applied even when the size of the transform block is larger than 32. In this case, the size of the transform block is not limited to being larger than 32. In addition, the intra prediction mode can be grouped into 2 or 3 groups instead of 5 groups for each size of the transform block. In addition to the transform pairs of Table 19 described above, the identity transform can be applied. In this case, one of the vertical or horizontal transforms can be one of DST 7, DCT 8, DCT 5, DST 4, and DST 1 as shown in Table 19, and the other can be the identity transform. In this case, the number of cases of the transform pair can increase. In addition, MTS may not be applied to the inverse transformation based on the fact that the image information does not include the MTS flag information and the value of the MTS index information is 0. In this case, the inverse transformation may be performed based on DCT-2 on the transform coefficients. In addition, when an identity transformation is applied to the above inverse transformation, any one of six transformation kernels may be applied as the identity transformation related kernel. In this case, information indicating whether the identity transformation is applied in the horizontal direction or the vertical direction may be signaled after the MTS index information. In addition, the image information may include flag information indicating whether the identity transformation is applied to the transform coefficients. Based on the value of the flag information being 1, information related to the direction in which the identity transformation is applied may be signaled. In addition, the image information may include information related to whether the identity transformation is applied to the transform coefficients in a horizontal direction and information related to whether the identity transformation is applied to the current block in a vertical direction. Here, the information indicating whether the identity transformation is applied may indicate a syntax element "idt_flag". In addition, based on the identity transformation being applied to the inverse transformation, the image information may include MTS inter available flag information or MTS intra available flag information. At this time, kernel index information representing six transformation kernels may be individually derived according to the direction in which the identity transformation is applied. For example, the kernel index information may be included at a lower level than the MTS inter available flag information or the MTS intra available flag information. At this time, the kernel index information may include horizontal direction kernel index information and / or vertical direction kernel index information. Here, the MTS inter available flag information may indicate a syntax element "mts_inter_enabled_flag", the MTS intra available flag information may indicate a syntax element "mts_intra_enabled_flag", the horizontal direction kernel index information may indicate "mts_horizontal_idx", and the vertical direction kernel index information may indicate "mts_vertical_idx". In addition, MTS can be applied even when the size of the transform block is larger than 32. In this case, the size of the transform block is not limited to being larger than 32. In addition, the intra prediction mode can be grouped into 2 or 3 groups instead of 5 groups for each size of the transform block. In addition to the transform pairs of Table 19 described above, the identity transform can be applied. In this case, one of the vertical or horizontal transforms can be one of DST 7, DCT 8, DCT 5, DST 4, and DST 1 as shown in Table 19, and the other can be the identity transform. In this case, the number of cases of the transform pair can increase. The decoding device generates restoration samples for the current block (S1930). For example, the decoding device may generate restoration samples for the current block based on the residual samples. At this time, the restoration samples for the current block may be generated based on the residual samples and the prediction samples for the current block. In addition, the decoding device may generate a restoration picture including the restoration samples, for example. As described above, the decoding device may then apply an in-loop filtering procedure, such as a deblocking filtering procedure and / or an SAO procedure, to the restoration picture to improve subjective / objective image quality, as needed. According to the above-described embodiment(s), in the case of IBC, SGPM, TIMD, palette mode, etc., the MTS transform set can be determined based on the intra prediction mode derived by applying DIMD, so that the MTS transform set can be adaptively determined in various intra prediction modes, thereby improving the efficiency of the transform. In addition, according to the above-described embodiment(s), a transformation pair can be determined by efficiently signaling MTS index information based on the number of transformation sets, signaling of MTS-related information, and / or application of identity transformation, thereby efficiently performing transformation. In the above-described embodiments, the methods are described based on a flow chart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will appreciate that the steps depicted in the flow chart are not exclusive, and other steps may be included or one or more of the steps in the flow chart may be deleted without affecting the scope of the embodiments of the present disclosure. The method according to the embodiments of the present disclosure described above may be implemented in the form of software, and the encoding device and / or the decoding device according to the present disclosure may be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc. The above-described embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable (program) instructions, such as program modules, that are executed by a computer. The modules may be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various means known in the art. The computer-readable medium may be any available medium that can be accessed by a computer and includes both volatile and nonvolatile media, removable and non-removable media. In addition, the computer-readable medium may include both computer storage media and communication media. The computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. The communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery media. In addition, the above-described embodiments of the present disclosure may be implemented as a computer program (or a computer program product) including instructions executable by a computer. The computer program includes programmable machine instructions processed by a processor, and may be implemented in a high-level programming language, an object-oriented programming language, an assembly language, or a machine language. In addition, the computer program may be recorded in a tangible computer-readable recording medium (e.g., a memory, a hard disk, a magnetic / optical medium, or a solid-state drive (SSD)). Accordingly, the above-described embodiment of the present disclosure can be implemented by executing the computer program as described above by a computing device. The computing device can include at least some of a processor, a memory, a storage device, a high-speed interface connecting the memory and a high-speed expansion port, and a low-speed interface connecting the low-speed bus and the storage device. Each of these components is connected to each other using various buses, and can be mounted on a common motherboard or mounted in another suitable manner. Here, the processor may process instructions within the computing device, such as instructions stored in a memory or storage device for displaying graphical information for providing a GUI (Graphical User Interface) on an external input / output device, such as a display connected to a high-speed interface. In other embodiments, multiple processors and / or multiple buses may be utilized, as appropriate, together with multiple memories and memory types. Additionally, the processor may be implemented as a chipset comprising multiple independent analog and / or digital processors. Memory also stores information within a computing device. For example, memory may be comprised of volatile memory units or a collection thereof. For another example, memory may be comprised of nonvolatile memory units or a collection thereof. Memory may also be another form of computer-readable media, such as, for example, a magnetic or optical disk. And the storage device can provide a large amount of storage space to the computing device. The storage device can be a computer-readable medium or a configuration including such a medium, and can include, for example, devices within a storage area network (SAN) or other configurations, and can be a floppy disk device, a hard disk device, an optical disk device, a tape device, flash memory, or other similar semiconductor memory device or array of devices. Additionally, the network can be implemented as a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or as various types of wireless networks such as a mobile radio communication network or a satellite communication network. The present disclosure discussed above has been described with reference to the embodiments illustrated in the drawings, but this is merely exemplary, and those skilled in the art will understand that various modifications and variations of the embodiments are possible from this. That is, the scope of the present disclosure is not limited to the above-described implementation examples, and various modifications and improvements made by those skilled in the art using the basic concept of the implementation examples defined in the following claims also fall within the scope of the implementation examples. Therefore, the true technical protection scope of the present disclosure should be determined by the technical idea of the appended claims.

Claims

1. In a video decoding method performed by a decoding device, A step of receiving image information including residual information for a current block; A step of deriving transformation coefficients for the current block based on the residual information; A step of performing inverse transformation based on the above transformation coefficients to derive residual samples for the current block; and A step of generating restoration samples for the current block based on the residual samples, The above image information includes MTS (Multiple Transform Selection) related information, Based on the above MTS-related information, one of the transformation pair candidates of the transformation set for the current block is determined, An image decoding method, characterized in that the number of candidates for the transformation pair of the transformation set is at least one.

2. In paragraph 1, The number of candidates for the above conversion pair is one of 1, 4 or 6, An image decoding method, characterized in that the above one transform pair is configured based on a DCT (discrete cosine transform) 2, DCT 5, DCT 8, DST (discrete sine transform) 4, DST 7, or an identity transform related kernel.

3. In paragraph 1, Based on the prediction mode for the current block being IBC (Intra Block Copy), SGPM (Spatial Geometric Partitioning mode), TIMD (Template-based Intra Mode Derivation) and palette mode, prediction samples are derived. Based on the above prediction samples, an intra prediction mode is derived using the DIMD (Decoder side Intra Mode Derivation) mode. An image decoding method, characterized in that one transform pair is determined based on the intra prediction mode.

4. In paragraph 1, The above MTS related information includes at least one of MTS available flag information, MTS flag information, or MTS index information, The above MTS flag information and the above MTS index information are included in the coding unit syntax or the residual coding syntax, A video decoding method, characterized in that the MTS available flag information is included in a higher level syntax than the MTS flag information and the MTS index information.

5. In paragraph 4, A video decoding method, characterized in that the value of the MTS flag information is 1 and the MTS index information is not included in the video information, based on the number of candidates in the above transformation set being 1.

6. In paragraph 5, Based on the above image information not including the above MTS index information, a specific transformation pair is used or the value of the above MTS index information is derived, An image decoding method, characterized in that the value of the above MTS index information is derived as 0 or 1.

7. In paragraph 4, An image decoding method, characterized in that, based on the number of candidates in the above transformation set being 1, either the first transformation pair among the transformation pair candidates of the four transformation sets or the first transformation pair among the transformation pair candidates of the six transformation sets is selected.

8. In paragraph 4, Based on the number of candidates in the above transformation set being 1, the value of the MTS flag information is 1, Based on the above MTS index information being 0, the first transformation pair among the transformation pair candidates of the four transformation sets is selected, An image decoding method, characterized in that the first transform pair among transform pair candidates of six transform sets is selected based on the above MTS index information being 1.

9. In paragraph 4, Based on the fact that the MTS flag information is not included in the above image information and the value of the MTS index information is 0, MTS is not applied to the inverse transformation. An image decoding method, characterized in that the inverse transformation is performed based on DCT 2 on the above transformation coefficients.

10. In paragraph 4, Based on the identity transformation being applied to the above inverse transformation, one of the six transformation kernels is applied as the identity transformation related kernel, A video decoding method, characterized in that information indicating whether an identity transformation is applied in a horizontal or vertical direction is signaled after the MTS index information.

11. In paragraph 10, The above image information includes flag information indicating whether the identity transformation is applied to the above transformation coefficients, A video decoding method, characterized in that information related to the direction in which the identity transformation is applied is signaled based on the value of the above flag information being 1.

12. In paragraph 11, An image decoding method, characterized in that the image information includes information related to whether an identity transformation is applied in a horizontal direction to the transform coefficients and information related to whether an identity transformation is applied in a vertical direction to the transform coefficients.

13. In paragraph 1, Based on the identity transformation being applied to the above inverse transformation, the image information includes MTS inter available flag information or MTS intra available flag information, The kernel index information representing the above six transformation kernels is individually derived according to the direction in which the identity transformation is applied. An image decoding method, characterized in that the kernel index information includes MTS horizontal index information and MTS vertical index information.

14. In a video encoding method performed by an encoding device, A step of deriving prediction samples for the current block; A step of deriving residual samples for the current block based on the above prediction samples; A step of performing transformation based on the residual samples to derive transformation coefficients for the current block; A step of generating residual information for the current block based on the above transformation coefficients; and Comprising a step of encoding image information including the above residual information, The above image information includes MTS (Multiple Transform Selection) related information, Based on the above MTS-related information, one of the transformation pair candidates of the transformation set for the current block is determined, An image encoding method, characterized in that the number of candidates for the transformation pair of the transformation set is at least one.

15. In paragraph 14, The number of candidates for the above conversion pair is one of 1, 4 or 6, An image encoding method, characterized in that the above one transform pair is configured based on DCT 2, DCT 5, DCT 8, DST 4, DST 7, or an identity transform related kernel.

16. In paragraph 14, Based on the prediction mode for the current block being IBC (Intra Block Copy), SGPM (Spatial Geometric Partitioning mode), TIMD (Template-based Intra Mode Derivation) and palette mode, prediction samples are derived. Based on the above prediction samples, an intra prediction mode is derived using the DIMD (Decoder side Intra Mode Derivation) mode. A video encoding method, characterized in that the one transform pair is determined based on the intra prediction mode.

17. In paragraph 14, The above MTS related information includes at least one of MTS available flag information, MTS flag information, or MTS index information, The above MTS flag information and the above MTS index information are signaled in the coding unit syntax or the residual coding syntax, A video encoding method, characterized in that the MTS available flag information is signaled in a higher level syntax than the MTS flag information and the MTS index information.

18. In paragraph 17, A video encoding method, characterized in that the value of the MTS flag information is signaled as 1 and the MTS index information is not signaled based on the number of candidates in the above transformation set being 1.

19. In Article 17, A video encoding method, characterized in that, based on the number of candidates in the above transformation set being 1, either the first transformation pair among the transformation pair candidates of the four transformation sets or the first transformation pair among the transformation pair candidates of the six transformation sets is selected.

20. A method for transmitting data for an image, comprising: obtaining a bitstream for the image, wherein the bitstream is generated based on a step of deriving prediction samples for a current block, a step of deriving residual samples for the current block based on the prediction samples, a step of performing transformation based on the residual samples to derive transform coefficients for the current block, a step of generating residual information for the current block based on the transform coefficients, and a step of encoding image information including the residual information; and Comprising a step of transmitting the data including the bitstream, The above image information includes MTS (Multiple Transform Selection) related information, Based on the above MTS-related information, one of the transformation pair candidates of the transformation set for the current block is determined, A transmission method, characterized in that the number of candidates for the conversion pair of the conversion set is at least one.

Citation Information

Patent Citations

  • Method and device for harmonizing between transformation skip mode and multiple transformation selection

    KR102591265B1

  • Transform method in picture block encoding, inverse transform method in picture block decoding, and apparatus

    US20210014492A1

  • KR20230169959A