Decoding device, encoding device, and data transmission device

Through the LFNST index encoding method, the transformation skip flag value and transformation coefficient are used to improve the image/video encoding efficiency, and solve the transmission and storage cost problems of high-resolution and high-quality images/video.

CN120075469APending Publication Date: 2025-05-30LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432428.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-11-13
Filing Date
2020-11-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit high-resolution, high-quality image/video data, especially in the transmission of virtual reality and artificial reality content, resulting in increased transmission and storage costs.

Method used

By applying the LFNST index coding method, the modified transformation coefficient is derived, and the transform skip flag value is used to improve encoding efficiency, including deriving variables to parse the LFNST index and selecting the optimal LFNST matrix for encoding.

Benefits of technology

Improves image/video compression efficiency, reduces transmission and storage costs, and is suitable for encoding and decoding of high-resolution and high-quality images/videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075469A_ABST
    Figure CN120075469A_ABST
Patent Text Reader

Abstract

The invention provides a decoding apparatus, an encoding apparatus, and a data transmission apparatus. An image decoding method according to the present document comprises a step of deriving a modified transform coefficient by applying LFNST to the transform coefficient, in which the step of deriving the modified transform coefficient comprises the steps of: deriving a variable indicating whether a significant coefficient exists in a DC component of a current block; and parsing the LFNST index based on variables indicating that significant coefficients exist at locations other than the DC component, where the variables may be derived based on respective transform skip flag values for color components of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the original application number 202080092623.X (International Application No.: PCT / KR2020 / 016004, Application Date: November 13, 2020, Invention Title: Transform-based Image Coding Method and Apparatus). Technical Field

[0002] The present disclosure relates to an image coding technology, and more particularly, to a method and apparatus for coding an image based on a transform in an image coding system. Background Art

[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been continuously increasing in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0004] In addition, nowadays, the interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.

[0005] Therefore, there is a need for an efficient image / video compression technology that can effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention

[0006] Technical Problem

[0007] One technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of image coding.

[0008] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of LFNST index coding.

[0009] Still another technical aspect of the present disclosure is to provide a method and apparatus for increasing the coding efficiency of LFNST indexes based on a transform skip flag.

[0010] Technical Solution

[0011] In one aspect, an image decoding method performed by a decoding device is provided. The method includes: deriving modified transform coefficients by applying LFNST to transform coefficients; wherein, deriving the modified transform coefficients includes: deriving a variable indicating whether valid coefficients exist in the DC component of the current block; and parsing the LFNST index based on the variable indicating that valid coefficients exist at positions other than the DC component, and the variable can be derived based on the respective transform skip flag values of the color components of the current block.

[0012] Based on the transform skip flag value of the color component being 0, the variable can indicate that valid coefficients exist at positions other than the DC component.

[0013] The variable can be initially set to 1 at the coding unit level of the current block. If the transform skip flag value is 0, the variable can be changed to 0 at the residual coding level, and the LFNST index can be parsed based on the variable being 0.

[0014] The transform skip flag of the current block can be signaled for each color component.

[0015] Deriving the modified transform coefficients can further include: setting a plurality of variables for LFNST based on whether the LFNST index is not 0 and whether the respective transform skip flag values of the color components are 0.

[0016] If the tree type of the current block is a single tree, the variable can be derived based on the transform skip flag value of the luminance component, the transform skip flag value of the chrominance Cb component, and the transform skip flag value of the chrominance Cr component.

[0017] If the tree type of the current block is a dual-tree luminance, the variable can be derived based on the transform skip flag value of the luminance component.

[0018] If the tree type of the current block is a dual-tree chrominance, the variable can be derived based on the transform skip flag value of the chrominance Cb component and the transform skip flag value of the chrominance Cr component.

[0019] According to an embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method includes: applying LFNST to derive modified transform coefficients from transform coefficients, applying a plurality of LFNST matrices to the transform coefficients to derive a variable indicating whether valid coefficients exist in the DC component of the current block; selecting an optimal LFNST matrix among the LFNST kernels based on the variable indicating that valid coefficients exist at positions other than the DC component, and deriving modified transform coefficients based on the selected LFNST matrix, and the variable can be derived based on the respective transform skip flag values of the color components of the current block.

[0020] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.

[0021] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and a bitstream to cause a decoding device to perform an image decoding method.

[0022] Technical Effects

[0023] According to the present disclosure, the overall image / video compression efficiency can be increased.

[0024] According to the present disclosure, the efficiency of LFNST index coding can be increased.

[0025] According to the present disclosure, the coding efficiency of the LFNST index can be increased based on a transform skip flag.

[0026] The effects that can be obtained through specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to those clearly described in the present disclosure and may include various effects that can be understood or derived based on the technical features of the present disclosure. Brief Description of the Drawings

[0027] Figure 1 An example of a video / image coding system to which the present disclosure can be applied is schematically illustrated.

[0028] Figure 2 is a diagram schematically illustrating the configuration of a video / image coding device to which the present disclosure can be applied.

[0029] Figure 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure can be applied.

[0030] Figure 4 is a diagram exemplarily illustrating the structural diagram of a content stream system to which the present disclosure is applied.

[0031] Figure 5 is a diagram schematically illustrating a multi-transform technique according to an embodiment of the present disclosure.

[0032] Figure 6 is a diagram schematically illustrating an intra-frame directional mode of 65 prediction directions.

[0033] Figure 7 is a diagram for describing the RST according to an embodiment of the present disclosure.

[0034] Figure 8 FIG. is an example of arranging the output data of the forward first transformation in the order of a one-dimensional vector according to the example.

[0035] Figure 9 FIG. is an example of arranging the output data of the forward second transformation in the order of a one-dimensional vector according to the example.

[0036] Figure 10 FIG. is an example of the wide-angle intra prediction mode according to an embodiment of the present disclosure.

[0037] Figure 11 FIG. is an example of the block shape to which LFNST is applied.

[0038] Figure 12 FIG. is an example of the arrangement of the output data of the forward LFNST according to the example.

[0039] Figure 13 FIG. is an example of zero clearing in the block where 4×4 LFNST is applied according to the example.

[0040] Figure 14 FIG. is an example of zero clearing in the block where 8×8 LFNST is applied according to the example.

[0041] Figure 15 FIG. is a flowchart illustrating the operation of a video decoding device according to an embodiment.

[0042] Figure 16 FIG. is a flowchart illustrating the operation of a video encoding device according to an embodiment. DETAILED DESCRIPTION

[0043] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the technical concept of the present disclosure. Unless the context clearly indicates otherwise, the singular forms may include the plural forms. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.

[0044] In addition, for the convenience of describing different characteristic functions from each other, each component in the drawings described herein is illustrated independently. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. The implementation manners in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.

[0045] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. In addition, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.

[0046] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).

[0047] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, these embodiments may be combined with each other and executed.

[0048] In this document, video may refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit representing an image of a specific time region, and a slice / tile is a unit that forms a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.

[0049] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample may mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it may mean a transform coefficient in the frequency domain.

[0050] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, terms such as unit and terms like block, region, etc. may be used interchangeably. Usually, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows.

[0051] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". Additionally, "A / B / C" may mean "at least one of A, B, and / or C".

[0052] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0053] In this disclosure, "at least one of A and B" may mean "only A", "only B", or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as "at least one of A and B".

[0054] Furthermore, in this disclosure, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0055] Additionally, the parentheses used in this disclosure may represent "for example". Specifically, when indicated as "prediction (intra prediction)", it may mean that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra prediction", and "intra prediction" is presented as an example of "prediction". Additionally, when indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".

[0056] The technical features described separately in one of the drawings in this disclosure may be implemented separately or may be implemented simultaneously.

[0057] Figure 1 An example of a video / image coding system to which this disclosure can be applied is schematically illustrated.

[0058] Reference Figure 1 , a video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transfer encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.

[0059] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0060] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0061] The encoding device may encode the input video / image. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0062] The transmitter may send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for sending via a broadcast / communication network. The receiver may receive / extract the bitstream and send the received / extracted bitstream to the decoding device.

[0063] The decoding device may decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.

[0064] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display.

[0065] Figure 2 FIG. is a diagram schematically illustrating a configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.

[0066] Referring to Figure 2 , the encoding device 200 may include an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0067] The image divider 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively divided according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, one coding unit may be divided into multiple coding units with a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is not further divided. In this case, based on the encoding efficiency according to the image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units with a deeper depth as needed, whereby the coding unit of the optimal size may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separated or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0068] Depending on the situation, the terms unit and terms such as block, region, etc. may be used in place of each other. Generally, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to the pixels or pels of a picture (or image).

[0069] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the block to be processed (hereinafter referred to as "current block") and can generate a prediction block including the prediction samples of the current block. The predictor 220 can determine whether to apply intra prediction or inter prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction such as prediction mode information and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0070] The intra predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include a variety of non - directional modes and a variety of directional modes. The non - directional modes can include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0071] The inter - frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted on a block, sub - block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as or different from each other. The temporal neighboring block can be referred to as a collocated reference block, collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, a residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.

[0072] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra - frame prediction or inter - frame prediction to the prediction of a block, and can also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode in order to perform prediction on a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games like screen content coding (SCC). Although IBC basically performs prediction within the current block, the way it performs is similar to inter - frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter - frame prediction techniques described in this disclosure.

[0073] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 may be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process may be applied to square pixel blocks of the same size, or may be applied to blocks of variable size rather than square blocks.

[0074] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients of a block type into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 can perform various coding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 can encode, together or separately, information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information can also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. Additionally, the video / image information can also include general constraint information. In the present disclosure, the information and / or syntax elements sent from the encoding device to / signaled to the decoding device can be included in the video / image information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be transmitted through a network or stored in a digital storage medium. Here, the network can include a broadcast network, a communication network, and / or the like, and the digital storage medium can include various storage media such as a USB, an SD, a CD, a DVD, a Blu-ray, an HDD, an SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 or a memory (not shown) that stores it can be configured as an internal / external element of encoding device 200, or the transmitter can be included in entropy encoder 240.

[0075] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual samples) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target picture, and as will be described later, can be used for inter-frame prediction of the next picture by filtering.

[0076] In addition, in the picture encoding and / or reconstruction processing, a luminance mapping with chroma scaling (LMCS) can be applied.

[0077] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can store the modified reconstructed picture in the memory 270, especially in the DPB of the memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As will be discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0078] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatches in the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve the encoding efficiency.

[0079] The memory 270 DPB can store the modified reconstructed picture so as to use it as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.

[0080] Figure 3It is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure can be applied.

[0081] Referring to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0082] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information that has been processed in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding by using the processing units applied in the encoding device. Thus, the decoded processing unit may be, for example, an encoding unit, which may be divided along a quadtree structure, a binary tree structure, and / or a ternary tree structure with a coding tree unit or a maximum coding unit. One or more transform units may be derived with the encoding unit. And, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproducer.

[0083] The decoding device 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. Additionally, the video / image information can also include conventional constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the conventional constraint information. In the present disclosure, the signaled / received information and / or syntax elements described subsequently can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. Among the information decoded in the entropy decoder 310, the information about prediction can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients) and associated parameter information for which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output by the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to the present disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0084] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.

[0085] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.

[0086] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically can determine the intra / inter prediction mode.

[0087] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to the prediction of a block, and can also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform intra-block copy (IBC) for the prediction of a block. Intra-block copy can be used for content image / video coding such as games like screen content coding (SCC). Although IBC basically performs prediction in the current block, the way it is performed is similar to inter prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0088] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0089] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the adjacent blocks can include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter-frame prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the mode of the inter-frame prediction for the current block.

[0090] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.

[0091] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for the intra-frame prediction of the next processing target block in the current block, and as described later, can be output through filtering or used for the inter-frame prediction of the next picture.

[0092] In addition, in the picture decoding process, a luminance mapping with chroma scaling (LMCS) can be applied.

[0093] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, especially to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0094] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - frame predictor 332. The memory 360 can store the motion information of blocks in the current picture from which the motion information has been derived (or decoded) and / or the motion information of blocks in the reconstructed pictures. The stored motion information can be sent to the inter - frame predictor 221 to be used as the motion information of neighboring blocks or the motion information of temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra - frame predictor 331.

[0095] In this specification, the examples described in the predictor 330, de - quantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, de - quantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.

[0096] As described above, prediction is performed to improve the compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for the current block, which is the block to be encoded, can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve the image encoding efficiency by signaling to the decoding device not the original sample values of the original block itself but the information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.

[0097] The residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the associated residual information to the decoding device (through the bitstream). Here, the residual information can include value information, position information, transformation technology, transformation kernel, quantization parameters, etc. of the quantized transform coefficients. The decoding device can perform a quantization / de - quantization process based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can de - quantize / inverse - transform the quantized transform coefficients to derive the residual block for use as a reference for inter - frame prediction of the next picture and can generate a reconstructed picture based on this.

[0098] Figure 4 The structure of a content - streaming system to which the present disclosure is applied is illustrated.

[0099] In addition, the content stream system applying the present disclosure may generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.

[0100] The encoding server is used to compress the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server may be omitted. The bitstream may be generated by applying the encoding method or the bitstream generation method of the present disclosure. And the streaming server may temporarily store the bitstream during the process of sending or receiving the bitstream.

[0101] The streaming server sends multimedia data to the user device via the web server based on the user's request, and the web server serves as a means for notifying the user of what services exist. When the user requests a service that the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content stream system may include a separate control server, and in this case, the control server is used to control the commands / responses between the corresponding devices in the content stream system.

[0102] The streaming server may receive content from the media storage device and / or the encoding server. For example, in the case of receiving content from the encoding server, the content may be received in real time. In this case, in order to smoothly provide the streaming service, the streaming server may store the bitstream for a predetermined time.

[0103] For example, the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a slate PC, a tablet PC, a superbook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. Each server in the content stream system may operate as a distributed server, and in this case, the data received by each server may be processed in a distributed manner.

[0104] Figure 5 The multi-transform technology according to an embodiment of the present disclosure is schematically illustrated.

[0105] Referring to Figure 5 , the transformer may correspond to the transformer in the encoding device described above Figure 2 , and the inverse transformer may correspond to the inverse transformer in the encoding device described above Figure 2 , or the inverse transformer in the decoding device described above Figure 3 .

[0106] The transformer can derive (a first) transform coefficient (S510) by performing a first transformation based on the residual samples (residual sample array) in the residual block. This first transformation can be referred to as the core transformation. In this document, the first transformation can be based on multi-transformation selection (MTS), and when multi-transformation is used as the first transformation, it can be referred to as multi-core transformation.

[0107] The multi-core transformation can represent a method of additionally performing transformations using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1. That is, the multi-core transformation can represent a transformation method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or first transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, from the perspective of the transformer, the first transform coefficients can be referred to as temporary transform coefficients.

[0108] In other words, when applying the conventional transformation method, transform coefficients can be generated by applying a transformation from the spatial domain to the frequency domain based on DCT type 2 to the residual signal (or residual block). In contrast, when applying the multi-core transformation, transform coefficients (or first transform coefficients) can be generated by applying a transformation from the spatial domain to the frequency domain based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1 to the residual signal (or residual block). In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transformation types, transform kernels, or transform cores. These DCT / DST transformation types can be defined based on basis functions.

[0109] When performing the multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transformation can be performed on the target block based on the vertical transform kernel, and a horizontal transformation can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transformation can indicate the transformation of the horizontal component of the target block, and the vertical transformation can indicate the transformation of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.

[0110] In addition, according to the example, if a transformation is performed by applying MTS, the mapping relationship of the transformation kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transformation kernel is represented as trTypeHor and the vertical transformation kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.

[0111] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transformation kernel sets. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.

[0112] In one example, the transformation kernel sets according to the MTS index information are shown in the following table.

[0113] [Table 1]

[0114] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2

[0115] The transformer may perform a secondary transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S520). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation refers to using the correlation existing between the (primary) transformation coefficients to transform into a more compact representation. The secondary transformation may include an inseparable transformation. In this case, the secondary transformation may be referred to as an inseparable secondary transformation (NSST) or a mode-dependent inseparable secondary transformation (MDNSST). The NSST may represent a transformation of performing a secondary transformation on the (primary) transformation coefficients derived through the primary transformation based on an inseparable transformation matrix to generate modified transformation coefficients (or secondary transformation coefficients) for the residual signal. Here, based on the inseparable transformation matrix, the transformation can be applied once to the (primary) transformation coefficients without separating the vertical transformation and the horizontal transformation (or independently applying the horizontal / vertical transformation). In other words, the NSST does not separately apply to the (primary) transformation coefficients in the vertical and horizontal directions, and may represent, for example, a transformation method of rearranging a two-dimensional signal (transformation coefficients) into a one-dimensional signal in a specific predetermined direction (e.g., row-major order or column-major order) and then generating modified transformation coefficients (or secondary transformation coefficients) based on the inseparable transformation matrix. For example, the row-major order is arranged in rows in the order of the first row, the second row, …, and the Nth row for M×N blocks, and the column-major order is arranged in columns in the order of the first column, the second column, …, and the Mth column for M×N blocks. The NSST may be applied to the upper left region of a block configured with (primary) transformation coefficients (hereinafter referred to as a transformation coefficient block). For example, when both the width W and the height H of the transformation coefficient block are 8 or more, an 8×8 NSST may be applied to the upper left 8×8 region of the transformation coefficient block. In addition, while both the width (W) and the height (H) of the transformation coefficient block are 4 or more, when the width (W) or the height (H) of the transformation coefficient block is less than 8, a 4×4 NSST may be applied to the upper left min(8, W)×min(8, H) region of the transformation coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width W or the height H of the transformation coefficient block is 4 or more is satisfied, a 4×4 NSST may be applied to the upper left end min(8, W)×min(8, H) region of the transformation coefficient block.

[0116] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transformation may be performed as follows.

[0117] The 4×4 input block X may be represented as follows.

[0118] [Equation 1]

[0119]

[0120] If X is represented in the form of a vector, the vector It can be expressed as follows.

[0121] [Formula 2]

[0122]

[0123] In Formula 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Formula 1 according to row-major order.

[0124] In this case, the inseparable quadratic transformation can be calculated as follows.

[0125] [Formula 3]

[0126]

[0127] In this formula, represents the transformation coefficient vector, and T represents a 16×16 (inseparable) transformation matrix.

[0128] Through the aforementioned Formula 3, a 16×1 transformation coefficient vector can be derived, and the vector can be reorganized into 4×4 blocks in a scanning order (such as horizontal, vertical, and diagonal, etc.). However, the above calculation is an example, and the hypercube-Givens transform (HyGT) etc. can also be used for the calculation of the inseparable quadratic transformation to reduce the computational complexity of the inseparable quadratic transformation.

[0129] In addition, in the inseparable quadratic transformation, the transformation kernel (or transformation core, transformation type) can be selected to be mode-dependent. In this case, the mode can include the intra prediction mode and / or the inter prediction mode.

[0130] As described above, the inseparable quadratic transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. The 8×8 transformation refers to a transformation that can be applied to the 8×8 region included in the transformation coefficient block when both W and H are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transformation coefficient block. Similarly, the 4×4 transformation refers to a transformation that can be applied to the 4×4 region included in the transformation coefficient block when both W and H are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transformation coefficient block. For example, the 8×8 transformation kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transformation kernel matrix can be a 16×16 / 8×16 matrix.

[0131] Here, in order to select a transform kernel related to a mode, for both the 8×8 transform and the 4×4 transform, two non-separable quadratic transform kernels can be configured for each transform set for the non-separable quadratic transform, and there can be four transform sets. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.

[0132] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transform kernels in each set can be k.

[0133] The transform sets can be referred to as NSST sets or LFNST sets. A specific set among the transform sets can be selected, for example, based on the intra prediction mode of the current block (CU or sub-block). The low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and represents a non-separable transform for low-frequency components.

[0134] As a reference, for example, the intra prediction mode can include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is an example, and this document can be applied even if the number of intra prediction modes is different. In addition, in some cases, the intra prediction mode numbered 67 can also be used, and the intra prediction mode numbered 67 can represent a linear model (LM) mode.

[0135] Figure 6 The intra-frame directional mode with 65 prediction directions is schematically shown.

[0136] Refer to Figure 6 , based on the intra prediction mode 34 with the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode with horizontal directivity and an intra prediction mode with vertical directivity. In Figure 6In this case, H and V respectively indicate the horizontal directionality and the vertical directionality, and the numbers -32 to 32 indicate displacements of 1 / 32 units at the sample grid positions. These numbers can represent offsets for the mode index values. Intra prediction modes 2 to 33 have horizontal directionality, and intra prediction modes 34 to 66 have vertical directionality. Strictly speaking, intra prediction mode 34 can be regarded as neither horizontal nor vertical, but can be classified as belonging to the horizontal directionality when determining the transform set of the secondary transform. This is because the input data is transposed for the vertically oriented mode symmetric to intra prediction mode 34, and the input data alignment method for the horizontal mode is used for intra prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data to N×M data. Intra prediction mode 18 and intra prediction mode 50 can respectively represent the horizontal intra prediction mode and the vertical intra prediction mode, and intra prediction mode 2 can be called the upper-right diagonal intra prediction mode because intra prediction mode 2 has a left reference pixel and performs prediction in the upper-right direction. Similarly, intra prediction mode 34 can be called the lower-right diagonal intra prediction mode, and intra prediction mode 66 can be called the lower-left diagonal intra prediction mode.

[0137] According to an example, four transform sets can be mapped according to the intra prediction mode, for example, as shown in the following table.

[0138] [Table 2]

[0139] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1

[0140] As shown in Table 2, any one of the four transform sets, that is, lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra prediction mode.

[0141] When determining a specific set for the non-separable transform, one of the k transform kernels in the specific set can be selected by the non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index indicating the specific transform kernel based on rate distortion (RD) checking, and can signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable secondary transform index. For example, the lfnst index value 0 can refer to the first non-separable secondary transform kernel, the lfnst index value 1 can refer to the second non-separable secondary transform kernel, and the lfnst index value 2 can refer to the third non-separable secondary transform kernel. Alternatively, the lfnst index value 0 can indicate that the first non-separable secondary transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate three transform kernels.

[0142] The transformer may perform an inseparable quadratic transform based on the selected transform kernel, and may obtain modified (quadratic) transform coefficients. As described above, the modified transform coefficients may be derived as the transform coefficients quantized by the quantizer, and may be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0143] In addition, as described above, if the quadratic transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform may be derived as the transform coefficients quantized by the quantizer as described above, and may be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0144] The inverse transformer may perform a series of processes in an order opposite to the order that has been performed in the above-mentioned transformer. The inverse transformer may receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing a quadratic (inverse) transform (S550), and may obtain a residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S560). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.

[0145] The decoding device may further include a quadratic inverse transform application determiner (or an element for determining whether to apply a quadratic inverse transform) and a quadratic inverse transform determiner (or an element for determining the quadratic inverse transform). The quadratic inverse transform application determiner may determine whether to apply a quadratic inverse transform. For example, the quadratic inverse transform may be NSST, RST, or LFNST, and the quadratic inverse transform application determiner may determine whether to apply a quadratic inverse transform based on the quadratic transform flag obtained by parsing the bitstream. In another example, the quadratic inverse transform application determiner may determine whether to apply a quadratic inverse transform based on the transform coefficients of the residual block.

[0146] The quadratic inverse transform determiner may determine the quadratic inverse transform. In this case, the quadratic inverse transform determiner may determine the quadratic inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra prediction mode. In an embodiment, the quadratic transform determination method may be determined depending on the primary transform determination method. Various combinations of the primary transform and the quadratic transform may be determined according to the intra prediction mode. In addition, in an example, the quadratic inverse transform determiner may determine the region to which the quadratic inverse transform is applied based on the size of the current block.

[0147] In addition, as described above, if the secondary (inverse) transformation is omitted, the (dequantized) transform coefficients can be received, a single (separable) inverse transformation can be performed, and a residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.

[0148] In addition, in the present disclosure, a reduced secondary transform (RST) in which the size of the transform matrix (kernel) is reduced can be applied in the concept of NSST, so as to reduce the amount of computation and storage required for the non-separable secondary transform.

[0149] In addition, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented by 8 bits. This can be a condition implemented in the decoding device and the encoding device, and compared with the existing 9 bits or 10 bits, it can reduce the storage amount required for storing the transform kernel, and can reasonably adapt to performance degradation. In addition, representing the kernel matrix by 8 bits can allow the use of small multipliers, and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.

[0150] In this specification, the term "RST" may refer to a transform performed on the residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing a reduction transform, since the size of the transform matrix is reduced, the amount of computation required for the transform can be reduced. That is, RST can be used to solve the computational complexity problems that occur in the transform of large-sized blocks or non-separable transforms.

[0151] RST can be referred to by various terms such as reduction transform, reduced secondary transform, scaled transform, simplified transform, and simple transform, and the names that RST can be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be referred to as a low-frequency non-separable transform (LFNST). The transform index can be referred to as the LFNST index.

[0152] In addition, when performing a secondary inverse transform based on RST, the inverse transformers 235 of the encoding device 200 and the inverse transformers 322 of the decoding device 300 can include: an inverse reduced secondary transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse primary transformer that derives the residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying the transform.

[0153] Figure 7FIG. is an illustration of an RST according to an embodiment of the present disclosure.

[0154] In the present disclosure, a "target block" may refer to a current block to be encoded, a residual block, or a transform block.

[0155] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduction transform matrix can be determined, where R is less than N. N may refer to the square of the length of the side of the block to which the transform is applied, or the total number of transform coefficients corresponding to the block to which the transform is applied, and the reduction factor may refer to the R / N value. The reduction factor may be referred to as a reduction factor, a shrinking factor, a simplification factor, a simple factor, or various other terms. In addition, R may be referred to as a reduction coefficient, but depending on the situation, the reduction factor may refer to R. In addition, depending on the situation, the reduction factor may refer to the N / R value.

[0156] In the example, the reduction factor or reduction coefficient may be signaled through a bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient may not be signaled separately.

[0157] The size of the reduction transform matrix according to the example may be R×N, which is less than N×N (the size of the conventional transform matrix), and may be defined as in Equation 4 below.

[0158] [Equation 4]

[0159]

[0160] Figure 7 The matrix T in the reduction transform block shown in (a) of may refer to the matrix T of Equation 4 R×N . As Figure 7 shown in (a) of, when the reduction transform matrix T R×N is multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.

[0161] In the example, if the size of the block to which the transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 7 (a) of can be represented as the matrix operation shown in Equation 5 below. In this case, the storage and multiplication calculations can be reduced to approximately 1 / 4 by the reduction factor.

[0162] In the present disclosure, matrix operation can be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix set on the left side of the column vector.

[0163] [Equation 5]

[0164]

[0165] In Equation 5, r 1 to r 64 can represent the residual samples of the target block, and specifically can be the transform coefficients generated by applying a single transform. As a result of the calculation of Equation 5, the transform coefficients c i of the target block can be derived, and the process of deriving c i can be as shown in Equation 6.

[0166] [Equation 6]

[0167]

[0168] As a result of the calculation of Equation 6, the transform coefficients c 1 to c R of the target block can be derived. That is, when R = 16, the transform coefficients c 1 to c 16 of the target block can be derived. If a conventional transform instead of RST is applied and a transform matrix of size 64×64 (N×N) is multiplied by residual samples of size 64×1 (N×1), then only 16 (R) transform coefficients are derived for the target block because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data sent from the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.

[0169] When considered from the perspective of the size of the transform matrix, the size of the conventional transform matrix is 64×64 (N×N), but the size of the reduced transform matrix is reduced to 16×64 (R×N). Therefore, compared with the case of performing a conventional transform, the storage utilization rate in the case of performing RST can be reduced by the ratio of R / N. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional transform matrix, using the reduced transform matrix can reduce the number of multiplication calculations (R×N) by the ratio of R / N.

[0170] In the example, the transformer 232 of the encoding device 200 can derive the transform coefficients of the target block by performing a single transform on the residual samples of the target block and a secondary transform based on RST. These transform coefficients can be transmitted to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive the modified transform coefficients based on the inverse reduced secondary transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse single transform for the modified transform coefficients.

[0171] Inverse RST matrix T according to the exampleN×R is of size N×R, which is smaller than the size of the conventional inverse transform matrix N×N, and is related to the reduced transform matrix T shown in Equation 4 R×N by a transpose relationship.

[0172] Figure 7 The matrix T in the reduced inverse transform block shown in (b) of t can refer to the inverse RST matrix T N×R T (the superscript T refers to transpose). As shown in Figure 7 (b) of N×R T when the inverse RST matrix T R×N T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. The inverse RST matrix T R×N can be expressed as (T T N×R ).

[0173] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.

[0174] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST in Figure 7 (b) of

[0175] [Equation 7]

[0176]

[0177] In Equation 7, c 1 to c 16 can represent the transform coefficients of the target block. As a result of the calculation in Equation 7, r i , which represents the modified transform coefficients of the target block or the residual samples of the target block, can be derived, and the process of deriving r i can be as shown in Equation 8.

[0178] [Equation 8]

[0179]

[0180] As a result of the calculation of Equation 8, r representing the modified transform coefficients of the target block or the residual samples of the target block can be derived 1 to r N Considering from the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduction transform matrix is reduced to 64×16 (R×N). Therefore, compared with the case of performing the conventional inverse transform, the storage utilization rate in the case of performing the inverse RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, using the inverse reduction transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.

[0181] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since according to the intra prediction mode, one transform set includes two or three transforms (kernels), it can be configured to select one of up to four transforms including the case without applying the secondary transform. Among the transforms without applying the secondary transform, applying the identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case of applying the identity matrix, that is, the case without applying the secondary transform), the transform index or lfnst index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, through the transform index, 8×8 NSST in the RST configuration can be specified, or 8×8 lfnst can be specified when LFNST is applied. 8×8 lfnst and 8×8 RST refer to the transforms that can be applied to the 8×8 region included in the transform coefficient block when both the W and H of the target block to be transformed are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to the transforms that can be applied to the 4×4 region included in the transform coefficient block when both the W and H of the target block are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block.

[0182] According to an embodiment of the present disclosure, for the transform in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when performing RST by applying an m×48 transform kernel matrix (m≤16) to an 8×8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 pieces of data constituting a region other than the lower right 4×4 region among the 8×8 region. Here, when performing matrix operations by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.

[0183] For the inverse transform in the decoding process, the transpose matrix of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST in the inverse transform process performed by the decoding device, the input coefficient data for applying inverse RST is configured in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged in a two-dimensional block according to a predetermined arrangement order.

[0184] In summary, in the transform process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on 48 transform coefficients in the upper left region, upper right region, and lower left region of the 8×8 region except for the lower right region with a 16×48 transform kernel matrix. For the matrix operations, 48 transform coefficients are input in a one-dimensional array. When performing the matrix operations, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the upper left region of the 8×8 region.

[0185] Conversely, in the inverse transformation process, when applying the inverse RST or LFNST to an 8×8 region, the 16 transform coefficients corresponding to the upper-left region of the 8×8 region among the transform coefficients in the 8×8 region can be input as a one-dimensional array according to the scanning order, and can undergo matrix operations with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix and can thus be represented as an n×1 column vector. In addition, * represents matrix multiplication. When performing the matrix operation, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper-left region, upper-right region, and lower-left region of the 8×8 region except for the lower-right region.

[0186] When the inverse quadratic transformation is based on RST, the inverse transformers 235 of the encoding device 200 and the inverse transformers 322 of the decoding device 300 can include an inverse reduced quadratic transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse primary transformer for deriving the residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In the present disclosure, deriving transform coefficients based on a transform can refer to deriving transform coefficients by applying the transform.

[0187] The non-separable transform (LFNST) described above will be described in detail below. The LFNST can include a forward transform performed by the encoding device and an inverse transform performed by the decoding device.

[0188] The encoding device receives the result (or a part of the result) derived after applying the primary (core) transform as input and applies the forward quadratic transform (quadratic transform).

[0189] [Equation 9]

[0190] y = G T x

[0191] In Equation 9, x and y are the input and output of the quadratic transform respectively, G is the matrix representing the quadratic transform, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of the forward LFNST, the transpose of the matrix G becomes the dimension of G T of.

[0192] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transform basis vectors sampled from the left sides of the [48×16] matrix and the [16×16] matrix, respectively.

[0193] On the other hand, for the forward LFNST, matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transform basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.

[0194] Therefore, in the case of the forward LFNST, a [48×1] vector or a [16×1] vector can be used as the input x, and a [16×1] vector or an [8×1] vector can be used as the output y. In video coding and decoding, the output of the forward single transform is two-dimensional (2D) data. Therefore, in order to construct a [48×1] vector or a [16×1] vector as the input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data that is the output of the forward transform.

[0195] Figure 8 is a diagram illustrating the order of arranging the output data of the forward single transform into a one-dimensional vector according to the example. Figure 8 The left diagrams of (a) and (b) of illustrate the order for constructing a [48×1] vector, and Figure 8 the right diagrams of (a) and (b) of illustrate the order for constructing a [16×1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by arranging the 2D data in the same order as in Figure 8 the (a) and (b) of.

[0196] The arrangement direction of the output data of the forward single transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 8 the (a) of, and when the intra prediction mode of the current block is in the vertical direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 8 the (b) of.

[0197] According to the example, an arrangement order different from Figure 8 the (a) and (b) of can be applied, and in order to derive and apply Figure 8The same result (y vector) for the arrangement orders of (a) and (b) can rearrange the column vectors of matrix G according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.

[0198] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data during the process of using the result of the forward quadratic transformation as input (for example, during quantization or residual coding), the output y vector of Equation 9 needs to be appropriately rearranged into 2D data again.

[0199] Figure 9 is a diagram illustrating the order of arranging the output data of the forward quadratic transformation into two-dimensional blocks according to the example.

[0200] In the case of LFNST, the output values can be arranged in a 2D block according to a predetermined scan order. Figure 9 (a) of shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions in the 2D block according to the diagonal scan order. Figure 9 (b) of shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions in the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 9 X in (b) indicates that it is filled with zeros.

[0201] According to another example, since the order of processing the output vector y can be preset during quantization or residual coding, the output vector y may not be arranged in a 2D block as shown in Figure 9 . However, in the case of residual coding, data coding can be performed in 2D block (e.g., 4×4) units (e.g., CG (coefficient group)), and in this case, the data is arranged according to a specific order in the diagonal scan order as shown in Figure 9 .

[0202] In addition, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output through the dequantization process according to a preset scan order for inverse transformation. The input vector y can be output as the output vector x through the following formula.

[0203] [Equation 10]

[0204] x = Gy

[0205] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.

[0206] The output vector x according toFigure 8 are arranged in the order shown in a two-dimensional block and are arranged as two-dimensional data, and this two-dimensional data becomes the input data (or a part of the input data) of the inverse first transform.

[0207] Therefore, the inverse second transform is overall the reverse of the forward second transform process, and in the case of the inverse transform, different from the forward direction, the inverse second transform is first applied, and then the inverse first transform is applied.

[0208] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transform matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.

[0209] In addition, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can consist of two matrices. Which transform set to use among the four transform sets is determined according to the intra-frame prediction mode, and more specifically, based on the value of the intra-frame prediction mode extended by considering wide-angle intra-frame prediction (WAIP). Which matrix to select from the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the transmitted index values, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra-frame prediction mode value.

[0210] Figure 10 FIG. is an illustration of a wide-angle intra-frame prediction mode according to an embodiment of this document.

[0211] The general intra-frame prediction mode value can have values from 0 to 66 and from 81 to 83, and the intra-frame prediction mode value extended due to WAIP can have values from -14 to 83 as shown. The values from 81 to 83 indicate the CCLM (Cross-Component Linear Model) mode, and the values from -14 to -1 and from 67 to 80 indicate the intra-frame prediction mode extended due to the application of WAIP.

[0212] When the width of the current prediction block is greater than the height, the upper reference pixel is generally closer to the position inside the block to be predicted. Therefore, predicting in the lower left direction can be more accurate than in the upper right direction. On the contrary, when the height of the block is greater than the width, the left reference pixel is generally closer to the position inside the block to be predicted. Therefore, predicting in the upper right direction can be more accurate than in the lower left direction. Therefore, it can be advantageous to apply remapping (i.e., mode index modification) to the index of the wide-angle intra-frame prediction mode.

[0213] When applying wide-angle intra prediction, information about existing intra prediction can be signaled, and after the information is parsed, the information can be remapped to an index of the wide-angle intra prediction mode. Thus, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) can remain unchanged, that is, the total number of intra prediction modes is 67, and the intra prediction mode coding for a specific block can remain unchanged.

[0214] Table 3 below shows the process of deriving the modified intra mode by remapping the intra prediction mode to the wide-angle intra prediction mode.

[0215] [Table 3]

[0216]

[0217] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not divided into sub-partitions by the intra sub-partition (ISP) technique currently adopted in the VVC standard, and the cIdx variable values of 0, 1, and 2 indicate the cases of the luminance component, Cb component, and Cr component, respectively. The log2 function shown in Table 3 returns the log value with base 2, and the Abs function returns the absolute value.

[0218] Variables such as the predModeIntra indicating the intra prediction mode and the height and width of the transform block are used as input values for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or coding block can be the height and width of the current block for the remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to width can be set to Abs(Log2(nW / nH)).

[0219] For non-square blocks, the intra prediction mode can be divided into two cases and modified.

[0220] First, if all of conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before modification is equal to or greater than 2, and (3) the intra prediction mode is less than the value derived as (8 + 2 * whRatio) when the variable whRatio is greater than 1 and less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra < (whRatio > 1)? (8 + 2 * whRatio) : 8), then the intra prediction mode is set to a value 65 greater than predModeIntra [predModeIntra is set to be equal to (predModeIntra + 65)].

[0221] If different from the above, i.e., if conditions (1) to (3) are satisfied: (1) the height of the current block is greater than the width, (2) the intra prediction mode before modification is less than or equal to 66, and (3) the intra prediction mode is greater than the value derived as (60 - 2 * whRatio) when whRatio is greater than 1 and greater than 60 when whRatio is less than or equal to 1 (predModeIntra > (whRatio > 1)? (60 - 2 * whRatio) : 60), then the intra prediction mode is set to a value 67 less than predModeIntra [predModeIntra is set to be equal to (predModeIntra - 67)].

[0222] Table 2 above shows how to select a transform set in LFNST based on the intra prediction mode values extended by WAIP. As Figure 9 shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same transform set is applied to modes located in symmetric directions, and this symmetry is also reflected in Table 2.

[0223] In addition, it is assumed that the forward LFNST input data of mode 54 is symmetric to the forward LFNST input data of mode 14. For example, for mode 14 and mode 54, the two-dimensional data is rearranged into one-dimensional data according to the arrangement order shown in (a) of Figure 8 and (b) of Figure 8 . Additionally, it can be seen that the patterns in the order shown in (a) of Figure 8 and (b) of Figure 8 are symmetric about the direction indicated by mode 34 (diagonal direction).

[0224] In addition, as described above, it is determined by the size and shape of the transform target block which of the [48×16] matrix and the [16×16] matrix is applied to LFNST.

[0225] Figure 11 is a diagram illustrating the block shapes to which LFNST is applied. Figure 11 (a) of Figure 11 shows a 4×4 block, Figure 11 (b) of Figure 11 shows 4×8 and 8×4 blocks, Figure 11 (c) of

[0226] shows 4×N or N×4 blocks, where N is 16 or greater, Figure 11Among them, the blocks with thick boundaries indicate the regions to which the LFNST is applied. For Figure 11 the blocks of (a) and (b), the LFNST is applied to the upper left 4×4 region, and for Figure 11 the block of (c), the LFNST is separately applied to two continuously arranged upper left 4×4 regions. In Figure 11 (a), (b) and (c), since the LFNST is applied in units of 4×4 regions, this LFNST will be referred to as "4×4 LFNST" hereinafter. Based on the matrix dimension of G, [16×16] or [16×8] matrices can be applied.

[0227] More specifically, the [16×8] matrix is applied to the 4×4 block (4×4 TU or 4×4 CU) of (a) of Figure 11 , and the [16×16] matrix is applied to the blocks in (b) and (c) of Figure 11 . This is to adjust the worst-case computational complexity to 8 multiplications per sample.

[0228] Regarding Figure 11 (d) and (e), the LFNST is applied to the upper left 8×8 region, and this LFNST is referred to as "8×8 LFNST" hereinafter. As the corresponding transformation matrix, [48×16] or [48×8] matrices can be applied. In the case of the forward LFNST, since the [48×1] vector (X vector in Equation 9) is input as input data, not all sample values in the upper left 8×8 region are used as input values for the forward LFNST. That is, as can be seen from the left order of (a) of Figure 8 or the left order of (b) of Figure 8 , a [48×1] vector can be constructed based on the samples belonging to the remaining 3 4×4 blocks while leaving the lower right 4×4 block unchanged.

[0229] [48×8] matrix can be applied to the 8×8 block (8×8 TU or 8×8 CU) in (d) of Figure 11 , and [48×16] matrix can be applied to the 8×8 block in (e) of Figure 11 . This is also to adjust the worst-case computational complexity to 8 multiplications per sample.

[0230] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (Y vector in Equation 9, [8×1] or [16×1] vector) are generated. In the forward LFNST, due to the characteristics of the matrix G T , the number of output data is equal to or less than the number of input data.

[0231] Figure 12 FIG. Figure 12 is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example, and shows the blocks in which the output data of the forward LFNST is arranged according to the block shape.

[0232] In Figure 12 the shaded area at the upper left of the shown block corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate the samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward transform remains unchanged.

[0233] As described above, since the size of the applied transform matrix varies according to the block shape, the number of output data also varies. As Figure 12 shown, the output data of the forward LFNST may not completely fill the upper left 4×4 block. In Figure 12 the cases of (a) and (d) of Figure 9 , a [16×8] matrix and an A [48×8] matrix are applied to the block indicated by the thick line or a partial area inside the block, respectively, and an [8×1] vector as the output of the forward LFNST is generated. That is, according to Figure 9 the scanning order shown in (b) of Figure 12 , only 8 output data can be filled as shown in (a) and (d) of Figure 12 , and 0 can be filled in the remaining 8 positions. In Figure 11 the case of the block to which the LFNST in (d) of Figure 11 is applied, as shown in (d) of Figure 12 , the two 4×4 blocks at the upper right and lower left adjacent to the upper left 4×4 block are also filled with the value 0.

[0234] As described above, basically, by signaling the LFNST index, it is specified whether the LFNST is applied and which transform matrix is to be applied. As Figure 12 shown, when the LFNST is applied, since the number of output data of the forward LFNST can be equal to or less than the number of input data, there are areas filled with zero values as follows.

[0235] 1) As Figure 12 shown in (a) of Figure 12 , the samples from the eighth position and the subsequent positions in the scanning order in the upper left 4×4 block, that is, the samples from the ninth to the sixteenth.

[0236] 2) As Figure 12 shown in (d) and (e) of Figure 12 , when a [48×16] matrix or a [48×8] matrix is applied, the two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.

[0237] Therefore, if there is non-zero data in the check regions 1) and 2), it is determined that LFNST is not applied, so that the signaling of the corresponding LFNST index can be omitted.

[0238] According to the example, for instance, in the case of LFNST adopted in the VVC standard, since the signaling of the LFNST index is performed after residual coding, the encoding device can know whether there is non-zero data (valid coefficients) at all positions within the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform the signaling regarding the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When there is no non-zero data in the regions specified in 1) and 2) above, the signaling of the LFNST index is performed.

[0239] In addition, for the adopted LFNST, the following simplified method can be applied.

[0240] (i) According to the example, the number of output data of the forward LFNST can be limited to a maximum of 16.

[0241] In Figure 11 case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions in the upper left respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), the 4×4 LFNST is only applied to one 4×4 region in the upper left, and the LFNST can be applied to Figure 11 all blocks only once. By this, the implementation manner of image coding can be simplified.

[0242] (ii) According to the example, the region to which LFNST is not applied can be additionally cleared. In this document, clearing can mean filling all positions belonging to a specific region with a value of 0. That is, clearing can be applied to the region that is not changed due to LFNST, and the result of the forward one-time transform is maintained. As described above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, the clearing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).

[0243] (ii)-(A) When the 4×4 LFNST is applied, the region to which the 4×4 LFNST is not applied can be cleared. Figure 13 is a diagram illustrating the clearing in the block where the 4×4 LFNST is applied according to the example.

[0244] As Figure 13 shown, regarding the block to which the 4×4 LFNST is applied, that is, for Figure 12For all blocks in (a), (b), and (c), the entire area where LFNST is not applied can be filled with zeros.

[0245] On the other hand, Figure 13 (d) of shows that when the maximum value of the number of output data of the forward LFNST according to the example is limited to 16, the remaining blocks where 4×4 LFNST is not applied are cleared.

[0246] (ii)-(B) When 8×8 LFNST is applied, the area where 8×8 LFNST is not applied can be cleared. Figure 14 is a diagram illustrating the clearing in the blocks where 8×8 LFNST is applied according to the example.

[0247] As Figure 14 shown, regarding the blocks where 8×8 LFNST is applied, that is, for all blocks in (d) and (e) of Figure 12 , the entire area where LFNST is not applied can be filled with zeros.

[0248] (iii) Due to the clearing presented in (ii) above, the area filled with zeros may not be the same as when LFNST is applied. Therefore, the clearing proposed in (ii) can be performed on a wider area according to the case of the compared LFNST to check for the presence of non-zero data. Figure 12 After checking for the presence of non-zero data in the area filled with zeros in (d) and (e) of

[0249] For example, when (ii)-(B) is applied, after checking for the presence of non-zero data in the area filled with zeros in (d) and (e) of Figure 12 , additionally checking for the presence of non-zero data in the area filled with 0 in Figure 14 , signaling for the LFNST index can be performed only when there is no non-zero data.

[0250] Of course, even when the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same manner as the existing LFNST index signaling. That is, after checking for the presence of non-zero data in the blocks filled with zeros in Figure 12 , the LFNST index signaling can be applied. In this case, the encoding device only performs the clearing and the decoding device does not assume the clearing, that is, only checking whether non-zero data exists only in the area clearly marked as 0 in Figure 12 , the LFNST index parsing can be performed.

[0251] Various embodiments of the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) for applying LFNST can be derived. Of course, the combination of the above simplified methods is not limited to the following embodiments, and any combination can be applied to LFNST.

[0252] Embodiment

[0253] - Limit the number of output data of the forward LFNST to a maximum of 16 → (i)

[0254] - When applying 4×4 LFNST, all areas where 4×4 LFNST is not applied are cleared → (II)-(A)

[0255] - When applying 8×8 LFNST, all areas where 8×8 LFNST is not applied are cleared → (II)-(B)

[0256] - After checking whether non-zero data also exists in the existing areas filled with zero values and the areas filled with zero due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only when there is no non-zero data → (iii).

[0257] In the case of the embodiment, when applying LFNST, the area where non-zero output data can exist is limited to the inside of the upper left 4×4 area. More specifically, in Figure 13 of (a) and Figure 14 of (a), the eighth position in the scanning order is the last position where non-zero data can exist. In Figure 13 of (b) and (c) and Figure 14 of (b), the sixteenth position in the scanning order (i.e., the position of the lower right edge of the upper left 4×4 block) is the last position where data other than 0 can exist.

[0258] Therefore, after applying LFNST, after checking whether non-zero data exists at positions where the residual coding process does not allow (at positions beyond the last position), it can be determined whether to signal the LFNST index.

[0259] In the case of the clearing method proposed in (ii), due to the number of data finally generated when both a transform and LFNST are applied once, the amount of computation required to perform the entire transform process can be reduced. That is, when LFNST is applied, since clearing is applied to the areas where the forward first transform output data exists and LFNST is not applied, there is no need to generate data for the areas that become cleared during the execution of the forward first transform. Therefore, the amount of computation required to generate the corresponding data can be reduced. The additional effects of the clearing method proposed in (ii) are summarized as follows.

[0260] First, as described above, reduce the amount of computation required to perform the entire transform process.

[0261] In particular, when applying (ii)-(B), the worst-case computational load is reduced, enabling the transformation process to be lightened. In other words, generally, a large amount of computation is required to perform a single transformation of a large size. By applying (ii)-(B), the number of data derived as a result of performing the forward LFNST can be reduced to 16 or less. Additionally, as the size of the entire block (TU or CU) increases, the effect of reducing the amount of transformation operations further increases.

[0262] Second, the computational load required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.

[0263] Third, the latency involved in the transformation process is reduced.

[0264] A secondary transformation such as LFNST adds computational load to the existing single transformation, thus increasing the overall latency time involved in performing the transformation. In particular, in the case of intra prediction, since the reconstructed data of adjacent blocks is used during the prediction process, during encoding, the increase in latency due to the secondary transformation leads to an increase in the latency until reconstruction. This can result in an increase in the overall latency of intra prediction coding.

[0265] However, if the zeroing proposed in (ii) is applied, the latency time for performing a single transformation can be greatly reduced when applying LFNST, maintaining or reducing the overall latency of the transformation, enabling a coding device to be implemented more simply.

[0266] In traditional intra prediction, the block to be currently encoded is regarded as one coding unit, and encoding is performed without division. However, intra-subpartition (ISP) coding means performing intra prediction coding by dividing the block to be currently encoded in the horizontal or vertical direction. In this case, reconstructed blocks can be generated by performing encoding / decoding in units of the divided blocks, and the reconstructed blocks can be used as reference blocks for the next divided blocks. According to an embodiment, in ISP coding, one coding block can be divided into two or four sub-blocks and encoded, and in ISP, within one sub-block, intra prediction is performed by referring to the reconstructed pixel values of the sub-blocks located adjacent to the left or adjacent to the upper side. Hereinafter, "encoding" can be used as a concept including both encoding performed by a coding device and decoding performed by a decoding device.

[0267] ISP divides a block predicted to be intra in the luminance frame into two or four sub-partitions in the vertical or horizontal direction according to the size of the block. For example, the minimum block size for which ISP can be applied is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into 4 sub-partitions.

[0268] When applying ISP, sub - blocks are encoded sequentially from left - to - right or top - to - bottom according to the division type (e.g., horizontally or vertically), and after performing reconstruction processing via inverse transform and intra - prediction for one sub - block, the encoding of the next sub - block can be performed. For the left - most or top - most sub - block, the reconstructed pixels of the already - encoded coded block are referred to, as in the conventional intra - prediction method. In addition, when each side of a subsequent internal sub - block is not adjacent to the previous sub - block, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already - encoded adjacent coded block are referred to, as in the conventional intra - prediction method.

[0269] In the ISP encoding mode, all sub - blocks can be encoded with the same intra - prediction mode, and flags indicating whether to use ISP encoding and indicating whether to divide in which direction (horizontal or vertical) can be signaled. At this time, the number of sub - blocks can be adjusted to 2 or 4 according to the shape of the block. When the size (width × height) of a sub - block is less than 16, it can be restricted so that division into the corresponding sub - block is not allowed or ISP encoding itself is not applied.

[0270] In the case of the ISP prediction mode, a coding unit is divided into two or four partitioned blocks (i.e., sub - blocks) for prediction, and the same intra - prediction mode is applied to the two or four partitioned blocks.

[0271] As described above, in the division direction, both the horizontal direction (when an M×N coding unit with horizontal length M and vertical length N is divided in the horizontal direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into M×(N / 2) blocks, and if the M×N coding unit is divided into four blocks, the M×N coding unit is divided into M×(N / 4) blocks)) and the vertical direction (when the M×N coding unit is divided in the vertical direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into (M / 2)×N blocks, and if the M×N coding unit is divided into four, the M×N coding unit is divided into (M / 4)×N blocks) are possible. When dividing the M×N coding unit in the horizontal direction, the partitioned blocks are encoded in the up - down order, and when dividing the M×N coding unit in the vertical direction, the partitioned blocks are encoded in the left - right order. In the case of horizontal (vertical) division, the reconstructed pixel values of the upper (left) partitioned block can be referred to for predicting the currently - encoded partitioned block.

[0272] A transform may be applied to the residual signals generated in block units by the ISP prediction method. The multi-transform selection (MTS) technology based on the DST-7 / DCT-8 combination and the existing DCT-2 may be applied to the forward-based primary transform (core transform), and the forward low-frequency non-separable transform (LFNST) may be applied to the transform coefficients generated according to the primary transform to generate the final modified transform coefficients.

[0273] That is, the LFNST may be applied to the blocks divided by applying the ISP prediction mode, and the same intra prediction mode is applied to the divided blocks, as described above. Therefore, when selecting the set of LFNSTs derived based on the intra prediction mode, the derived set of LFNSTs may be applied to all the blocks. That is, since the same intra prediction mode is applied to all the blocks, the same set of LFNSTs may be applied to all the blocks.

[0274] According to an embodiment, the LFNST may be applied only to transform blocks having both a horizontal length and a vertical length of 4 or greater. Therefore, when the horizontal length or the vertical length of the divided block according to the ISP prediction method is less than 4, the LFNST is not applied and the LFNST index is not signaled. In addition, when applying the LFNST to each block, the corresponding block may be regarded as one transform block. When the ISP prediction method is not applied, the LFNST may be applied to the coding block.

[0275] A method of applying the LFNST to each block will be described in detail.

[0276] According to an embodiment, after applying the forward LFNST to each block, only up to 16 (8 or 16) coefficients are left in the upper left 4×4 region in the transform coefficient scanning order, and then zeroing may be applied, where the remaining positions and regions are all filled with 0.

[0277] Alternatively, according to an embodiment, when the length of one side of the block is 4, the LFNST is applied only to the upper left 4×4 region, and when the length of all sides of the block (i.e., width and height) is 8 or greater, the LFNST may be applied to the remaining 48 coefficients in the upper left 8×8 region except for the lower right 4×4 region.

[0278] Alternatively, according to an embodiment, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each block is 4×4 or 8×8, only 8 transform coefficients may be output after applying the forward LFNST. That is, when the block is 4×4, an 8×16 matrix may be used as the transform matrix, and when the block is 8×8, an 8×48 matrix may be used as the transform matrix.

[0279] In the current VVC standard, the LFNST index signaling is performed on a coding unit basis. Therefore, in the ISP prediction mode and when applying LFNST to all sub-blocks, the same LFNST index value can be applied to the corresponding sub-blocks. That is to say, when the LFNST index value is sent once at the coding unit level, the corresponding LFNST index can be applied to all sub-blocks in the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates the case where LFNST is not applied, and 1 and 2 indicate two transform matrices present in an LFNST set when LFNST is applied.

[0280] As described above, the LFNST set is determined by the intra prediction mode, and in the case of the ISP prediction mode, since all sub-blocks in the coding unit are predicted in the same intra prediction mode, the sub-blocks can refer to the same LFNST set.

[0281] As another example, the LFNST index signaling is still performed on a coding unit basis, but in the case of the ISP prediction mode, it is not determined whether LFNST is uniformly applied to all sub-blocks, and for each sub-block, it can be determined whether to apply the LFNST index value signaled at the coding unit level and whether to apply LFNST through a separate condition. Here, the separate condition can be signaled in the bitstream in the form of a flag for each sub-block, and when the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST may not be applied.

[0282] In the following, a method for maintaining the worst-case computational complexity when applying LFNST to the ISP mode will be described.

[0283] In the case of the ISP mode, when applying LFNST, in order to keep the number of multiplications per sample (or per coefficient, per position) at a certain value or less, the application of LFNST may be restricted. Depending on the size of the sub-block, by applying LFNST as follows, the number of multiplications per sample (or per coefficient, per position) can be kept at 8 or less.

[0284] 1. When both the horizontal length and the vertical length of the sub-block are 4 or greater, the same method as the worst-case computational complexity control method for LFNST in the current VVC standard can be applied.

[0285] That is to say, when the divided block is a 4×4 block, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix instead of the 16×16 matrix can be applied in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns from the 16×16 matrix can be applied in the reverse direction. In addition, when the divided block is an 8×8 block, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the upper 8 rows from the 16×48 matrix is applied, and in the reverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns from the 48×16 matrix can be applied.

[0286] In the case of a 4×N or N×4 (N>4) block, when performing the forward transform, the 16 coefficients generated after applying the 16×16 matrix only to the upper left 4×4 block can be set in the upper left 4×4 region, and the other regions can be filled with the value 0. Additionally, when performing the inverse transform, the 16 coefficients located in the upper left 4×4 block are set in scan order to form an input vector, and then 16 output data can be generated by multiplying by the 16×16 matrix. The generated output data can be set in the upper left 4×4 region, and the remaining regions except the upper left 4×4 region can be filled with the value 0.

[0287] In the case of an 8×N or N×8 (N>8) block, when performing the forward transform, the 16 coefficients generated after applying the 16×48 matrix to the ROI region (the remaining region except for the lower right 4×4 block in the upper left 8×8 block) within only the upper left 8×8 block can be set in the upper left 4×4 region, and all other regions can be filled with the value 0. Moreover, when performing the inverse transform, the 16 coefficients located in the upper left 4×4 region are set in scan order to form an input vector, and then 48 output data can be generated by multiplying by the 48×16 matrix. The generated output data can be filled in the ROI region, and all other regions can be filled with the value 0.

[0288] As another example, in order to keep the number of multiplications per sample (or per coefficient, per position) at a certain value or less, the number of multiplications per sample (or per coefficient, per position) based on the ISP coding unit size rather than the size of the ISP divided block can be kept at 8 or less. When only one block among the ISP divided blocks meets the condition for applying LFNST, the worst-case complexity calculation of LFNST can be applied based on the corresponding coding unit size rather than the size of the divided block. For example, when the luminance coding block of a certain coding unit is divided into four divided blocks of size 4×4 and encoded by ISP, and there are no non-zero transform coefficients for two of the four divided blocks, it can be configured such that for the other two divided blocks (based on the encoder), 16 transform coefficients instead of eight transform coefficients are generated.

[0289] In the following, a method of signaling the LFNST index in the ISP mode will be described.

[0290] As described above, the LFNST index can have any one of the values 0, 1, and 2, where the value 0 indicates that LFNST is not applied, and the values 1 and 2 indicate two LFNST kernel matrices included in the selected LFNST set, respectively. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method for sending the LFNST index in the current VVC standard will be described as follows.

[0291] 1. The LFNST index can be sent once for each coding unit (CU), and in the case of the dual-tree type, separate LFNST indices can be signaled for the luminance block and the chrominance block.

[0292] 2. When the LFNST index is not signaled, the value of the LFNST index is set (inferred) to the default value 0. The cases where the value of the LFNST index is inferred to be 0 are as follows:

[0293] A. The case of using a mode that does not apply transformation (e.g., transform skip, BDPCM, lossless coding, etc.).

[0294] B. The case where a single transformation is not DCT-2 (DST7 or DCT8), i.e., the transformation in the horizontal direction or the transformation in the vertical direction is not DCT-2.

[0295] C. The case where the horizontal length or the vertical length of the luminance block of the coding unit exceeds the maximum luminance transformable size. For example, when the maximum luminance transform size is 64, LFNST is not applied when the luminance block size of the coding block is 128×16.

[0296] In the case of the dual-tree type, for each of the coding units of the luminance component and the coding units of the chrominance component, it is determined whether the maximum luminance transformable size is exceeded. That is, for the luminance block, it is checked whether the maximum luminance transformable size is exceeded, and for the chrominance block, it is checked whether the horizontal / vertical length of the corresponding luminance block for the color format and the maximum luminance transformable size are exceeded. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luminance block is twice the horizontal / vertical length of the chrominance block, and the transform size of the corresponding luminance block is twice the transform size of the chrominance block. In another example, when the color format is 4:4:4, the horizontal / vertical length and the transform size of the corresponding luminance block are the same as those of the chrominance block.

[0297] A 64-length transform or a 32-length transform refers to a transform applied horizontally or vertically with lengths of 64 or 32 respectively, and "transform size" may refer to the corresponding length of 64 or 32.

[0298] In the case of the single-tree type, it is possible to check whether the horizontal length or the vertical length of a luminance block exceeds the maximum transform size of a luminance transform block, and when the horizontal length or the vertical length of the luminance block exceeds the maximum transform size of the luminance transform block, signaling of the LFNST index may be omitted.

[0299] D. It is possible to send the LFNST index only when both the horizontal length and the vertical length of a coding unit are each greater than or equal to 4.

[0300] In the case of the double-tree type, it is possible to signal the LFNST index only when both the horizontal length and the vertical length of the corresponding component (i.e., the luminance component or the chrominance component) are each greater than or equal to 4.

[0301] In the case of the single-tree type, it is possible to signal the LFNST index when both the horizontal length and the vertical length of the luminance component are each greater than or equal to 4.

[0302] E. In the case where the last non-zero coefficient position is not the DC position (which is the position at the upper left of the block), if for a luminance block of the double-tree type the last non-zero coefficient position is not the DC position, the LFNST index is sent. In the case of a chrominance block of the double-tree type, if any one of the last non-zero coefficient positions of Cb and the last non-zero coefficient position of Cr is not the DC position, the corresponding LNFST index is sent.

[0303] In the case of the single-tree type, the LFNST index is sent when the last non-zero coefficient position of any one of the luminance component, the Cb component, and the Cr component is not the DC position.

[0304] Here, when the value of the coding block flag (CBF) indicating whether transform coefficients of a transform block exist is 0, the last non-zero coefficient position of the corresponding transform block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, so the position of the last non-zero coefficient may be not considered when checking the conditions for LFNST index signaling.

[0305] For example, 1) in the case of the luminance component of the dual-tree type, if the corresponding CBF value is 0, the LFNST index is not signaled; 2) in the case of the chrominance component of the dual-tree type, if the CBF value of Cb is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked and the corresponding LFNST index is sent; 3) in the case of the single-tree type, the last non-zero coefficient position of any component with a CBF value of 1 among the luminance component, the Cb component, and the Cr component is checked.

[0306] F. When a transform coefficient is found at a position other than the position where the LFNST transform coefficient is allowed to exist, the signaling of the LFNST index can be omitted. In the case of 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 8 positions starting from the DC position and all the remaining positions are filled with 0. Additionally, in cases other than 4×4 transform blocks and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficient can exist in 16 positions starting from the DC position and all the remaining positions are filled with 0.

[0307] Therefore, when there is a non-zero transform coefficient in the area that should be filled with 0 after residual coding, the LFNST index signaling can be omitted.

[0308] In addition, the ISP mode can be applied only to luminance blocks, or can be applied to both luminance blocks and chrominance blocks. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four sub-blocks and predicted, and a transform can be applied to each corresponding sub-block. Therefore, when determining the conditions for signaling the LFNST index on a coding unit basis, the fact that LFNST can be applied to each of the corresponding sub-blocks needs to be considered. Additionally, when the ISP prediction mode is applied only to a specific component (e.g., luminance blocks), the fact that only the corresponding component is divided into sub-blocks should be taken into account when signaling the LFNST index. The available LFNST index signaling methods in the ISP mode can be summarized as follows.

[0309] 1. The LFNST index can be sent once for each coding unit (CU), and in the case of the dual-tree type, separate LFNST indices can be signaled for luminance blocks and chrominance blocks respectively.

[0310] 2. When signaling the LFNST index, the value of the LFNST index is set (inferred) to the default value 0. The cases where the LFNST index value is inferred to be 0 are as follows.

[0311] A. Cases where a mode that does not apply a transform (e.g., transform skip, BDPCM, lossless coding, etc.) is used.

[0312] B. In the case where the horizontal length or vertical length of the luminance block of the coding unit exceeds the maximum luminance transformable size, for example, when the maximum transformable luminance size is 64, LFNST cannot be applied when the size of the luminance block of the coding block is 128×16.

[0313] Whether to signal the LFNST index can be determined based on the size of the partitioned block rather than the size of the coding unit. That is, when the horizontal length or vertical length of the partitioned block corresponding to the luminance block exceeds the maximum luminance transformable size, the LFNST index signaling can be omitted, and the value of the LFNST index can be inferred as 0.

[0314] In the case of the dual-tree type, it is determined for each of the coding units or partitioned blocks of the luminance component and the coding units or partitioned blocks of the chrominance component whether it exceeds the maximum transform block size. That is, the horizontal length and vertical length of the coding unit or partitioned block for luminance are compared with the maximum luminance transformable size, and if either the horizontal length or the vertical length is greater than the maximum luminance transformable size, LFNST is not applied. And in the case of the coding unit or partitioned block for chrominance, the horizontal / vertical length of the corresponding luminance block for the color format is compared with the maximum luminance transformable size. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luminance block is twice the horizontal / vertical length of the chrominance block, and the transform size of the corresponding luminance block is twice the transform size of the chrominance block. In another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luminance block are the same as those of the chrominance block.

[0315] In the case of the single-tree type, it can be checked whether the horizontal length or vertical length of the luminance block (coding unit or partitioned block) exceeds the maximum transformable size of the luminance transform block, and if so, the LFNST index signaling can be omitted.

[0316] C. When applying the LFNST included in the current VVC standard, the LFNST index can be sent only when both the horizontal length and vertical length of the partitioned block are each greater than or equal to 4.

[0317] When applying the LFNST for 2×M (1×M) or M×2 (M×1) blocks in addition to the LFNST included in the current VVC standard, the LFNST index can be sent only when the size of the partitioned block is equal to or greater than 2×M (1×M) or M×2 (M×1) blocks. Here, if the P×Q block is equal to or greater than the R×S block, this means P≥R and Q≥S.

[0318] In summary, the LFNST index may be sent only when the partition block is equal to or larger than the minimum size to which LFNST can be applied. In the case of the dual-tree type, the LFNST index may be signaled only when the partition block of the luma or chroma component is equal to or larger than the minimum size to which LFNST can be applied. In the case of the single-tree type, the LFNST index may be signaled only when the partition block of the luma component is equal to or larger than the minimum size to which LFNST can be applied.

[0319] In this document, if the M×N block is equal to or larger than the K×L block, this means that M is equal to or larger than K and N is equal to or larger than L. If the M×N block is larger than the K×L block, this means that M is equal to or larger than K, N is equal to or larger than L, and M is larger than K or N is larger than L. If the M×N block is less than or equal to the K×L block, this means that M is less than or equal to K and N is less than or equal to L, and if the M×N block is less than the K×L block, this means that M is less than or equal to K, N is less than or equal to L, and M is less than K or N is less than L.

[0320] D. In the case where the last non-zero coefficient position is not the DC position (which is the position at the upper left of the block), if for a luma block of the dual-tree type, the last non-zero coefficient position in any one of all the partition blocks is not the DC position, the LFNST index may be sent. In the case of a chroma block of the dual-tree type, if the last non-zero coefficient position of all the partition blocks for Cb (when the ISP mode is not applied to the chroma component, the number of partition blocks is considered to be 1) and the last non-zero coefficient position of all the partition blocks for Cr (when the ISP mode is not applied to the chroma component, the number of partition blocks is considered to be 1) in any one of them is not the DC position, the corresponding LNFST index may be sent.

[0321] In the case of the single-tree type, when the last non-zero coefficient position in any one of all the partition blocks of the luma component, Cb component, and Cr component is not the DC position, the corresponding LFNST index may be sent.

[0322] Here, when the coding block flag (CBF) value indicating whether there are transform coefficients for each partition block is 0, the last non-zero coefficient position of the corresponding partition block is not checked to determine whether to signal the LFNST index. That is, when the corresponding CBF value is 0, the transform is not applied to the corresponding block, so the last non-zero coefficient position of the corresponding partition block is not considered when checking the conditions for LFNST index signaling.

[0323] For example, 1) in the case of the luminance component of the dual-tree type, when the corresponding CBF value of each partition block is 0, the corresponding partition block is excluded when determining whether to signal the LFNST index; 2) in the case of the chrominance component of the dual-tree type, when the CBF value of Cb of each partition block is 0 and the CBF value of Cr is 1, only the last non-zero coefficient position of Cr is checked when determining whether to signal the LFNST index; and 3) in the case of the single-tree type, it is possible to determine whether to signal the LFNST index by checking only the last non-zero coefficient position of the blocks with CBF value of 1 for all partition blocks of the luminance component, Cb component, and Cr component.

[0324] In the case of the ISP mode, the image information can be configured such that the last non-zero coefficient position is not checked, and its implementation is as follows.

[0325] i. In the case of the ISP mode, it is possible to allow the LFNST index signaling without checking the last non-zero coefficient position of both the luminance block and the chrominance block. That is, even when the last non-zero coefficient position of all partition blocks is at the DC position or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.

[0326] ii. In the case of the ISP mode, it is possible to omit the check of the last non-zero coefficient position only for the luminance block and check the last non-zero coefficient position of the chrominance block in the above manner. For example, in the case of the luminance block of the dual-tree type, the LFNST index signaling is allowed without checking the last non-zero coefficient position, while in the case of the chrominance block of the dual-tree type, it is possible to determine whether to signal the corresponding LFNST index by checking whether there is a DC position of the last non-zero coefficient position in the above manner.

[0327] iii. In the case of the ISP mode and the single-tree type, the above method i or the above method ii can be applied. That is, when method i is applied to the single-tree type in the ISP mode, the check of the last non-zero coefficient position can be omitted for both the luminance block and the chrominance block, and the LFNST index signaling is allowed. Alternatively, when method ii is applied, the check of the last non-zero coefficient position can be omitted for the partition blocks of the luminance component, and for the partition blocks of the chrominance component (when the ISP is not applied to the chrominance component, the number can be considered to be 1), the last non-zero coefficient position can be checked in the above manner to determine whether to signal the corresponding LFNST index.

[0328] E. When it is found that there are transform coefficients in a position other than the position where the LFNST transform coefficients can exist for even one of all partition blocks, the LFNST index signaling can be omitted.

[0329] For example, in the case of 4×4 and 8×8 blocks, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficients can exist in 8 positions starting from the DC position, and all the remaining positions are filled with 0. Additionally, when the block is equal to or larger than 4×4 and is neither a 4×4 block nor an 8×8 block, according to the transform coefficient scanning order in the VVC standard, the LFNST transform coefficients can exist in 16 positions starting from the DC position, and all the remaining positions are filled with 0.

[0330] Therefore, when there are non-zero transform coefficients in the area filled with 0 after performing residual coding, the LFNST index signaling can be omitted.

[0331] On the other hand, in the case of the ISP mode, according to the current VVC standard, the horizontal and vertical directions are independently regarded as length conditions, and DST-7 is applied instead of DCT-2 without signaling the MTS index. It is determined whether the horizontal length or the vertical length is greater than or equal to 4 and greater than or equal to 16, and a single transform kernel is determined based on the determined result. Therefore, for the cases where LFNST can be applied in the ISP mode, the following transform combinations are feasible.

[0332] 1. When the LFNST index is 0 (including the case where the LFNST index is inferred to be 0), the single transform determination conditions in the ISP mode, which are included in the current VVC standard, can be followed. That is, if the length conditions (which are the conditions of being equal to or greater than 4 and equal to or less than 16) are separately and independently satisfied for the horizontal and vertical directions; if so, DST-7 can be applied instead of DCT-2, and if not, DCT-2 can be applied.

[0333] 2. When the LFNST index is greater than 0, the following two configurations can be used as the single transform.

[0334] A. DCT-2 can be applied to both the horizontal and vertical directions.

[0335] B. The conditions for determining the single transform in the ISP mode, which are included in the current VVC standard, can be followed. In other words, it is checked whether the length conditions (which are the conditions of being equal to or greater than 4 and equal to or less than 16) are separately and independently satisfied for the horizontal and vertical directions; if so, DST-7 can be applied instead of DCT-2, and if not, DCT-2 can be applied.

[0336] In the ISP mode, the image information can be configured such that: for each partition block instead of each coding unit, the LFNST index is sent. In this case, in the above LFNST index signaling method, considering that there is only one partition block in the unit where the LFNST index is sent, it can be determined whether to signal the LFNST index.

[0337] In addition, the signaling of the LFNST index and the MTS index will be described below.

[0338] The following table gives the coding unit syntax table, transform unit syntax table, and residual coding syntax table related to the signaling of the LFNST index and the MTS index according to the example. According to Table 4, the MTS index moves from the transform unit level to the coding unit level syntax and is signaled after the LFNST index signaling. In addition, the constraint that LFNST is not allowed when ISP is applied to the coding unit has been removed. When ISP is applied to the coding unit, the constraint that LFNST is not allowed is removed so that LFNST can be applied to all intra-prediction blocks. In addition, both the MTS index and the LFNST index are conditionally signaled at the end part of the coding unit level.

[0339] [Table 4]

[0340]

[0341] [Table 5]

[0342]

[0343] [Table 6]

[0344]

[0345] The meanings of the main variables in the table are as follows.

[0346] 1. cbWidth, cbHeight: The width and height of the current coding block

[0347] 2. log2TbWidth, log2TbHeight: The base-2 logarithms of the width and height of the current transform block, and reflect zeroing to reduce to the upper left region where non-zero coefficients can exist.

[0348] 3. sps_lfnst_enabled_flag: It is a flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the sequence parameter set (SPS).

[0349] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the position (x0, y0). chType can have values of 0 and 1, where 0 represents the luma component, and 1 represents the chroma component. The position (x0, y0) indicates a position on the picture, and MODE_INTRA (intra prediction) and MODE_INTER (inter prediction) can have the CuPredMode[chType][x0][y0] value.

[0350] 5. IntraSubPartitionsSplit[x0][y0]: The content at the position (x0, y0) is the same as in item 4. It indicates which ISP partition is applied at the position (x0, y0). ISP_NO_SPLIT indicates that the coding unit corresponding to the position (x0, y0) is not divided into sub-blocks.

[0351] 6. intra_mip_flag[x0][y0]: The content at the position (x0, y0) is the same as in item 4 above. intra_mip_flag is a flag indicating whether the matrix-based intra prediction (MIP) prediction mode is applied. The flag value 0 indicates that MIP is not enabled, and the flag value 1 indicates that MIP is applied.

[0352] 7. cIdx: The value 0 indicates luma, and the values 1 and 2 indicate Cb and Cr of the chroma components respectively.

[0353] 8. treeType: It indicates single tree, dual tree, etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luma component, DUAL_TREE_CHROMA: dual tree for chroma component)

[0354] 9. lastSubBlock: It indicates the position of the sub-block (coefficient group (CG)) where the last non-zero coefficient is located in the scan order. 0 indicates the sub-block containing the DC component, and if it is greater than 0, it is not the sub-block containing the DC component.

[0355] 10. lastScanPos: It indicates where the last valid coefficient in a sub-block is located in the scan order. If a sub-block consists of 16 positions, values from 0 to 15 can be used.

[0356] 11. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If it is not parsed, it is inferred to be the value 0. That is, the default value is set to 0, indicating that LFNST is not applied.

[0357] 12. LastSignificantCoeffX, LastSignificantCoeffY: It indicates the x - coordinate and y - coordinate where the last significant coefficient in the transform block is located. The x - coordinate starts from 0 and increases from left to right, and the y - coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means the last significant coefficient is located at DC.

[0358] 13. cu_sbt_flag: It is a flag that indicates whether sub - block transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.

[0359] 14. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: They are flags that respectively indicate whether explicit MTS is applied to inter - frame CUs and intra - frame CUs. If the corresponding flag value is 0, it indicates that MTS cannot be applied to the inter - frame CU or intra - frame CU, and if it is 1, it indicates that MTS is applicable.

[0360] 15. tu_mts_idx[x0][y0]: It is the MTS index syntax element to be parsed. If it is not parsed, it is inferred to have a value of 0. That is, the default value is set to 0, which indicates that DCT - 2 is applied to both the horizontal and vertical directions.

[0361] As shown in Table 4, several conditions need to be checked when encoding mts_idx[x0][y0], and tu_mts_idx[x0][y0] is signaled only when the lfnst_idx[x0][y0] value is 0.

[0362] In addition, tu_cbf_luma[x0][y0] is a flag that indicates whether there are valid coefficients for the luminance component. The value 0 indicates that there are no valid coefficients in the corresponding transform block of the luminance component, while 1 indicates that there are valid coefficients in the corresponding transform block of the luminance component.

[0363] According to Table 4, when both the width and height of the coding unit of the luminance component are 32 or less, mts_idx[x0][y0] is signaled (Max(cbWidth, cbHeight)<=32), that is, whether to apply MTS is determined by the width and height of the coding unit of the luminance component.

[0364] Additionally, according to Table 4, even in the ISP mode (IntraSubPartitionsSplitType != ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to signal, and the same LFNST index value can be applied to all ISP partition blocks.

[0365] On the other hand, mts_idx[x0][y0] can be signaled only when not in the ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT).

[0366] In the process of determining log2ZoTbWidth and log2ZoTbHeight as shown in Table 6 (where log2ZoTbWidth and log2ZoTbHeight respectively represent the base-2 logarithms of the width and height of the remaining upper-left region after zeroing), the part of checking the value of mts_idx[x0][y0] can be omitted.

[0367] Furthermore, according to the example, when determining log2ZoTbWidth and log2ZoTbHeight in residual coding, the condition for checking sps_mts_enable_flag can be added.

[0368] If there are valid coefficients at the zeroing position when applying LFNST, the variable LfnstZeroOutSigCoeffFlag in Table 4 is 0, otherwise it is 1. The variable LfnstZeroOutSigCoeffFlag can be set according to several conditions shown in Table 6.

[0369] According to the example, the variable LfnstDcOnly in Table 4 becomes 1 when all the last valid coefficients are located at the DC position (upper-left position) of the transform block with a corresponding coding block flag (CBF) having a value of 1 (1 if there is at least one valid coefficient in the block, otherwise 0), otherwise it becomes 0. More specifically, in the case of dual-tree luminance, the position of the last valid coefficient is checked for one luminance transform block, and in the case of dual-tree chrominance, the positions of the last valid coefficients are checked for both the Cb transform block and the Cr transform block. In the case of single-tree, the positions of the last valid coefficients can be checked for the luminance, Cb, and Cr transform blocks.

[0370] In Table 4, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 6. If there are significant coefficients in the area filled with zeros due to zeroing (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, as shown in Table 4, the MTS index is not signaled.

[0371] In addition, as shown in Table 4, when tu_cbf_luma[x0][y0] is 0, the mts_idx[x0][y0] coding can be omitted. That is, if the CBF value of the luminance component is 0, since no transform is applied, there is no need to signal the MTS index, so the MTS index coding can be omitted.

[0372] According to the example, the technical feature can be implemented with another conditional syntax. For example, after performing MTS, a variable indicating whether there are significant coefficients in the area other than the DC area of the current block can be derived. If this variable indicates that there are significant coefficients in the area other than the DC area, the MTS index can be signaled. That is, the presence of significant coefficients in the area other than the DC area of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, the MTS index can be signaled.

[0373] This variable can be represented as MtsDcOnly. After the variable MtsDcOnly is initially set to 1 at the coding unit level, its value can change to 0 when the residual coding level indicates that there are significant coefficients in the area other than the DC area of the current block. When the variable MtsDcOnly is 0, the image information can be configured to signal the MTS index.

[0374] If tu_cbf_luma[x0][y0] is 0, the variable MtsDcOnly remains its initial value of 1 because the residual coding syntax is not called at the transform unit level of Table 5. In this case, since the variable MtsDcOnly does not change to 0, the image information can be configured so that the MTS index is not signaled. That is, the MTS index is not parsed and signaled.

[0375] In addition, the decoding device can determine the color index cIdx of the transform coefficient to derive the variable MtsZeroOutSigCoeffFlag in Table 6. A color index cIdx of 0 represents the luminance component.

[0376] According to the example, since MTS can be applied only to the luminance component of the current block, the decoding device can determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag for determining whether to parse the MTS index.

[0377] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether zeroing is performed when applying MTS. It indicates whether there are transform coefficients in the area outside the upper-left area where the last valid coefficient can exist due to zeroing after performing MTS (i.e., in the area other than the upper-left 16×16 area). As shown in Table 4, the variable MtsZeroOutSigCoeffFlag is initially set to 1 (MtsZeroOutSigCoeffFlag = 1) at the coding unit level, and if there are transform coefficients in the area other than the 16×16 area, this value changes from 1 to 0 at residual coding, and as shown in Table 6, it can be changed (MtsZeroOutSigCoeffFlag = 0). If the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.

[0378] As shown in Table 6, at the residual coding level, depending on whether zeroing accompanying MTS is performed, a non-zeroing area where non-zero transform coefficients can exist can be set. Even in this case, when the color index (cIdx) is 0, the non-zeroing area can be set to the upper-left 16×16 area of the current block.

[0379] In this way, when deriving the variable for determining whether to parse the MTS index, it is determined whether the color component is luminance or chrominance. However, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color component is not determined when deriving the variable for determining whether to parse the LFNST index.

[0380] For example, Table 4 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate that zeroing is performed when applying LFNST. The variable LfnstZeroOutSigCoeffFlag indicates whether there are valid coefficients in the second area of the current block other than the first area located in the upper left. This value is initially set to 1, and if there are valid coefficients in the second area, this value can be changed to 0. The LFNST index can be parsed only when the value of the initially set variable LfnstZeroOutSigCoeffFlag remains 1. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, since LFNST can be applied to both the luminance component and the chrominance component of the current block, the color index of the current block is not determined.

[0381] In addition, the syntax table for signaling the coding unit of the LFNST index according to the example is as follows.

[0382] [Table 7]

[0383]

[0384] In Table 7, lfnst_idx represents the LFNST index and can have values 0, 1, and 2 as described above. As shown in Table 7, lfnst_idx is signaled only when the condition (!intra_mip_flag[x0][y0] || Min(lfnstWidth, lfnstHeight) >= 16) is satisfied. Here, intra_mip_flag[x0][y0] is a flag indicating whether the matrix-based intra prediction (MIP) mode is applied to the luma block belonging to the (x0, y0) coordinates. If the MIP mode is applied to the luma block, the value is 1, and if not, the value is 0.

[0385] lfnstWidth and lfnstHeight indicate the width and height to which the LFNST is applied for the current coding block being encoded (including both the luma coding block and the chroma coding block). When the ISP is applied to the coding block, it can indicate the width and height of each sub-block partitioned into two or four.

[0386] In addition, in the above condition, when Min(lfnstWidth, lfnstHeight) >= 16 is equal to or greater than a 16×16 block (e.g., both the width and height of the luma coding block to which the MIP is applied are equal to or greater than 16) when the MIP is applied, it indicates that the LFNST can be applied. The meanings of the main variables included in Table 7 that are not repeated in the description of Table 4 are briefly introduced below.

[0387] 1. IntraSubPartitionsSplitType: It indicates how the ISP partitions are formed for the current coding unit, and ISP_NO_SPLIT indicates that the corresponding coding unit is not a coding unit partitioned into sub-blocks. ISP_VER_SPLIT indicates vertical splitting, and ISP_HOR_SPLIT indicates horizontal splitting. For example, when a W×H (width W, height H) block is horizontally split into n sub-blocks, it is split into W×(H / n) blocks, and when a W×H (width W, height H) block is vertically split into n sub-blocks, it is split into (W / n)×H blocks.

[0388] 2. SubWidthC, SubHeightC: SubWidthC and SubHeightC are values set according to the color format (or chroma format, e.g., 4:2:0, 4:2:2, 4:4:4), and more specifically, they respectively indicate the ratios of the widths and heights of the luminance component and the chroma components. (See the table below)

[0389] [Table 8]

[0390] Chrominance format SubWidthC SubHeightC Monochrome 1 1 4:2:0 2 2 4:2:2 2 1 4:4:4 1 1 4:4:4 1 1

[0391] 3. NumIntraSubPartitions: It indicates how many sub-blocks are divided when ISP is applied. That is, it indicates that the partition is divided into NumIntraSubPartitions sub-blocks.

[0392] 4. LfnstDcOnly: For all transform blocks belonging to the current coding unit, each last non-zero coefficient position is the DC position (i.e., the upper left position within the corresponding transform block) or when there are no valid coefficients (i.e., when the corresponding CBF value is 0), the value of the LnfstDCOnly variable becomes 1.

[0393] In the case of a luminance single tree or a luminance double tree, the value of the LfnstDcOnly variable is determined by checking the condition only for the transform blocks corresponding to the luminance component in the corresponding coding unit, and in the case of a chroma single tree or a chroma double tree, the value of the LfnstDcOnly variable can be determined by checking the condition only for the transform blocks corresponding to the chroma components (Cb, Cr) in the corresponding coding unit. In the case of a single tree, the value of the LfnstDcOnly variable can be determined by checking the above conditions for all transform blocks corresponding to the luminance component and the chroma components (Cb, Cr) in the corresponding coding unit.

[0394] 5. LfnstZeroOutSigCoeffFlag: When applying LFNST, if there are valid coefficients only in the regions where valid coefficients can exist, it can be set to 1; otherwise, it can be set to 0.

[0395] In the case of a 4×4 transform block or an 8×8 transform block, up to 8 valid coefficients can be located starting from the (0,0) position (top left) in the corresponding transform block according to the scan order, and the remaining positions in the corresponding transform block are set to zero. In the case of a transform block that is not 4×4 and 8×8 and whose width and height are each equal to or greater than 4 (i.e., a transform block to which LFNST can be applied), 16 valid coefficients can be located starting from the (0,0) position (top left) in the corresponding transform block according to the scan order (i.e., the valid coefficients can be located only within the top left 4×4 block), and the remaining positions in the corresponding transform block can be set to zero.

[0396] In addition, as shown in Table 7, when being encoded in a split tree or a dual tree, when signaling the LFNST index, it is not checked whether MIP is applied to the chrominance component. In this way, LFNST can be appropriately applied to the chrominance component.

[0397] As shown in Table 7, when the condition (treeType == DUAL_TREE_CHROMA ||!intra_mip_flag[x0][y0] || Min(lfnstWidth, lfnstHeight) >= 16)) is satisfied, the LFNST index is signaled. This means that when the tree type is the dual tree chroma type (treeType == DUAL_TREE_CHROMA), the MIP mode is not applied (!intra_mip_flag[x0][y0]), or the smaller of the width and height of the block to which LFNST is applied is 16 or greater (Min(lfnstWidth, lfnstHeight) >= 16)), the LFNST index is signaled. That is, when the coded block is dual tree chroma, the LFNST index is signaled without determining whether the MIP mode is applied or the width and height of the block to which LFNST is applied.

[0398] In addition, the above condition can be interpreted as if the coded block is not dual tree chroma and MIP is not applied, the LFNST index is signaled without determining the width and height of the block to which LFNST is applied.

[0399] In addition, when the coded block is not dual tree chroma and MIP is applied, it can be interpreted that when the smaller of the width and height of the block to which LFNST is applied is 16 or greater, the LFNST index can be signaled.

[0400] On the other hand, the LFNST index is signaled only when transform skip is not applied to the luminance component as shown in Table 7 (i.e., when the condition transform_skip_flag[x0][y0][0] == 0 is satisfied).

[0401] Here, x0 and y0 represent the (x0, y0) coordinates when, for the luminance component, the top-left position in the picture is (0, 0), the horizontal X coordinate increases from left to right, and the vertical Y coordinate increases from top to bottom.

[0402] (x0, y0) are the coordinates based on the luminance component, but can also be used for the chrominance phase component. In this case, the actual position indicated by the (x0, y0) coordinates can be scaled based on the picture for the chrominance component. For example, when the chrominance format is 4:2:0, the actual position of the chrominance component indicated by (x0, y0) on the picture can be (x0 / 2, y0 / 2). For example, when the chrominance format is 4:2:0, the actual position of the chrominance component indicated by (x0, y0) on the picture can be (x0 / 2, y0 / 2).

[0403] In transform_skip_flag[x0][y0][0], the last index 0 refers to the luminance component. More specifically, in transform_skip_flag[x0][y0][cIdx], cIdx refers to the component it is for, and if the cIdx value is 0, a cIdx value of 0 indicates luminance, while a cIdx greater than 0 (1 or 2) indicates chrominance.

[0404] In addition, the variable LfnstDcOnly is initialized to the value 1, as shown in Table 7, and can be set to 0 according to the conditions in the parsing function used for residual coding, as shown in the following table.

[0405] [Table 9]

[0406]

[0407]

[0408] As shown in Table 9, the LfnstDcOnly value can only be set to 0 when the transform_skip_flag[x0][y0][cIdx] value is 0 (that is, only when transform skip is not applied to the component indicated by cIdx). If it is not in ISP mode, as shown in Table 7, the LFNST index is signaled only when the LfnstDcOnly value is 0, and when the LFNST index is not signaled, the LFNST index value can be inferred to be 0.

[0409] For reference, when calling the residual coding function presented in Table 9 during the call of transform_tree in Table 7, and for a single tree, the residual coding functions for luminance (cIdx = 0) and chrominance (cIdx = 1 or 2, corresponding to the Cb and Cr components) are all called, and for a dual tree, only the residual coding function for luminance (cIdx = 0) is called in the case of the dual-tree luminance (DUAL_TREE_LUMA), while in the case of the dual-tree chrominance (DUAL_TREE_CHROMA), only the residual coding function for chrominance (cIdx = 1 or 2, corresponding to the Cb and Cr components) is called.

[0410] The conditions for signaling the LFNST index for cases not in the ISP mode are summarized as follows (here, it can be assumed that other conditions for signaling the LFNST index are satisfied, for example, it is assumed that the condition Max(cbWidth, cbHeight) <= MaxTbSizeY is satisfied).

[0411] 1. When transform_skip_flag[x0][y0][0] is 1

[0412] - The LFNST index is inferred to be 0 without signaling.

[0413] 2. When transform_skip_flag[x0][y0][0] is 0

[0414] 2-A. When transform_skip_flag[x0][y0][1] is 0 and transform_skip_flag[x0][y0][2] is 0

[0415] - In Table 9, for all cIdx (for cIdx 0, 1, 2), the LfnstDcOnly value can be set to 0.

[0416] - If the LfnstDcOnly value is 0, then the LFNST index is signaled; otherwise, the LFNST index is not signaled and the value is inferred to be 0.

[0417] 2-B. When transform_skip_flag[x0][y0][1] is 0 and transform_skip_flag[x0][y0][2] is 1

[0418] - In Table 9, the LfnstDcOnly value can be set to 0 only when cIdx is 0 and 1.

[0419] - If the LfnstDcOnly value is 0, then signal the LFNST index; otherwise, do not signal the LFNST index and the value is inferred to be 0.

[0420] 2-C. When transform_skip_flag[x0][y0][1] is 1 and transform_skip_flag[x0][y0][2] is 0

[0421] - In Table 9, the LfnstDcOnly value can be set to 0 only when cIdx is 0 and 2

[0422] - If the LfnstDcOnly value is 0, then signal the LFNST index; otherwise, do not signal the LFNST index and the value is inferred to be 0.

[0423] 2-D. When transform_skip_flag[x0][y0][1] is 1 and transform_skip_flag[x0][y0][2] is 1

[0424] - In Table 9, the LfnstDcOnly value can be set to 0 only when cIdx is 0

[0425] - If the LfnstDcOnly value is 0, then signal the LFNST index; otherwise, do not signal the LFNST index and the value is inferred to be 0.

[0426] In the case of a single tree, check the values of transform_skip_flag[x0][y0][0], transform_skip_flag[x0][y0][1], transform_skip_flag[x0][y0][2] for the above cases. In the case of a luma double tree, only check transform_skip_flag[x0][y0][0], while in the case of a chroma double tree, check the values of transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2].

[0427] In the case of the ISP mode (the IntraSubPartitionsSplitType!= ISP_NO_SPLIT condition in Table 7, i.e., horizontal partitioning or vertical partitioning), as shown in Table 7, do not check the LfnstDcOnly variable and signal the LFNST index.

[0428] Therefore, in the case of the ISP mode under the dual-tree and single-tree of luminance, regardless of the value of the LfnstDcOnly variable, when the value of transform_skip_flag[x0][y0][0] is 0 (transform skip is not applied to the luminance component), the LFNST index is signaled (when the LFNST index is not signaled, the LFNST index value can be inferred as 0).

[0429] In the case of the chrominance dual-tree, based on the fact that ISP prediction is only applied to luminance in the current VVC standard, it is considered that ISP is not applied to chrominance, and the LFNST index can be signaled by checking the LfnstDcOnly variable in the above method, and as shown in Table 9, the LfnstDcOnly variable can be set to 0 only when the value of transform_skip_flag[x0][y0][cIdx] is 0.

[0430] Of course, the application of the ISP mode to luminance even affects the chrominance dual-tree. Therefore, even in the case of the chrominance dual-tree, regardless of the LfnstDcOnly variable, the LFNST index is signaled when the value of transform_skip_flag[x0][y0][0] is 0.

[0431] The conditions for signaling the LFNST index when the ISP mode is applied and the value of transform_skip_flag[x0][y0][0] is 0 are summarized as follows. If the value of transform_skip_flag[x0][y0][0] is 1, the LFNST index is not signaled and the LFNST index is inferred as 0. Of course, it can be assumed that other conditions required for signaling the LFNST index in Table 7 are satisfied. For example, it can be assumed that conditions such as Max(cbWidth, cbHeight)<=MaxTbSizeY are satisfied.

[0432] 1. In the case of a single tree

[0433] - Signal the LFNST index regardless of the value of the LfnstDcOnly variable

[0434] 2. In the case of a dual tree

[0435] 2-A. In the case of the luminance dual-tree

[0436] - Signal the LFNST index regardless of the value of the LfnstDcOnly variable

[0437] 2-B. In the case of the chrominance dual-tree

[0438] - According to the values of transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2], when the value of the LfnstDcOnly variable is set to 0, that is, when the cIdx value in transform_skip_flag[x0][y0][cIdx] is 1, the LfnstDcOnly variable value can be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0, and when the cIdx value is 2, the LfnstDcOnly variable value can be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0.

[0439] - If the LfnstDcOnly value is 0, signal the LFNST index; otherwise, do not signal the LFNST index and the value is inferred to be 0.

[0440] In the above cases, the situation of the chroma dual-tree is the same as the case where ISP is not applied.

[0441] In addition, according to the example, although transform skip is allowed for chroma components in the current VVC standard, as shown in the following table, transform skip flags corresponding to each chroma component are added.

[0442] [Table 10]

[0443]

[0444] In Table 10, except for the case of the luma dual-tree, it can be confirmed that transform_skip_flag[xC][yC][1] corresponding to whether transform skip is applied to Cb and transform_skip_flag[xC][yC][2] corresponding to whether transform skip is applied to Cr can be signaled. If the value of transform_skip_flag[xC][yC][1] is 1, transform skip is applied to Cb, and if it is 0, transform skip is not applied to Cb, and if the value of transform_skip_flag[xC][yC][2] is 1, transform skip is applied to Cr, and if it is 0, transform skip is not applied to Cr.

[0445] Therefore, even when the LFNST index value is greater than 0 (i.e., when LFNST is applied), the value of each transform_skip_flag[x0][y0][cIdx] for the luminance component (Y component) and chrominance components (Cb component and Cr component) can be different. According to Table 7, since the LFNST index value can be greater than 0 only when the transform_skip_flag[x0][y0][0] value is 0, when the LFNST index value is greater than 0, transform_skip_flag[x0][y0][0] is always zero.

[0446] Therefore, the cases where LFNST can be applied according to the transform_skip_flag[x0][y0][cIdx] value are summarized as follows. Here, the LFNST index is greater than 0, and the transform_skip_flag[x0][y0][0] value is 0. It can be assumed that other conditions for applying LFNST are satisfied, for example, both the width and height of the corresponding block can be greater than or equal to 4.

[0447] 1. Single tree

[0448] - LFNST is applied to the luminance component

[0449] - If the transform_skip_flag[x0][y0][1] value is 0, then LFNST is applied to the Cb component, and if it is 1, then LFNST is not applied to the Cb component.

[0450] - If the transform_skip_flag[x0][y0][2] value is 0, then LFNST is applied to the Cr component, and if it is 1, then LFNST is not applied to the Cr component.

[0451] 2. Luminance dual tree

[0452] - LFNST is applied to the luminance component

[0453] 3. Chrominance dual tree

[0454] - If the transform_skip_flag[x0][y0][1] value is 0, then LFNST is applied to the Cb component, and if it is 1, then LFNST is not applied to the Cb component.

[0455] - If the transform_skip_flag[x0][y0][2] value is 0, then LFNST is applied to the Cr component, and if it is 1, then LFNST is not applied to the Cr component.

[0456] As described above, in order to selectively apply LFNST according to the value of transform_skip_flag[x0][y0][cIdx], the following conditions should be added to the specification text of LFNST.

[0457] [Table 11]

[0458]

[0459]

[0460]

[0461] As shown in Table 11, when the value of the LFNST index (lfnst_idx) is not 0 (i.e., when LFNST is applied), by checking the value of transform_skip_flag[xTbY][yTbY][cIdx] of the component specified by cIdx (when lfnst_idx is not equal to 0 and transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0 and both nTbW and nTbH are greater than or equal to 4, the following is applied), it can be configured such that the subsequent encoding process is only executed when the value of transform_skip_flag[xTbY][yTbY][cIdx] is 0 (i.e., LFNST is applied).

[0462] In addition, the LFNST index can be signaled according to whether the transform is skipped for each color component.

[0463] As an example, when compared with Table 7, in Table 12, the condition for sending the LFNST index can be removed only when the value of transform_skip_flag[x0][y0][0] is 0.

[0464] [Table 12]

[0465]

[0466] However, since the method of setting the value of the LfnstDcOnly variable described in Table 12 is the same as that in Table 9, the setting of the value of the LfnstDcOnly variable is changed according to the value of transform_skip_flag[x0][y0][cIdx], and finally whether the LFNST index is signaled will also change.

[0467] When the ISP mode is not applied, how the LFNST index is signaled by the transform_skip_flag[x0][y0][cIdx] value is summarized as follows. It can be assumed that other conditions for signaling the LFNST index are already met, for example, conditions such as Max(cbWidth, cbHeight) <= MaxTbSizeY are met.

[0468] 1. In the case of a single tree

[0469] - When the transform_skip_flag[x0][y0][0] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0470] - When the transform_skip_flag[x0][y0][1] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0471] - When the transform_skip_flag[x0][y0][2] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0472] - When the LfnstDcOnly value is 0, the LFNST index can be signaled. If the LFNST index is not signaled, it can be inferred as 0.

[0473] 2. In the case of a dual tree for the luminance component

[0474] - When the transform_skip_flag[x0][y0][0] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0475] - When the LfnstDcOnly value is 0, the LFNST index can be signaled. If the LFNST index is not signaled, it can be inferred as 0.

[0476] 3. In the case of a dual tree for the chrominance component

[0477] - When the transform_skip_flag[x0][y0][1] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0478] - When the transform_skip_flag[x0][y0][2] value is 0, the LfnstDcOnly variable value can be set to 0 according to the method shown in Table 9.

[0479] - When the LfnstDcOnly value is 0, the LFNST index can be signaled. If the LFNST index is not signaled, it can be inferred as 0.

[0480] As shown in Table 12, the LfnstDcOnly value is initialized to 1, and in the case of a dual tree, the LFNST index corresponding to the luminance dual tree and the LFNST index corresponding to the chrominance dual tree can be signaled separately. This means that different LFNST kernels can be applied to luminance and chrominance.

[0481] In addition, the dual trees shown in Tables 7 to 12 can include DUAL_TREE_LUMA (corresponding to the luminance component) and DUAL_TREE_CHROMA (corresponding to the chrominance component) that appear in the current VVC specification document, and they can include a syntax parsing tree for luminance and a syntax parsing tree for chrominance that are different due to conditions such as the size condition of the coding unit. For example, the case of a single tree can be included.

[0482] When the ISP mode is applied, transform_skip_flag[x0][y0][0] is not signaled, and transform_skip_flag[x0][y0][0] is inferred as 0, as shown in Table 10. That is, as shown in Table 10, transform_skip_flag[x0][y0][0] is signaled only when the condition IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT (which is the case where the ISP mode is not applied) is satisfied.

[0483] In addition, as shown in Table 10, transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2] can be signaled regardless of whether the ISP mode is applied.

[0484] Therefore, the LFNST signaling in the case of applying the ISP mode can be summarized as follows. It can be assumed that other conditions for sending the LFNST index have been satisfied, for example, conditions such as Max(cbWidth, cbHeight) <= MaxTbSizeY can be satisfied.

[0485] 1. In the case of a single tree

[0486] - The LFNST index can be signaled. If the LFNST index is not signaled, it can be inferred as 0.

[0487] 2. In the case of a dual tree for the luminance component

[0488] - It is possible to signal the LFNST index. If the LFNST index is not signaled, it can be inferred as 0.

[0489] 3. In the case of the chrominance component dual-tree

[0490] - When the value of transform_skip_flag[x0][y0][1] is 0, the value of the LfnstDcOnly variable can be set to 0 according to the method shown in Table 9.

[0491] - When the value of transform_skip_flag[x0][y0][2] is 0, the value of the LfnstDcOnly variable can be set to 0 according to the method shown in Table 9.

[0492] - When the value of LfnstDcOnly is 0, it is possible to signal the LFNST index. If the LFNST index is not signaled, it can be inferred as 0.

[0493] When the ISP mode is applied, the LfnstDcOnly condition is not checked, as shown in Table 12. Therefore, in the case of the first single-tree and the second luminance component dual-tree, the LFNST index can be signaled without checking the LfnstDcOnly condition. In the case of the chrominance dual-tree, the LFNST index can be signaled according to the same conditions as in the case of not applying the ISP mode in Case 3 above. That is, the LFNST index can be signaled according to the LfnstDcOnly condition.

[0494] According to the example, since the value of transform_skip_flag[x0][y0][cIdx] can be assigned to the luminance component and the two chrominance components respectively, when the LFNST index value is greater than 0, that is, even when LFNST is applied, the LFNST can be applied to the component indicated by cIdx only when the value of transform_skip_flag[x0][y0][cIdx] is 0. The corresponding changed specification text is the same as Table 11.

[0495] According to the example, if the condition of checking whether the value of transform_skip_flag[x0][y0][0] is 0 only for the dual-tree case compared to Table 7 is removed, the LFNST index signaling can be configured as shown in Table 13. The LfnstDcOnly variable shown in Table 13 can be set to 0 according to the conditions shown in Table 9.

[0496] [Table 13]

[0497]

[0498] When configured as shown in Table 13, in the case of a single tree, the LFNST index signaling methods shown in Tables 7 to 9 can be applied, and in the case of a double tree, the method shown in Table 12 can be applied. Additionally, since the transform_skip_flag[x0][y0][cIdx] values are assigned to the luminance component and the two chrominance components respectively, as shown in Tables 7 to 9, even when the LFNST index value is greater than 0 (i.e., when LFNST is applied), it can be configured to apply LFNST to the component indicated by cIdx only when the transform_skip_flag[x0][y0][cIdx] value is 0. The corresponding changed specification text is the same as that in Table 11.

[0499] The following drawings are created to illustrate specific examples of this specification. Since the names of specific devices described in the drawings or the names of specific signals / messages / fields are presented by way of example, the technical features of this specification are not limited to the specific names used in the following drawings.

[0500] Figure 15 is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.

[0501] Figure 15 Each step disclosed in is based on some of the content described above Figures 5 to 14 described in. Therefore, the detailed descriptions that are repetitive with the content described above in Figure 3 and Figures 5 to 14 described in will be omitted or simplified.

[0502] The decoding device 300 according to an embodiment can receive information on an intra prediction mode, residual information, and an LFNST index from a bitstream (S1510).

[0503] More specifically, the decoding device 300 can decode information on the quantized transform coefficients of a current block from the bitstream and derive the quantized transform coefficients of a target block based on the information on the quantized transform coefficients of the current block. The information on the quantized transform coefficients of the target block can be included in a sequence parameter set (SPS) or a slice header and can include at least one of information on a reduction factor, information on a minimum transform size to which a reduced transform is applied, information on a maximum transform size to which a reduced transform is applied, and information on a transform index indicating either a transform kernel matrix included in a transform set or a reduced inverse transform size.

[0504] In addition, the decoding device may also receive information about the intra prediction mode of the current block and information about whether ISP is applied to the current block. The decoding device may derive whether the current block is divided into a predetermined number of sub-partition transform blocks by receiving and parsing flag information indicating whether ISP coding or ISP mode is applied. Here, the current block may be a coding block. In addition, the decoding device may derive the size and number of the divided sub-partition blocks through flag information indicating in which direction the current block is to be divided.

[0505] The decoding device 300 may derive the transform coefficients (S1520) by performing inverse quantization on the residual information (i.e., quantized transform coefficients) about the current block.

[0506] The derived transform coefficients may be arranged in units of 4×4 blocks in an inverse diagonal scan order, and the transform coefficients within the 4×4 block may also be arranged in an inverse diagonal scan order. That is, the transform coefficients for which inverse quantization has been performed may be arranged according to the inverse scan order applied in a video codec (such as in VVC or HEVC).

[0507] The decoding device may derive modified transform coefficients by applying LFNST to the transform coefficients.

[0508] Unlike the first transform that separates and transforms the transform target coefficients in the vertical or horizontal direction, LFNST is a non-separable transform that applies the transform without separating the coefficients in a specific direction. This non-separable transform may be a low-frequency non-separable transform that applies the forward transform only to the low-frequency region rather than the entire block region.

[0509] The LFNST index information may be received as syntax information, and the syntax information may be received as a binarized bin string including 0 and 1.

[0510] The syntax element of the LFNST index according to the present embodiment may indicate whether to apply inverse LFNST or inverse non-separable transform and any one of the transform kernel matrices included in the transform set, and when two transform kernel matrices are included in the transform set, the syntax element of the transform index may have three values.

[0511] That is, according to the embodiment, the syntax element values of the LFNST index may include: 0, which indicates the case where inverse LFNST is not applied to the target block; 1, which indicates the first transform kernel matrix among the transform kernel matrices; and 2, which indicates the second transform kernel matrix among the transform kernel matrices.

[0512] The intra prediction mode information and the LFNST index information may be signaled at the coding unit level.

[0513] The decoding device may derive a variable indicating whether there is a valid coefficient in the DC component of the current block to determine whether to parse the LFNST index for the current block, and may derive the variable based on the respective transform skip flag values of the color components of the current block (S1530).

[0514] The variable indicating whether there is a valid coefficient in the DC component of the current block may be represented as the variable LfnstDcOnly, and becomes 0 when there is a non-zero coefficient at a non-DC component for at least one transform block in a coding unit, and becomes 1 when there is no non-zero coefficient at positions other than the DC component for all transform blocks in a coding unit. In the present disclosure, the DC component refers to the upper left position or (0, 0) which is the position reference of the 2D component.

[0515] There may be several transform blocks within one coding unit. For example, in the case of the chrominance component, there may be transform blocks for Cb and Cr, and in the case of the single-tree type, there may be transform blocks for luminance, Cb, and Cr. According to the example, when a non-zero coefficient is found at a position other than the DC component position even in one of the transform blocks constituting the current coding block, the value of the variable LnfstDcOnly may be set to 0.

[0516] In addition, since residual coding is not performed on the corresponding transform block if there is no non-zero coefficient in the transform block, the value of the variable LfnstDcOnly is not changed by the corresponding transform block. Therefore, if there is no non-zero coefficient in the non-DC component of the transform block, the value of the variable LfnstDcOnly does not change and remains the previous value. For example, when the coding unit is coded according to the single-tree type and the value of the variable LfnstDcOnly is changed to 0 due to the luminance transform block, the value of the variable LfnstDcOnly remains 0 when a non-zero coefficient exists only in the DC component of the Cb transform block or when there is no non-zero coefficient in the Cb transform block. The value of the variable LfnstDcOnly is initially initialized to 1. If no component in the current coding unit updates the value of the variable LfnstDcOnly to 0, it remains the value 1 as it is, and when one of the transform blocks constituting the coding unit sets the value of the variable LfnstDcOnly to 0, it finally remains 0.

[0517] Furthermore, the variable LfnstDcOnly can be derived based on the respective transform skip flag values of the color components of the current block. The transform skip flag of the current block can be signaled for each color component, and if the tree type of the current block is a single tree and the transform skip flag value of the luminance component, then the variable LfnstDcOnly can be derived based on the transform skip flag value of the luminance component, the transform skip flag value of the chrominance Cb component, and the transform skip flag value of the chrominance Cr component. Alternatively, if the tree type of the current block is a dual-tree luminance, then the variable LfnstDcOnly is derived based on the transform skip flag value of the luminance component, and if the tree type of the current block is a dual-tree chrominance, then the variable LfnstDcOnly can be derived based on the transform skip flag value of the chrominance Cb component and the transform skip flag value of the chrominance Cr component.

[0518] According to an example, based on the transform skip flag values of the color components being 0, the variable LfnstDcOnly can indicate the presence of valid coefficients at positions other than the DC component. That is, if the tree type of the current block is a single tree and the transform skip flag value of the luminance component, then the variable LfnstDcOnly can be derived as 0 based on the fact that at least one of the transform skip flag value of the luminance component, the transform skip flag value of the chrominance Cb component, and the transform skip flag value of the chrominance Cr component is 0. Alternatively, if the tree type of the current block is a dual-tree luminance, then the variable LfnstDcOnly is derived based on the transform skip flag value of the luminance component, and if the tree type of the current block is a dual-tree chrominance, then the variable LfnstDcOnly can be derived based on the transform skip flag value of the chrominance Cb component and the transform skip flag value of the chrominance Cr component.

[0519] As described above, the variable LfnstDcOnly can be initially set to 1 at the coding unit level of the current block, and if the transform skip flag value is 0, then the variable LfnstDcOnly can be changed to 0 at the residual coding level.

[0520] The decoding device can parse the LFNST index (S1540) based on the variable LfnstDcOnly indicating the presence of valid coefficients at positions that are not the DC component, i.e., based on the variable LfnstDcOnly being 0.

[0521] Furthermore, in the case of a luminance block where the intra sub-partition (ISP) mode can be applied, the LFNST index can be parsed without deriving the variable LfnstDcOnly.

[0522] Specifically, when the ISP mode is applied and the transform skip flag of the luminance component (i.e., the value of transform_skip_flag[x0][y0][0]) is 0, and the tree type of the current block is a single tree or a luminance dual tree, the LFNST index can be signaled, regardless of the value of the variable LfnstDcOnly.

[0523] On the other hand, in the case of the chrominance component without the ISP mode applied, the value of the variable LfnstDcOnly can be set to 0 according to transform_skip_flag[x0][y0][1] which is the transform skip flag of the chrominance Cb component and transform_skip_flag[x0][y0][2] which is the transform skip flag of the chrominance Cr component. That is, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0, and when the cIdx value is 2, the value of transform_skip_flag[x0][y0][2] can be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. If the value of the variable LfnstDcOnly is 0, the decoding device can parse the LFNST index, otherwise the LFNST index can be not signaled and can be inferred as the value 0.

[0524] Thereafter, the decoding device can derive the modified transform coefficients from the transform coefficients based on the LFNST index and the LFNST matrix for the LFNST (S1550).

[0525] The decoding device can set multiple variables for the LFNST based on whether the LFNST index is not 0 (i.e., whether the LFNST index is greater than 0) and the respective transform skip flag values of the color components are 0.

[0526] For example, in the step of applying the LFNST after parsing the LFNST index, the decoding device can determine again whether the respective transform skip flag values of the color components are 0, and can set various variables for applying the LFNST. For example, the intra prediction mode for selecting the LFNST set, the number of transform coefficients output after applying the LFNST, the size of the block to which the LFNST is applied, etc. can be set.

[0527] In the case of a block encoded with BDPCM, the transform skip flag may be automatically set to 1, and in this case, even if the LFNST index is not 0, the transform skip flag may be 1. Therefore, when actually applying LFNST, the transform skip flag value for each color component can be checked again.

[0528] Alternatively, according to an example, when the flag value indicating whether there are valid encoding coefficients in the transform block is 0, there may be a case where the transform skip flag value is not checked. Also in this case, since it cannot be guaranteed that the transform skip flag value is 0 just because the LFNST index is not 0, when actually applying LFNST, the transform skip flag value for each color component can be checked again.

[0529] That is to say, the decoding device can check the transform skip flag value for each color component in the LFNST index parsing step, and can check the transform skip flag value for each color component again when actually applying LFNST.

[0530] The decoding device can determine an LFNST set including an LFNST matrix based on the intra prediction mode derived from the intra prediction mode information, and select any one of the multiple LFNST matrices based on the LFNST set and the LFNST index.

[0531] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks divided from the current block. That is to say, since the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can be equally applied to all sub-partition transform blocks. In addition, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks divided from the current block.

[0532] As described above, the transform set can be determined according to the intra prediction mode of the transform block to be transformed, and the inverse LFNST can be performed based on any one of the transform kernel matrices (i.e., LFNST matrices) included in the transform set indicated by the LFNST index. The matrix applied to the inverse LFNST can be referred to as an inverse LFNST matrix or an LFNST matrix, and such a matrix can have any name as long as it has a transpose relationship with the matrix used for the forward LFNST.

[0533] In one example, the inverse LFNST matrix can be a non-square matrix in which the number of columns is less than the number of rows.

[0534] The decoding device can derive the residual samples of the current block (S1560) based on a single inverse transform of the modified transform coefficients.

[0535] In this case, as the inverse first transform, a conventional separation transform can be used, and the above MTS can be used.

[0536] Subsequently, the decoding device 300 can generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block.

[0537] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices shown in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.

[0538] Figure 16 is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.

[0539] Figure 16 Each step disclosed in is based on some of the above Figures 5 to 14 described content. Therefore, the detailed description that duplicates the above in Figure 2 and Figures 5 to 14 will be omitted or simplified.

[0540] The encoding device 200 according to an embodiment can derive predicted samples of the current block based on the intra prediction mode applied to the current block.

[0541] The encoding device can perform prediction for each sub-partition transform block when applying ISP to the current block.

[0542] The encoding device can determine whether to apply ISP encoding or ISP mode to the current block (i.e., the encoding block), determine in which direction the current block is to be divided according to the determination result, and derive the size and number of the divided sub-blocks.

[0543] The same intra prediction mode is applied to the sub-partition transform blocks divided from the current block, and the encoding device can derive predicted samples for each sub-partition transform block. That is, the encoding device performs intra prediction sequentially, for example, horizontally or vertically, from left to right or from top to bottom, according to the division form of the sub-partition transform blocks. For the leftmost or topmost sub-blocks, the reconstructed pixels of the already encoded encoding blocks are referred to in the conventional intra prediction method. Additionally, for each side of the subsequent internal sub-partition transform blocks, when it is not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to the corresponding side, the already encoded adjacent encoding blocks refer to the reconstructed pixels as in the conventional intra prediction method.

[0544] The encoding device 200 can derive the residual samples of the current block based on the predicted samples (S1610).

[0545] The encoding device 200 may derive the transform coefficients of the current block by applying at least one of LFNST and MTS to the residual samples, and may arrange the transform coefficients according to a predetermined scan order.

[0546] The encoding device may derive the transform coefficients of the current block based on a single transform of the residual samples (S1620).

[0547] The single transform may be performed by multiple transform kernels as in MTS, and in this case, the transform kernel may be selected based on the intra prediction mode.

[0548] The encoding device 200 may determine whether to perform a secondary transform or a non-separable transform (specifically, LFNST) on the transform coefficients of the current block, and apply LFNST to the transform coefficients to derive modified transform coefficients.

[0549] Unlike the first transform that separates and transforms the transform target coefficients in the vertical or horizontal direction, LFNST is a non-separable transform that applies the transform without separating the coefficients in a specific direction. The non-separable transform may be a low-frequency non-separable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.

[0550] The encoding device may apply multiple LFNST matrices to the transform coefficients to derive a variable indicating whether there are valid coefficients in the DC component of the current block, and may derive the variable based on the respective transform skip flag values of the color components of the current block (S1630).

[0551] The encoding device may derive the variable after applying LFNST to each LFNST matrix candidate, or in the state where LFNST is not applied without applying LFNST.

[0552] Specifically, the encoding device may apply multiple LFNST candidates (i.e., LFNST matrices) to exclude the corresponding LFNST matrix in which the valid coefficients of all transform blocks exist only in the DC position (of course, when CBF is 0, the corresponding variable is excluded from the process), and compare the RD values only among the LFNST matrices in which the variable LfnstDcOnly value is 0. For example, when LFNST is not applied, it is included in the comparison process because it is independent of the variable LfnstDcOnly value (in this case, since LFNST is not applied, the variable LfnstDcOnly value may be determined based on the transform coefficients obtained as a result of the single transform), and the LFNST matrix with the corresponding LfnstDcOnly value of 0 is also included in the RD value comparison process.

[0553] A variable indicating whether there are valid coefficients in the DC component of the current block can be represented as the variable LfnstDcOnly, and becomes 0 when there are non-zero coefficients at non-DC components for at least one transform block in a coding unit, and becomes 1 when there are no non-zero coefficients at positions other than the DC component for all transform blocks in a coding unit.

[0554] There may be several transform blocks in a coding unit. For example, in the case of chrominance components, there may be transform blocks for Cb and Cr, and in the case of single-tree type, there may be transform blocks for luminance, Cb, and Cr. According to the example, when a non-zero coefficient is found at a position other than the DC component position in even one of the transform blocks constituting the current coding block, the value of the variable LnfstDcOnly can be set to 0.

[0555] In addition, since residual coding is not performed on the corresponding transform block if there are no non-zero coefficients in the transform block, the value of the variable LfnstDcOnly is not changed by the corresponding transform block. Therefore, if there are no non-zero coefficients in the non-DC component of the transform block, the value of the variable LfnstDcOnly does not change and remains the previous value. For example, when the coding unit is coded according to the single-tree type and the value of the variable LfnstDcOnly is changed to 0 due to the luminance transform block, the value of the variable LfnstDcOnly remains 0 when non-zero coefficients exist only in the DC component of the Cb transform block or when there are no non-zero coefficients in the Cb transform block. The value of the variable LfnstDcOnly is initially initialized to 1, and if no component in the current coding unit updates the value of the variable LfnstDcOnly to 0, it remains the value 1 as it is, and when one of the transform blocks constituting the coding unit sets the value of the variable LfnstDcOnly to 0, it finally remains 0.

[0556] In addition, the variable LfnstDcOnly can be derived based on the respective transform skip flag values of the color components of the current block. The transform skip flag of the current block can be signaled for each color component, and if the tree type of the current block is single-tree, the transform skip flag value of the luminance component, then the variable LfnstDcOnly can be derived based on the transform skip flag value of the luminance component, the transform skip flag value of the chrominance Cb component, and the transform skip flag value of the chrominance Cr component. Alternatively, if the tree type of the current block is dual-tree luminance, the variable LfnstDcOnly is derived based on the transform skip flag value of the luminance component, and if the tree type of the current block is dual-tree chrominance, the variable LfnstDcOnly can be derived based on the transform skip flag value of the chrominance Cb component and the transform skip flag value of the chrominance Cr component.

[0557] According to the example, the transform skip flag value based on the color component is 0, and the variable LfnstDcOnly can indicate the existence of valid coefficients at positions other than the DC component. That is, if the tree type of the current block is a single tree and the transform skip flag value of the luminance component is such that the variable LfnstDcOnly can be derived as 0 based on the fact that at least one of the transform skip flag values of the luminance component, the chrominance Cb component, and the chrominance Cr component is 0. Alternatively, if the tree type of the current block is a dual-tree luminance, the variable LfnstDcOnly is derived based on the transform skip flag value of the luminance component, and if the tree type of the current block is a dual-tree chrominance, the variable LfnstDcOnly can be derived based on the transform skip flag values of the chrominance Cb component and the chrominance Cr component.

[0558] As described above, the variable LfnstDcOnly can be initially set to 1 at the coding unit level of the current block, and if the transform skip flag value is 0, the variable LfnstDcOnly can be changed to 0 at the residual coding level.

[0559] The encoding device can select an optimal LFNST matrix based on the variable indicating the existence of valid coefficients at positions other than the DC component, and can derive modified transform coefficients based on the selected LFNST matrix (S1640).

[0560] The encoding device can set multiple variables for LFNST based on whether the respective transform skip flag values of the color components are 0 in the step of deriving the modified transform coefficients.

[0561] For example, after determining whether to apply LFNST, the encoding device can again determine whether the respective transform skip flag values of the color components are 0 in the step of applying LFNST, and can set various variables for applying LFNST. For example, the intra prediction mode for selecting an LFNST set, the number of transform coefficients output after applying LFNST, the size of the block to which LFNST is applied, etc. can be set.

[0562] In the case of a block encoded by BDPCM, since the transform skip flag can be automatically set to 1, when actually applying LFNST, the transform skip flag values of each color component can be checked again.

[0563] Alternatively, according to the example, when the flag value indicating the existence of encoded valid coefficients in the transform block is 0, there may be a case where the transform skip flag value is not checked. Similarly in this case, since it cannot be guaranteed that the transform skip flag value is 0 just because the LFNST index is not 0, when actually applying LFNST, the transform skip flag values of each color component can be checked again.

[0564] That is to say, the encoding device can check the transform skip flag value of each color component in the step of determining whether to apply LFNST, and can check the transform skip flag value of each color component again when actually applying LFNST.

[0565] In addition, in the case of a luminance block where the intra sub-partition (ISP) mode can be applied, LFNST can be applied without deriving the variable LfnstDcOnly.

[0566] Specifically, when the ISP mode is applied and the transform skip flag for the luminance component (i.e., the value of transform_skip_flag[x0][y0][0]) is 0, and the tree type of the current block is a single tree or a luminance double tree, LFNST can be applied regardless of the value of the variable LfnstDcOnly.

[0567] On the other hand, in the case of a chrominance component where the ISP mode is not applied, the value of the variable LfnstDcOnly can be set to 0 according to transform_skip_flag[x0][y0][1] which is the transform skip flag of the chrominance Cb component and transform_skip_flag[x0][y0][2] which is the transform skip flag of the chrominance Cr component. That is to say, in transform_skip_flag[x0][y0][cIdx], when the cIdx value is 1, the value of the variable LfnstDcOnly can be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0, and when the cIdx value is 2, the value of transform_skip_flag[x0][y0][2] can be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. If the value of the variable LfnstDcOnly is 0, the encoding device can apply LFNST, otherwise LFNST is not applied.

[0568] The encoding device 200 can determine the LFNST set based on the mapping relationship according to the intra prediction mode applied to the current block, and perform LFNST based on one of the two LFNST matrices included in the LFNST set, that is, the non-separable transform.

[0569] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks partitioned from the current block. That is, since the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can also be equally applied to all sub-partition transform blocks. In addition, since the LFNST index is encoded in units of coding units, the same LFNST matrix can be applied to the sub-partition transform blocks partitioned from the current block.

[0570] As described above, the transform set can be determined according to the intra prediction mode of the transform block to be transformed. The matrix applied to the LFNST has a transpose relationship with the matrix used for the inverse LFNST.

[0571] In one example, the LFNST matrix can be a non-square matrix in which the number of rows is less than the number of columns.

[0572] The encoding device can construct the image information such that the LFNST index indicating the LFNST matrix applied to the LFNST is parsed based on the fact that the variable LfnstDcOnly is initially set to 1 at the coding unit level of the current block, and the variable LfnstDcOnly is changed to 0 at the residual coding level when the transform skip flag value is 0, and the variable LfnstDcOnly is 0 (S1650).

[0573] The encoding device can perform quantization based on the modified transform coefficients of the current block to derive quantized transform coefficients, and encode the LFNST index.

[0574] That is, the encoding device can generate residual information including information about the quantized transform coefficients. The residual information can include the above-mentioned transform-related information / syntax elements. The encoding device can encode the image / video information including the residual information and output the encoded image / video information in the form of a bitstream.

[0575] More specifically, the encoding device 200 can generate information about the quantized transform coefficients and encode the information about the generated quantized transform coefficients.

[0576] The syntax element of the LFNST index according to the present embodiment can indicate whether (inverse) LFNST is applied and any one of the LFNST matrices included in the LFNST set, and when the LFNST set includes two transform kernel matrices, the syntax element of the LFNST index can have three values.

[0577] According to an embodiment, when the partition tree structure of the current block is of a double-tree type, the LFNST index can be encoded for each of the luminance block and the chrominance block.

[0578] According to an embodiment, the syntax element value of the transform index can be derived as 0, 1, and 2. 0 indicates the case where (inverse) LFNST is not applied to the current block, 1 indicates the first LFNST matrix in the LFNST matrix, and 2 indicates the second LFNST matrix in the LFNST matrix.

[0579] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation can be omitted. When quantization / dequantization is omitted, the quantized transform coefficients can be referred to as transform coefficients. When transformation / inverse transformation is omitted, the transform coefficients can be referred to as coefficients or residual coefficients, or can still be referred to as transform coefficients for the sake of consistency in expression.

[0580] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information can include information about the transform coefficients, and the information about the transform coefficients can be signaled through the residual coding syntax. The transform coefficients can be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. The residual samples can be derived based on the inverse transform (transformation) of the scaled transform coefficients. These details can also be applied / expressed in other parts of the present disclosure.

[0581] In the above embodiment, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step can be executed in an order or steps different from the above order or steps, or a certain step can be executed concurrently with other steps. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and without affecting the scope of the present disclosure, another step can be incorporated or one or more steps in the flowchart can be deleted.

[0582] The above method according to the present disclosure can be implemented in software form, and the encoding device and / or decoding device according to the present disclosure can be included in devices for image processing such as televisions, computers, smart phones, set-top boxes, and display devices.

[0583] When the embodiments in the present disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor can include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.

[0584] In addition, the decoding device and the encoding device applying the present disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.

[0585] In addition, the processing method applying the present disclosure can be produced in the form of a program executable by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in the computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices for storing computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in the computer-readable recording medium or transmitted through a wired or wireless communication network. In addition, the embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer according to the embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.

[0586] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features of the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or executed in a method.

Claims

1. An apparatus, the apparatus comprises: a memory; and at least one processor connected to the memory, the at least one processor being configured to: receive residual information through a bitstream, derive transform coefficients for a current block based on the residual information, derive residual samples for the current block based on a non-separable transform for the transform coefficients, and generate a reconstructed picture based on the residual samples, wherein a variable indicating whether valid coefficients exist only in the upper-left sample position within the current block is derived based on a first transform skip flag for a chrominance Cb component in the current block and a second transform skip flag for a chrominance Cr component in the current block, wherein the first transform skip flag for the chrominance Cb component and the second transform skip flag for the chrominance Cr component are signaled from the bitstream respectively based on the tree type of the current block being a dual-tree chrominance, wherein, based on the variable indicating that the valid coefficients exist at positions other than the upper-left sample position, a non-separable transform index is parsed, wherein, based on the non-separable transform index being equal to 0, the non-separable transform is not applied to the transform coefficients for the chrominance Cb component, wherein, based on the non-separable transform index not being equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 0, the non-separable transform is applied to the transform coefficients for the chrominance Cb component, and wherein, based on the non-separable transform index not being equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 1, the non-separable transform is not applied to the transform coefficients for the chrominance Cb component.

2. The apparatus according to claim 1, wherein the variable indicating that the valid coefficients exist at the positions other than the upper-left sample position is derived based on at least one of the first transform skip flag for the chrominance Cb component or the second transform skip flag for the chrominance Cr component being equal to 0.

3. The apparatus according to claim 2, wherein the variable is initially set to 1 at the coding unit level of the current block, wherein, based on at least one of the first transform skip flag for the chrominance Cb component or the second transform skip flag for the chrominance Cr component being equal to 0, the variable is changed to 0 at the residual coding level, and wherein the non-separable transform index is parsed based on the variable being equal to 0.

4. The apparatus according to claim 1, wherein the at least one processor is further configured to set a plurality of variables for the non-separable transform based on whether the non-separable transform index is not equal to 0 and whether the first transform skip flag and the second transform skip flag are equal to 0.

5. An apparatus, the apparatus comprises: a memory; and at least one processor connected to the memory, the at least one processor being configured to: derive prediction samples for a current block, Derive a residual sample for the current block based on the prediction sample, Derive transform coefficients for the current block from the residual sample based on an inseparable transform, Generate residual information based on the transform coefficients, and Encode image information including the residual information, wherein a variable indicating whether valid coefficients exist only in the top-left sample position within the current block is derived based on a first transform skip flag for the chrominance Cb component in the current block and a second transform skip flag for the chrominance Cr component in the current block, wherein the first transform skip flag for the chrominance Cb component and the second transform skip flag for the chrominance Cr component are respectively encoded into the bitstream based on the tree type of the current block being a dual-tree chrominance, wherein an inseparable transform index is encoded based on the variable indicating the existence of the valid coefficients at positions other than the top-left sample position, wherein based on the inseparable transform index being equal to 0, the inseparable transform is not applied to the transform coefficients for the chrominance Cb component, wherein based on the inseparable transform index not being equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 0, the inseparable transform is applied to the transform coefficients for the chrominance Cb component, and wherein based on the inseparable transform index not being equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 1, the inseparable transform is not applied to the transform coefficients for the chrominance Cb component.

6. The apparatus according to claim 5, wherein, the variable indicating the existence of the valid coefficients at the positions other than the top-left sample position is derived based on at least one of the first transform skip flag for the chrominance Cb component or the second transform skip flag for the chrominance Cr component being equal to 0.

7. The apparatus according to claim 6, wherein, the variable is initially set to 1 at the coding unit level of the current block, wherein based on at least one of the first transform skip flag for the chrominance Cb component or the second transform skip flag for the chrominance Cr component being equal to 0, the variable is changed to 0 at the residual coding level, and wherein the inseparable transform index is encoded based on the variable being equal to 0.

8. The apparatus according to claim 5, wherein, the at least one processor is further configured to set a plurality of variables for the inseparable transform based on whether the first transform skip flag and the second transform skip flag are equal to 0.

9. An apparatus, the apparatus comprising: at least one processor, the at least one processor being configured to: Derive a prediction sample for a current block, Derive a residual sample for the current block based on the prediction sample, Derive transform coefficients for the current block from the residual sample based on an inseparable transform, Generate residual information based on the transform coefficients, and Encode image information including the residual information to generate a bitstream; and A transmitter configured to transmit the bitstream, wherein a variable indicating whether valid coefficients exist only in the upper-left sample position within the current block is derived based on a first transform skip flag for the chrominance Cb component in the current block and a second transform skip flag for the chrominance Cr component in the current block, wherein the first transform skip flag for the chrominance Cb component and the second transform skip flag for the chrominance Cr component are respectively encoded into the bitstream based on the tree type of the current block being a dual-tree chrominance, wherein an inseparable transform index is encoded based on the variable indicating the existence of the valid coefficients at positions other than the upper-left sample position, wherein based on the inseparable transform index being equal to 0, the inseparable transform is not applied to the transform coefficients for the chrominance Cb component, wherein based on the inseparable transform index being not equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 0, the inseparable transform is applied to the transform coefficients for the chrominance Cb component, and wherein based on the inseparable transform index being not equal to 0 and the first transform skip flag for the chrominance Cb component being equal to 1, the inseparable transform is not applied to the transform coefficients for the chrominance Cb component.