Image compilation method and device based on transformation

By applying LFNST transform technology in the image encoding and decoding process, optimizing transform coefficient processing and quantization of chrominance components, the problem of efficient compression of high-resolution images and videos is solved, and transmission and storage costs are reduced.

CN115066904BActive Publication Date: 2025-09-16LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180013799.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-10
Filing Date
2021-01-11
Publication Date
2025-09-16
Estimated Expiration
2041-01-11

AI Technical Summary

Technical Problem

As the demand for high-resolution and high-quality images and videos increases, existing transmission and storage costs increase, and existing compression technologies find it difficult to effectively compress image and video data of various characteristics.

Method used

The LFNST transform technology is used to process the transform coefficients in the image encoding and decoding process. By determining the tree type and the application of the scaling list, the quantization process of the chrominance component is optimized to improve the image coding efficiency.

Benefits of technology

The compression efficiency of images and videos is improved, especially the quantization efficiency of chrominance components in single-tree types, which reduces transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115066904B_ABST
    Figure CN115066904B_ABST
Patent Text Reader

Abstract

The image decoding method according to the present document includes the following steps: applying LFNST to a transform coefficient so as to derive a modified transform coefficient; and deriving a residual sample for a target block based on an inverse primary transform of the modified transform coefficient, wherein the step of deriving the transform coefficient includes: determining whether a scaling list is applied to the current block based on a tree type of the current block and whether LFNST is applied; and deriving the transform coefficient for the current block from the residual information based on the determination result, and when the tree type of the current block is a single tree and a chroma component, the scaling list can be applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to image coding technology, and more particularly, to a method and device for coding an image based on transformation in an image coding system. Background Art

[0002] Recently, the demand for high-resolution and high-quality images and videos such as ultra-high-definition (HUD) images and 4K or 8K or larger videos has been increasing in various fields. As image and video data becomes higher in resolution and higher in quality, the amount of information or the number of bits transmitted increases compared to existing image and video data. Therefore, if a medium such as an existing wired or wireless broadband line is used to transmit image data or an existing storage medium is used to store image and video data, the transmission cost and storage cost increase.

[0003] In addition, there has been a recent increase in interest and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms, and the broadcasting of images and videos having image characteristics different from those of real images, such as game images, has increased.

[0004] Therefore, in order to efficiently compress and transmit or store and play back information of high-resolution and high-quality images and videos having such various characteristics, efficient image and video compression technology is required. Summary of the Invention

[0005] Technical issues

[0006] The technical aspect of the present disclosure is to provide a method and apparatus for increasing image coding efficiency.

[0007] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing quantization efficiency.

[0008] Yet another technical aspect of the present disclosure is to provide a method and apparatus for increasing quantization efficiency of chroma components in a single tree type.

[0009] Technical Solution

[0010] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method may include: deriving transform coefficients for a current block based on residual information received from a bitstream, deriving modified transform coefficients by applying LFNST to the transform coefficients, and deriving residual samples for a target block based on an inverse primary transform of the modified transform coefficients, wherein the deriving of the transform coefficients may include determining whether a scaling list is applied to the current block based on whether LFNST is applied and a tree type of the current block, deriving the transform coefficients for the current block from the residual information based on the determination result, and applying the scaling list when the tree type of the current block is a single tree and the current block is a chroma component.

[0011] LFNST may not be applied to the chroma components of the current block.

[0012] When the tree type of the current block is a single tree and LFNST is performed on the current block, the scaling list may not be applied to the luma component of the current block.

[0013] The method may further include receiving flag information indicating whether the scaling list is available when performing the LFNST.

[0014] When the flag information indicates that the scaling list is unavailable and the LFNST index is greater than 0, the scaling list may not be applied to the luma component.

[0015] When the flag information indicates that the scaling list is unavailable and the LFNST index is greater than 0, if the tree type of the current block is dual-tree chroma, the scaling list may not be applied to the chroma component.

[0016] When the flag information indicates that the scaling list is unavailable and the LFNST index is greater than 0, if the tree type of the current block is dual-tree luma, the scaling list may not be applied to the luma component.

[0017] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method may include deriving transform coefficients of a current block from residual samples of the current block based on a transform process, determining whether a scaling list is applied to the current block based on whether LFNST is performed in the transform process and a tree type of the current block, and quantizing the transform coefficients based on the determination, wherein the scaling list may be applied when the tree type of the current block is a single tree and the current block is a chroma component.

[0018] According to yet another embodiment of the present disclosure, a digital storage medium storing image data including encoded image information and a bit stream generated according to an image encoding method performed by an encoding device may be provided.

[0019] According to yet another embodiment of the present disclosure, a digital storage medium may be provided that stores image data including encoded image information and a bit stream, so that a decoding device performs an image decoding method.

[0020] Beneficial effects

[0021] According to the present disclosure, the overall image / video compression efficiency can be increased.

[0022] According to the present disclosure, quantization efficiency can be increased.

[0023] According to the present disclosure, it is possible to increase the quantization efficiency of chroma components in a single tree type.

[0024] The effects that can be achieved through the specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from the present disclosure by a person of ordinary skill in the relevant field. Therefore, the specific effects of the present disclosure are not limited to the effects explicitly described in the present disclosure, and can include various effects that can be understood or derived from the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 An example of a video / image coding system to which the present disclosure is applicable is schematically illustrated.

[0026] Figure 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure is applicable.

[0027] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.

[0028] Figure 4 A multiple transformation scheme according to an embodiment of this document is schematically illustrated.

[0029] Figure 5 The intra direction mode of 65 prediction directions is exemplarily shown.

[0030] Figure 6 is a diagram for explaining RST according to an embodiment of the present disclosure.

[0031] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example.

[0032] Figure 8 is a diagram illustrating an order in which output data of a forward sub-transform is arranged into two-dimensional blocks according to an example.

[0033] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.

[0034] Figure 10 is a diagram illustrating a block shape to which LFNST is applied.

[0035] Figure 11 is a diagram illustrating an arrangement of output data of a forward LFNST according to an example.

[0036] Figure 12 is a diagram illustrating limiting the number of output data for forward LFNST to a maximum of 16 according to an example.

[0037] Figure 13 is a diagram illustrating zeroing in a block to which 4×4 LFNST is applied according to an example.

[0038] Figure 14 is a diagram illustrating clearing in a block to which 8×8 LFNST is applied according to an example.

[0039] Figure 15 is a diagram illustrating an image decoding method according to an example.

[0040] Figure 16 is a diagram illustrating an image encoding method according to an example.

[0041] Figure 17 The structure of a content streaming system to which the present disclosure is applied is illustrated. DETAILED DESCRIPTION

[0042] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used to describe specific embodiments, rather than to limit the technical spirit of this document. Unless otherwise clearly indicated in the context, singular expressions include plural expressions. Terms such as "including" or "having" in this specification should be understood to indicate the presence of characteristics, numbers, steps, operations, elements, components, or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more characteristics, numbers, steps, operations, elements, components, or combinations thereof.

[0043] In addition, to facilitate the description of different feature functions, the elements in the drawings described in this document are illustrated independently. This does not mean that each element is implemented as separate hardware or separate software. For example, at least two elements can be combined to form a single element, or a single element can be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included in the scope of the rights of this document unless it deviates from the essence of this document.

[0044] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.

[0045] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the Versatile Video Coding (VVC) standard (ITU-T Recommendation H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Recommendation H.265), the Essential Video Coding (EVC) standard, the AVS2 standard, etc.).

[0046] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, the embodiments may be combined with each other and performed.

[0047] In this document, video can mean a collection of a series of images over time. Generally, a picture means a unit of an image representing a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.

[0048] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a value of a pixel, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or when this pixel value is converted into a frequency domain, it may refer to a transform coefficient in a frequency domain.

[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include a luma block and two chroma (e.g., CB, CR) blocks. Depending on the situation, terms such as unit, block, and region may be used interchangeably. In general, an MxN block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0050] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".

[0051] Additionally, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) "only A," 2) "only B," and / or 3) "both A and B." In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."

[0052] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present disclosure, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.

[0053] In addition, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”

[0054] In addition, the brackets used in this disclosure may mean "for example." Specifically, when it is indicated as "prediction (intra-frame prediction)", this may mean that "intra-frame prediction" is proposed as an example of "prediction." That is, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, when it is indicated as "prediction (i.e., intra-frame prediction)", this may also mean that "intra-frame prediction" is proposed as an example of "prediction".

[0055] Technical features described individually in one drawing in the present disclosure may be implemented individually or may be implemented simultaneously.

[0056] Figure 1 An example of a video / image coding system to which embodiments of this document may be applied is schematically illustrated.

[0057] refer to Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit the coded video / image information or data to the receive device in the form of a file or stream transmission via a digital storage medium or a network.

[0058] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0059] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.

[0060] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0061] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver in a receiving device via a digital storage medium or network in the form of a file or streaming. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.

[0062] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0063] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.

[0064] Figure 2 The figure schematically illustrates the configuration of a video / image encoding device to which the present document can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.

[0065] refer to Figure 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be composed of one or more hardware components (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0066] The image partitioner 210 partitions the input image (or picture, or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding units may be recursively partitioned according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units of increasing depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree structure and / or the ternary tree structure. Alternatively, the binary tree structure may be applied first. The coding process according to this document may be performed based on the final coding unit that has not been further partitioned. In this case, based on coding efficiency according to image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively partitioned into coding units of increasing depth as needed, so that the optimally sized coding unit may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit described above. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal based on the transform coefficient.

[0067] Depending on the situation, the terms such as unit and block, region, etc. can be used interchangeably. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to a pixel or pel of a picture (or image).

[0068] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block including prediction samples of the current block. The predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to the prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0069] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.

[0070] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal cannot be sent. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector prediction item, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0071] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can also apply intra prediction and inter prediction at the same time. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform prediction on the block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding such as screen content coding (SCC) for games. Although IBC basically performs prediction in the current block, it can be performed similarly to inter prediction in terms of deriving a reference block in the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0072] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstruction signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size or can be applied to blocks of variable size instead of square blocks.

[0073] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signals (information about the quantized transform coefficients) and output the encoded signals in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information necessary for video / image reconstruction, in addition to the quantized transform coefficients (e.g., syntax element values, etc.), together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of the network abstraction layer (NAL). The video / image information may also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), and the like. In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded by the above-mentioned encoding process and included in the bitstream. The bitstream may be sent over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage device (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.

[0074] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 250 adds the reconstructed residual signal to the prediction signal output from the predictor 220, so that a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample or reconstructed sample array) can be generated. When there is no residual in the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and as described later, can be used for inter-frame prediction of the next picture through filtering.

[0075] Meanwhile, during picture encoding and / or reconstruction, luminance mapping and chroma scaling (LMCS) may be applied.

[0076] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 290. The information about filtering can be encoded in the entropy encoder 290 and output in the form of a bitstream.

[0077] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter-frame predictor 280. This allows the encoding device to avoid prediction mismatches in the encoding device 200 and the decoding device when applying inter-frame prediction, and improves coding efficiency.

[0078] The memory 270DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of blocks in the current picture from which motion information has been derived (or encoded), and / or motion information of blocks in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information of neighboring blocks or motion information of temporally neighboring blocks. The memory 270 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.

[0079] Figure 3 This is a diagram schematically illustrating the configuration of a video / image decoding device to which this document can be applied.

[0080] refer to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be composed of one or more hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0081] When a bit stream including video / image information is input, the decoding device 300 can Figure 2 The image is reconstructed correspondingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the information related to the block segmentation obtained from the bit stream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided into a quadtree structure, a binary tree structure and / or a ternary tree structure using a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. In addition, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.

[0082] The decoding device 300 can receive the data from the Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In the present disclosure, the information and / or syntax elements sent / received using the signal to be described later can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values ​​of the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoded target syntax element information and decoding information of neighboring and decoding target blocks, or information about the symbol / bin decoded in the previous step, to determine a context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bin to generate the symbol corresponding to each syntax element value. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the symbol / bin information decoded by the context model for the next symbol / bin. Information about prediction from the information decoded in the entropy decoder 310 can be provided to the predictor 330, and information about the residual, i.e., quantized transform coefficients and associated parameter information, which have been entropy decoded in the entropy decoder 310, can be input to the dequantizer 321. In addition, information about filtering from the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiver (not shown) that receives the signal output from the encoding device can further constitute the decoding device 300 as an internal / external element, and the receiver can be a component of the entropy decoder 310. Meanwhile, the decoding device according to the present disclosure may be referred to as a video / image / picture encoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, a predictor 330, an adder 340, a filter 350, and a memory 360.

[0083] The dequantizer 321 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0084] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing inverse transformation on the transformation coefficients.

[0085] The predictor may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically may determine the intra / inter prediction mode.

[0086] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can also apply intra prediction and inter prediction at the same time. This can be called combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copying (IBC) to predict the block. Intra block copying can be used for content image / video coding of games, etc., such as screen content coding (SCC). Although IBC basically performs prediction in the current block, it can be performed similarly to inter prediction in terms of deriving a reference block in the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0087] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located in the neighborhood of the current block or far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.

[0088] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the inter-frame prediction mode used for the current block.

[0089] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual when processing the target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.

[0090] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter prediction of the next picture.

[0091] In addition, luma mapping and chroma scaling (LMCS) can be applied to the picture decoding process.

[0092] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0093] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block in the current picture from which the motion information has been derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 260 to be used as the motion information of the neighboring block or the motion information of the temporally neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.

[0094] In this specification, the examples described in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 may be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.

[0095] As described above, when performing video coding, prediction is performed to improve compression efficiency. A prediction block including prediction samples for a current block, i.e., a target coding block, can be generated by prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived similarly in the encoding device and the decoding device. The encoding device can improve image coding efficiency by signaling information (residual information) about the residual between the original block, rather than the original sample value itself of the original block, and the prediction block to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and can generate a reconstructed picture including the reconstructed block.

[0096] Residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, can derive a transform coefficient by performing a transformation process on the residual samples (residual sample array) included in the residual block, can derive a quantized transform coefficient by performing a quantization process on the transform coefficient, and can send the relevant residual information to the decoding device (via a bitstream) with a signal. In this case, the residual information may include information such as value information, position information, a transformation scheme, a transform kernel, and a quantization parameter of the quantized transform coefficient. The decoding device can perform an inverse quantization / inverse transformation process based on the residual information and can derive residual samples (or residual blocks). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can derive a residual block for inter-frame prediction reference of a subsequent picture by performing inverse quantization / inverse transformation on the quantized transform coefficient, and can generate a reconstructed picture.

[0097] Figure 4 The multi-transformation technology according to an embodiment of the present disclosure is schematically illustrated.

[0098] refer to Figure 4 , the converter can correspond to the above Figure 2 The converter in the encoding device of , and the inverse converter may correspond to the above Figure 2 The inverse transformer in the encoding device, or corresponding to Figure 3 An inverse transformer in a decoding device.

[0099] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S410). This primary transform may be referred to as a core transform. In this document, the primary transform may be based on a multi-transform selection (MTS), and when a multi-transform is applied as the primary transform, it may be referred to as a multi-core transform.

[0100] Multi-core transform may refer to a method of transforming using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1. In other words, multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. Herein, the primary transform coefficient may be referred to as a time transform coefficient from the perspective of the transformer.

[0101] That is, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. However, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. Here, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. These DCT / DST types may be defined based on basis functions.

[0102] If a multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for the target block may be selected from the transform kernels, a vertical transform may be performed on the target block based on the vertical transform kernel, and a horizontal transform may be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform may refer to a transform for the horizontal component of the target block, and the vertical transform may refer to a transform for the vertical component of the target block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on the prediction mode and / or transform index of the target block (CU or sub-block) including the residual block.

[0103] Furthermore, according to one example, if the primary transform is performed by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transform or the horizontal transform. For example, when the horizontal transform kernel is expressed as trTypeHor and the vertical transform kernel is expressed as trTypeVer, a trTypeHor or trTypeVer value of 0 can be set to DCT2, a trTypeHor or trTypeVer value of 1 can be set to DCT-7, and a trTypeHor or trTypeVer value of 2 can be set to DCT-8.

[0104] In this case, MTS index information may be encoded and signaled to a decoding device to indicate any one of a plurality of transform kernel sets. For example, MTS index 0 may indicate that both trTypeHor and trTypeVer values ​​are 0, MTS index 1 may indicate that both trTypeHor and trTypeVer values ​​are 1, MTS index 2 may indicate that both trTypeHor and trTypeVer values ​​are 2 and trTypeVer values ​​are 1, MTS index 3 may indicate that both trTypeHor and trTypeVer values ​​are 1 and 2, and MTS index 4 may indicate that both trTypeHor and trTypeVer values ​​are 2.

[0105] In one example, the transformation kernel set according to the MTS index information is illustrated in the following table.

[0106] [Table 1]

[0107] tu_mts_idx[x0][y0] 0 1 2 3 4 trTvpeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2

[0108] The transformer can derive modified (secondary) transform coefficients by performing a secondary transform based on the (primary) transform coefficients (S420). The primary transform is a transform from the spatial domain to the frequency domain, while the secondary transform refers to a transform performed with a more compressed expression by using the correlation between the (primary) transform coefficients. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a pattern-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent the generation of modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix. At this time, the vertical transform and the horizontal transform may not be applied separately to the (primary) transform coefficients (or the horizontal transform and the vertical transform may not be applied independently), but the transform may be applied at one time based on the non-separable transform matrix. In other words, the non-separable secondary transform can refer to a transform method in which the vertical and horizontal directions of the (primary) transform coefficients are not applied separately, and for example, a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal in a certain direction (e.g., a row-first direction or a column-first direction), and then the modified transform coefficient (or secondary transform coefficient) is generated based on the non-separable transform matrix. For example, according to the row-first order, M×N blocks are arranged in a row in the order of the first row, the second row, ... and the Nth row. According to the column-first order, M×N blocks are arranged in a row in the order of the first column, the second column, ... and the Nth column. The non-separable secondary transform can be applied to the upper left area of ​​the block configured with the (primary) transform coefficients (hereinafter, may be referred to as the transform coefficient block). For example, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 8, an 8×8 non-separable secondary transform can be applied to the upper left 8×8 area of ​​the transform coefficient block. In addition, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 4, and the width (W) or height (H) of the transform coefficient block is less than 8, a 4×4 non-separable sub-transform may be applied to the upper left min(8,W)×min(8,H) area of ​​the transform coefficient block. However, this embodiment is not limited thereto, and for example, even if the condition that only the width (W) or height (H) of the transform coefficient block is equal to or greater than 4 is satisfied, a 4×4 non-separable sub-transform may be applied to the upper left min(8,W)×min(8,H) area of ​​the transform coefficient block.

[0109] Specifically, for example, if a 4×4 input block is used, the non-separable sub-transform may be performed as follows.

[0110] A 4×4 input block X can be represented as follows.

[0111] [Formula 1]

[0112]

[0113] If X is represented as a vector, then the vector It can be expressed as follows.

[0114] [Formula 2]

[0115]

[0116] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.

[0117] In this case, the non-separable secondary transform can be calculated as follows.

[0118] [Formula 3]

[0119]

[0120] In this formula, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.

[0121] By using the above formula 3, the 16×1 transform coefficient vector can be derived And the vector can be scanned in order (horizontally, vertically, diagonally, etc.) Reorganized into 4×4 blocks. However, the above calculations are examples, and hypercube-Givens transform (HyGT) or the like may also be used for the calculation of the inseparable secondary transform in order to reduce the computational complexity of the inseparable secondary transform.

[0122] Furthermore, in a non-separable secondary transform, the transform kernel (or transform core, transform type) may be selected to be mode-dependent. In this case, the mode may include an intra prediction mode and / or an inter prediction mode.

[0123] As described above, an inseparable secondary transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 area can be the upper left 8×8 area in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 area can be the upper left 4×4 area in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0124] Here, in order to select mode-dependent transform kernels, two inseparable sub-transform kernels per transform set can be configured for inseparable sub-transforms for both the 8×8 transform and the 4×4 transform, and four transform sets can exist. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.

[0125] However, as the size of the transform (ie, the size of the region to which the transform is applied) may be other than 8×8 or 4×4, for example, the number of sets may be n, and the number of transform kernels in each set may be k.

[0126] The transform set may be referred to as an NSST set or a LFNST set. A specific set among the transform sets may be selected, for example, based on the intra prediction mode of the current block (CU or subblock). A low-frequency non-separable transform (LFNST) may be an example of a reduced non-separable transform, which will be described later and represents a non-separable transform for low-frequency components.

[0127] For reference, for example, the intra prediction mode may include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction mode may include a plane intra prediction mode No. 0 and a DC intra prediction mode No. 1, and the directional intra prediction mode may include 65 intra prediction modes No. 2 to No. 66. However, this is an example, and this document may be applied even if the number of intra prediction modes is different. In addition, in some cases, intra prediction mode No. 67 may also be used, and intra prediction mode No. 67 may represent a linear model (LM) mode.

[0128] Figure 5 The intra directional mode for 65 prediction directions is schematically shown.

[0129] refer to Figure 5 , based on the intra prediction mode 34 having the upper left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode having a horizontal directionality and an intra prediction mode having a vertical directionality. Figure 5In

[15] , H and V represent horizontal and vertical directivities, respectively, and the numbers -32 to 32 indicate displacements of 1 / 32 units on the sample grid position. These numbers can represent offsets for the mode index values. Intra-frame prediction modes 2 to 33 have horizontal directivities, and intra-frame prediction modes 34 to 66 have vertical directivities. Strictly speaking, intra-frame prediction mode 34 can be considered neither horizontal nor vertical, but can be classified as belonging to horizontal directivity when determining the transform set of the secondary transform. This is because the input data is transposed for a vertically oriented mode that is symmetrical based on intra-frame prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-frame prediction mode 34. Transposing the input data means switching the rows and columns of two-dimensional M×N block data to N×M data. Intra-frame prediction mode 18 and intra-frame prediction mode 50 can represent horizontal intra-frame prediction mode and vertical intra-frame prediction mode, respectively, and intra-frame prediction mode 2 can be called upper right diagonal intra-frame prediction mode because intra-frame prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra-prediction mode 34 may be referred to as a bottom-right diagonal intra-prediction mode, and intra-prediction mode 66 may be referred to as a bottom-left diagonal intra-prediction mode.

[0130] According to an example, four transform sets according to intra prediction modes may be mapped, for example, as shown in the following table.

[0131] [Table 2]

[0132] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=80 1

[0133] As shown in Table 2, any one of four transform sets, ie, lfnstTrSetIdx, may be mapped to any one of four indexes (ie, 0 to 3) according to the intra prediction mode.

[0134] When it is determined that a specific set is used for an inseparable transform, one of the k transform cores in the specific set can be selected by an inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can send the inseparable secondary transform index to a decoding device with a signal. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an lfnst index value 0 can refer to a first inseparable secondary transform core, an lfnst index value 1 can refer to a second inseparable secondary transform core, and an lfnst index value 2 can refer to a third inseparable secondary transform core. Alternatively, an lfnst index value 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst index values ​​1 to 3 can indicate three transform cores.

[0135] The transformer can perform a non-separable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer, and can be encoded and sent to a decoding device with a signal, and transmitted to a dequantizer / inverse transformer in the encoding device.

[0136] In addition, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and sent to the decoding device with a signal, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0137] The inverse transformer may perform a series of processes in the reverse order of the order already performed in the above-mentioned transformer. The inverse transformer may receive the (dequantized) transform coefficients and derive the (primary) transform coefficients by performing a secondary (inverse) transform (S450), and may obtain the residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients may be referred to as modified transform coefficients. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.

[0138] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner may determine whether to apply a secondary inverse transform. For example, the secondary inverse transform may be NSST, RST, or LFNST, and the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a transform coefficient of a residual block.

[0139] The secondary inverse transform determiner may determine the secondary inverse transform. In this case, the secondary inverse transform determiner may determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-frame prediction mode. In an embodiment, the secondary transform determination method may be determined depending on the primary transform determination method. Various combinations of the primary transform and the secondary transform may be determined according to the intra-frame prediction mode. In addition, in an example, the secondary inverse transform determiner may determine the area to which the secondary inverse transform is applied based on the size of the current block.

[0140] In addition, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) inverse transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.

[0141] Furthermore, in the present disclosure, reduced secondary transform (RST) in which the size of a transformation matrix (kernel) is reduced may be applied in the concept of NSST in order to reduce the amount of computation and storage required for an inseparable secondary transform.

[0142] In addition, the transformation kernel, transformation matrix, and coefficients constituting the transformation kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented in 8 bits. This can be a condition for implementation in decoding devices and encoding devices, and compared with existing 9 bits or 10 bits, the amount of storage required to store the transformation kernel can be reduced, and performance degradation can be reasonably adapted. In addition, representing the kernel matrix in 8 bits can allow the use of small multipliers and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.

[0143] In this specification, the term "RST" may refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. When performing a downscale transform, the amount of computation required for the transform can be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to address the computational complexity issues that arise when transforming large blocks or non-separable transforms.

[0144] RST may be referred to by various terms such as reduced transform, reduced sub-transform, downscaled transform, simplified transform, and simple transform, and the names that RST may be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in a low-frequency region including non-zero coefficients in a transform block, it may be referred to as a low-frequency non-separable transform (LFNST). The transform index may be referred to as an LFNST index.

[0145] Meanwhile, when performing a secondary inverse transform based on an RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include an inverse downscaled secondary transformer for deriving modified transform coefficients based on an inverse RST of the transform coefficients, and an inverse primary transformer for deriving residual samples for the target block based on an inverse primary transform of the modified transform coefficients. The inverse primary transform refers to an inverse transform of a primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.

[0146] Figure 6is a diagram illustrating an RST according to an embodiment of the present disclosure.

[0147] In this disclosure, a “target block” may refer to a current block to be encoded, a residual block, or a transform block.

[0148] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduced transformation matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a shrinkage factor, a simplification factor, a simple factor, or various other terms. In addition, R can be referred to as a reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to an N / R value.

[0149] In an example, the reduction factor or the reduction coefficient may be signaled through a bitstream, but the example is not limited thereto. For example, a predefined value for the reduction factor or the reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or the reduction coefficient may not be signaled separately.

[0150] The size of the reduced transform matrix according to an example may be R×N, which is smaller than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.

[0151] [Formula 4]

[0152]

[0153] Figure 6 The matrix T in the reduced transform block shown in (a) may refer to the matrix T of Equation 4 R×N .like Figure 6 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual samples for the target block, the transform coefficients for the current block may be derived.

[0154] In an example, if the size of the block to which the transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (a) can be expressed as a matrix operation shown in the following Equation 5. In this case, the storage and multiplication calculations can be reduced to about 1 / 4 by a reduction factor.

[0155] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix provided on the left side of the column vector.

[0156] [Formula 5]

[0157]

[0158] In formula 5, r1 to r 64 ℓ may represent a residual sample for the target block, and specifically may be a transform coefficient generated by applying the primary transform. As a result of the calculation of Equation 5, a transform coefficient ci of the target block may be derived, and the process of deriving ci may be as shown in Equation 6.

[0159] [Formula 6]

[0160]

[0161] As a result of the calculation of Equation 6, the transform coefficients c1 to c R That is, when R=16, the transform coefficients c1 to c 16 If a normal transform is applied instead of RST, and a transform matrix of 64×64 (N×N) size is multiplied by a residual sample of 64×1 (N×1) size, although 64 (N) transform coefficients are derived for the target block, only 16 (R) transform coefficients are derived for the target block because RST is applied. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.

[0162] When considering the size of the transformation matrix, the size of the conventional transformation matrix is ​​64×64 (N×N), but the size of the reduced transformation matrix is ​​reduced to 16×64 (R×N). Therefore, compared with the case of performing the conventional transformation, the memory usage when performing RST can be reduced by the R / N ratio. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional transformation matrix, the number of multiplication calculations (R×N) can be reduced by the R / N ratio when using the reduced transformation matrix.

[0163] In an example, the transformer 232 of the encoding device 200 may derive transform coefficients for the target block by performing a primary transform and an RST-based secondary transform on the residual samples of the target block. These transform coefficients may be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) for the transform coefficients, and may derive residual samples for the target block based on an inverse primary transform for the modified transform coefficients.

[0164] According to the example, the inverse RST matrix T N×R The size of is N×R which is smaller than the size of the conventional inverse transform matrix N×N, and is the same as the reduced transform matrix T shown in Equation 4.R×N Has a transposition relationship.

[0165] Figure 6 (b) The matrix T in the reduced inverse transform block shown t It can refer to the inverse RST matrix T N×R T (The superscript T refers to transposition). Figure 6 As shown in (b), when the inverse RST matrix T RxN T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. RxN T It can be expressed as (T RxN T ) NxR .

[0166] More specifically, when the inverse RST is used as the secondary inverse transform, when the inverse RST matrix T RxN T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as the inverse primary transform, and in this case, when the inverse RST matrix T RxN T When multiplied by the transform coefficients of the target block, residual samples of the target block can be derived.

[0167] In an example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (ie, R / N=16 / 64=1 / 4), then according to Figure 6 The RST of (b) can be expressed as a matrix operation shown in the following equation 7.

[0168] [Formula 7]

[0169]

[0170] In formula 7, c1 to c 16 As a result of the calculation of Equation 7, r representing the modified transform coefficient of the target block or the residual sample of the target block can be derived. i , and derive r i The process can be shown as formula 8.

[0171] [Formula 8]

[0172]

[0173] As a result of the calculation of Equation 8, r1 to r2 representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N. From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is ​​64×64 (N×N), but the size of the inverse reduced transform matrix is ​​reduced to 64×16 (R×N), so the storage usage in the case of performing inverse RST can be reduced by the R / N ratio compared to the case of performing the conventional inverse transform. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, the use of the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.

[0174] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since one transform set includes two or three transforms (kernels) according to the intra prediction mode, it can be configured to select one of up to four transforms, including the case where the secondary transform is not applied. In the transform where the secondary transform is not applied, the application of the identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is ​​applied, that is, the case where the secondary transform is not applied), the transform index or LFNST index can be signaled as a syntax element for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, the 8×8 NSST in the RST configuration can be specified by the transform index, or the 8×8 LFNST can be specified when LFNST is applied. 8×8lfnst and 8×8RST refer to transforms that can be applied to an 8×8 region included in a transform coefficient block when both W and H of a target block to be transformed are equal to or greater than 8, and the 8×8 region may be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4lfnst and 4×4RST refer to transforms that can be applied to a 4×4 region included in a transform coefficient block when both W and H of a target block are equal to or greater than 4, and the 4×4 region may be the upper left 4×4 region in the transform coefficient block.

[0175] According to an embodiment of the present disclosure, for the transformation in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 area. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m×48 transform kernel matrix (m≤16) to an 8×8 area, 48 pieces of data are input and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming an 8×8 area can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on the 48 pieces of data constituting the area other than the lower right 4×4 area in the 8×8 area. Here, when the matrix operation is performed by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.

[0176] For the inverse transform in the decoding process, a transposed matrix of the aforementioned transform kernel matrix may be used. That is, when inverse RST or LFNST is performed in the inverse transform process performed by the decoding device, input coefficient data to which inverse RST is applied is arranged in a one-dimensional vector according to a predetermined arrangement order, and a modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged into a two-dimensional block according to the predetermined arrangement order.

[0177] In summary, during the transform process, when RST or LFNST is applied to an 8×8 region, the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region, are subjected to a matrix operation with the 16×48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input as a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived and arranged in the upper left region of the 8×8 region.

[0178] On the contrary, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 area, 16 transform coefficients corresponding to the upper left area of ​​the 8×8 area among the transform coefficients in the 8×8 area can be input in a one-dimensional array according to the scanning order, and can undergo a matrix operation with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix)*(16×1 transform coefficient vector)=(48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and can therefore be expressed as an n×1 column vector. In addition, * represents matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left area, the upper right area, and the lower left area of ​​the 8×8 area except the lower right area.

[0179] When the secondary inverse transform is based on an RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse downscaled secondary transformer for deriving modified transform coefficients based on the inverse RST for the transform coefficients; and an inverse primary transformer for deriving residual samples for the target block based on the inverse primary transform for the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.

[0180] The non-separable transform (LFNST) described above will be described in detail as follows: The LFNST may include a forward transform performed by an encoding device and an inverse transform performed by a decoding device.

[0181] The encoding device receives as input a result (or a portion of the result) derived after applying a primary (core) transform, and applies a forward secondary transform (Secondary transform).

[0182] [Formula 9]

[0183] y=G T x

[0184] In Equation 9, x and y are the input and output of the secondary transform, respectively, G is a matrix representing the secondary transform, and the transform basis vectors consist of column vectors. In the case of inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of the matrix G becomes G T dimension.

[0185] For inverse LFNST, the dimensions of the matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transformation basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.

[0186] On the other hand, for forward LFNST, the matrix G T The dimensions are [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper parts of the [16×48] matrix and the [16×16] matrix, respectively.

[0187] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or a [8×1] vector can be used as output y. In video coding and decoding, the output of the forward primary transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data as the output of the forward transform.

[0188] Figure 7 is a diagram illustrating an order of arranging output data of a forward primary transform into a one-dimensional vector according to an example. Figure 7 The left figures of (a) and (b) show the order for constructing a [48×1] vector, and Figure 7 The right figures of (a) and (b) show the order for constructing a [16×1] vector. In the case of LFNST, the 2D data can be constructed by Figure 7 Arrange them in the same order as in (a) and (b) to obtain a one-dimensional vector x.

[0189] The arrangement direction of the output data of the forward primary transform may be determined according to the intra-frame prediction mode of the current block. For example, when the intra-frame prediction mode of the current block is in the horizontal direction relative to the diagonal direction, the output data of the forward primary transform may be arranged in the following manner: Figure 7 The output data of the forward primary transform are arranged in the order of (a) and when the intra prediction mode of the current block is in a vertical direction relative to the diagonal direction, the output data of the forward primary transform can be arranged in the order of (a) Figure 7 (b) The output data of the forward primary transform are arranged in sequence.

[0190] According to the example, different Figure 7 The order of (a) and (b) is determined, and in order to derive and apply Figure 7If the order of (a) and (b) is the same as the result (y vector), the column vectors of matrix G can be rearranged according to the order of arrangement. In other words, the column vectors of G can be rearranged so that each element that makes up the x vector is always multiplied by the same transformation basis vector.

[0191] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data in a process using the result of the forward secondary transform as input (for example, in the process of performing quantization or residual coding), the output y vector of Equation 9 needs to be properly arranged into 2D data again.

[0192] Figure 8 is a diagram illustrating an order in which output data of a forward sub-transform is arranged into two-dimensional blocks according to an example.

[0193] In the case of LFNST, the output values ​​can be arranged in 2D blocks according to a predetermined scanning order. Figure 8 (a) shows that when the output y is a [16×1] vector, the output values ​​are arranged in 16 positions of a 2D block according to a diagonal scanning order. Figure 8 (b) shows that when the output y is an [8×1] vector, the output values ​​are arranged in 8 positions of the 2D block according to the diagonal scanning order, and the remaining 8 positions are filled with zeros. Figure 8 The X in (b) indicates that it is filled with zeros.

[0194] According to another example, since the order of processing the output vector y when performing quantization or residual coding can be preset, the output vector y may not be arranged in the order shown in FIG. Figure 8 However, in the case of residual coding, data coding can be performed in 2D block (e.g., 4×4) units (e.g., CG (coefficient group)), and in this case, according to Figure 8 The data is arranged in a specific order in the diagonal scan order of .

[0195] Meanwhile, the decoding apparatus may configure a one-dimensional input vector y by arranging two-dimensional data output through the dequantization process according to a preset scanning order for inverse transformation. The input vector y may be output as an output vector x through the following equation.

[0196] [Equation 10]

[0197] x=Gy

[0198] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y as a [16×1] vector or a [8×1] vector by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.

[0199] The output vector x is based on Figure 7 The order shown in is arranged in a two-dimensional block and is arranged as two-dimensional data, and the two-dimensional data becomes input data (or a part of input data) of the inverse primary transform.

[0200] Thus, the inverse secondary transform is overall the reverse of the forward secondary transform process and in the case of an inverse transform, unlike in the forward direction, the inverse secondary transform is applied first and then the inverse primary transform.

[0201] In inverse LFNST, one of 8 [48×16] matrices and 8 [16×16] matrices can be selected as the transformation matrix G. Whether the [48×16] matrix or the [16×16] matrix is ​​applied depends on the size and shape of the block.

[0202] In addition, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can be composed of two matrices. Which transform set to use among the four transform sets is determined according to the intra prediction mode, and more specifically, the transform set is determined based on the value of the intra prediction mode extended by considering wide-angle intra prediction (WAIP). Which matrix is ​​selected from the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be sent as index values, 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra prediction mode value.

[0203] Figure 9 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of this document.

[0204] The general intra prediction mode value may have values ​​from 0 to 66 and from 81 to 83, and the intra prediction mode value extended due to WAIP may have values ​​from -14 to 83 as shown. The values ​​from 81 to 83 indicate the CCLM (Cross Component Linear Model) mode, and the values ​​from -14 to -1 and the values ​​from 67 to 80 indicate the intra prediction mode extended due to WAIP application.

[0205] When predicting a current block whose width is greater than its height, the upper reference pixel is typically closer to a position inside the block to be predicted. Therefore, prediction in the lower left direction can be more accurate than in the upper right direction. Conversely, when the block's height is greater than its width, the left reference pixel is typically closer to a position inside the block to be predicted. Therefore, prediction in the upper right direction can be more accurate than in the lower left direction. Therefore, applying remapping (i.e., mode index modification) to the index of the Wide Intra prediction mode can be advantageous.

[0206] When wide-angle intra prediction is applied, information about existing intra predictions may be signaled, and after the information is parsed, the information may be remapped to the index of the wide-angle intra prediction mode. Therefore, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) may not be changed, that is, the total number of intra prediction modes is 67, and the intra prediction mode coding for the specific block may not be changed.

[0207] Table 3 below shows a process of deriving a modified intra mode by remapping the intra prediction mode to the wide-angle intra prediction mode.

[0208] [Table 3]

[0209]

[0210] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not divided into sub-partitions by the intra sub-partitioning (ISP) technology currently adopted in the VVC standard, and the cIdx variable values ​​​​of 0, 1, and 2 indicate the cases of the luma component, Cb component, and Cr component, respectively. The log2 function shown in Table 3 returns a log value with a base of 2, and the Abs function returns an absolute value.

[0211] The variable predModeIntra indicating the intra prediction mode and the height and width of the transform block are used as input values ​​for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or coding block can be the height and width of the current block used for remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to width can be set to Abs(Log2(nW / nH)).

[0212] For non-square blocks, the intra prediction mode can be divided into two cases and modified.

[0213] First, if all of conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before modification is equal to or greater than 2, and (3) the intra prediction mode is less than a value derived from (8+2*whRatio) when the variable whRatio is greater than 1 and is less than 8 when the variable whRatio is less than or equal to 1 [predModeIntra is less than (whRatio>1)? (8+2*whRatio):8], the intra prediction mode is set to a value 65 greater than the intra prediction mode [predModeIntra is set equal to (predModeIntra+65)].

[0214] If different from the above, that is, if the following conditions (1) to (3) are satisfied, (1) the height of the current block is greater than the width, (2) the intra prediction mode before modification is less than or equal to 66, and (3) the intra prediction mode is greater than a value derived from (60-2*whRatio) when the variable whRatio is greater than 1 and is greater than 60 when the variable whRatio is less than or equal to 1 [predModeIntra greater than (whRatio>1)? (60-2*whRatio):60], then the intra prediction mode is set to a value that is 67 less than the intra prediction mode [predModeIntra is set equal to (predModeIntra-67)].

[0215] Table 2 above shows how to select a transform set based on the intra prediction mode value extended by WAIP in LFNST. Figure 9 As shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same set of transforms is applied to modes located in mutually symmetric directions, and this symmetry is also reflected in Table 2.

[0216] At the same time, it is assumed that the forward LFNST input data of mode 54 is symmetric with the forward LFNST input data of mode 14. For example, for mode 14 and mode 54, according to Figure 7 (a) and Figure 7 The arrangement order shown in (b) rearranges the two-dimensional data into one-dimensional data. In addition, it can be seen that Figure 7 (a) and Figure 7 The pattern of the sequence shown in (b) is symmetrical about the direction indicated by the pattern 34 (diagonal direction).

[0217] Meanwhile, as described above, which transform matrix of the [48×16] matrix and the [16×16] matrix is ​​applied to the LFNST is determined by the size and shape of the transform target block.

[0218] Figure 10 is a diagram illustrating a block shape to which LFNST is applied. Figure 10 (a) shows a 4×4 block, (b) shows a 4×8 block and an 8×4 block, (c) shows a 4×N block or an N×4 block, where N is 16 or greater, (d) shows an 8×8 block, and (e) shows an M×N block, where M≥8, N≥8, and N>8 or M>8.

[0219] exist Figure 10 In , blocks with thick borders indicate the areas where LFNST is applied. Figure 10For the blocks (a) and (b), LFNST is applied to the top left 4×4 region, and for Figure 10 In block (c), LFNST is applied separately to the two upper left 4×4 regions that are arranged consecutively. Figure 10 In (a), (b), and (c), since LFNST is applied in units of 4×4 regions, this LFNST will be referred to as “4×4 LFNST” hereinafter. A [16×16] or [16×8] matrix may be applied based on the matrix dimension of G used in Equations 9 and 10.

[0220] More specifically, the [16×8] matrix is ​​applied to Figure 10 (a) 4×4 block (4×4TU or 4×4CU), and the [16×16] matrix is ​​applied to Figure 10 This is to adjust the worst-case computational complexity to 8 multiplications per sample.

[0221] about Figure 10 In (d) and (e), LFNST is applied to the upper left 8×8 region, and this LFNST is hereinafter referred to as "8×8LFNST". As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values ​​of the upper left 8×8 region are used as input values ​​of the forward LFNST. That is, as can be seen from Figure 7 (a) the left order or Figure 7 As can be seen from the left sequence of (b), a [48×1] vector can be constructed based on samples belonging to the remaining three 4×4 blocks while leaving the lower right 4×4 block as is.

[0222] The [48×8] matrix can be applied to Figure 10 The 8×8 block (8×8TU or 8×8CU) in (d) and the [48×16] matrix can be applied to Figure 10 The 8×8 block in (e) is used to adjust the worst-case computational complexity to 8 multiplications per sample.

[0223] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data are generated (Y vector in Equation 9, [8×1] or [16×1] vector). In the forward LFNST, since the matrix G T The characteristic is that the amount of output data is equal to or less than the amount of input data.

[0224] Figure 11is a diagram illustrating arrangement of output data of a forward LFNST according to an example, and shows blocks in which the output data of the forward LFNST is arranged according to block shapes.

[0225] exist Figure 11 The shaded area on the upper left of the block shown corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate samples filled with 0 values, and the remaining areas represent areas that are not changed by the forward LFNST. In the areas that are not changed by the LFNST, the output data of the forward primary transform remains unchanged.

[0226] As described above, since the dimension of the applied transformation matrix varies according to the shape of the block, the amount of output data also varies. Figure 11 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. Figure 11 In the cases of (a) and (d), the [16×8] matrix and the [48×8] matrix are applied to the block indicated by the bold line or the partial area inside the block, respectively, and the [8×1] vector is generated as the output of the forward LFNST. That is, according to Figure 8 The scanning order shown in (b) can only fill 8 output data, such as Figure 11 As shown in (a) and (d), the remaining 8 positions can be filled with 0. Figure 10 (d) The case of LFNST application blocks, such as Figure 11 As shown in (d), the two 4×4 blocks on the upper right and lower left adjacent to the upper left 4×4 block are also filled with zero values.

[0227] As described above, basically, by signaling the LFNST index, whether to apply LFNST and the transformation matrix to be applied are specified. Figure 11 As shown, when LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, an area filled with zero values ​​occurs as follows.

[0228] 1) If Figure 11 As shown in (a), samples of the upper left 4×4 block are located at positions subsequent to the eighth position in the scanning order, that is, samples from the ninth to the sixteenth position.

[0229] 2) If Figure 11 As shown in (d) and (e), when the [48×16] matrix or the [8×8] matrix is ​​applied, two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.

[0230] Therefore, if non-zero data exists by checking areas 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.

[0231] Meanwhile, for the adopted LFNST, the following simplified method can be applied.

[0232] (i) According to an example, the number of output data of the forward LFNST may be limited to 16 at most.

[0233] exist Figure 10 In the case of (c), 4×4 LFNST can be applied to two 4×4 regions adjacent to the upper left, respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), 4×4 LFNST is applied only to one 4×4 region in the upper left, and LFNST can be applied only to Figure 10 By doing this, the implementation of image compilation can be simplified.

[0234] Figure 12 FIG. 4 shows an example in which the number of output data for the forward LFNST is limited to a maximum of 16. Figure 12 , when LFNST is applied to the upper-leftmost 4x4 region in a 4xN or Nx4 block where N is 16 or greater, the output data of the forward LFNST becomes 16.

[0235] (ii) According to an example, clearing can be additionally applied to areas to which LFNST is not applied. In this document, clearing can mean filling all positions belonging to a specific area with a value of 0. That is, clearing can be applied to areas that are not changed by LFNST and maintain the results of the forward primary transform. As described above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, clearing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.

[0236] (ii)-(A) When 4×4 LFNST is applied, a region to which 4×4 LFNST is not applied may be cleared. Figure 13 is a diagram illustrating zeroing in a block to which 4×4 LFNST is applied according to an example.

[0237] like Figure 13 As shown, regarding the block to which 4×4 LFNST is applied, that is, for Figure 11 For all blocks in (a), (b), and (c), the entire region to which LFNST is not applied can be filled with zeros.

[0238] on the other hand, Figure 13 (d) shows that when Figure 12When the maximum value of the number of output data of the forward LFNST shown in is limited to 16, clearing is performed on the remaining blocks to which the 4×4 LFNST is not applied.

[0239] (ii)-(B) When 8×8 LFNST is applied, an area to which 8×8 LFNST is not applied may be cleared. Figure 14 is a diagram illustrating clearing in a block to which 8×8 LFNST is applied according to an example.

[0240] like Figure 14 As shown, for the block where 8×8 LFNST is applied, i.e., for Figure 11 For all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.

[0241] (iii) Due to the zeroing presented in (ii) above, the area filled with zeros may not be the same when LFNST is applied. Figure 11 In the case of LFNST, the wider region performs the zeroing proposed in (ii) to check whether there is non-zero data.

[0242] For example, when (ii)-(B) is applied, when checking Figure 11 After the areas filled with zero values ​​in (d) and (e) have non-zero data, additionally check Figure 14 Whether there is non-zero data in the area filled with 0, the signaling of the LFNST index can be performed only when there is no non-zero data.

[0243] Of course, even if the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. Figure 11 After checking whether there is non-zero data in the block filled with zeros in , LFNST index signaling can be applied. In this case, the encoding device only performs zero clearing and the decoding device does not assume zero clearing, that is, it only checks whether non-zero data exists only in Figure 11 In the regions explicitly marked as 0 in , LFNST index parsing can be performed.

[0244] Various embodiments of combinations of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) for applying LFNST can be derived. Of course, the combination of the above simplified methods is not limited to the following embodiments, and any combination can be applied to LFNST.

[0245] Example

[0246] -Limit the number of output data of the forward LFNST to a maximum of 16 → (i)

[0247] - When 4×4 LFNST is applied, all regions where 4×4 LFNST is not applied are cleared → (ii)-(A)

[0248] - When 8×8 LFNST is applied, all regions where 8×8 LFNST is not applied are cleared → (ii)-(B)

[0249] - After checking whether non-zero data also exists in the existing areas filled with zero values ​​and the areas filled with zeros due to the additional clearing ((ii)-(A), (ii)-(B)), the LFNST index is signaled only when no non-zero data exists → (iii).

[0250] In the case of an embodiment, when LFNST is applied, the area where non-zero output data can exist is limited to the interior of the upper left 4×4 area. In more detail, in Figure 13 (a) and Figure 14 In the case of (a), the eighth position in the scanning order is the last position where non-zero data can exist. Figure 13 (b) and (c) and Figure 14 In the case of (b), the sixteenth position in the scanning order (ie, the position of the lower right edge of the upper left 4×4 block) is the last position where data other than 0 may exist.

[0251] Therefore, when LFNST is applied, after checking whether non-zero data exists at a position where the residual coding process is not allowed (at a position beyond the last position), it may be determined whether to signal the LFNST index.

[0252] In the case of the zeroing method proposed in (ii), since the amount of data ultimately generated when both the primary transform and LFNST are applied is reduced, the amount of computation required to perform the entire transform process can be reduced. That is, when LFNST is applied, since zeroing is applied to the forward primary transform output data in areas where LFNST is not applied, there is no need to generate data for areas that were zeroed during the forward primary transform. Therefore, the amount of computation required to generate the corresponding data can be reduced. Additional effects of the zeroing method proposed in (ii) are summarized below.

[0253] First, as mentioned above, the amount of computation required to perform the entire transformation process is reduced.

[0254] In particular, when (ii)-(B) is applied, the worst-case computational effort is reduced, making it possible to lightweight the transform process. In other words, generally speaking, a large amount of computation is required to perform a large-scale primary transform. By applying (ii)-(B), the amount of data derived as a result of performing forward LFNST can be reduced to 16 or less. In addition, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transform operations further increases.

[0255] Second, the amount of computation required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.

[0256] Third, the delay involved in the transformation process is reduced.

[0257] Secondary transforms such as LFNST add computational complexity to the existing primary transform, thus increasing the overall latency involved in performing the transform. In particular, in the case of intra prediction, since reconstructed data from neighboring blocks is used in the prediction process, the increased latency due to the secondary transform during encoding results in an increased latency until reconstruction. This can lead to an increase in the overall latency of intra prediction encoding.

[0258] However, if the zeroing suggested in (ii) is applied, the delay time for performing the primary transform can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transform, so that the encoding device can be implemented more simply.

[0259] Meanwhile, in conventional intra prediction, the coded target block is regarded as one coding unit, and coding is performed without partitioning. However, ISP (Intra Sub Partitioning) coding refers to performing intra prediction coding by partitioning the coded target block in a horizontal direction or a vertical direction. In this case, a reconstructed block can be generated by performing encoding / decoding in units of partitioned blocks, and the reconstructed block can be used as a reference block for the next partitioned block. According to an example, in ISP coding, one coding block can be partitioned into two or four sub-blocks and coded, and in ISP, intra prediction is performed on one sub-block by referring to the reconstructed pixel value of a sub-block located adjacent to its left or upper side. Hereinafter, the term "coding" may be used as a concept including both coding performed by an encoding device and decoding performed by a decoding device.

[0260] Hereinafter, an embodiment in which LFNST is applied only to the luma component in a single tree is described.

[0261] The following shows a coding unit syntax table related to signaling of LFNST index and MTS index according to an example.

[0262] [Table 4]

[0263]

[0264] The meanings of the main variables in the above table are as follows.

[0265] 1.cbWidth, cbHeight: width and height of the current compilation block

[0266] 2. log2TbWidth, log2TbHeight: The base 2 logarithms of the width and height of the current transform block. By applying zeroing, the size of the current transform block can be reduced to the upper left area where non-zero coefficients may exist.

[0267] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. A flag value equal to 0 indicates that LFNST is not enabled, and a flag value equal to 1 indicates that LFNST is enabled. This flag is defined in a sequence parameter set (SPS).

[0268] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the position (x0, y0). chType can have values ​​0 and 1, where 0 represents the luma component and 1 represents the chroma component. The position (x0, y0) indicates the position on the picture, and with the value of CuPredMode[chType][x0][y0], MODE_INTRA (intra-frame prediction) and MODE_INTER (inter-frame prediction) are possible.

[0269] 5. IntraSubPartitionsSplitType: indicates which ISP is applied to the current coding unit, and ISP_NO_SPLIT indicates that the coding unit is not split into partition blocks.

[0270] 6. intra_mip_flag[x0][y0]: Position (x0, y0) is as described above in 4. intra_mip_flag is a flag indicating whether matrix-based intra prediction (MIP) mode is applied. A flag value equal to 0 indicates that MIP is not applicable, and a flag value equal to 1 indicates that MIP is applied.

[0271] 7. cIdx: A value of 0 indicates luma, while values ​​1 and 2 indicate chroma components Cb and Cr, respectively.

[0272] 8.treeType: Indicates single tree, dual tree, etc. (SINGLE_TREE: single tree, DUAL_TREE_LUMA: dual tree for luminance component, DUAL_TREE_CHROMA: dual tree for chrominance component)

[0273] 9. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, this element is inferred to have a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.

[0274] The description of the preceding syntax elements can be applied to the syntax elements shown in the following table.

[0275] In Table 4, transform_skip_flag[x0][y0][0]==0 is one condition for determining whether to signal an lfnst index with respect to a luma component according to whether transform is skipped.

[0276] Therefore, according to an example, the following coding unit syntax table is proposed in order to remove the dependency between the signaling of the transform skip flag for the luma component and the signaling of the LFNST index for the chroma components.

[0277] [Table 5]

[0278]

[0279] In the embodiment shown in Table 5, the signaling of the LFNST index for the luma component depends only on the transform skip flag of the luma component for both the dual-tree type and the single-tree partition mode. In the dual-tree mode, the LFNST index of the chroma component can be signaled only based on the transform skip flag of the chroma component. In the single-tree partition mode, LFNST is not applied to the chroma component to reduce the worst-case delay.

[0280] The variable LfnstTransformNotSkipFlag shown in Table 5 is set according to the tree type of the current block and the transform skip flag value for the color component, and only when the value of the variable is 1, the LFNST index may be signaled.

[0281] When the tree type is not dual-tree chroma (treeType!=DUAL_TREE_CHROMA), that is, when the tree type is single-tree or dual-tree luma, if the transform skip flag value for the luma component is 0, the variable LfnstTransformNotSkipFlag may be set to 1 ( treeType! =DUAL_TREE_CHROMA? transform_ skip_flag[x0][y0][0]==0:(transform_skip_flag[x0][y0][1]==0||transform_skip_flag[x0][y0][2]==0)), and when the tree type is dual-tree chroma, if the transform skip flag value (transform_skip_flag[x0][y0][1]) for the chroma component Cb is 0 or the transform skip flag value (transform_skip_flag[x0][y0][1]) for the chroma component Cr is 0, the variable LfnstTransformNotSkipFlag may be set to 1 ( treeType! =DUAL_TREE_CHROMA? transform_skip_flag[x0][y0][0]==0: (transform_skip_flag[x0][y0][1]==0||transform_skip_flag[x0][y0][2]==0) ).

[0282] In this disclosure, the operator "x?y:z" indicates that if x is true, then x is y, otherwise x is z (if x is true, then evaluates to the value of y, otherwise evaluates to the value of z).

[0283] Considering the specification text of the transformation process of Table 5, it is as follows.

[0284] [Table 6]

[0285]

[0286]

[0287] If there are valid coefficients at zeroed positions when LFNST is applied, the variable LfnstZeroOutSigCoeffFlag in Table 4 is 0, otherwise it is 1. The variable LfnstZeroOutSigCoeffFlag can be set according to multiple conditions shown in Table 11 below.

[0288] The variable LfnstZeroOutSigCoeffFlag indicates whether a significant coefficient exists in a second region other than the upper left first region of the current block. The variable's value is initially set to 1 and can be changed to 0 if a significant coefficient exists in the second region. The LFNST index can be parsed only while maintaining the initial value of the variable LfnstZeroOutSigCoeffFlag. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, LFNST can be applied to the luma component or all chroma components of the current block, thus unspecifying the color index of the current block.

[0289] According to an example, when all the last significant coefficients of a transform block whose corresponding coding block flag (CBF, which is 1 when there is at least one significant coefficient in the corresponding block and 0 otherwise) is 1 are located at the DC position (upper left position), the variable LfnstDcOnly in Table 4 is 1, otherwise it is 0. Specifically, the position of the last significant coefficient is checked relative to a luma transform block in the dual-tree luma, and the position of the last significant coefficient is checked relative to both the Cb transform block and the Cr transform block in the dual-tree chroma. In a single tree, the position of the last significant coefficient can be checked relative to the transform blocks of luma, Cb, and Cr.

[0290] A syntax table of a coding unit signaling an LFNST index according to another example is as follows.

[0291] [Table 7]

[0292]

[0293] In Table 7, the variable transform_skip_flag[x0][y0][cIdx] indicates whether transform skipping is applied to the coding block for the component indicated by cIdx. cIdx can have values ​​0, 1, and 2, where 0 indicates a luma component, and 1 and 2 indicate a Cb component and a Cr component, respectively. A value of transform_skip_flag[x0][y0][cIdx] equal to 1 indicates that transform skipping is applied, while 0 indicates that transform skipping is not applied.

[0294] In Table 7, the variable LfnstNotSkipFlag may be set to 1 only when transform skipping is not applied to all components (all coding blocks) forming the current coding unit, and may be set to 0 in other cases, and only when LfnstNotSkipFlag is 1, the LFNST index (lfnst_idx in Table 7) is signaled.

[0295] When the current coding unit is coded in a single tree structure (when treeType is SINGLE_TREE in Table 7), all components include Y, Cb, and Cr, when the current coding unit is coded in a separate tree structure for luma (when treeType is DUAL_TREE_LUMA in Table 7), all components include only Y, and when the current coding unit is coded in a separate tree structure for chroma (when treeType is DUAL_TREE_CHROMA in Table 7), all components include Cb and Cr.

[0296] In other words, when even one of the components consisting of the current coding unit is skipped for coding by transform, the LFNST index is not signaled, and the value of the LFNST index is inferred to be 0. That is, LFNST is not applied.

[0297] In a structure where LFNST is restricted from being applied when even one of the components (Y, Cb, and Cr) is coded by transform skip as in Table 7, when a plurality of components (Y, Cb, and Cr) are continuously coded as in a single tree and it is determined during parsing that the components are subject to transform skip (for example, when transform skip is determined for Cb and Cr), the corresponding transform coefficients may be configured not to be additionally buffered with respect to the corresponding components coded by transform skip until the value of the LFNST index is parsed.

[0298] For example, when any one component is found to be coded by transform skipping, it is determined not to apply LFNST, and thus inverse quantization, inverse transform, etc. may be immediately performed subsequently.

[0299] Instead of Table 7, the syntax table may be concisely described as in Table 8.

[0300] [Table 8]

[0301]

[0302] When determining whether to signal an LFNST index by checking only whether transform skipping is applied to the luma component in a single tree, a syntax table for a coding unit may be configured as follows.

[0303] [Table 9]

[0304]

[0305] In Table 9, LfnstNotSkipFlag is determined to be the same as in Table 7 or Table 8 in the separate tree (that is, in the luma separate tree, when transform skipping is not applied to the luma component, LfnstNotSkipFlag is set to 1, otherwise it is set to 0, and in the chroma separate tree, when transform skipping is not applied to both the Cb component and the Cr component, LfnstNotSkipFlag is set to 1, otherwise it is set to 0). In the single tree, when no transform skipping is applied to only the luma component, LfnstNotSkipFlag is set to 1, otherwise it is set to 0. Instead of Table 8, the syntax table shown in Table 9 may be applied.

[0306] [Table 10]

[0307]

[0308] An embodiment of determining an LFNST index signaling condition when LFNST is applied only to the luma component in a single tree is described below.

[0309] In the coding unit syntax tables of Tables 4, 5, 7, 8, 9, and 10, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are used as conditions for signaling LFNST indexes. Basically, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are initialized to 1, as shown in Table 7, and their values ​​can be updated to 0 in the syntax table for residual coding, as shown in Table 11. For reference, when coding a component (which can be Y, Cb, or Cr) using transform skipping, a different syntax table (transform_ts_coding) is imported instead of the residual coding in Table 11. Therefore, when parsing the LFNST index for the component, the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are not updated.

[0310] [Table 11]

[0311]

[0312]

[0313] In Table 11, lastSubBlock indicates the position of a subblock (coefficient group (CG)) where the last significant (non-zero) coefficient is located in the scanning order. 0 indicates a subblock including a DC component, while a value greater than 0 indicates a subblock not including a DC component.

[0314] lastScanPos indicates the position of the last significant coefficient within a subblock in scan order. When a subblock includes 16 positions, values ​​from 0 to 15 are possible.

[0315] LastSignificantCoeffX and LastSignificantCoeffY indicate the x- and y-coordinates of the last significant coefficient in the transform block. The x-coordinate starts at 0 and increases from left to right, and the y-coordinate starts at 0 and increases from top to bottom. A value of 0 for both variables means the last significant coefficient is at DC.

[0316] Basically, Table 11 is applied to the embodiment of Table 5 to determine the values ​​of the variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag. When Table 11 is applied and the coding block is coded in a single tree, the residual coding shown in Table 11 can be introduced for all components. For example, when all Y, Cb, and Cr components are not coded using transform skipping, residual coding can be performed on each component.

[0317] Therefore, when Table 11 is applied and the current coding block is coded by a single tree, if the last non-zero coefficient of even one component is located at a position other than the DC position (the upper left position of the transform block), the value of the variable LfnstDcOnly can be updated to 0, and when the position of the last non-zero coefficient of even one component is located in an area where the transform coefficient cannot be located when LFNST is applied (i.e., in an area other than the first to eighth positions in a 4×4 transform block or an 8×8 transform block according to the forward transform coefficient scanning order, or in an area other than the upper left 4×4 area in a transform block to which LFNST is applicable in the current VVC standard), the value of the variable LfnstZeroOutSigCoeffFlag can be updated to 0. As shown in the compilation unit syntax tables of Tables 5, 7, 8, 9, and 10, the LFNST index may be signaled only when the value of the variable LfnstZeroOutSigCoeffFlag is 1, and in modes other than the ISP mode, the LFNST index may be signaled only when the value of the variable LfnstDcOnly is 0.

[0318] However, when LFNST is applied only to the luma component in a single tree, when residual coding is performed on a component to which LFNST is not applied (chroma component Cb or Cr), updating of the variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag may not be allowed. This is because it is logically inappropriate for the arrangement or distribution of transform coefficients of the component to which LFNST is not applied to determine whether to signal the LFNST index (i.e., whether LFNST is applied).

[0319] Table 12 limits the update of the variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag to only the luma component in a single tree. In cases other than a single tree, the variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag may be updated for all components (as shown in Table 11).

[0320] [Table 12]

[0321]

[0322]

[0323] In the residual coding presented in Table 12, since treeType is added as a parameter compared to Table 11, the syntax table for the transform unit may be modified as shown in Table 13.

[0324] [Table 13]

[0325]

[0326] Based on the details of Table 5, some details may be replaced with the embodiments of Tables 7 to 10, or the details of Table 11 or Table 12 may be applied. The following possible combinations may be configured based on Tables 7 to 10, Table 11, or Table 12.

[0327] 1. Table 7 (or Table 8) + Table 11

[0328] 2. Table 7 (or Table 8) + Table 12

[0329] 3. Table 9 (or Table 10) + Table 11

[0330] 4. Table 9 (or Table 10) + Table 12

[0331] Hereinafter, a method for applying a scaling list to chroma components when LFNST is applied only to luma components in a single tree is described.

[0332] Currently, in VVC WD, a syntax element called scaling_matrix_for_lfnst_disabled_flag is defined. When scaling_matrix_for_lfnst_disabled_flag is 1, the scaling list is not applied when LFNST is applied, and when scaling_matrix_for_lfnst_disabled_flag is 0, the scaling list can be applied when LFNST is applied.

[0333] Here, the scaling list is a matrix for specifying specific weights (weighting values) for each transform coefficient position in a transform block, and dequantization or quantization is achieved by multiplying the weight for each transform coefficient, thereby enabling different dequantization or quantization to be applied depending on the importance of the transform coefficient.

[0334] In a single tree, as in the embodiment of Table 5, LFNST may be applied only to the luma component, and when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied while coding a coding block in a single tree, the scaling list is not applied to the luma component. Here, the scaling list may be applied to the chroma components to which LFNST is not applied.

[0335] Table 14 shows an example of a dequantization process (scaling process) that realizes the above situation.

[0336] [Table 14]

[0337]

[0338]

[0339]

[0340]

[0341] In Table 14, treeType indicates a tree type of a coding unit to which a currently processed transform block belongs, and SINGLE_TREE, DUAL_TREE_LUMA, and DUAL_TREE_CHROMA indicate a single tree, a separate tree for luma, and a separate tree for chroma (dual-tree chroma), respectively.

[0342] In this embodiment, since LFNST can be applied only to the luma component in a single tree, when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied (lfnst_idx[xTby][yTby] value is greater than 0), the scaling list is not applied to the luma component (when the cIdx value is 0).

[0343] However, for chroma components (when the cIdx value is greater than 0), it may be determined whether to apply the scaling list by further checking different conditions (eg, checking transform_skip_flag[xTby][yTby][cIdx]).

[0344] In separate trees, such as in the case of luma component in a single tree, when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST is applied (lfnst_idx[xTby][yTby] value is greater than 0), the scaling list is not applied to luma and chroma components.

[0345] Furthermore, in separate trees, such as in the case of chroma components in a single tree, whether to apply the scaling list may be determined by further checking different conditions (eg, checking transform_skip_flag[xTby][yTby][cIdx]).

[0346] Therefore, when the value of scaling_matrix_for_lfnst_disabled_flag is 1 and LFNST may be applied only to the luma component in a single tree, the scaling list may not be applied to the luma component and may be applied to the chroma components.

[0347] According to an example, a combination of Table 14 and the above embodiments may be applied (based on the details of Table 5, some details may be replaced with the embodiments of Tables 7 to 10 or a combination of the details of Table 11 or Table 12 may be applied).

[0348] In this case, as in the normative text of “Transform process for scaling transform coefficients” in Table 6, LFNST may be configured to be applied only to the luma component in a single tree.

[0349] The following drawings are provided to describe specific examples of the present disclosure. Since specific terms used for devices or specific terms used for signals / messages / fields shown in the drawings are provided for illustration, the technical features of the present disclosure are not limited to the specific terms used in the following drawings.

[0350] Figure 15 is a flowchart illustrating the operation of a video decoding apparatus according to an embodiment of the present disclosure.

[0351] Figure 15 Each process disclosed in the reference is based on Figures 4 to 14 Therefore, with reference to Figures 3 to 14 Descriptions of those overlapping specific details will be omitted or will be made schematically.

[0352] The decoding apparatus 300 according to an embodiment may receive flag information indicating whether a scaling list is available when performing LFNST, an LFNST index for a current block, and residual information from a bitstream ( S1510 ).

[0353] Specifically, the decoding device 300 can decode information about the quantized transform coefficients for the current block from the bitstream, and can derive the quantized transform coefficients of the target block based on the information about the quantized transform coefficients for the current block. The information about the quantized transform coefficients of the target block may be included in a sequence parameter set (SPS) or a slice header, and may include information about whether to apply RST, information about a reduction factor, information about a minimum transform size for applying RST, information about a maximum transform size for applying RST, an inverse RST size, and at least one of information about a transform index indicating any one transform kernel matrix included in the transform set.

[0354] The decoding device may further receive information about the intra prediction mode for the current block and information about whether ISP is applied to the current block. The decoding device may receive and parse flag information indicating whether ISP coding or ISP mode is applied, thereby deriving whether the current block is partitioned into a predetermined number of sub-partition transform blocks. Here, the current block may be a coding block. In addition, the decoding device may derive the size and number of the partitioned sub-partition blocks using flag information indicating the direction in which the current block is partitioned.

[0355] The LFNST index is a value for specifying an LFNST matrix when LFNST is applied as an inverse secondary inseparable transform, and may have a value ranging from 0 to 2. For example, an LFNST index value of 0 may indicate that LFNST is not applied to the current block, an LFNST index value of 1 may indicate a first LFNST matrix, and an LFNST index value of 2 may indicate a second LFNST matrix.

[0356] Information about ISP and LFNST indexes may be received in a coding unit level.

[0357] Flag information received by the decoding device indicating whether the scaling list is available when LFNST is performed can be represented by scaling_matrix_for_lfnst_disabled_flag or sps_scaling_matrix_for_lfnst_disabled_flag and can be signaled in the sequence parameter set. A value of 1 for this flag indicates that the scaling list is not applied when LFNST is applied, and a value of 0 for this flag indicates that the scaling list is applicable when LFNST is applied. The scaling list is a matrix for specifying a specific weight (weight value) for each transform coefficient position in the transform block, and can be dequantized or quantized by multiplying the weight for each transform coefficient, thereby enabling differential dequantization or quantization to be applied according to the importance of the transform coefficient.

[0358] The decoding apparatus 300 may determine whether a scaling list is applied to the current block based on whether LFNST is applied and a tree type of the current block, in order to dequantize a transform coefficient for the current block ( S1520 ).

[0359] Whether to apply the scaling list may be determined based on the flag information and the value of the LFNST index.

[0360] When the tree type of the current block is single-tree, the color components of the current block may include a luma component, a first chroma component indicating chroma Cb, and a second chroma component indicating chroma Cr. When the tree type of the current block is dual-tree luma, the current block may include a luma component. When the tree type of the current block is dual-tree chroma, the color components of the current block may include a first chroma component and a second chroma component.

[0361] Here, the current block may be a transform block, which is a transform unit, and when the tree type of the current block is single-tree, the current block may include a transform block for a luma component, a transform block for a first chroma component, and a transform block for a second chroma component. When the tree type of the current block is dual-tree luma, the current block may include a transform block for a luma component, and when the tree type of the current block is dual-tree chroma, the current block may include a transform block for a first chroma component and a transform block for a second chroma component.

[0362] According to an example, when the type of the current block is a single tree, LFNST can be applied only to the luma component, and when the current block is coded in a single tree, the value of scaling_matrix_for_lfnst_disabled_flag is 1, and LFNST is applied, the scaling list is not applied to the luma component. However, the scaling list can be applied to the chroma components to which LFNST is not applied.

[0363] In summary, when the flag information about the scaling list indicates that the scaling list is not available and the LFNST index is greater than 0 (i.e., LFNST is applied), when the tree type of the current block is a single tree and the current block is a luminance component, the scaling list may not be applied, and when the tree type of the current block is a single tree and the current block is a chrominance component, the scaling list may be applied.

[0364] According to an example, if the flag information about the scaling list indicates that the scaling list is not available and the LFNST index is greater than 0, when the tree type of the current block is dual-tree chroma, LFNST can be applied to the current block, and thus the scaling list is not applied to the chroma component.

[0365] According to an example, when flag information about the scaling list indicates that the scaling list is unavailable and the LFNST index is greater than 0, when the tree type of the current block is dual-tree luma, LFNST is applicable to the current block and thus the scaling list is not applied to the luma component.

[0366] Subsequently, the decoding apparatus derives a transform coefficient for the current block from the residual information based on the determination result ( S1530 ).

[0367] The derived transform coefficients can be arranged in 4×4 block units according to the reverse diagonal scanning order, and the transform coefficients in the 4×4 block can also be arranged according to the reverse diagonal scanning order. That is, the dequantized transform coefficients can be arranged according to the reverse scanning order applied in the video codec, such as in VVC or HEVC.

[0368] The decoding apparatus may derive a modified transform coefficient from the transform coefficient based on the LFNST matrix and the LFNST index for LFNST, that is, by applying LFNST ( S1540 ).

[0369] LFNST is a non-separable transform in which a transform is applied to coefficients without separating the coefficients in a specific direction, unlike a primary transform in which the coefficients to be transformed are separated vertically or horizontally and then transformed. This non-separable transform may be a low-frequency non-separable transform in which the forward transform is applied only to the low-frequency region rather than the entire region of the block.

[0370] The decoding device may derive various variables to apply LFNST, and may determine whether to apply LFNST based on the tree type and size of the current block.

[0371] The decoding device can derive a first variable (variable LfnstDcOnly) indicating whether there is a valid coefficient at a position outside the DC component in the current block and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether there is a transformation coefficient in a second area outside the upper left first area of ​​the current block.

[0372] The first variable and the second variable are initially set to 1, wherein the first variable may be updated to 0 when the significant coefficient exists at a position other than the position of the DC component in the current block, and the second variable may be updated to 0 when the transform coefficient exists in the second region.

[0373] When the first variable is updated to 0 and the second variable remains 1, LFNST may be applied to the current block.

[0374] For the luma component to which the intra subpartitioning (ISP) mode is applicable, the LFNST index may be parsed without deriving the variable LfnstDcOnly.

[0375] Specifically, when the ISP mode is applied and the transform skip flag for the luma component (i.e., transform_skip_flag[x0][y0][0]) is 0, the LFNST index can be signaled when the tree type of the current block is a single tree or dual tree for luma, regardless of the value of the variable LfnstDcOnly.

[0376] However, for the chroma component to which the ISP mode is not applied, the value of the variable LfnstDcOnly may be set to 0 according to the value of the transform skip flag for the chroma component Cb (transform_skip_flag[x0][y0][1]) and the value of the transform skip flag for the chroma component Cr (transform_skip_flag[x0][y0][2]). That is, when the value of cIdx in transform_skip_flag[x0][y0][cIdx] is 1, the value of the variable LfnstDcOnly may be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0, and when the value of cIdx is 2, the value of the variable LfnstDcOnly may be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. When the value of the variable LfnstDcOnly is 0, the decoding device may parse the LFNST index; otherwise, the LFNST index may be inferred to be 0 without signaling.

[0377] The second variable may be a variable LfnstZeroOutSigCoeffFlag, which may indicate that zeroing is performed when LFNST is applied. The second variable may be initially set to 1 and may be changed to 0 when a significant coefficient exists in the second region.

[0378] When the index of the subblock where the last non-zero coefficient exists is greater than 0 and the width and height of the transform block are both equal to or greater than 4, or when the position of the last non-zero coefficient in the subblock where the last non-zero coefficient exists is greater than 7 and the size of the transform block is 4×4 or 8×8, the variable LfnstZeroOutSigCoeffFlag can be derived as 0. A subblock refers to a 4×4 block used as a coding unit in residual coding and can be called a coefficient group (CG). Subblock index 0 refers to the top left 4×4 subblock.

[0379] That is, when a non-zero coefficient is derived in an area other than the upper left area (where LFNST transform coefficients may exist in the transform block), or when a non-zero coefficient exists at a position other than the eighth position in the scanning order for a 4×4 block or an 8×8 block, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0380] The decoding apparatus may determine an LFNST set including an LFNST matrix based on the intra prediction mode derived from the information about the intra prediction mode, and may select any one of a plurality of LFNST matrices based on the LFNST set and the LFNST index.

[0381] Here, the same LFNST set and the same LFNST index can be applied to the sub-partitioned transform blocks into which the current block is divided. That is, since the same intra prediction mode is applied to the sub-partitioned transform blocks, the LFNST set determined based on the intra prediction mode can also be equally applied to all sub-partitioned transform blocks. In addition, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partitioned transform blocks into which the current block is divided.

[0382] As described above, a transform set may be determined according to the intra prediction mode for the transform block to be transformed, and inverse LFNST may be performed based on the transform kernel matrix (i.e., any one of the LFNST matrices) included in the transform set indicated by the LFNST index. The matrix used for inverse LFNST may be referred to as an inverse LFNST matrix or an LFNST matrix, and may be referred to by any term as long as the matrix is ​​the transpose of the matrix used for forward LFNST.

[0383] In an example, the inverse LFNST matrix may be a non-square matrix, where the number of columns is less than the number of rows.

[0384] The decoding apparatus may derive residual samples for the current block based on the primary inverse transform for the modified transform coefficient ( S1550 ).

[0385] Here, as the primary inverse transform, a general separable transform may be used, or the aforementioned MTS may be used.

[0386] Subsequently, the decoding apparatus 300 may generate reconstructed samples based on the residual samples for the current block and the prediction samples for the current block.

[0387] The following figures are provided to describe specific examples of the present disclosure. Because the specific terms of the devices or specific terms of signals / messages / fields illustrated in the figures are provided for explanation, the technical features of the present disclosure are not limited to the specific terms used in the following figures.

[0388] Figure 16 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of the present disclosure.

[0389] exist Figure 16 Each process disclosed in the reference is based on Figures 4 to 14 Therefore, with reference to Figure 2 and Figures 4 to 14 Descriptions of those overlapping specific details will be omitted or will be made schematically.

[0390] The encoding apparatus 200 according to an embodiment may derive a prediction sample for a current block based on an intra prediction mode applied to the current block.

[0391] When the ISP is applied to the current block, the encoding apparatus may perform prediction through each sub-partitioned transform block.

[0392] The encoding device can determine whether to apply ISP coding or ISP mode to the current block, that is, the coding block, and can determine the direction in which the current block is partitioned, and can derive the size and number of partitioned sub-blocks according to the determination result.

[0393] The same intra prediction mode can be applied to the sub-partition transform blocks into which the current block is partitioned, and the encoding device can derive prediction samples for each sub-partition transform block. That is, the encoding device performs intra prediction sequentially according to the partitioning form of the sub-partition transform blocks, for example, horizontally or vertically, or from left to right or from top to bottom. For the leftmost or topmost sub-block, the reconstructed pixels of the already coded coding block are referred to in the conventional intra prediction method. Further, for each side of a subsequent internal sub-partition transform block that is not adjacent to the previous sub-partition transform block, in order to derive reference pixels adjacent to the side, reference is made to the reconstructed pixels of the already coded adjacent coding block as in the conventional intra prediction method.

[0394] The encoding apparatus 200 may derive residual samples for the current block based on the prediction samples ( S1610 ).

[0395] The encoding apparatus 200 may derive transform coefficients for the current block by applying at least one of LFNST or MTS to residual samples, and may arrange the transform coefficients according to a predetermined scanning order.

[0396] The encoding device may derive a transform coefficient for the current block based on a transform process such as a primary transform and / or a secondary transform on the residual sample, may apply LFNST when the current block is a single tree type and a luma component, and may not apply LFNST when the current block is a single tree type and a chroma component (S1620).

[0397] The primary transform may be performed by multiple transform cores, as in MTS, in which case the transform core may be selected based on the intra prediction mode.

[0398] The encoding apparatus 200 may determine whether to perform a secondary transform or an inseparable transform, specifically, LFNST, on a transform coefficient for a current block, and may derive a modified transform coefficient by applying LFNST to the transform coefficient.

[0399] LFNST is a non-separable transform in which a transform is applied to coefficients without separating the coefficients in a specific direction, unlike a primary transform that separates the coefficients to be transformed vertically or horizontally and transforms them. This non-separable transform may be a low-frequency non-separable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.

[0400] The encoding device may derive various variables to apply LFNST, and may determine whether to apply LFNST based on a tree type and size of a current block.

[0401] The encoding device can derive a first variable (variable LfnstDcOnly) indicating whether there is a valid coefficient at a position other than the position of the DC component in the current block and a second variable (variable LfnstZeroOutSigCoeffFlag) indicating whether there is a transformation coefficient in a second area other than the upper left first area of ​​the current block.

[0402] The first variable and the second variable are initially set to 1, wherein the first variable may be updated to 0 when the significant coefficient exists at a position other than the position of the DC component in the current block, and the second variable may be updated to 0 when the transform coefficient exists in the second region.

[0403] When the first variable is updated to 0 and the second variable remains 1, LFNST may be applied to the current block.

[0404] For the luma component to which the intra subpartitioning (ISP) mode is applicable, LFNST may be applied without deriving the variable LfnstDcOnly.

[0405] Specifically, when the ISP mode is applied and the transform skip flag for the luma component (i.e., transform_skip_flag[x0][y0][0]) is 0, LFNST can be applied when the tree type of the current block is a single tree or dual tree for luma, regardless of the value of the variable LfnstDcOnly.

[0406] However, for chroma components to which the ISP mode is not applied, the value of the variable LfnstDcOnly may be set to 0 according to the value of the transform skip flag for the chroma component Cb (transform_skip_flag[x0][y0][1]) and the value of the transform skip flag for the chroma component Cr (transform_skip_flag[x0][y0][2]). That is, when the value of cIdx in transform_skip_flag[x0][y0][cIdx] is 1, the value of the variable LfnstDcOnly may be set to 0 only when the value of transform_skip_flag[x0][y0][1] is 0, and when the value of cIdx is 2, the value of the variable LfnstDcOnly may be set to 0 only when the value of transform_skip_flag[x0][y0][2] is 0. When the value of the variable LfnstDcOnly is 0, the encoding device may apply LFNST, and otherwise, the encoding device may not apply LFNST.

[0407] The second variable may be a variable LfnstZeroOutSigCoeffFlag, which may indicate that zeroing is performed when LFNST is applied. The second variable may be initially set to 1 and may be changed to 0 when a significant coefficient exists in the second region.

[0408] When the index of the subblock where the last non-zero coefficient exists is greater than 0 and the width and height of the transform block are both equal to or greater than 4, or when the position of the last non-zero coefficient in the subblock where the last non-zero coefficient exists is greater than 7 and the size of the transform block is 4×4 or 8×8, the variable LfnstZeroOutSigCoeffFlag can be derived as 0. A subblock refers to a 4×4 block used as a coding unit in residual coding and can be called a coefficient group (CG). Subblock index 0 refers to the top left 4×4 subblock.

[0409] That is, when a non-zero coefficient is derived in an area other than the upper left area (where LFNST transform coefficients may exist in the transform block), or when a non-zero coefficient exists at a position other than the eighth position in the scanning order for a 4×4 block or an 8×8 block, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0410] The encoding apparatus may determine an LFNST set including an LFNST matrix based on the intra prediction mode derived from the information about the intra prediction mode, and may select any one of a plurality of LFNST matrices.

[0411] Here, the same LFNST set and the same LFNST index can be applied to the sub-partitioned transform blocks into which the current block is divided. That is, since the same intra prediction mode is applied to the sub-partitioned transform blocks, the LFNST set determined based on the intra prediction mode can also be equally applied to all sub-partitioned transform blocks. In addition, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partitioned transform blocks into which the current block is divided.

[0412] As described above, a transform set may be determined based on the intra prediction mode for the transform block to be transformed, and LFNST may be performed based on the transform kernel matrix (i.e., any one of the LFNST matrices) included in the LFNST transform set. A matrix applied to LFNST may be referred to as an LFNST matrix and may be referred to by any term as long as the matrix is ​​the transpose of the matrix used for inverse LFNST.

[0413] In an example, the LFNST matrix may be a non-square matrix having a smaller number of rows than columns.

[0414] The encoding apparatus may determine whether the scaling list is applied to the current block based on whether LFNST is performed in the transformation process and the tree type of the current block ( S1630 ).

[0415] The scaling list is a matrix used to specify a specific weight (weight value) for each transform coefficient position in the transform block, and can be dequantized or quantized by multiplying the weight used for each transform coefficient, thereby enabling differential dequantization or quantization according to the importance of the transform coefficient.

[0416] According to an example, when the tree type of the current block is a single tree and the current block is a luma component, the encoding device may not apply the scaling list, and when the tree type of the current block is a single tree and the current block is a chroma component, the encoding device may apply the scaling list.

[0417] When the tree type of the current block is single-tree, the color components of the current block may include a luma component, a first chroma component indicating chroma Cb, and a second chroma component indicating chroma Cr. When the tree type of the current block is dual-tree luma, the current block may include a luma component. When the tree type of the current block is dual-tree chroma, the color components of the current block may include a first chroma component and a second chroma component.

[0418] Here, the current block may be a transform block, which is a transform unit, and when the tree type of the current block is single-tree, the current block may include a transform block for a luma component, a transform block for a first chroma component, and a transform block for a second chroma component. When the tree type of the current block is dual-tree luma, the current block may include a transform block for a luma component, and when the tree type of the current block is dual-tree chroma, the current block may include a transform block for a first chroma component and a transform block for a second chroma component.

[0419] According to an example, when the current block is a single tree, the encoding device may apply LFNST only to the luma component, and when LFNST is applied, the encoding device does not apply the scaling list to the luma component. However, the encoding device may apply the scaling list to the chroma components to which LFNST is not applied.

[0420] In summary, when the LFNST index is greater than 0 (i.e., LFNST is applied), when the tree type of the current block is a single tree and the current block is a luminance component, the encoding device may not apply the scaling list, and when the tree type of the current block is a single tree and the current block is a chrominance component, the encoding device may apply the scaling list.

[0421] According to an example, in case the LFNST index is greater than 0, when the tree type of the current block is dual-tree chroma, LFNST may be applied to the current block, and thus the encoding apparatus does not apply the scaling list.

[0422] According to an example, in case the LFNST index is greater than 0, when the tree type of the current block is dual-tree luma, LFNST may be applied to the current block, and thus the encoding apparatus does not apply the scaling list.

[0423] The encoding apparatus may quantize the transform coefficient based on the determination, that is, whether the scaling list is applied to the current block ( S1640 ).

[0424] That is, the encoding apparatus may quantize the transform coefficients of the transform block to which LFNST is not applied using the scaling list, and may quantize the transform coefficients of the transform block to which LFNST is applied without using the scaling list.

[0425] The encoding apparatus may encode and output residual information and flag information indicating whether a scaling list is available when performing LFNST ( S1650 ).

[0426] Flag information indicating whether the scaling list is applicable when LFNST is performed can be represented by scaling_matrix_for_lfnst_disabled_flag or sps_scaling_matrix_for_lfnst_disabled_flag and can be signaled in the sequence parameter set. A value of 1 for this flag indicates that the scaling list is not applied when LFNST is applied, and a value of 0 for this flag indicates that the scaling list is applicable when LFNST is applied.

[0427] When the LFNST index is greater than 0 and the current block is a single tree, LFNST may be applied to the luma component, and thus the encoding apparatus may encode the value of the flag as 1.

[0428] However, when the LFNST index is greater than 0 and the current block is a single tree, LFNST is not applied to the chroma component, and thus the encoding apparatus may construct image information such that the scaling list may be applied.

[0429] When the LFNST index is greater than 0 and the tree type of the current block is dual-tree chroma, LFNST is applicable to the current block, and thus the encoding apparatus may encode the value of the flag as 1 so that the scaling list is not applied to the chroma component.

[0430] According to an example, when the LFNST index is greater than 0 and the tree type of the current block is dual-tree luma, LFNST is applicable to the current block, and thus the encoding apparatus may encode the value of the flag as 1 so that the scaling list is not applied to the luma component.

[0431] The encoding apparatus may derive a quantized transform coefficient by quantizing the modified transform coefficient for the current block, and may encode the LFNST index.

[0432] The encoding device may generate residual information including information about the quantized transform coefficients. The residual information may include the aforementioned transform-related information / syntax elements. The encoding device may encode the image / video information including the residual information and may output the image / video information in the form of a bitstream.

[0433] Specifically, the encoding apparatus 200 may generate information about the quantized transform coefficient and may encode the information about the quantized transform coefficient.

[0434] The syntax element of the LFNST index according to the present embodiment may indicate whether to apply (inverse) LFNST and any one of the LFNST matrices included in the LFNST set, and when the LFNST set includes two transform kernel matrices, the syntax element of the LFNST index may have three values.

[0435] According to an example, when the partitioned tree structure of the current block is a dual tree type, an LFNST index may be encoded for each of the luma block and the chroma block.

[0436] According to an embodiment, the value of the syntax element of the transform index may include: 0, which indicates that no (inverse) LFNST is applied to the current block; 1, which indicates the first LFNST matrix among the LFNST matrices; and 2, which indicates the second LFNST matrix among the LFNST matrices.

[0437] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.

[0438] In addition, in the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled through residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.

[0439] In the above embodiments, the method is described based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a step may be performed in an order or step different from the above order or step or simultaneously with another step. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps of the flowchart may be removed without affecting the scope of the present disclosure.

[0440] The above-mentioned method according to the present disclosure may be implemented in a software form, and the encoding device and / or decoding device according to the present disclosure may be included in an image processing device such as a TV, a computer, a smart phone, a set-top box, a display device, etc.

[0441] When the embodiments of the present disclosure are specifically implemented by software, the above methods can be embodied as modules (processes, functions, etc.) for performing the above functions. The modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, the embodiments described in the present disclosure can be specifically implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in each figure can be specifically implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.

[0442] In addition, the decoding device and the encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), and the like.

[0443] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include all kinds of storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media embodied in the form of carrier waves (e.g., transmission via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or sent via a wired or wireless communication network. Additionally, the embodiments of the present invention can be embodied as a computer program product through program code, and the program code can be executed on a computer through the embodiments of the present invention. The program code can be stored on a computer-readable carrier.

[0444] Figure 17 The structure of a content streaming system to which the present disclosure is applied is illustrated.

[0445] Furthermore, the content streaming system to which the present disclosure is applied may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0446] The encoding server compresses the content input from a multimedia input device such as a smart phone, camera, or camcorder into digital data to generate a bitstream, and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method of the embodiments of this document. The streaming server can also temporarily store the bitstream during the process of sending or receiving the bitstream.

[0447] The streaming server transmits multimedia data to user devices via a network server based on user requests. The network server serves as a medium for informing users of services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, which then delivers the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices in the content streaming system.

[0448] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a stable streaming service, the streaming server can store the bitstream for a predetermined time.

[0449] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each of the servers in the content streaming system may be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.

[0450] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and device claims that can be combined can be combined to be implemented or performed in a device, and the technical features of the method claims and device claims that can be combined can be implemented or performed in a method.

Claims

1. An image decoding method performed by a decoding device, comprising: deriving transform coefficients for a current block based on residual information received from a bitstream; as well as deriving residual samples for the current block based on an inverse transform for the transform coefficients, Wherein, deriving the transform coefficients includes: determining whether a scaling list is applied to the current block based on whether a low frequency non-separable transform (LFNST) is applied and a tree type of the current block; and deriving the transform coefficients for the current block from the residual information based on the determination, and wherein, based on the tree type of the current block being a single tree and the color component of the current block being a chroma component, the LFNST is not applied to the chroma component of the current block and the scaling list is applied to the chroma component of the current block, and Wherein, based on flag information indicating that the scaling list is unavailable for the block to which the LFNST is applied, the tree type of the current block is dual-tree chroma, and LFNST is applied to the chroma components of the current block, the scaling list is not applied to the chroma components of the current block.

2. The image decoding method according to claim 1, wherein: Based on the tree type of the current block being a single tree and the LFNST being applied to the luma component of the current block, the scaling list is not applied to the luma component of the current block.

3. The image decoding method according to claim 1, further comprising: The flag information indicating whether the scaling list is available for the block to which the LFNST is applied is received.

4. The image decoding method according to claim 3, wherein: Based on the flag information indicating that the scaling list is unavailable, the LFNST index is greater than 0, and the tree type of the current block is a single tree, the scaling list is not applied to the luma component of the current block.

5. The image decoding method according to claim 3, wherein: Based on the flag information indicating that the scaling list is unavailable, the LFNST index is greater than 0, and the tree type of the current block is dual-tree luma, the scaling list is not applied to the luma component of the current block.

6. An image encoding method performed by an image encoding device, comprising: deriving transform coefficients for the current block from residual samples for the current block based on a transform process; determining whether a scaling list is applied to the current block based on whether a low frequency non-separable transform (LFNST) is performed in the transform process and a tree type of the current block; as well as Based on the determination, quantizing the transform coefficients, wherein, based on the tree type of the current block being a single tree and the color component of the current block being a chrominance component, applying the scaling list to the chrominance component of the current block, and Wherein, based on the tree type of the current block being dual-tree chroma and the LFNST being performed on the current block, the scaling list is not applied to the chroma component of the current block.

7. The image encoding method according to claim 6, wherein: Based on the tree type of the current block being the single tree and the color component of the current block being the chroma component, the LFNST is not performed on the chroma component of the current block.

8. The image encoding method according to claim 6, wherein: Based on the tree type of the current block being a single tree and the LFNST being performed on the current block, the scaling list is not applied to a luma component of the current block.

9. The image encoding method according to claim 6, wherein: Based on the tree type of the current block being dual-tree luma and the LFNST being performed on the current block, the scaling list is not applied to a luma component of the current block.

10. The image encoding method according to claim 6, further comprising: Flag information indicating whether the scaling list is available for the block to which the LFNST is applied is encoded and output.

11. The image encoding method according to claim 6, wherein: The current block includes a transform block.

12. A method for transmitting data for image information, comprising: generating a bitstream for the image information, wherein the bitstream is generated based on: deriving transform coefficients for the current block from residual samples for the current block based on a transform process, determining whether a scaling list is applied to the current block based on whether a low frequency non-separable transform (LFNST) is performed in the transform process and a tree type of the current block, quantizing the transform coefficients based on the determination, and encoding residual information related to the quantized transform coefficients; as well as sending said data comprising said bitstream, wherein, based on the tree type of the current block being a single tree and the color component of the current block being a chrominance component, applying the scaling list to the chrominance component of the current block, and Wherein, based on the tree type of the current block being dual-tree chroma and the LFNST being performed on the current block, the scaling list is not applied to the chroma component of the current block.