Transform-based image coding method and device therefor

The image coding method employing LFNST addresses the need for efficient compression of high-resolution images/videos by enhancing coding efficiency and transform index coding efficiency, effectively handling various image characteristics.

JP2025083584AActive Publication Date: 2025-05-30LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025044301
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-23
Filing Date
2025-03-19
Publication Date
2025-05-30
Estimated Expiration
2040-09-21

AI Technical Summary

Technical Problem

There is a need for highly efficient image/video compression technologies to effectively compress, transmit, store, and reproduce high-resolution and high-quality images/videos with various characteristics, especially for applications like VR, AR, and game images.

Method used

The proposed solution involves an image coding method and apparatus that utilizes Low-Frequency Non-Separable Transform (LFNST) to enhance coding efficiency. This method includes deriving modified transform coefficients by parsing an LFNST index and applying it to sub-partition blocks, which improves the efficiency of transform index coding.

Benefits of technology

The implementation of LFNST in the image coding method significantly increases the overall compression efficiency of images/videos and enhances the efficiency of transform index coding, addressing the challenges of high-resolution and high-quality image/video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083584000001_ABST
    Figure 2025083584000001_ABST
Patent Text Reader

Abstract

To provide an image coding technique.SOLUTION: An image decoding method according to the present document comprises a step of deriving a corrected transform coefficient, wherein the step of deriving the corrected transform coefficient comprises the steps of: determining whether the transform coefficient exists in a second region excluding the upper left first region of the current block; parsing a LFNST index on the basis of the determined result; and deriving the corrected transform coefficient on the basis of the LFNST index and a LFNST matrix, wherein the LFNST index can be parsed on the basis of the current block being divided into a plurality of sub-partition blocks, and the absence of the transform coefficient in all of the individual second regions for the plurality of sub-partition blocks.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to image coding technology, and more particularly, to an image coding method and apparatus based on transform in an image coding system.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Accordingly, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and apparatus for enhancing the coding efficiency of images.

[0006] Another technical problem of this document is to provide a method and apparatus for enhancing the efficiency of transform index coding.

[0007] Still another technical problem of this document is to provide an image coding method and apparatus utilizing LFNST.

[0008] Still another technical problem of this document is to provide a video coding method and apparatus applying LFNST to sub - partition blocks.

Means for Solving the Problem

[0009] According to one embodiment of this document, an image decoding method performed by a decoding apparatus is provided. The method may include a step of deriving a modified transform coefficient. The step of deriving the modified transform coefficient includes a step of determining whether the transform coefficient exists in a second region excluding a first region at the upper left end of the current block, a step of parsing an LFNST index based on the determination result, and a step of deriving the modified transform coefficient based on the LFNST index and an LFNST matrix. The current block is divided into a plurality of sub - partition blocks, and the LFNST index can be parsed based on the fact that the transform coefficient does not exist in all of the individual second regions for the plurality of sub - partition blocks.

[0010] If the current block is not divided into the plurality of sub - partition blocks and the transform coefficient does not exist in the second region, the LFNST index can be parsed.

[0011] The current block is a coding block. If the width and height of an individual sub - partition block are 4 or more, the LFNST index for the current block can be applied to the plurality of sub - partition blocks.

[0012] If the divided sub - partition block is a 4×4 block or an 8×8 block, the LFNST can be applied to the conversion coefficients from the upper - left corner of the current block to the eighth in the scanning direction.

[0013] The step of deriving the modified conversion coefficient further includes the step of deriving a first variable indicating whether the conversion coefficient exists in the region excluding the DC position of the current block. The LFNST index can be parsed when the first variable indicates that the conversion coefficient exists in the region excluding the DC position.

[0014] Based on the current block being divided into a plurality of sub - partition blocks, the LFNST index can be parsed without deriving the first variable.

[0015] If the sub - partition block is not a 4×4 block or an 8×8 block, the LFNST can be applied to the conversion coefficients of the 4×4 region at the upper - left corner of the sub - partition block.

[0016] According to an embodiment of the present document, an image encoding method performed by an encoding device is provided. The method includes the steps of deriving conversion coefficients for the current block based on a primary conversion for residual samples,

[0017] deriving modified conversion coefficients for the current block based on the conversion coefficients of the first region at the upper - left corner of the current block and a predetermined LFNST matrix, zeroing out a second region of the current block where the modified conversion coefficients do not exist, configuring image information such that the LFNST index is signaled based on the current block being divided into a plurality of sub - partition blocks and the zeroing - out being performed for all of the plurality of sub - partition blocks, and outputting the image information including the residual information derived through quantization of the modified conversion coefficients and the LFNST index.

[0018] According to still another embodiment of the present document, a digital storage medium storing encoded image information generated according to an image encoding method performed by an encoding device and image data including a bitstream may be provided.

[0019] According to still another embodiment of the present document, a digital storage medium storing encoded image information that causes an image decoding method to be performed by a decoding device and image data including a bitstream may be provided.

Advantages of the Invention

[0020] According to this document, the overall compression efficiency of images / videos can be increased.

[0021] According to this document, the efficiency of transform index coding can be increased.

[0022] Still another technical problem of this document is to provide an image coding method and apparatus utilizing LFNST.

[0023] Still another technical problem of this document is to provide a video coding method and apparatus for applying LFNST to subpartition blocks.

[0024] The effects obtained through a specific example in this specification are not limited to the effects listed above. For example, there may be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this specification. Accordingly, the specific effects of this specification are not limited to those explicitly described in this specification, and may include various effects that can be understood or induced from the technical features of this specification.

Brief Description of the Drawings

[0025]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

DETAILED DESCRIPTION OF THE INVENTION

[0026] This document can be modified in various ways and can have various embodiments. However, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not intended to limit the technical idea in this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof in advance.

[0027] On the one hand, each component in the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and it does not mean that each component is implemented by separate hardware or separate software. For example, among the components, two or more components may be combined to form one component, or one component may be divided into multiple components. Embodiments in which the components are integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0028] Hereinafter, with reference to the accompanying drawings, the preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0029] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be related to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0030] In this document, various embodiments related to video / image coding are presented, and unless otherwise mentioned, the embodiments may be executed in combination with each other.

[0031] In this document, "video" can mean a set of a series of "images" over time. "Picture" generally means a unit indicating one image in a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.

[0032] "Pixel" or "pel" can mean the smallest unit that constitutes one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample generally may indicate a pixel or a pixel value, may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component. Alternatively, a sample may mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it may mean a conversion coefficient in the frequency domain.

[0033] "Unit" can indicate the basic unit of image processing. A unit can include at least one of a specific region of a picture and information regarding the region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit may be used interchangeably with terms such as "block" or "area" as the case may be. In a general case, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0034] In this document, “ / ” and “,” shall be interpreted as “and / or.” For example, “A / B” shall be interpreted as “A and / or B,” and “A, B” shall be interpreted as “A and / or B.” Further, “A / B / C” means “at least one of A, B, and / or C.” Also, “A, B, C” also means “at least one of A, B, and / or C.”

[0035] Further, in this document, “or” shall be interpreted as “and / or.” For example, “A or B” may mean 1) only “A,” 2) only “B,” or 3) both “A and B.” In other words, “or” in this document may mean “additionally or alternatively.”

[0036] In this specification, "at least one of A and B" may mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" may be interpreted in the same way as "at least one of A and B".

[0037] Also, in this specification, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0038] Also, the parentheses used in this specification may mean "for example". Specifically, when displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when displayed as "prediction (i.e., intra prediction)", "intra prediction" may be proposed as an example of "prediction".

[0039] The technical features individually described within one drawing in this specification may be realized individually or simultaneously.

[0040] FIG. 1 is a drawing schematically illustrating the configuration of a video / image encoding apparatus to which this document is applicable. Hereinafter, the video encoding apparatus can include an image encoding apparatus.

[0041] Referring to FIG. 1, the encoding apparatus 100 can be configured to include an image partitioner 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 can include an inter predictor 121 and an intra predictor 122. The residual processor 130 can include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 can further include a subtractor 131. The adder 150 can be referred to as a reconstructor or a reconstructed block generator. The above-described image partitioner 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 170 can include a DPB (decoded picture buffer) and can be configured by a digital storage medium. The hardware component can further include the memory 170 as an internal / external component.

[0042] The image segmentation unit 110 can divide an input image (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or the binary-tree structure can also be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this document can be performed. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of a deeper depth, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0043] The unit can be used interchangeably with terms such as block or area depending on the context. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can generally represent pixels or pixel values, and can represent only the pixels / pixel values of the luma component, or only the pixels / pixel values of the chroma component. A sample can be used as a term corresponding to a pixel or pel for one picture (or image).

[0044] The encoding device 100 can subtract a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 121 or the intra prediction unit 122 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 132. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 100 can be called the subtraction unit 131. The prediction unit can perform prediction on a processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 140 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.

[0045] The intra prediction unit 122 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a distance from it depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 122 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0046] The inter prediction unit 121 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called by names such as a collocated reference block and a collocated CU (col CU), and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, the inter prediction unit 121 can configure a motion information candidate list based on the peripheral block, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be executed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 121 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, different from the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.

[0047] The prediction unit 120 can generate a prediction signal based on various prediction methods described later. For example, for the prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on the information regarding the palette table and the palette index.

[0048] The prediction signal generated via the prediction unit (including the inter-prediction unit 121 and / or the intra-prediction unit 122) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 132 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from the graph when the relationship information between pixels is represented by the graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of variable size that is not square.

[0049] The quantization unit 133 quantizes the transform coefficients and transmits them to the entropy encoding unit 140. The entropy encoding unit 140 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 133 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 140 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 140 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document can be included in the video / image information. The video / image information can be encoded through the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 140 can be configured as an internal / external element of the encoding device 100 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 140.

[0050] The quantized transform coefficients output from the quantization unit 133 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 134 and the inverse transform unit 135, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 121 or the intra prediction unit 122. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 150 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.

[0051] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.

[0052] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering and transmit it to the entropy encoding unit 140 as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.

[0053] The modified restored picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 121. The encoding device can avoid a prediction mismatch between the encoding device 100 and the decoding device when inter prediction is applied through this, and can also improve the encoding efficiency.

[0054] The DPB of the memory 170 can store the modified restored picture for use as a reference picture in the inter prediction unit 121. The memory 170 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 121 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 122.

[0055] FIG. 2 is a diagram schematically explaining the configuration of a video / image decoding apparatus to which this document is applicable.

[0056] Referring to FIG. 2, the decoding apparatus 200 can be configured to include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filtering unit 250, and a memory 260. The predictor 230 can include an inter-prediction unit 231 and an intra-prediction unit 232. The residual processor 220 can include a dequantizer 221 and an inverse transformer 222. The entropy decoder 210, the residual processor 220, the predictor 230, the adder 240, and the filtering unit 250 described above can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 260 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 260 as an internal / external component.

[0057] When a bitstream including video / image information is input, the decoding device 200 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 2. For example, the decoding device 200 can derive units / blocks based on the information regarding block division obtained from the bitstream. The decoding device 200 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit, and the coding unit can be divided according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 200 can be reproduced via a reproducing device.

[0058] The decoding device 200 can receive the signal output from the encoding device in FIG. 1 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the surrounding and decoded information of the block to be decoded, or the information of symbols / bins decoded in the previous step, predicts the occurrence probability of a bin according to the determined context model, performs arithmetic decoding of the bin, and generates a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction unit (inter prediction unit 232 and intra prediction unit 231), and the residual value obtained by performing entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 220. The residual processing unit 220 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 250. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 200, or the receiving unit can be a component of the entropy decoding unit 210. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 221, inverse transform unit 222, addition unit 240, filtering unit 250, memory 260, inter prediction unit 232, and intra prediction unit 231.

[0059] In the inverse quantization unit 221, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 221 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the scan order of the coefficients performed in the encoding device. The inverse quantization unit 221 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0060] In the inverse conversion unit 222, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0061] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode.

[0062] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, for prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, for prediction of a block, the prediction unit can be based on the intra block copy (IBC) prediction mode, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.

[0063] The Intra prediction unit 231 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood (neighbor) of the current block or away from it depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The Intra prediction unit 231 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0064] The Inter prediction unit 232 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At that time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the Inter prediction unit 232 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0065] The adder 240 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 232 and / or the intra prediction unit 231). When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0066] The adder 240 may be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture, and may be output after filtering, as described later, or may be used for inter prediction of the next picture.

[0067] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0068] The filtering unit 250 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 250 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 260, specifically, to the DPB of the memory 260. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0069] The (corrected) restored picture stored in the DPB of the memory 260 can be used as a reference picture in the inter prediction unit 232. The memory 260 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 232 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 260 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 231.

[0070] In this document, the embodiments described in the filtering unit 160, inter prediction unit 121, and intra prediction unit 122 of the encoding device 100 can be applied to the filtering unit 250, inter prediction unit 232, and intra prediction unit 231 of the decoding device 200 in the same or corresponding manner, respectively.

[0071] As described above, when performing video coding, prediction is performed to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived identically in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling to the decoding device information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0072] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between an original block and a predicted block, perform a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, perform a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, and quantization parameter of the quantized conversion coefficients. The decoding device can perform an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. The encoding device can further inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.

[0073] FIG. 3 schematically shows the multiple conversion techniques according to this document.

[0074] Referring to FIG. 3, the conversion unit may correspond to the conversion unit in the encoding device of FIG. 1 described above, and the inverse conversion unit may correspond to the inverse conversion unit in the encoding device of FIG. 1 described above or the inverse conversion unit in the decoding device of FIG. 3.

[0075] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S310). Such a primary conversion can be referred to as a core transform. Here, the primary conversion may be based on Multiple Transform Selection (MTS), and when multiple conversion is applied as the primary conversion, it can be referred to as a multiple core transform.

[0076] The multiple-core transform can indicate a method of further performing the transform using the Discrete Cosine Transform (DCT) type 2, the Discrete Sine Transform (DST) type 7, the DCT type 8, and / or the DST type 1. That is, the multiple-core transform can indicate a transform method of converting a residual signal (or a residual block) in the spatial domain into a transform coefficient (or a primary transform coefficient) in the frequency domain based on a plurality of selected transform kernels among the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary transform coefficient may be called a temporary transform coefficient from the perspective of the transform unit.

[0077] In other words, when an existing transform method is applied, based on the DCT type 2, the conversion from the spatial domain to the frequency domain for the residual signal (or the residual block) is applied to generate a transform coefficient. Different from this, when the multiple-core transform is applied, based on the DCT type 2, the DST type 7, the DCT type 8, and / or the DST type 1, etc., the conversion from the spatial domain to the frequency domain for the residual signal (or the residual block) is applied to generate a transform coefficient (or a primary transform coefficient). Here, the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1, etc. may be called a transform type, a transform kernel, or a transform core. Such DCT / DST transform types can be defined based on basis functions.

[0078] When the multi-core transformation is executed, among the transformation kernels, a vertical transformation kernel and a horizontal transformation kernel for a target block can be selected, a vertical transformation for the target block can be executed based on the vertical transformation kernel, and a horizontal transformation for the target block can be executed based on the horizontal transformation kernel. Here, the horizontal transformation can indicate a transformation for a horizontal component of the target block, and the vertical transformation can indicate a transformation for a vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on a prediction mode and / or a transformation index of a target block (CU or sub-block) including a residual block.

[0079] Also, according to an example, when applying MTS to execute a primary transformation, specific basis functions are set to predetermined values, and a mapping relationship for the transformation kernel can be set by combining which basis functions are applied when it is a vertical transformation or a horizontal transformation. For example, when the horizontal transformation kernel is indicated by trTypeHor and the vertical transformation kernel is indicated by trTypeVer, the value 0 of trTypeHor or trTypeVer can be set to DCT2, the value 1 of trTypeHor or trTypeVer can be set to DST7, and the value 2 of trTypeHor or trTypeVer can be set to DCT8.

[0080] In this case, in order to indicate any one of a number of conversion kernel sets, MTS index information can be encoded and signaled to the decoding device. For example, if the MTS index is 0, it indicates that all the values of trTypeHor and trTypeVer are 0; if the MTS index is 1, it indicates that all the values of trTypeHor and trTypeVer are 1; if the MTS index is 2, it indicates that the value of trTypeHor is 2 and the value of trTypeVer is 1; if the MTS index is 3, it indicates that the value of trTypeHor is 1 and the value of trTypeVer is 2; if the MTS index is 4, it can indicate that all the values of trTypeHor and trTypeVer are 2.

[0081] By way of example, the conversion kernel sets based on the MTS index information are shown in a table as follows.

[0082] [Table 1]

[0083] The conversion unit performs a secondary conversion based on the (primary) conversion coefficient to derive a corrected (secondary) conversion coefficient (S320). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion means converting to a more compressed representation by utilizing the correlation existing between the (primary) conversion coefficients. The secondary conversion includes a non-separable transform. In this case, the secondary conversion may be referred to as a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The non-separable secondary transform indicates a conversion that performs a secondary conversion on the (primary) conversion coefficient derived by the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, the conversion can be applied at once without separating the vertical conversion and the horizontal conversion (or independently applying the horizontal-vertical conversion) to the (primary) conversion coefficient based on the non-separable transform matrix. In other words, the non-separable secondary conversion is not separately applied to the vertical direction and the horizontal direction with respect to the (primary) conversion coefficient. For example, after rearranging a two-dimensional signal (conversion coefficient) into a one-dimensional signal in a specific determined direction (e.g., row-first direction or column-first direction), it indicates a conversion method for generating a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix. For example, the row-first order is to arrange the first row, the second row,..., the Nth row of the M×N block in a column in this order, and the column-first order is to arrange the first column, the second column,..., the Mth column of the M×N block in a column in this order. The non-separable secondary conversion can be applied to the top-left region of a block composed of (primary) conversion coefficients (hereinafter referred to as a conversion coefficient block). For example, when both the width (W) and the height (H) of the conversion coefficient block are 8 or more, an 8×8 non-separable secondary conversion can be applied to the top-left 8×8 region of the conversion coefficient block.Also, when both the width (W) and height (H) of the conversion coefficient block are 4 or more, and either the width (W) or height (H) of the conversion coefficient block is less than 8, a 4×4 non-separable second-order conversion can be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block. However, the embodiment is not limited thereto. For example, even if only the condition that both the width (W) and height (H) of the conversion coefficient block are 4 or more is satisfied, the 4×4 non-separable second-order conversion can also be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block.

[0084] Specifically, for example, when a 4×4 input block is used, the non-separable second-order conversion can be executed as follows.

[0085] The 4×4 input block X can be shown as follows.

[0086]

Number

[0087] When showing X in the form of a vector, the vector JPEG2025083584000004.jpg1289 can be shown as follows.

[0088]

Number

[0089] As in Equation 2, the vector JPEG2025083584000006.jpg1086 rearranges the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.

[0090] In this case, the second-order non-separable conversion can be calculated as follows.

[0091]

Number

[0092] Here, JPEG2025083584000008.jpg11100 represents a transform coefficient vector, and T represents a 16×16 (non-separable) transform matrix.

[0093] Through the above formula 3, a 16×1 transform coefficient vector JPEG2025083584000009.jpg1096 can be derived, and the JPEG2025083584000010.jpg1191 can be re-organized in 4×4 blocks through a scan order (horizontal, vertical, diagonal, etc.). However, the above calculations are illustrative, and for reducing the computational complexity of non-separable quadratic transforms, HyGT (Hypercube-Givens Transform), etc. can also be used for the calculation of non-separable quadratic transforms.

[0094] On the other hand, for the non-separable quadratic transform, a transform kernel (or transform core, transform type) can be selected in a mode-dependent manner. Here, the mode can include an intra prediction mode and / or an inter prediction mode.

[0095] As described above, the non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included inside the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 area can be the upper left 8×8 area inside the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included inside the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 area can be the upper left 4×4 area inside the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0096] At that time, for the selection of the mode-based transform kernel, two non-separable second-order transform kernels per transform set for the non-separable second-order transform can be configured for both the 8×8 transform and the 4×4 transform, and the number of transform sets can be four. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and in this case, each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.

[0097] However, the size of the transform, that is, the size of the area to which the transform is applied, can be a size other than 8×8 or 4×4 as an example, the number of the sets can be n, and the number of transform kernels in each set can be k.

[0098] The transformation set can be referred to as an NSST set or an LFNST set. The selection of a specific set among the transformation sets can be performed, for example, based on the intra prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) can be an example of a reduced non-separable transform described later, and represents a non-separable transform for low-frequency components.

[0099] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is an example, and this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can represent the LM (linear model) mode.

[0100] FIG. 4 exemplarily shows the intra-directional modes of 65 prediction directions.

[0101] Referring to FIG. 4, intra prediction modes having horizontal directionality and intra prediction modes having vertical directionality can be distinguished centering around the 34th intra prediction mode having a predicted direction of the lower right diagonal. H and V in FIG. 4 respectively represent horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in units of 1 / 32 on the sample grid position. This can indicate an offset with respect to the mode index value. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. On the other hand, strictly speaking, the 34th intra prediction mode can be regarded as neither horizontal directionality nor vertical directionality, but from the perspective of determining the conversion set of the secondary conversion, it can be classified as belonging to horizontal directionality. This is because for the vertical mode symmetric with respect to the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the alignment method of the input data for the horizontal mode is used. Transposing the input data means that for the data MxN of the two-dimensional block, the rows become columns and the columns become rows, constituting data of NxM. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. Since the 2nd intra prediction mode predicts in the upper right direction with the left reference pixel, it can be called the upper right diagonal intra prediction mode. In the same context, the 34th intra prediction mode can be called the lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called the lower left diagonal intra prediction mode.

[0102] By way of example, the mapping of four conversion sets by the intra prediction mode can be shown, for example, as in the following table.

[0103]

Table 2

[0104] As shown in Table 2, it can be mapped to any one of the four conversion sets according to the intra prediction mode, that is, lfnstTrSetIdx can be mapped to any one of 0 to 3, that is, any of the four.

[0105] On the other hand, when it is determined that a specific set is used for the non-separable conversion, one of the k conversion kernels in the specific set can be selected through the non-separable second-order conversion index. The encoding device can derive a non-separable second-order conversion index indicating a specific conversion kernel based on the rate-distortion (RD) check, and can signal the non-separable second-order conversion index to the decoding device. The decoding device can select one of the k conversion kernels in the specific set based on the non-separable second-order conversion index. For example, the index value 0 of lfnst can indicate the first non-separable second-order conversion kernel, the index value 1 of lfnst can indicate the second non-separable second-order conversion kernel, and the index value 2 of lfnst can indicate the third non-separable second-order conversion kernel. Alternatively, the index value 0 of lfnst can indicate that the first non-separable second-order conversion is not applied to the target block, and the index values 1 to 3 of lfnst can indicate the three conversion kernels.

[0106] The conversion unit can perform the non-separable second-order conversion based on the selected conversion kernel and obtain the modified (second-order) conversion coefficient. The modified conversion coefficient can be derived from the conversion coefficient quantized through the quantization unit as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization / inverse conversion unit in the encoding device.

[0107] On the one hand, when the secondary conversion is omitted as described above, the (primary) conversion coefficients that are the output of the primary (separation) conversion can be derived from the conversion coefficients quantized through the quantization unit as described above, encoded, and transmitted to the decoding device for signaling and to the inverse quantization / inverse conversion unit within the encoding device.

[0108] The inverse conversion unit can execute a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficients, executes a secondary (inverse) conversion to derive the (primary) conversion coefficients (S350), executes a primary (inverse) conversion on the (primary) conversion coefficients, and can obtain a residual block (residual samples) (S360). Here, the primary conversion coefficients can be called modified conversion coefficients from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0109] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). The secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion can be NSST, RST, or LFNST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed from the bitstream. As another example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficients of the residual block.

[0110] The second inverse transform determination unit can determine the second inverse transform. At this time, the second inverse transform determination unit can determine the second inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra prediction mode. Also, as an example, the second transform determination method can be determined depending on the first transform determination method. Various combinations of the first transform and the second transform can be determined by the intra prediction mode. Also, as an example, the second inverse transform determination unit can also determine the area to which the second inverse transform is applied based on the size of the current block.

[0111] On the other hand, as described above, when the second (inverse) transform is omitted, the (inverse quantized) transform coefficients can be received and the first (separated) inverse transform can be executed to obtain a residual block (residual sample). As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0112] On the other hand, in this document, in order to reduce the amount of calculation and memory requirements associated with the non-separable second transform, an RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced with the concept of NSST can be applied.

[0113] On the other hand, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in this document, that is, the kernel coefficients or matrix coefficients, can be represented by 8 bits. This can be one condition for implementation in the decoding device and the encoding device, and can reduce the memory requirement for storing the transform kernel while accompanying a reasonably acceptable performance degradation compared to existing 9 bits or 10 bits. Also, by representing the kernel matrix with 8 bits, a small multiplier can be used, which may be suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.

[0114] In this specification, RST can mean a transform performed on residual samples for a target block based on a transform matrix whose size is reduced by a simplification factor. When performing a simplified transform, the reduction in the size of the transform matrix can reduce the amount of computation required during the transform. That is, RST can be used to solve the issue of the complexity of operations that occur during the transform of large blocks or non-separable transforms.

[0115] RST can be referred to by various terms such as reduced transform, reduction transform, simplified transform, simple transform, etc. The names by which RST can be referred are not limited to the exemplified ones. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it may also be referred to as LFNST (Low-Frequency Non-Separable Transform). The transform index can be named the LFNST index.

[0116] On the other hand, when the inverse secondary transform is performed based on RST, the inverse transform unit 135 of the encoding device 100 and the inverse transform unit 222 of the decoding device 200 can include an inverse RST unit that derives a modified transform coefficient based on the inverse RST for the transform coefficient, and an inverse primary transform unit that derives the residual sample for the target block based on the inverse primary transform for the modified transform coefficient. The inverse primary transform means the inverse transform of the primary transform applied to the residual. In this document, deriving a transform coefficient based on a transform can mean deriving the transform coefficient by applying the transform.

[0117] FIG. 5 is a diagram for explaining RST according to an embodiment of this document.

[0118] In this specification, "target block" can mean the current block or residual block or transform block on which coding is performed.

[0119] In an RST according to one embodiment, an N-dimensional vector is mapped to an R-dimensional vector located in a different space, and a reduced transform matrix can be determined, where R is smaller than N. N can mean the square of the length of one side of the block to which the transform is applied, or the total number of transform coefficients corresponding to the block to which the transform is applied, and the simplification factor can mean the R / N value. The simplification factor can be referred to by various terms such as reduction factor, reduced factor, reduced coefficient, reduction coefficient, simplified factor, simple factor, etc. On the other hand, R can be referred to as a reduced coefficient, but in some cases, the simplification factor may mean R. Also, in some cases, the simplification factor may mean the N / R value.

[0120] In one embodiment, the simplification factor or reduced coefficient can be signaled via a bitstream, but the embodiments are not limited thereto. For example, there may be predefined values for the simplification factor or reduced coefficient stored in each encoding device 100 and decoding device 200, and in this case, the simplification factor or reduced coefficient may not be signaled separately.

[0121] The size of the simplified transform matrix according to one embodiment is RxN, which is smaller than the size NxN of the normal transform matrix, and can be defined as in Equation 4 below.

[0122]

Equation

[0123] The matrix T in the Reduced Transform block shown in Fig. 5(a) can represent the matrix T of Equation 4 RxN As shown in Fig. 5(a), when the simplified transformation matrix T RxN is multiplied by the residual samples for the target block, the transformation coefficients for the target block can be derived.

[0124] In one embodiment, when the size of the block to which the transformation is applied is 8x8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to Fig. 5(a) can be expressed by the matrix operation of Equation 5 below. In this case, the memory and multiplication operations can be reduced by a simplification factor to approximately 1 / 4.

[0125] In this document, matrix operation can be understood as an operation of placing a matrix on the left side of a column vector and multiplying the matrix by the column vector to obtain a column vector.

[0126]

Number

[0127] In Equation 5, r 1 to r 64 can represent the residual samples for the target block, and more specifically, can be the transformation coefficients generated by applying a first-order transformation. The operation result of Equation 5, the transformation coefficients c i for the target block can be derived, and the derivation process of c i is as shown in Equation 6.

[0128]

Number

[0129] The operation result of Equation 6, the transformation coefficients c 1 to c R for the target block can be derived. That is, when R = 16, the transformation coefficients c 1 to c 16can be derived. If, instead of RST, a normal (regular) conversion is applied and a conversion matrix of size 64x64 (NxN) is multiplied by a residual sample of size 64x1 (Nx1), 64 (N) conversion coefficients for the target block may be derived. However, since RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 100 to the decoding device 200 decreases, the transmission efficiency between the encoding device 100 and the decoding device 200 can increase.

[0130] From the perspective of the size of the conversion matrix, the size of the normal conversion matrix is 64x64 (NxN), but the size of the simplified conversion matrix decreases to 16x64 (RxN). Therefore, when compared with the case of performing a normal conversion, the use of memory when performing RST can be reduced at a ratio of R / N. Also, when compared with the number of multiplication operations NxN when using a normal conversion matrix, when using a simplified conversion matrix, the number of multiplication operations can be reduced (RxN) at a ratio of R / N.

[0131] In one embodiment, the conversion unit 132 of the encoding device 100 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual sample for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 200. The inverse conversion unit 222 of the decoding device 200 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and based on an inverse primary conversion for the modified conversion coefficients, can derive the residual sample for the target block.

[0132] Inverse RST matrix T according to one embodiment NxR has a size of NxR, which is smaller than the size NxN of the normal inverse conversion matrix, and the simplified conversion matrix T shown in Equation 4 RxNis related to transpose.

[0133] The matrix T within the Reduced Inv. Transform block shown in Fig. 5(b) t is the inverse RST matrix T RxN T can be meant by (the superscript T means transpose). As shown in Fig. 5(b), when the inverse RST matrix T RxN T is multiplied by the transform coefficients for the target block, the corrected transform coefficients for the target block or the residual samples for the target block can be derived. The inverse RST matrix T RxN T can also be expressed as (T RxN ) T NxR in some cases.

[0134] More specifically, when inverse RST is applied to the second - order inverse transform, when the inverse RST matrix T RxN T is multiplied by the transform coefficients for the target block, the corrected transform coefficients for the target block can be derived. On the other hand, inverse RST can be applied to the inverse first - order transform. In this case, when the inverse RST matrix T RxN T is multiplied by the transform coefficients for the target block, the residual samples for the target block can be derived.

[0135] In one embodiment, when the size of the block to which the inverse transform is applied is 8x8 and R = 16 (i.e., when R / N = 16 / 64 = 1 / 4), the RST according to Fig. 5(b) can be expressed by the matrix operation as shown in Equation 7 below.

[0136]

Equation

[0137] In Equation 7, c 1 to c 16can indicate the conversion coefficient for the target block. r indicating the operation result of Equation 7, the corrected conversion coefficient for the target block, or the residual sample for the target block j can be derived, and r j is derived as shown in Equation 8.

[0138]

Equation

[0139] r indicating the operation result of Equation 8, the corrected conversion coefficient for the target block, or the residual sample for the target block 1 to r N can be derived. Considering from the perspective of the size of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64x64 (NxN), but the size of the simplified inverse transformation matrix is reduced to 64x16 (NxR). Therefore, when comparing with the time of performing the normal inverse transformation, the memory usage when performing the inverse RST can be reduced at a ratio of R / N. Also, comparing with the number of multiplication operations NxN when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced at a ratio of R / N (NxR).

[0140] On the other hand, for an 8x8 RST as well, the configuration of the conversion set as shown in Table 2 can be applied. That is, the 8x8 RST can be applied by the conversion set in Table 2. Since one conversion set is composed of two or three conversions (kernels) depending on the prediction mode within the screen, it can be configured to select one out of a maximum of four conversions, including the case where no second-order conversion is applied. The conversion when no second-order conversion is applied can be regarded as the one to which the identity matrix is applied. When indices 0, 1, 2, and 3 are assigned to the four conversions respectively (for example, the 0th index can be assigned to the identity matrix, that is, the case where no second-order conversion is applied), a syntax element called the conversion index or the index of lfnst can be signaled for each block of the conversion coefficients to specify the applied conversion. That is, through the conversion index, for the 8x8 upper-left block, the 8x8 RST can be specified in the RST configuration, or when LFNST is applied, the 8x8 lfnst can be specified. The 8x8 lfnst and the 8x8 RST refer to the conversions that can be applied to the 8x8 region included inside the block of the conversion coefficients when all of the W and H of the target block to be converted are equal to or greater than 8, and the 8x8 region can be the upper-left 8x8 region inside the block of the conversion coefficients. Similarly, the 4x4 lfnst and the 4x4 RST refer to the conversions that can be applied to the 4x4 region included inside the block of the conversion coefficients when all of the W and H of the target block are equal to or greater than 4, and the 4x4 region can be the upper-left 4x4 region inside the block of the conversion coefficients.

[0141] On the one hand, in one embodiment of this document, in the conversion of the encoding process, for the 64 data constituting the 8x8 region, instead of a 16x64 conversion kernel matrix, only 48 data are selected, and a maximum 16x48 conversion kernel matrix can be applied. Here, "maximum" means that for an mx48 conversion kernel matrix that can generate m coefficients, the maximum value of m is 16. That is, when performing RST by applying an mx48 conversion kernel matrix (m ≤ 16) to an 8x8 region, an input of 48 data can be received and m coefficients can be generated. When m is 16, an input of 48 data is received and 16 coefficients are generated. That is, assuming that 48 data form a 48x1 vector, multiplying the 16x48 matrix and the 48x1 vector in order can generate a 16x1 vector. At that time, the 48 data forming the 8x8 region can be appropriately arranged to form a 48x1 vector. At that time, when performing matrix operations by applying a maximum 16x48 conversion kernel matrix, 16 modified conversion coefficients are generated, and the 16 modified conversion coefficients can be arranged in the upper left 4x4 region according to the scanning order, and the upper right 4x4 region and the lower left 4x4 region can be filled with 0.

[0142] For the inverse conversion in the decoding process, the transposed matrix of the above-described conversion kernel matrix can be used. That is, when inverse RST or LFNST is performed in the inverse conversion process executed by the decoding device, the input coefficient data to which the inverse RST is applied is composed of a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the matrix of the inverse RST on the left side of the one-dimensional vector can be arranged in a two-dimensional block according to a predetermined arrangement order.

[0143] Upon arrangement, in the conversion process, when RST or LFNST is applied to an 8x8 region, among the conversion coefficients of the 8x8 region, matrix operations are performed between the 48 conversion coefficients in the upper left, upper right, and lower left regions excluding the lower right region of the 8x8 region and a 16x48 conversion kernel matrix. For the matrix operations, the 48 conversion coefficients are input into a one-dimensional array. When such matrix operations are performed, 16 modified conversion coefficients are derived, and the modified conversion coefficients can be arranged in the upper left region of the 8x8 region.

[0144] Conversely, in the inverse conversion process, when inverse RST or LFNST is applied to an 8x8 region, among the conversion coefficients of the 8x8 region, the 16 conversion coefficients corresponding to the upper left side of the 8x8 region can be input in the form of a one-dimensional array according to the scanning order and matrix-operated with a 48x16 conversion kernel matrix. That is, the matrix operations in such a case can be represented as (48x16 matrix) * (16x1 conversion coefficient vector) = (48x1 modified conversion coefficient vector). Here, since the nx1 vector can be interpreted in the same sense as an nx1 matrix, it may also be denoted as an nx1 column vector. Also, * means the matrix multiplication operation. When such matrix operations are performed, 48 modified conversion coefficients can be derived, and the 48 modified conversion coefficients can be arranged in the upper left, upper right, and lower left regions excluding the lower right region of the 8x8 region.

[0145] On the other hand, when the secondary inverse conversion is performed based on RST, the inverse conversion unit 135 of the encoding device 100 and the inverse conversion unit 222 of the decoding device 200 can include an inverse RST unit that derives modified conversion coefficients based on the inverse RST for the conversion coefficients, and an inverse primary conversion unit that derives residual samples for the target block based on the inverse primary conversion for the modified conversion coefficients. The inverse primary conversion means the inverse conversion of the primary conversion applied to the residue. In this document, deriving conversion coefficients based on a conversion can mean deriving conversion coefficients by applying the conversion.

[0146] Looking specifically at the aforementioned non-separable transform, LFNST, it is as follows. LFNST can include a forward transform by an encoding device and an inverse transform by a decoding device.

[0147] The encoding device applies a forward secondary transform using, as input, the result (or a part of the result) derived after applying a forward primary transform.

[0148]

Number

[0149] In the above formula (9), x and y are the input and output of the secondary transform respectively, G is a matrix representing the secondary transform, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of the forward LFNST, taking the transpose of the matrix G gives the dimension of G T .

[0150] In the case of the inverse LFNST, the dimensions of the matrix G are [48x16], [48x8], [16x16], [16x8], and the [48x8] matrix and the [16x8] matrix are submatrices obtained by sampling 8 transform basis vectors from the left side of the [48x16] matrix and the [16x16] matrix respectively.

[0151] On the other hand, in the case of the forward LFNST, the dimension of the matrix G T is [16x48], [8x48], [16x16], [8x16], and the [8x48] matrix and the [8x16] matrix are submatrices obtained by sampling 8 transform basis vectors from above the [16x48] matrix and the [16x16] matrix respectively.

[0152] Therefore, in the case of the forward LFNST, the input x can be a [48x1] vector or a [16x1] vector, and the output y can be a [16x1] vector or an [8x1] vector. Since the output of the forward first-order transform in video coding and decoding is two-dimensional (2D) data, in order to form a [48x1] vector or a [16x1] vector as the input x, the 2D data that is the output of the forward transform must be appropriately arranged to form a one-dimensional vector.

[0153] FIG. 6 shows, by way of example, a diagram illustrating the order of arranging the output data of the forward first-order transform into a one-dimensional vector. The left diagrams of FIGS. 6(a) and 6(b) show the order for creating a [48x1] vector, and the right diagrams of FIGS. 6(a) and 6(b) show the order for creating a [16x1] vector. In the case of LFNST, the 2D data is sequentially arranged in the order as shown in FIGS. 6(a) and 6(b) to obtain a one-dimensional vector x.

[0154] The array direction of the output data of such a forward first-order transform can be determined by the intra prediction mode of the current block. For example, if the intra prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of FIG. 6(a), and if the intra prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of FIG. 6(b).

[0155] By way of example, an array order different from the array orders of FIGS. 6(a) and 6(b) can be applied. In order to derive the same result (y vector) as when applying the array orders of FIGS. 6(a) and 6(b), the column vectors of the matrix G can be rearranged according to the said array order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transform basis vector.

[0156] Since the output y derived through Equation 9 is a one-dimensional vector, if a configuration that processes the result of the forward quadratic transformation as input, for example, a configuration that performs quantization or residual coding, requires two-dimensional data as input data, the output y vector of Equation 9 must be appropriately arranged into 2D data again.

[0157] Figure 7 is a diagram showing the order of arranging the output data of the forward quadratic transformation in a two-dimensional block by way of an example.

[0158] In the case of LFNST, it can be arranged in a 2D block according to a determined scan order. (a) of Figure 7 shows that when the output y is a [16x1] vector, the output values are arranged in the 16 positions of the two-dimensional block according to the diagonal scan order. (b) of Figure 7 shows that when the output y is an [8x1] vector, the output values are arranged in the 8 positions of the two-dimensional block according to the diagonal scan order, and the remaining 8 positions are filled with 0s. X in (b) of Figure 7 indicates that it is filled with 0.

[0159] According to another example, in the configuration that performs quantization or residual coding, since the order in which the output vector y is processed can be executed according to a preset order, the output vector y may not be arranged in a 2D block as shown in Figure 7. However, in the case of residual coding, data coding can be executed in units of 2D blocks such as CG (Coefficient Group) (for example, 4x4), and in this case, the data can be arranged according to a specific order such as the diagonal scan order in Figure 7.

[0160] On the other hand, the decoding device can list the two-dimensional data output through the inverse quantization process or the like for the inverse transformation according to a preset scan order to form the one-dimensional input vector y. The input vector y can be output to the input vector x by the following equation.

[0161]

Equation

[0162] In the case of the inverse-direction LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16x1] vector or an [8x1] vector, by the G matrix. In the case of the inverse-direction LFNST, the output vector x can be a [48x1] vector or a [16x1] vector.

[0163] The output vector x is arranged in 2D blocks and arrayed into 2D data according to the order shown in FIG. 6, and such 2D data becomes the input data (or a part of the input data) for the inverse-direction primary transform.

[0164] Therefore, the inverse-direction secondary transform is overall the opposite of the forward-direction secondary transform process. In the case of the inverse transform, different from the forward direction, the inverse-direction secondary transform is applied first, and then the inverse-direction primary transform is applied.

[0165] In the inverse-direction LFNST, one of eight [48x16] matrices and eight [16x16] matrices can be selected as the transform matrix G. Whether to apply a [48x16] matrix or a [16x16] matrix is determined by the size and shape of the block.

[0166] Also, the eight matrices can be derived from four transformation sets as shown in Table 2 described above, and each transformation set can be composed of two matrices. Among the four transformation sets, which transformation set to use is determined by the intra prediction mode. More specifically, considering up to the Wide Angle Intra Prediction (WAIP), the transformation set is determined based on the extended intra prediction mode value. Among the two matrices constituting the selected transformation set, which matrix to select is derived through index signaling. More specifically, as the index value to be transmitted, 0, 1, or 2 is possible. 0 indicates not to apply LFNST, and 1 and 2 can indicate either of the two transformation matrices constituting the transformation set selected based on the intra prediction mode value.

[0167] On the other hand, as described above, among the [48x16] matrix and the [16x16] matrix, whether to apply LFNST to which transformation matrix is determined by the size and shape of the transformation target block.

[0168] FIG. 8 is a diagram showing the shape of a block to which LFNST is applied. (a) of FIG. 8 shows a 4x4 block, (b) shows 4x8 and 8x4 blocks, (c) shows a 4xN or Nx4 block where N is 16 or more, (d) shows an 8x8 block, and (e) shows an MxN block where M≧8, N≧8, and N>8 or M>8.

[0169] In FIG. 8, the block with the thick frame indicates the area to which LFNST is applied. For the blocks in FIGS. 8(a) and 8(b), LFNST is applied to the top-left 4x4 area, and for the block in FIG. 8(c), LFNST is applied to two consecutively arranged top-left 4x4 areas respectively. In FIGS. 8(a), 8(b), and 8(c), since LFNST is applied in units of 4x4 areas, such LFNST will be hereinafter referred to as "4x4 LFNST", and as the transformation matrix, a matrix with dimensions based on G in Equations 9 and 10 [16x16] or [16x8] matrix can be applied.

[0170] More specifically, for the 4x4 block (4x4 TU or 4x4 CU) in FIG. 8(a), a [16x8] matrix is applied, and for the blocks in FIGS. 8(b) and 8(c), a [16x16] matrix is applied. This is to match the computational complexity for the worst case to 8 multiplications per sample.

[0171] For FIGS. 8(d) and 8(e), LFNST is applied to the top-left 8x8 area, and such LFNST will be hereinafter referred to as "8x8 LFNST". As the transformation matrix, a [48x16] or [48x8] matrix can be applied. In the case of the forward LFNST, since a [48x1] vector (the x vector in Equation 9) is input as the input data, not all the sample values of the top-left 8x8 area are used as the input values of the forward LFNST. That is, as seen in the left-side order of FIG. 6(a) or the left-side order of FIG. 6(b), the bottom-right 4x4 block is left as it is, and based on the samples belonging to the remaining three 4x4 blocks, a [48x1] vector can be constructed.

[0172] In the 8x8 block (8x8 TU or 8x8 CU) in (d) of FIG. 8, a [48x8] matrix is applied, and a [48x16] matrix can be applied to the 8x8 block in (e) of FIG. 8. This is also to adjust the computational complexity for the worst case to 8 multiplications per sample.

[0173] According to the block shape, when the corresponding forward LFNST (4x4 LFNST or 8x8 LFNST) is applied thereto, 8 or 16 output data (y vector in Equation 9, [8x1] or [16x1] vector) are generated. In the forward LFNST, due to the characteristics of matrix G T the number of output data is equal to or less than the number of input data.

[0174] FIG. 9 is a drawing showing the arrangement of the output data of the forward LFNST by way of an example, and shows a block in which the output data of the forward LFNST are arranged along the block shape.

[0175] The shaded area processed on the upper left side of the block shown in FIG. 9 corresponds to the area where the output data of the forward LFNST are located. The positions indicated by 0 represent samples filled with 0 values, and the remaining areas represent areas not changed by the forward LFNST. In the areas not changed by the LFNST, the output data of the forward first-order transformation exist unchanged.

[0176] As described above, since the dimension of the transformation matrix applied according to the shape of the block changes, the number of output data also changes. As shown in FIG. 9, the output data of the forward LFNST may not fill all of the upper left 4x4 blocks. In the cases of FIGS. 11(a) and (d), a [16x8] matrix and a [48x8] matrix are applied to the blocks indicated by thick lines or partial regions inside the blocks, respectively, to generate an [8x1] vector in the output of the forward LFNST. That is, according to the scan order shown in FIG. 7(b), only 8 output data are filled as shown in FIGS. 9(a) and (d), and the remaining 8 positions can be filled with 0. In the case of the LFNST application block in FIG. 8(d), the two 4x4 blocks on the upper right and lower left adjacent to the upper left 4x4 block as shown in FIG. 9(d) are also filled with 0 values.

[0177] As described above, basically, the LFNST index is signaled to specify whether the LFNST is applicable and which transformation matrix is to be applied. As shown in FIG. 9, when the LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, regions filled with 0 values occur as follows.

[0178] 1) As shown in FIG. 9(a), positions after the 8th position in the scan order within the upper left 4x4 block, that is, samples from the 9th to the 16th

[0179] 2) As shown in FIGS. 9(d) and (e), two 4x4 blocks adjacent to the upper left 4x4 block or the second and third 4x4 blocks in the scan order to which a [16×48] matrix or an [8×48] matrix is applied

[0180] Therefore, when checking the regions of 1) and 2) above and finding that there are non-zero data, it is certain that the LFNST is not applied, so the signaling of the LFNST index can be omitted.

[0181] In one example, for instance, in the case of LFNST adopted in the VVC standard, since the signaling of the LFNST index is performed after residual coding, the encoding device can know the presence or absence of non-zero data (valid coefficients) for all positions inside the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can determine whether to parse the LFNST index. If there is no non-zero data in the regions specified in the above 1) and 2), signaling of the LFNST index will be performed.

[0182] Since a truncated unary code is applied to the LFNST index in a binary method, the LFNST index is composed of a maximum of two bins, and as binary codes for the possible LFNST index values of 0, 1, and 2, 0, 10, and 11 are assigned respectively. In the case of LFNST adopted in the current VVC, context-based CABAC coding (regular coding) is applied to the first bin, and bypass coding is applied to the second bin. The total number of contexts for the first bin is two. When a primary transform pair (DCT-2, DCT-2) is applied in the horizontal and vertical directions and the luma component and chroma component are coded in a dual-tree type, one context is assigned, and for the remaining cases, another context is applied. The coding of such an LFNST index is shown in the following table.

[0183]

Table 3

[0184] For one selected LFNST, the following simplification method can be applied.

[0185] (i) By way of example, the number of output data for the forward LFNST can be limited to a maximum of 16.

[0186] In the case of (c) in FIG. 8, 4x4 LFNSTs can be applied to two adjacent 4x4 regions on the upper left side, and at that time, a maximum of 32 LFNST output data can be generated. If the number of output data for the forward LFNST is limited to a maximum of 16, for a 4xN / Nx4 (N≥16) block (TU or CU), the 4x4 LFNST is only applied to one 4x4 region existing on the upper left side, and for all blocks in FIG. 8, the LFNST can be applied only once. Through this, the implementation for image coding is simplified.

[0187] FIG. 10 shows, by way of example, that the number of output data for the forward LFNST is limited to a maximum of 16. As shown in FIG. 10, when the LFNST is applied to the uppermost left 4x4 region in a 4xN or Nx4 block where N is 16 or more, the output data of the forward LFNST becomes 16.

[0188] (ii) By way of example, zero-out can be further applied to the regions where the LFNST is not applied. Zero-out in this document can be meant to fill the values of all positions belonging to a specific region with 0. That is, zero-out can be applied to the regions that maintain the result of the forward first-order transformation without being changed by the LFNST. As described above, since the LFNST is classified into 4x4 LFNST and 8x8 LFNST, zero-out can be classified into the following two types ((ii)-(A) and (ii)-(B)).

[0189] (ii)-(A) When the 4x4 LFNST is applied, the regions where the 4x4 LFNST is not applied can be zeroed out. FIG. 11 is a diagram showing zero-out in a block where the 4x4 LFNST is applied, by way of example.

[0190] As shown in FIG. 11, for the block to which 4x4 LFNST is applied, that is, the regions in FIGS. 9(a), (b), and (c) where LFNST is not applied can all be filled with 0s.

[0191] On the other hand, FIG. 11(d) shows that when the maximum number of output data of the forward LFNST is limited to 16 as shown in FIG. 12, zeroing out is performed on the remaining blocks where 4×4 LFNST is not applied.

[0192] (ii)-(B) When 8x8 LFNST is applied, the regions where 8x8 LFNST is not applied can be zeroed out. FIG. 12 is a diagram showing zeroing out in the block where 8x8 LFNST is applied by way of an example.

[0193] As shown in FIG. 12, for the block to which 8x8 LFNST is applied, that is, the regions in FIGS. 9(d) and (e) where LFNST is not applied can all be filled with 0s up to the regions where LFNST is not applied.

[0194] (iii) Due to the zeroing out presented in (ii) above, when LFNST is applied, the regions filled with 0s can change. Therefore, it is possible to check whether there is non-zero data in a wider region than in the case of the LFNST in FIG. 9 with respect to whether there is non-zero data due to the zeroing out proposed in (ii).

[0195] For example, when applying (ii)-(B), after checking whether there is non-zero data up to the regions further filled with 0s in FIG. 12 in addition to the regions filled with 0s in FIGS. 9(d) and (e), signaling for the LFNST index can be executed only when there is no non-zero data.

[0196] Of course, even if the zero-out proposed in (ii) above is applied, it is possible to check whether there is non-zero data, in the same way as the signaling of the existing LFNST index. That is, for the blocks filled with 0 in FIG. 9, it is possible to check whether there is non-zero data and apply the signaling of the LFNST index. In such a case, the zero-out is executed only in the encoding device, and the decoding device does not assume the zero-out. That is, it is possible to check only whether there is non-zero data for the regions explicitly denoted by 0 in FIG. 9 and execute the parsing of the LFNST index.

[0197] Alternatively, as another example, zero-out can also be executed as shown in FIG. 13. FIG. 13 is a diagram showing zero-out in a block to which 8x8 LFNST is applied according to another example.

[0198] As shown in FIGS. 11 and 12, zero-out can be applied to all regions other than the regions to which LFNST is applied, and it is also possible to apply zero-out only to partial regions as shown in FIG. 13. Zero-out is applied only to the regions other than the upper left 8x8 region in FIG. 13, and it may not be necessary to apply zero-out to the lower right 4x4 block inside the upper left 8x8 region.

[0199] Various embodiments are derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) for the LFNST. Of course, the combinations for the simplification methods are not limited to the following embodiments, and any combination can be applied to the LFNST.

[0200] Embodiment

[0201] - Limit the number of output data for the forward LFNST to a maximum of 16 → (i)

[0202] - When a 4x4 LFNST is applied, zero out all regions where the 4x4 LFNST is not applied → (ii)-(A)

[0203] - When an 8x8 LFNST is applied, zero out all regions where the 8x8 LFNST is not applied → (ii)-(B)

[0204] - After checking whether there is non-zero data in the regions filled with existing 0 values and the additional zeroed-out regions ((ii)-(A), (ii)-(B)), signaling of the LFNST indexing is performed only when there is no non-zero data → (iii)

[0205] In the case of the above embodiment, when the LFNST is applied, the region where non-zero output data can exist is limited to the inside of the upper left 4×4 region. More specifically, in the cases of (a) in FIG. 11 and (a) in FIG. 12, the 8th position in the scan order is the last position where non-zero data can exist, and in the cases of (b) and (d) in FIG. 11 and (b) in FIG. 12, the 16th position in the scan order (i.e., the position at the lower right of the upper left 4×4 block) is the last position where non-zero data can exist.

[0206] Therefore, after checking whether there is non-zero data at a position where the residual coding process is not allowed (a position beyond the last position) when the LFNST is applied, it is possible to determine whether signaling of the LFNST index is possible.

[0207] In the case of the zeroing-out method proposed in (ii), since the number of finally generated data decreases when both the first-order transformation and LFNST are applied, the amount of computation required for the entire transformation process can be reduced. That is, when LFNST is applied, zeroing-out is also applied to the forward first-order transformation output data existing in the area where LFNST is not applied, so it is not necessary to generate data for the area that will be zeroed out from the time of performing the forward first-order transformation. Therefore, the amount of computation required for the data generation can be saved. Summarizing the additional effects of the zeroing-out method proposed in (ii), it is as follows.

[0208] First, as described above, the amount of computation required for the execution of the entire transformation process is reduced.

[0209] In particular, when applying (ii)-(B), the amount of computation for the worst case decreases, and the transformation process can be lightened. To elaborate, generally, a large amount of computation is required for executing a first-order transformation of a large size. However, when applying (ii)-(B), the number of data derived as the execution result of the forward LFNST can be reduced to 16 or less, and the larger the size of the entire block (TU or CU), the greater the reduction effect of the transformation computation amount.

[0210] Second, the amount of computation required for the entire transformation process decreases, and the power consumption required for the transformation execution can be reduced.

[0211] Third, the latency associated with the transformation process is reduced.

[0212] Second-order transformations such as LFNST add computation to the existing first-order transformation, thus increasing the overall latency associated with the transformation execution. In particular, in the case of intra prediction, since the restored data of adjacent blocks is used in the prediction process, the increase in the latency due to the second-order transformation during encoding leads to an increase in the latency until reconstruction, which may lead to an increase in the overall latency of the intra prediction encoding.

[0213] However, when applying the zero-out presented in (ii), the delay time of the primary conversion execution can be significantly reduced when applying LFNST. Therefore, the delay time for the entire conversion execution will either be maintained or reduced, and the encoding device can be realized more easily.

[0214] On the other hand, conventional intra prediction encoded without division considering the block to be currently encoded as one encoding unit. However, ISP (Intra Sub-Paritions) coding means performing intra prediction coding by dividing the block to be currently encoded horizontally or vertically. At this time, the blocks restored by performing encoding / decoding in units of the divided blocks are generated, and the restored blocks are used as reference blocks for the next divided blocks. By way of example, one coding block may be divided into two or four sub-blocks and coded during ISP coding. In ISP, one sub-block performs intra prediction by referring to the restored pixel values of the sub-blocks located adjacent to the left or adjacent above. Hereinafter, the "coding" used includes all concepts of encoding performed in the encoding device and decoding performed in the decoding device.

[0215] Table 4 shows the number of sub-blocks divided according to the block size when applying ISP, and the sub-partitions divided by ISP may be called transform blocks (TUs).

[0216]

Table 4

[0217] The ISP is to divide the block predicted by luma intra according to the block size into two or four sub - partitionings in the vertical or horizontal direction. For example, the minimum block size to which the ISP can be applied is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into four sub - partitionings.

[0218] FIG. 14 and FIG. 15 show an example of sub - blocks into which one coding block is divided. More specifically, FIG. 14 is an illustration of the division when the coding block (width (W) × height (H)) is a 4×8 block or an 8×4 block, and FIG. 15 shows an illustration of the division when the coding block is not a 4×8 block, an 8×4 block, or a 4×4 block.

[0219] When the ISP is applied, the sub - blocks are coded sequentially, for example, horizontally (Horizontal) or vertically (Verticial), from left to right or from top to bottom according to the form of the division. After the inverse transformation and intra - prediction for one sub - block are performed up to the restoration process, the coding for the next sub - block is performed. For the left - most or top - most sub - block, the restored pixels of the coding block that has already been coded in the same way as the normal intra - prediction method are referred to. Also, when a side of a subsequent internal sub - block is not adjacent to the previous sub - block, the restored pixels of the adjacent coding block that has already been coded in the same way as the normal intra - prediction method are referred to in order to derive the reference pixels adjacent to that side.

[0220] In the ISP coding mode, all sub-blocks may be coded in the same intra prediction mode, and a flag indicating whether to use ISP coding and a flag indicating in which direction (horizontal or vertical) to divide are signaled. As shown in FIGS. 14 and 15, the number of sub-blocks can be adjusted to two or four according to the block shape. When the size (width × height) of one sub-block is less than 16, it is possible to disallow the division into the corresponding sub-block or to limit the application of ISP coding itself.

[0221] On the other hand, in the case of the ISP prediction mode, one coding unit is divided into two or four partition blocks, that is, sub-blocks, and the same in-picture prediction mode is applied to the two or four divided partition blocks for prediction.

[0222] As described above, the division directions include the horizontal direction (when an M×N coding unit with horizontal and vertical lengths of M and N respectively is divided horizontally, it is divided into two M×(N / 2) blocks when divided into two, and into four M×(N / 4) blocks when divided into four), and the vertical direction (when an M×N coding unit is divided vertically, it is divided into two (M / 2)×N blocks when divided into two, and into four (M / 4)×N blocks when divided into four), both are possible. When divided horizontally, the partition blocks are coded in the order from the upper side to the lower side, and when divided vertically, the partition blocks are coded in the order from the left side to the right side. When the currently coded partition block is a horizontal (vertical) direction division, it can be predicted by referring to the restored pixel values of the upper (left) partition block.

[0223] The conversion can be applied to the residual signal generated by the ISP prediction method in units of partition blocks. In the forward direction, not only the existing DCT-2 but also the DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technology is applied to the primary transform (core transform or primary transform), and the forward LFNST (Low Frequency Non-Separable Transform) is applied to the transform coefficients generated by the primary transform to generate the final corrected transform coefficients.

[0224] That is, LFNST can also be applied to the partition blocks divided by applying the ISP prediction mode. As described above, the same intra prediction mode is applied to the divided partition blocks. Therefore, when selecting the LFNST set derived based on the intra prediction mode, the LFNST set derived for all partition blocks can be applied. That is, since the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.

[0225] On the other hand, by way of example, LFNST can only be applied to transform blocks where both the horizontal and vertical lengths are 4 or more. Therefore, when the vertical or horizontal length of the partition block divided according to the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled either. Also, when applying LFNST to each partition block, the partition block can be regarded as one transform block. Of course, when the ISP prediction method is not applied, LFNST is applied to the coding block.

[0226] Specifically explaining the application of LFNST to each partition block is as follows.

[0227] In one example, after applying the forward LFNST to individual partition blocks, only up to 16 (8 or 16) coefficients are left in the upper left 4×4 region according to the conversion coefficient scan order, and then zero-out is applied to fill the remaining positions and regions with all 0 values.

[0228] Alternatively, in one example, when the length of one side of the partition block is 4, LFNST is applied only to the upper left 4×4 region, and when the lengths of all sides of the partition block, i.e., the width and height, are 8 or more, LFNST can be applied to the remaining 48 coefficients excluding the lower right 4×4 region inside the upper left 8×8 region.

[0229] Alternatively, in one example, in order to match the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4×4 or 8×8, only 8 conversion coefficients can be output after applying the forward LFNST. That is, when the partition block is 4×4, an 8×16 matrix is applied as the conversion matrix, and when the partition block is 8×8, an 8×48 matrix is applied as the conversion matrix.

[0230] On the other hand, in the current VVC standard, LFNST index signaling is performed on a coding unit basis. Therefore, in the ISP prediction mode, when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the partition block. That is, once the LFNST index value is transmitted at the coding unit level, the corresponding LFNST index can be applied to all partition blocks inside the coding unit. As described above, the LFNST index value has 0, 1, 2 values, where 0 indicates the case where LFNST is not applied, and 1 and 2 indicate the two conversion matrices existing in one LFNST set when LFNST is applied.

[0231] As described above, the LFNST set is determined by the intra prediction mode. In the case of the ISP prediction mode, since all partition blocks within a coding unit are predicted with the same intra prediction mode, the partition blocks can refer to the same LFNST set.

[0232] As another example, although the LFNST index signaling is still performed on a coding unit basis, in the case of the ISP prediction mode, instead of uniformly determining whether to apply LFNST to all partition blocks, it is determined whether to apply the LFNST index value signaled at the coding unit level to each partition block according to separate conditions, or not to apply LFNST. Here, the separate conditions are signaled in flag form for each partition block via the bitstream. If the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and if the flag value is 0, LFNST is not applied.

[0233] On the other hand, in a coding unit to which the ISP mode is applied, when the length of one side of a partition block is less than 4, an example of applying LFNST is described as follows.

[0234] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper left M×2 (2×M) region (where M≦N). For example, when M = 8, since the upper left region is 8×2 (2×8), the region where 16 residual signals exist becomes the input of the forward LFNST, and an R×16 (R≦16) forward transform matrix can be applied.

[0235] Here, the forward LFNST matrix can be a separate additional matrix rather than the matrix currently included in the VVC standard. Also, for worst-case complexity adjustment, an 8×16 matrix obtained by sampling only the upper eight row vectors of the 16×16 matrix is used for the transformation. The method of complexity adjustment will be described in detail later.

[0236] Second, when the size of the partition block is N×1 (1×N), LFNST can be applied to the upper left M×1 (1×M) region (where M≦N). For example, when M = 16, the upper left region is 16×1 (1×16), so the region where 16 residual signals exist becomes the input of the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.

[0237] Here, the forward LFNST matrix can be a separate additional matrix rather than the matrix currently included in the VVC standard. Also, for worst-case complexity adjustment, an 8×16 matrix obtained by sampling only the upper eight row vectors of the 16×16 matrix can be used for the transformation. The method of complexity adjustment will be described in detail later.

[0238] The first embodiment and the second embodiment may be applied simultaneously, or only one of the two embodiments may be applied. In particular, in the case of the second embodiment, by considering a one-dimensional transformation for LFNST, it has been experimentally observed that the improvement in compression performance obtained in the existing LFNST is not relatively large compared to the LFNST index signaling cost. However, in the case of the first embodiment, an improvement in compression performance similar to that obtained in the existing LFNST has been observed, that is, in the case of ISP, it has been experimentally confirmed that the application of LFNST for 2×N and N×2 contributes to the actual compression performance.

[0239] In the current LFNST in VVC, the symmetry between intra prediction modes is applied. The same LFNST set is applied to the bi-directional modes centered around Mode 34 (prediction in the 45-degree diagonal direction to the lower right), for example, the same LFNST set is applied to Mode 18 (horizontal prediction mode) and Mode 50 (vertical prediction mode). However, when applying the forward LFNST to Modes 35 to 66, after transposing the input data, the LFNST is applied.

[0240] On the other hand, VVC supports the Wide Angle Intra Prediction (WAIP) mode, and the LFNST set is derived based on the intra prediction mode modified considering the WAIP mode. For the modes extended by WAIP, the symmetry is utilized to determine the LFNST set in the same way as the general intra prediction direction modes. For example, since Mode -1 is symmetric to Mode 67, the same LFNST set is applied, and since Mode -14 is symmetric to Mode 80, the same LFNST set is applied. For Modes 67 to 80, after transposing the input data before applying the forward LFNST, the LFNST transformation is applied.

[0241] In the case of the LFNST applied to the upper left M×2 (M×1) block, the symmetry for the aforementioned LFNST cannot be applied because the block to which the LFNST is applied is non-square. Therefore, instead of applying the symmetry based on the intra prediction mode like the LFNST in Table 2, the symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.

[0242] Figure 16 is a diagram showing the symmetry between the M×2 (M×1) block and the 2×M (1×M) block by way of an example.

[0243] As shown in FIG. 16, since the second mode in the M×2 (M×1) block is symmetric to the 66th mode in the 2×M (1×M) block, the same set of LFNSTs can be applied to the 2×M (1×M) block and the M×2 (M×1) block.

[0244] At this time, in order to apply the LFNST set applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on the second mode instead of the 66th mode. That is, before applying the forward LFNST, after transposing the input data of the 2×M (1×M) block, the LFNST can be applied.

[0245] FIG. 17 is a drawing showing an example of transposing a 2×M block by way of example.

[0246] FIG. 17(a) is a diagram explaining that the input data can be read in column-first order for the 2×M block and the LFNST can be applied, and FIG. 17(b) is a diagram explaining that the input data can be read in row-first order for the M×2 (M×1) block and the LFNST can be applied. Organizing the method of applying the LFNST to the upper left M×2 (M×1) or 2×M (M×1) block, it is as follows.

[0247] 1. First, as shown in FIGS. 17(a) and 17(b), the input data is arranged to form the input vector of the forward LFNST. For example, referring to FIG. 16, for the M×2 block predicted in the second mode, following the order in FIG. 17(b), and for the 2×M block predicted in the 66th mode, after arranging the input data according to the order in FIG. 17(a), the LFNST set for the second mode can be applied.

[0248] For a 2×M (M×2) block, an LFNST set is determined based on a modified intra prediction mode considering WAIP. As described above, a pre - set mapping relationship is established between the intra prediction mode and the LFNST set, which can be shown as a mapping table as in Table 2.

[0249] For a 2×M (1×M) block, after obtaining a mode symmetric about the prediction mode in the 45 - degree diagonal direction downward to the right (mode 34 in the case of the VVC standard) from the modified intra prediction mode considering WAIP, an LFNST set is determined based on the symmetric mode and the mapping table. The mode (y) symmetric about mode 34 is derived by the following formula. The mapping table will be described more specifically below.

[0250]

Equation

[0251] When applying the forward LFNST, the input data prepared in the first process is multiplied by the LFNST kernel to derive the transformation coefficients. The LFNST kernel is selected from the LFNST set determined in the second process and a pre - specified LFNST index.

[0252] For example, when M = 8 and a 16×16 matrix is applied as the LFNST kernel, the matrix is multiplied by 16 input data to generate 16 transformation coefficients. The generated transformation coefficients are arranged in the upper - left 8×2 or 2×8 region according to the scan order used in the VVC standard.

[0253] Figure 18 is a diagram showing the scan order for an 8×2 or 2×8 region by way of an example.

[0254] For regions other than the upper left 8×2 or 2×8 region, they may be filled with all 0 values (zero-out), or the existing conversion coefficients to which a primary conversion has been applied may be maintained as they are. The pre-specified LFNST index can be any one of the LFNST index values (0, 1, 2) that are tried when calculating the RD cost while changing the LFNST index value during the encoding process.

[0255] In the case of a configuration that keeps the computational complexity for the worst case below a certain level (e.g., 8 multiplications / sample), for example, after multiplying an 8×16 matrix obtained by taking only the upper 8 rows of the 16×16 matrix to generate only 8 conversion coefficients, the 8 conversion coefficients are arranged according to the scan order as shown in FIG. 18, and zero-out may be applied to the remaining coefficient regions. The adjustment of the worst-case complexity will be described later.

[0256] When applying the inverse LFNST, a set number (e.g., 16) of conversion coefficients are set as the input vector, and after selecting the LFNST kernel (e.g., a 16×16 matrix) derived from the LFNST set obtained in the second process and the parsed LFNST index, the LFNST kernel is multiplied by the input vector to derive the output vector.

[0257] For the M×2 (M×1) block, the output vector is arranged in row-major order as shown in FIG. 17(b), and for the 2×M (1×M) block, the output vector is arranged in column-major order as shown in FIG. 17(a).

[0258] For the remaining regions excluding the regions where the output vector is arranged inside the upper left M×2 (M×1) or 2×M (M×2) region, and for the regions other than the upper left M×2 (M×1) or 2×M (M×2) region within the partition block, they are all configured to be filled with all 0 values (zero-out) or to maintain the conversion coefficients restored in the residual coding and inverse quantization process as they are.

[0259] Similar to the third example, when constructing the input vector, the input data is arranged according to the scan order in FIG. 18, and the number of input data can be reduced (for example, from 16 to 8) to construct the input vector in order to keep the computational complexity for the worst case below a certain level.

[0260] For example, when M = 8 and using eight input data, only the left 16×8 matrix can be taken from the 16×16 matrix for multiplication, and then 16 output data can be obtained. The adjustment of the complexity for the worst case will be described later.

[0261] In the above embodiment, when applying LFNST, the case of applying symmetry between the M×2 (M×1) block and the 2×M (1×M) block is presented. However, different LFNST sets can also be applied to the shapes of the two blocks according to other examples.

[0262] Hereinafter, various examples regarding the LFNST set configuration for the ISP mode and the mapping method using the intra prediction mode will be described.

[0263] In the case of the ISP mode, the LFNST set configuration is different from the existing LFNST set. In other words, a kernel different from the existing LFNST kernel may be applied, or a mapping table different from the mapping table between the intra prediction mode index currently applied to the VVC standard and the LFNST set may be applied. The mapping table currently applied to the VVC standard is as shown in Table 2.

[0264] In Table 2, the preModeIntra value means the intra prediction mode value changed considering WAIP, and the lfnstTrSetIdx value is the index value indicating a specific LFNST set. Each LFNST set is composed of two LFNST kernels.

[0265] When the ISP prediction mode is applied, if both the horizontal length and the vertical length of each partition block are greater than or equal to 4, the same kernel as the LFNST kernel applied in the current VVC standard may be applied, and the mapping table may also be applied as it is. Of course, a different LFNST kernel and a different mapping table from the current VVC standard may also be applied.

[0266] When the ISP prediction mode is applied, if the horizontal length or the vertical length of each partition block is less than 4, a different LFNST kernel and a different mapping table from those in the current VVC standard may be applied. Tables 5 to 7 below show the mapping tables between the intra prediction mode values (intra prediction mode values changed considering WAIP) and the LFNST sets that can be applied to M×2 (M×1) blocks or 2×M (1×M) blocks.

[0267]

Table 5

[0268]

Table 6

[0269]

Table 7

[0270] The first mapping table in Table 5 is composed of 7 LFNST sets, the mapping table in Table 6 is composed of 4 LFNST sets, and the mapping table in Table 7 is composed of 2 LFNST sets. As another example, when composed of 1 LFNST set, the lfnstTrSetIdx value is fixed to 0 for the preModeIntra value.

[0271] Hereinafter, a method for maintaining the computational complexity in the worst case when applying LFNST to the ISP mode will be described.

[0272] When in ISP mode, the application of LFNST can be restricted to maintain the multiplication per sample (or per coefficient, per position) below a certain value when applying LFNST. Depending on the size of the partition block, LFNST can be applied as follows to maintain the multiplication per sample (or per coefficient, per position) at 8 or less.

[0273] 1. When both the horizontal and vertical lengths of the partition block are 4 or more, the same method as the worst-case computational complexity adjustment method for LFNST in the current VVC standard can be applied.

[0274] That is, when the partition block is a 4×4 block, instead of a 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from the 16×16 matrix can be applied in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns from the 16×16 matrix can be applied in the reverse direction. Also, when the partition block is an 8×8 block, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the upper 8 rows from the 16×48 matrix can be applied, and in the reverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns from 48×16 can be applied.

[0275] In the case of a 4×N or N×4 (N>4) block, when performing the forward transform, a 16×16 matrix is applied only to the upper left 4×4 block. After that, the 16 generated coefficients are arranged in the upper left 4×4 area, and the other areas are filled with 0 values. Also, when performing the reverse transform, after arranging the 16 coefficients located in the upper left 4×4 block in scan order to form an input vector, a 16×16 matrix can be multiplied to generate 16 output data. The generated output data is arranged in the upper left 4×4 area, and the remaining areas except the upper left 4×4 area are filled with 0.

[0276] In the case of an 8×N or N×8 (N>8) block, when performing the forward transform, a 16×48 matrix is applied only to the ROI region (the remaining region excluding the lower-right 4×4 block from the upper-left 8×8 block) inside the upper-left 8×8 block. After that, the 16 generated coefficients are arranged in the upper-left 4×4 region, and all other regions are filled with 0 values. Also, when performing the inverse transform, after arranging the 16 coefficients located in the upper-left 4×4 block in the scan order to form an input vector, a 48×16 matrix is multiplied to generate 48 output data. The generated output data is filled in the ROI region, and all other regions are filled with 0 values.

[0277] 2. When the size of the partition block is N×2 or 2×N and the LFNST is applied to the upper-left M×2 or 2×M region (M≦N), a matrix sampled according to the N value can be applied.

[0278] When M = 8, for a partition block with N = 8, that is, an 8×2 or 2×8 block, in the case of the forward transform, instead of a 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix is applied. In the case of the inverse transform, instead of a 16×16 matrix, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix is applied.

[0279] When N is greater than 8, in the case of the forward transform, after applying a 16×16 matrix to the upper-left 8×2 or 2×8 block, the 16 generated output data are arranged in the upper-left 8×2 or 2×8 block, and the remaining regions are filled with 0 values. In the case of the inverse transform, after arranging the 16 coefficients located in the upper-left 8×2 or 2×8 block in the scan order to form an input vector, the corresponding 16×16 matrix is multiplied to generate 16 output data. The generated output data are arranged in the upper-left 8×2 or 2×8 block, and all other regions are filled with 0 values.

[0280] 3. When the size of the partition block is N×1 or 1×N and the LFNST is applied to the upper left M×1 or 1×M area (M≦N), a matrix sampled according to the N value is applied.

[0281] When M = 16, for a partition block where N = 16, that is, a 16×1 or 1×16 block, in the case of forward transformation, instead of a 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix is applied. In the case of inverse transformation, instead of a 16×16 matrix, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix is applied.

[0282] When N is greater than 16, in the case of forward transformation, after applying a 16×16 matrix to the upper left 16×1 or 1×16 block, the 16 generated output data are arranged in the upper left 16×1 or 1×16 block, and the remaining area is filled with 0 values. In the case of inverse transformation, after arranging the 16 coefficients located in the upper left 16×1 or 1×16 block in scan order to form an input vector, the corresponding 16×16 matrix is multiplied to generate 16 output data. The generated output data are arranged in the upper left 16×1 or 1×16 block, and the entire remaining area is filled with 0 values.

[0283] As another example, in order to keep the multiplication per sample (or per coefficient, per position) below a certain value, the multiplication per sample (or per coefficient, per position) is kept below 8 based on the size of the ISP coding unit rather than the size of the ISP partition block. If there is only one block among the ISP partition blocks that satisfies the condition where LFNST is applied, the complexity calculation for the worst case of LFNST is applied based on the size of the coding unit that is not the size of the partition block. For example, if the luma coding block for a certain coding unit is divided into 4 partition blocks of 4×4 size and coded by ISP, and there are no non-zero transform coefficients for 2 of the partition blocks, it can be set so that 16 transform coefficients, each not 8 (based on the encoder standard), are generated for the other 2 partition blocks.

[0284] Hereinafter, when in the ISP mode, a method for signaling the LFNST index will be described.

[0285] As described above, the LFNST index has values of 0, 1, and 2. 0 indicates that LFNST is not applied, and 1 and 2 indicate any one of the two LFNST kernel matrices included in the selected set of LFNSTs. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. Describing the transmission method of the LFNST index in the current VVC standard, it is as follows.

[0286] 1. The LFNST index can be transmitted once for each coding unit (CU). In the case of a dual-tree, individual LFNST indexes are signaled for the luma block and the chroma block respectively.

[0287] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, the following applies.

[0288] A. When in a mode where no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0289] B. When the first transformation is not DCT-2 (DST7 or DCT8), i.e., when the horizontal or vertical transformation is not DCT-2

[0290] C. When the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma transformation that can be applied. For example, when the size of the maximum luma transformation that can be applied is 64, and the size of the luma block of the coding block is 128×16, LFNST cannot be applied.

[0291] In the case of a dual tree, for each of the coding unit for the luma component and the coding unit for the chroma component, it is determined whether the size exceeds the size of the maximum luma transformation. That is, it is checked whether the size exceeds the size of the maximum luma transformation that can be applied to the luma block, and it is checked whether the vertical / horizontal length of the corresponding luma block for the color format and the size of the maximum transformation that can be applied exceed the size of the maximum luma transformation that can be applied to the chroma block. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice that of the chroma block, and the size of the corresponding luma block transformation is twice that of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length and the transformation size of the corresponding luma block are the same as those of the corresponding chroma block.

[0292] 64-length conversion or 32-length conversion means a conversion applied horizontally or vertically that has a length of 64 or 32 respectively, and "conversion size" means 64 or 32 which is the said length.

[0293] In the case of a single tree, after checking whether the horizontal or vertical length of the luma block exceeds the size of the maximum luma conversion block that can be converted, if it exceeds, the LFNST index signaling may be omitted.

[0294] The LFNST index can be transmitted only when both the horizontal length and the vertical length of the coding unit are 4 or more.

[0295] In the case of a dual tree, the LFNST index can be signaled only when both the horizontal length and the vertical length of the corresponding component (i.e., luma or chroma component) are 4 or more.

[0296] In the case of a single tree, the LFNST index can be signaled when both the horizontal length and the vertical length of the luma component are 4 or more.

[0297] E. If the position of the last non-zero coefficient is not the DC position (the upper left position of the block), for a dual-tree type luma block, when the position of the last non-zero coefficient is not the DC position, the LFNST index is transmitted. For a dual-tree type chroma block, if the position of the last non-zero coefficient for Cb or Cr is not the DC position, the corresponding LNFST index is transmitted.

[0298] In the case of a single-tree type, if the position of the last non-zero coefficient of any one of the luma component, Cb component, and Cr component is not the DC position, the LFNST index is transmitted.

[0299] Here, if the CBF (coded block flag) value indicating the presence or absence of a transform coefficient for a transform block is 0, the position of the last non-zero coefficient for the transform block is not checked to determine whether to perform LFNST index signaling. That is, when the CBF value is 0, since no transform is applied to the block, the position of the last non-zero coefficient does not need to be considered when checking the conditions for LFNST index signaling.

[0300] For example, 1) in the dual-tree type and for the luma component, if the CBF value is 0, no LFNST index is signaled; 2) in the dual-tree type and for the chroma component, if the CBF value for Cb is 0 and the CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is transmitted; 3) in the single-tree type, the position of the last non-zero coefficient is checked only for components where the CBF value for all of luma, Cb, and Cr is 1.

[0301] If it is confirmed that a transform coefficient exists at a position where an F.LFNST transform coefficient cannot exist, LFNST index signaling can be omitted. In the case of 4×4 and 8×8 transform blocks, according to the transform coefficient scan order in the VVC standard, there are LFNST transform coefficients at 8 positions starting from the DC position, and the remaining positions are all filled with 0. Also, in the case of transform blocks other than 4×4 and 8×8, according to the transform coefficient scan order in the VVC standard, there are LFNST transform coefficients at 16 positions starting from the DC position, and the remaining positions are all filled with 0.

[0302] Therefore, after performing residual coding, if there is a non-zero transform coefficient in the area where the 0 value should be filled, LFNST index signaling can be omitted.

[0303] On the one hand, the ISP mode is applied only when it is a luma block, or it may be applied to both luma and chroma blocks. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks for prediction, and the transformation is also applied to the corresponding partition blocks respectively. Therefore, when determining the conditions for signaling the LFNST index in units of coding units, the fact that LFNST can be applied to each corresponding partition block must also be considered. Also, when the ISP prediction mode is applied only to a specific component (e.g., a luma block), the LFNST index must be signaled considering the fact that it is divided into partition blocks only for that component. When in the ISP mode, sorting out the possible LFNST index signaling methods is as follows.

[0304] 1. The LFNST index can be transmitted once for each coding unit (CU), and in the case of a dual-tree, individual LFNST indexes can be signaled for the luma and chroma blocks respectively.

[0305] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.

[0306] A. In the case of a mode where transformation is not applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0307] B. When the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma transformation that can be transformed. For example, when the size of the maximum luma transformation that can be transformed is 64, and the size of the luma block of the coding block is the same as 128×16, LFNST cannot be applied.

[0308] It is also possible to determine whether to perform signaling of the LFNST index based on the size of the partition block instead of the coding unit. That is, if the horizontal or vertical length of the partition block for the luma block exceeds the size of the maximum luma transform that can be converted, the LFNST index signaling is omitted, and the LFNST index value can be analogized to 0.

[0309] In the case of a dual tree, it is determined whether each of the coding unit or partition block for the luma component and the coding unit or partition block for the chroma component exceeds the maximum transform block size. That is, when comparing the vertical and horizontal lengths of the coding unit or partition block for luma with the maximum luma transform size respectively, if either one is larger than the maximum luma transform size, LFNST is not applied. In the case of the coding unit or partition block for chroma, the horizontal / vertical length of the corresponding luma block for the color format and the size of the maximum luma transform that can be converted are compared. For example, when the color format is 4:2:0, the horizontal / vertical lengths of the corresponding luma blocks are each twice that of the chroma block, and the transform size of the corresponding luma block is twice that of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length and transform size of the corresponding luma block are the same as those of the corresponding chroma block.

[0310] In the case of a single tree, after checking whether the horizontal or vertical length of the luma block (coding unit or partition block) exceeds the maximum luma transform block size that can be converted, if it exceeds, the LFNST index signaling may be omitted.

[0311] C. If the LFNST included in the current VVC standard is applied, the LFNST index can be transmitted only when both the horizontal and vertical lengths of the partition block are 4 or more.

[0312] If, in addition to the LFNST currently included in the VVC standard, it is applied to the LFNST for 2×M (1×M) or M×2 (M×1) blocks, the LFNST index can be transmitted only when the size of the partition block is greater than or equal to 2×M (1×M) or M×2 (M×1) blocks. Here, the meaning that a P×Q block is greater than or equal to an R×S block means that P≧R and Q≧S.

[0313] After sorting, the LFNST index can be transmitted only when the partition block is greater than or equal to the minimum size to which the LFNST is applicable. In the case of a dual tree, the LFNST index can be signaled only when the partition block for the luma or chroma component is greater than or equal to the minimum size to which the LFNST is applicable. In the case of a single tree, the LFNST index can be signaled only when the partition block for the luma component is greater than or equal to the minimum size to which the LFNST is applicable.

[0314] In this document, the fact that an M×N block is greater than or equal to a K×L block means that M is greater than or equal to K and N is greater than or equal to L. The fact that an M×N block is greater than a K×L block means that M is greater than or equal to K and N is greater than or equal to L, while M is greater than K or N is greater than L. The fact that an M×N block is less than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L, and the fact that an M×N block is less than a K×L block means that M is less than or equal to K and N is less than or equal to L, while M is less than K or N is less than L.

[0315] D. If the position of the last non-zero coefficient (last non-zero coefficient position) is not the DC position (the upper left corner position of the block), for a dual-tree type luma block, if the position of the last non-zero coefficient in any one of all the partition blocks is not the DC position, the LFNST can be transmitted. For a dual-tree type chroma block, if the position of the last non-zero coefficient in all the partition blocks for Cb (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) and the position of the last non-zero coefficient in all the partition blocks for Cr (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) are not the DC position in any one of them, the LNFST index can be transmitted.

[0316] In the case of the single-tree type, if the position of the last non-zero coefficient in any one of all the partition blocks for the luma component, the Cb component, and the Cr component is not the DC position, the corresponding LFNST index can be transmitted.

[0317] Here, when the CBF (coded block flag) value indicating whether there is a transform coefficient for each partition block is 0, in order to determine whether to perform LFNST index signaling, the position of the last non-zero coefficient for the partition block is not checked. That is, when the CBF value is 0, since no transform is applied to the block, the position of the last non-zero coefficient for the partition block is not considered when checking the conditions related to LFNST index signaling.

[0318] For example, 1) in the case of the dual-tree type and the luma component, when determining whether to perform LFNST index signaling, if the corresponding CBF value for each partition block is 0, the corresponding partition block is excluded; 2) in the case of the dual-tree type and the chroma component, when the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, only the position of the last non-zero coefficient for Cr is checked to determine whether to perform the corresponding LFNST index signaling; 3) in the case of the single-tree type, for all partition blocks of the luma component, Cb component, and Cr component, only for blocks where the CBF value is 1, the position of the last non-zero coefficient is checked to determine whether to perform LFNST index signaling.

[0319] In the case of the ISP mode, the video information may be configured so as not to check the position of the last non-zero coefficient, and the embodiments related thereto are as follows.

[0320] i. In the case of the ISP mode, the check regarding the position of the last non-zero coefficient is omitted for both the luma block and the chroma block, and LFNST index signaling is allowed. That is, for all partition blocks, even if the position of the last non-zero coefficient is the DC position or the corresponding CBF value is 0, the LFNST index signaling is allowed.

[0321] ii. In the case of the ISP mode, the check regarding the position of the last non-zero coefficient is omitted only for the luma block, and in the case of the chroma block, the check regarding the position of the last non-zero coefficient in the above-described manner is performed. For example, in the case of the dual-tree type and the luma block, LFNST index signaling is allowed without checking the position of the last non-zero coefficient, and in the case of the dual-tree type and the chroma block, the presence or absence of the DC position for the position of the last non-zero coefficient in the above-described manner is checked to determine whether to perform the signaling of the corresponding LFNST index.

[0322] iii. If it is in ISP mode and of single-tree type, apply the method of item i or ii above. That is, when it is in ISP mode and single-tree type and item i is applied, for both the luma block and the chroma block, check on the position of the non-zero coefficient at the end is omitted, and LFNST index signaling is allowed. Or, when item ii is applied, for the partition block for the luma component, check on the position of the non-zero coefficient at the end is omitted, and for the partition block for the chroma component (when ISP is not applied to the chroma component, it is regarded as having 1 partition block), check on the position of the non-zero coefficient at the end is performed in the above-mentioned manner to determine whether to perform the corresponding LFNST index signaling.

[0323] E. If it is confirmed that a conversion coefficient exists at a position where there should be no LFNST conversion coefficient for even one of all the partition blocks, LFNST index signaling can be omitted.

[0324] For example, in the case of 4×4 and 8×8 partition blocks, according to the conversion coefficient scan order in the VVC standard, LFNST conversion coefficients exist at 8 positions starting from the DC position, and the remaining positions are all filled with 0. Also, when it is equal to or larger than 4×4 but not a 4×4 or 8×8 partition block, according to the conversion coefficient scan order in the VVC standard, LFNST conversion coefficients exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0.

[0325] Therefore, after performing residual coding, if there is a non-zero conversion coefficient in the area where the 0 value should be filled, LFNST index signaling can be omitted.

[0326] If the LFNST can also be applied to the case where the partition block is 2×M (1×M) or M×2 (M×1), the area where the LFNST conversion coefficient can be located can be specified as follows. The area outside the area where the conversion coefficient can be located is filled with 0. If there is a non-zero conversion coefficient in the area that must be filled with 0 when assuming that the LFNST is applied, the LFNST index signaling can be omitted.

[0327] i. When the LFNST can be applied to a 2×M or M×2 block and M = 8, only 8 LFNST conversion coefficients are generated for a 2×8 or 8×2 partition block. When the conversion coefficients are arranged in the scan order as shown in FIG. 18, 8 conversion coefficients are arranged in the scan order from the DC position, and the remaining 8 positions are filled with 0.

[0328] For a 2×N or N×2 (N>8) partition block, 16 LFNST conversion coefficients are generated. When the conversion coefficients are arranged in the scan order as shown in FIG. 18, 16 conversion coefficients are arranged in the scan order from the DC position, and the remaining area is filled with 0. That is, in a 2×N or N×2 (N>8) partition block, the area other than the upper left 2×8 or 8×2 block is filled with 0. For a 2×8 or 8×2 partition block, 16 conversion coefficients are generated instead of 8 LFNST conversion coefficients. In this case, no area that must be filled with 0 occurs. As described above, when the LFNST is applied, if it is detected that there is a non-zero conversion coefficient in the area that is determined to be filled with 0 even in one partition block, the LFNST index signaling can be omitted and the LFNST index can be assumed to be 0.

[0329] ii. The LFNST can be applied to a 1×M or M×1 block. When M = 16, only 8 LFNST transform coefficients are generated for a 1×16 or 16×1 partition block. When the transform coefficients are arranged in the scan order from left to right or from top to bottom, 8 transform coefficients are arranged in the corresponding scan order starting from the DC position, and the remaining 8 positions are filled with 0.

[0330] For a 1×N or N×1 (N>16) partition block, 16 LFNST transform coefficients are generated. When the transform coefficients are arranged in the scan order from left to right or from top to bottom, 16 transform coefficients are arranged in the corresponding scan order starting from the DC position, and the remaining area is filled with 0. That is, in a 1×N or N×1 (N>16) partition block, the area other than the upper left 1×16 or 16×1 block is filled with 0.

[0331] Even for a 1×16 or 16×1 partition block, 16 transform coefficients are generated instead of 8 LFNST transform coefficients, and in this case, there is no area that needs to be filled with 0. As described above, when the LFNST is applied, if it is detected that there is a non-zero transform coefficient in the area defined to be filled with 0 in a partition block, the LFNST index signaling can be omitted, and the LFNST index can be inferred as 0.

[0332] On the other hand, in the ISP mode, in the current VVC standard, DST-7 is applied instead of DCT-2 without signaling for the MTS index by looking at the length conditions independently for the horizontal and vertical directions. It is determined whether the vertical length or the horizontal length is greater than 4 or equal to and less than 16, and the primary transform kernel is determined according to the determination result. Therefore, for the case where the LFNST can be applied while in the ISP mode, the following conversion combination configuration is possible.

[0333] 1. When the LFNST index is 0 (including the case where the LFNST index is analogized to 0), follow the decision conditions for the primary transformation when it is the ISP included in the current VVC standard. That is, check whether the length conditions (greater than 4 or equal to less than 16) are satisfied independently for the horizontal and vertical directions respectively. If satisfied, apply DST-7 instead of DCT-2 for the primary transformation; if not satisfied, apply DCT-2.

[0334] 2. When the LFNST index is greater than 0, the following two configurations are possible for the primary transformation.

[0335] A. DCT-2 can be applied to both the horizontal and vertical directions.

[0336] B. It is possible to follow the decision conditions for the primary transformation when it is the ISP included in the current VVC standard. That is, check whether the length conditions (greater than 4 or equal to less than 16) are satisfied independently for the horizontal and vertical directions respectively. If satisfied, apply DST-7 instead of DCT-2; if not satisfied, apply DCT-2.

[0337] In the case of the ISP mode, since the LFNST index is not transmitted for each coding unit, the video information can be configured to be transmitted for each partition block. In such a case, consider that there is only one partition block within the unit in which the LFNST index is transmitted in the aforementioned LFNST index signaling method, and it is possible to determine whether to perform LFNST index signaling.

[0338] Summarizing the embodiments in which LFNST is applied in the aforementioned ISP mode, it is as follows.

[0339] (1) When LFNST is applied in the ISP mode, the transformed unit to be divided must have a size of at least 4×4 or more.

[0340] (2) The same LFNST kernel as the existing LFNST kernel applied to the coding unit to which the ISP mode is not applied can be used.

[0341] (3) All conversion units must satisfy the maximum last position value condition (the position condition of the valid coefficient that is not 0 at the end). If one or more conversion units do not satisfy the maximum last position value condition, the LFNST is not used and the LFNST index is not parsed.

[0342] (4) When the ISP mode is applied, the setting that the LFNST cannot be applied if the valid coefficient does not exist outside the DC position may be ignored.

[0343] (5) When the LFNST is applied, DCT-2 is used for the primary conversion of the conversion unit to which the ISP is applied.

[0344] The following Table 8 shows the syntax elements including the above content.

[0345]

Table 8

[0346] Table 8 sets the width and height of the area to which the LFNST is applied according to the tree type, and shows the conditions for transmitting the LFNST index. The syntax elements in Table 8 are signaled at the coding unit (CU) level. In the case of a dual-tree, individual LFNST indexes are signaled for the luma block and the chroma block respectively.

[0347] First, when the tree type of the coding unit is dual - tree chroma, the width (lfnstWidth) of the area to which LFNST is applied can be set to the width reflected by the color format from the width of the coding unit ((treeType == DUAL_TREE_CHROMA)? cbWidth / SubWidthC).

[0348] On the contrary, when the tree type of the coding unit is not dual - tree chroma, that is, it is dual - tree luma or single - tree, the width (lfnstWidth) of the area to which LFNST is applied is set to the value obtained by dividing the coding unit by the number of sub - partitions or the width of the coding unit depending on whether the coding unit is split by ISP ((IntraSubPartitionsSplitType == ISP_VER_SPLIT)? cbWidth / NumIntraSubPartitions: cbWidth). That is, when the coding unit is split vertically by ISP (IntraSubPartitionsSplitType == ISP_VER_SPLIT), the width of the area to which LFNST is applied is set to the value obtained by dividing the coding unit by the number of sub - partitions (cbWidth / NumIntraSubPartitions), and when it is not split, it is set to the width of the coding unit (cbWidth).

[0349] Similarly, when the tree type of the coding unit is dual - tree chroma, the height (lfnstHeight) of the area to which LFNST is applied is set to the height reflected by the color format from the height of the coding unit ((treeType == DUAL_TREE_CHROMA)? cbHeight / SubHeightC).

[0350] If the tree type of the coding unit is not dual tree chroma, i.e., it is dual tree luma or single tree, the height (lfnstHeight) of the area to which LFNST is applied is set to the value obtained by dividing the coding unit by the number of sub - partitions or the height of the coding unit depending on whether the coding unit is split by ISP ((IntraSubPartitionsSplitType == ISP_HOR_SPLIT)? cbHeight / NumIntraSubPartitions: cbHeight). That is, when the coding unit is split horizontally by ISP (IntraSubPartitionsSplitType == ISP_HOR_SPLIT), the height of the area to which LFNST is applied is set to the value obtained by dividing the coding unit by the number of sub - partitions (cbHeight / NumIntraSubPartitions), and when it is not split, it is set to the height of the coding unit (cbHeight).

[0351] In order for LFNST to be applied in this way, the width and height of the area to which the LFNST is applied must be 4 or more (Min(lfnstWidth, lfnstHeight) ≥ 4). That is, in the case of a coding dual tree, the LFNST index is signaled only when both the horizontal and vertical lengths of the component (i.e., luma or chroma component) are 4 or more, and in the case of a single tree, the LFNST index is signaled when both the horizontal and vertical lengths of the luma component are 4 or more.

[0352] When ISP is applied to the coding unit, the LFNST index is transmitted only when both the horizontal and vertical lengths of the partition block are 4 or more.

[0353] Also, when the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma transform block that can be transformed (when the condition Max(cbWidth, cbHeight) ≦ MaxTbSizeY is not satisfied), LFNST cannot be applied and the LFNST index is not transmitted.

[0354] Also, the LFNST index is signaled only when the position of the last non-zero coefficient (last non-zero coefficient position) is not the DC position (the upper left position of the block).

[0355] In the case of a dual-tree type luma block, if the position of the last non-zero coefficient is not the DC position, the LFNST index is transmitted. In the case of a dual-tree type chroma block, if either the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position, the LFNST index is transmitted. In the case of a single-tree type, if the position of the last non-zero coefficient is not the DC position for any one of the luma component, Cb component, and Cr component, the LFNST index is transmitted.

[0356] On the other hand, when ISP is applied to the coding unit, the LFNST index can be signaled without checking the position of the last non-zero coefficient (IntraSubPartitionsSplitType!= ISP_NO_SPLIT || LfnstDcOnly == 0). That is, even if the position of the last non-zero coefficient for all partition blocks is the DC position, LFNST index signaling can be allowed. The DC position indicates the upper left position of the block.

[0357] Finally, when it is confirmed that there is a transform coefficient at a position where the LFNST transform coefficient cannot exist, the LFNST index signaling can be omitted (LfnstZeroOutSigCoeffFlag == 1).

[0358] When ISP is applied to a coding unit, if it is confirmed that a transform coefficient exists at a position where there should be no transform coefficient for one partition block among all partition blocks, LFNST index signaling can be omitted.

[0359] In the following, an embodiment in which an LFNST kernel sampled from 8×8 LFNST is applied in the ISP mode will be described.

[0360] By way of an example, LFNST kernels (A) applicable to the upper left 4×4 block, LFNST kernels (B) applicable to the upper left 4×8 block, and LFNST kernels (C) applicable to the upper left 8×4 block can be derived by sampling kernel data from 8×8 LFNST (a 16×48 matrix for the forward LFNST, for example, the 8×8 LFNST included in the current VVC standard).

[0361] The derived kernels can be used as LFNST kernels when in the ISP mode and when LFNST is applied. For example, (A) can be applied to a 4×4 ISP partition block, (B) can be applied to an N×4 ISP partition block (N≥8), and (C) can be applied to a 4×N ISP partition block (N≥8). For an ISP partition with both the horizontal length and the vertical length being the same as or larger than 8, the existing 8×8 LFNST (for example, the 8×8 LFNST included in the current VVC standard) can be applied.

[0362] In order to unify the LFNST computational complexity and reduce memory usage, a 16×48 LFNST kernel for LFNST has been proposed. For example, a block with a small size such as 4×N or N×4 can be regarded as a part of an 8×N or N×8 block. For this reason, the overlapping part with the 16×48 LFNST kernel, that is, a part of the 16×48 LFNST kernel can be used as an LFNST kernel.

[0363] FIG. 19 is a diagram for explaining a sampled LFNST kernel in the case of an ISP mode according to an example.

[0364] In the case of the forward LFNST, when a 16×48 LFNST kernel is applied to a 4×N or N×4 block, a matrix overlapping the 16×48 LFNST kernel can be used for the secondary transformation.

[0365] FIG. 19 shows a 16×48 LFNST kernel that can be applied to 4×4, 8×4, 4×8, and 16×4 regions. (a) in FIG. 19 shows that when applying the 16×48 LFNST kernel to a 4×4 region, only the kernel part overlapping the 4×4 region is used; (b) in FIG. 19 shows that when applying the 16×48 LFNST kernel to an 8×4 region, only the kernel part overlapping the 8×4 region is used; (c) in FIG. 19 shows that when applying the 16×48 LFNST kernel to a 4×8 region, only the kernel part overlapping the 4×8 region is used. (d) in FIG. 19 shows that when applying the 16×48 LFNST kernel to a 16×4 region, a 16×32 matrix overlapping the 16×48 LFNST kernel in the 16×4 region can be used.

[0366] In the case of the reverse LFNST, 8 or 16 coefficients can be input, and 16 or 48 coefficients can be output as a result. On the other hand, operations for transformation can be performed only on the samples within the block size region, that is, on the coefficients, and the remaining samples are not calculated. For example, in the case of a 4×4 transformation unit, if 8 coefficients and a 16×48 LFNST kernel are given, only the 4×4 region in the upper left corner is output and calculated, and the coefficients outside the 4×4 region are not operated on.

[0367] In order to maintain the number of multiplication operations in the worst case per coefficient, in the case of 8×4 and 4×8 blocks, only 8 coefficients are calculated, which is consistent with the number of operations of the LFNST applied to the existing 4×4 and 8×8 transformation units.

[0368] The following drawings are created to illustrate a specific example of this specification. Since the names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited by the specific names used in the following drawings.

[0369] FIG. 20 is a flowchart showing the operation of a video decoding apparatus according to an embodiment of this document.

[0370] Each step disclosed in FIG. 20 is based on a part of the content described above in FIGS. 2 to 19. Therefore, specific content overlapping with the content described above in FIGS. 2 to 19 is omitted or simplified in the description.

[0371] A decoding apparatus 200 according to an embodiment can receive residual information from a bitstream (S2010).

[0372] More specifically, the decoding apparatus 200 can decode information regarding quantized transform coefficients for a current block from the bitstream, and based on the information regarding quantized transform coefficients for the current block, can derive quantized transform coefficients for a target block. The information regarding quantized transform coefficients for the target block may be included in an SPS (Sequence Parameter Set) or a slice header, and may include at least one of information regarding whether simplified transform (RST) is applied, information regarding a simplification factor, information regarding the minimum transform size to which simplified transform is applied, information regarding the maximum transform size to which simplified transform is applied, a simplified inverse transform size, and information regarding a transform index indicating any one of transform kernel matrices included in a transform set.

[0373] Further, the decoding device can further receive information regarding the intra prediction mode for the current block and information regarding whether ISP is applied to the current block. The decoding device can derive whether the current block is divided into a predetermined number of sub-partition conversion blocks by receiving and parsing flag information indicating whether to apply ISP coding or the ISP mode. Here, the current block can be a coding block. Also, the decoding device can derive the size and number of sub-partition blocks into which the current block is divided via flag information indicating the direction in which the current block is divided.

[0374] For example, as shown in FIG. 14, when the size (width × height) of the current block is 8 × 4, the current block is divided vertically into two sub-blocks, and when the size (width × height) of the current block is 4 × 8, the current block is divided horizontally into two sub-blocks. Or, as shown in FIG. 15, when the size (width × height) of the current block is larger than 4 × 8 or 8 × 4, that is, when the size of the current block is 1) 4 × N or N × 4 (N ≧ 16), or 2) M × N (M ≧ 8, N ≧ 8), the current block is divided into four sub-blocks in the horizontal or vertical direction.

[0375] In the current block, the same intra prediction mode is applied to the sub - partition blocks that have been split, and the decoding device derives prediction samples for each sub - partition block. That is, the decoding device performs intra prediction sequentially, for example, horizontally or vertically, from left to right or from top to bottom, according to the split form of the sub - partition block. For the left - most or top - most sub - block, the restored pixels of the coded block that has already been coded in the same way as the normal intra prediction method are referred to. Also, when each side of a subsequent internal sub - partition block is not adjacent to the previous sub - partition block, in order to derive the reference pixels adjacent to that side, the restored pixels of the adjacent coded block that has already been coded in the same way as the normal intra prediction method are referred to.

[0376] The decoding device 200 performs inverse quantization on the residual information for the current block, that is, the quantized transform coefficients, to derive the transform coefficients (S2020).

[0377] The derived transform coefficients are arranged in the reverse diagonal scan order in units of 4×4 blocks, and the transform coefficients within the 4×4 block are also arranged in the reverse diagonal scan order. That is, the transform coefficients after inverse quantization are arranged according to the reverse scan order applied in video codecs such as VVC and HEVC.

[0378] The transform coefficients derived based on such residual information are the transform coefficients inverse - quantized as described above and may also be the quantized transform coefficients. That is, the transform coefficients may be any data that can check whether the data in the current block is non - zero data regardless of whether quantization is performed.

[0379] The decoding device can determine whether the conversion coefficient exists in the second region excluding the first region at the upper left end of the current block. If the conversion coefficient does not exist in the second region, the LFNST index can be parsed. Further, the decoding device can determine whether the conversion coefficient does not exist in all of the individual second regions for a plurality of sub-partition blocks when the current block is divided into a plurality of sub-partition blocks (S2030).

[0380] The decoding device can check whether zeroing out has been performed on the second region by deriving a first variable indicating whether a valid coefficient exists in the second region excluding the first region at the upper left end of the current block.

[0381] The first variable can be the variable LfnstZeroOutSigCoeffFlag which can indicate that zeroing out has been performed when LFNST is applied. The first variable is initially set to 1, and if a valid coefficient exists in the second region, the second variable can be changed to 0.

[0382] The variable LfnstZeroOutSigCoeffFlag can be derived to be 0 when the index of the sub-block where the last non-zero coefficient exists is greater than 0 and both the width and height of the transform block are the same as or greater than 4, or when the position of the last non-zero coefficient within the sub-block where the last non-zero coefficient exists is greater than 7 and the size of the transform block is 4×4 or 8×8. A sub-block means a 4×4 block used as a coding unit in residual coding and can also be named CG (Coefficient Group). The index of the sub-block being 0 refers to the upper left 4×4 sub-block.

[0383] That is, if a non-zero coefficient is derived in a region other than the upper left region where the LFNST conversion coefficient can exist in the conversion block, or if a non-zero coefficient exists at a position other than the 8th position in the scan procedure for 4×4 blocks and 8×8 blocks, the variable LfnstZeroOutSigCoffFlag is set to 0.

[0384] For example, when ISP is applied to a coding unit, if it is confirmed that a conversion coefficient exists at a position where the LFNST conversion coefficient cannot exist even for one sub-partition block among all sub-partition blocks, the LFNST index signaling can be omitted. That is, if there are valid coefficients in the second region without zeroing out in one sub-partition block, the LFNST index is not signaled.

[0385] On the other hand, the first region is derived based on the size of the current block.

[0386] For example, if the size of the current block is 4×4 or 8×8, the first region is from the upper left side of the current block to the 8th sample position in the scan direction. When the current block is divided, if the size of the sub-partition block is 4×4 or 8×8, the first region is from the upper left side of the sub-partition block to the 8th sample position in the scan direction.

[0387] If the size of the current block is 4x4 or 8x8, 8 pieces of data are output via the forward LFNST, so the 8 conversion coefficients received by the decoding device can be arranged from the upper left side of the current block to the 8th sample position in the scan direction as shown in FIGS. 11(a) and 12(a).

[0388] Also, in the remaining cases where the current block size is not 4x4 or 8x8, the first region can be a 4x4 region in the upper left corner of the current block. If the current block size is not 4x4 or 8x8, 16 data are output via the forward LFNST, so the 16 conversion coefficients received by the decoding device can be arranged in the 4x4 region in the upper left corner of the current block, as shown in FIGS. 11(b) to (d) and FIG. 12(b).

[0389] On the other hand, the conversion coefficients that can be arranged in the first region can be arranged along the diagonal scan direction as shown in FIG. 7.

[0390] As described above, when the current block is divided into sub-partition blocks, if there are no conversion coefficients in all of the individual second regions for the plurality of sub-partition blocks, the decoding device parses the LFNST index. If there are conversion coefficients in the second region for any one of the sub-partition blocks, the LFNST index is not parsed.

[0391] As described above, LFNST is applied to sub-partition blocks having a width and height of 4 or more, and the LFNST index for the current block, which is a coding block, can be applied to a plurality of sub-partition blocks.

[0392] On the other hand, since the zero-out with LFNST reflection (including all zero-outs associated with the application of LFNST) is also applied as it is to the sub-partition blocks, it is applied identically to both the first region and the sub-partition blocks. That is, when the divided sub-partition block is a 4×4 block or an 8×8 block, LFNST is applied to the conversion coefficients up to the 8th in the scan direction from the upper left corner of the sub-partition block, and when the sub-partition block is not a 4×4 block or an 8×8 block, LFNST is applied to the conversion coefficients in the upper left 4×4 region of the sub-partition block.

[0393] On the one hand, according to an example, in order for the decoding device to determine whether the LFNST index can be parsed, the decoding device can derive a second variable indicating whether there is the conversion coefficient, that is, a valid coefficient, in the area excluding the DC position of the current block.

[0394] The second variable can be the variable LfnstDcOnly that can be derived in the residual coding process. If the index of the sub-block including the last valid coefficient in the current block is 0 and the position of the last valid coefficient in the sub-block is greater than 0, the second variable can be derived to be 0. If the second variable is 0, the LFNST index can be parsed. A sub-block means a 4×4 block used as a coding unit in residual coding, and can also be named as CG (Coefficient Group). The index of the sub-block being 0 refers to the left-upper 4×4 sub-block.

[0395] The second variable can be initially set to 1, and can be maintained as 1 or changed to 0 depending on whether there is a valid coefficient in the area excluding the DC position.

[0396] The variable LfnstDcOnly indicates whether there is a non-zero coefficient at a position other than the DC component for at least one conversion block in one coding unit. If there is a non-zero coefficient at a position other than the DC component for at least one conversion block in one coding unit, it becomes 0, and if there is no non-zero coefficient at a position other than the DC component for all conversion blocks in one coding unit, it can become 1.

[0397] The decoding device can parse the LFNST index based on the derivation result (S2040).

[0398] That is, when the current block is divided into a plurality of sub - partition blocks, if there are no transform coefficients in all of the individual second regions for the plurality of sub - partition blocks, the LFNST index can be parsed to perform LFNST.

[0399] The LFNST index information is received as syntax information, and the syntax information can be received as a binary string including 0 and 1.

[0400] The syntax element of the LFNST index according to this embodiment can indicate whether inverse LFNST or inverse non - separable transform is applied and any one of the transform kernel matrices included in the transform set. When the transform set includes two transform kernel matrices, the value of the syntax element of the transform index can be three.

[0401] That is, according to one embodiment, the syntax element value for the LFNST index can include 0 indicating that inverse LFNST is not applied to the target block, 1 indicating the first transform kernel matrix among the transform kernel matrices, and 2 indicating the second transform kernel matrix among the transform kernel matrices.

[0402] The intra - prediction mode information and the LFNST index information can be signaled at the coding unit level.

[0403] On the other hand, the decoding device can parse the LFNST index without deriving the second variable based on the fact that the current block is divided into a plurality of sub - partition blocks.

[0404] In one example, when the current block is not divided into a plurality of sub - partition blocks and the second variable indicates that there are transform coefficients in the region excluding the DC position, the decoding device can parse the LFNST index. If, by chance, the current block is divided into a plurality of sub - partition blocks, the second variable is not checked, or the first variable value is ignored, and the LFNST index can be parsed.

[0405] That is, when ISP is applied to the current block, even if the position of the last non - zero coefficient for all sub - partition blocks is located at the DC position, LFNST index signaling can be allowed.

[0406] The decoding device can derive a corrected transform coefficient from the transform coefficient based on the LFNST index and the LFNST matrix for LFNST (S2050).

[0407] Unlike the first - order transform that separates the coefficients to be transformed in the vertical or horizontal direction and then transforms them, LFNST is a non - separable transform that applies the transform without separating the coefficients in a specific direction. Such a non - separable transform can be a low - frequency non - separable transform that applies the forward transform only to the low - frequency region that is not the entire block region.

[0408] The decoding device can determine an LFNST set including an LFNST matrix based on the intra - prediction mode derived from the intra - prediction mode information, and can select any one of the plurality of LFNST matrices based on the LFNST set and the LFNST index.

[0409] At this time, the same LFNST set and the same LFNST index are applied to the sub-partition conversion blocks divided in the current block. That is, since the same intra prediction mode is applied to the sub-partition conversion blocks, the LFNST set determined based on the intra prediction mode is also applied identically to all the sub-partition conversion blocks. Also, since the LFNST index is signaled at the coding unit level, the same LFNST matrix is applied to the sub-partition conversion blocks divided in the current block.

[0410] On the other hand, as described above, a conversion set is determined according to the intra prediction mode of the conversion block to be transformed, and the inverse LFNST is performed based on the conversion kernel matrix included in the conversion set indicated by the LFNST index, that is, any one of the matrices of the LFNST. The matrix applied to the inverse LFNST is named the inverse LFNST matrix or the LFNST matrix, and it doesn't matter what the name is as long as such a matrix has a transpose relationship with the matrix used for the forward LFNST.

[0411] In one example, the matrix of the inverse LFNST can be a non-square matrix whose number of columns is less than the number of rows.

[0412] On the other hand, the conversion coefficients that are the output data of the LFNST are derived as a predetermined number based on the size of the current block or the sub-partition conversion block. For example, when the height and width of the current block or the sub-partition conversion block are 8 or more, 48 conversion coefficients as shown on the left side of FIG. 6 are derived, and when the width and height of the sub-partition conversion block are not 8 or more, that is, when the width and height of the sub-partition conversion block are 4 or more but the width or height of the sub-partition conversion block is less than 8, 16 conversion coefficients as shown on the right side of FIG. 6 are derived.

[0413] As shown in FIG. 6, the 48 transform coefficients are arranged in the upper left, upper right, and lower left 4×4 regions in the upper left 8×8 region of the sub-partition transform block, and the 16 transform coefficients are arranged in the upper left 4×4 region of the sub-partition transform block.

[0414] The 48 transform coefficients and the 16 transform coefficients are arranged vertically or horizontally according to the intra prediction mode of the sub-partition transform block. For example, when the intra prediction mode is in the horizontal direction (modes 2 to 34 in FIG. 4) with respect to the diagonal direction (mode 34 in FIG. 4), the transform coefficients are arranged in the horizontal direction as shown in FIG. 6(a), that is, in the row-first order. When the intra prediction mode is in the vertical direction (modes 35 to 66 in FIG. 4) with respect to the diagonal direction, the transform coefficients are arranged in the horizontal direction as shown in FIG. 6(b), that is, in the column-first order.

[0415] The decoding device derives residual samples for the current block based on the first-order inverse transform for the modified transform coefficients (S2060).

[0416] At this time, for the inverse first-order transform, a normal separable transform can be used, or the above-described MTS can also be used.

[0417] Subsequently, the decoding device 200 can generate restored samples based on the residual samples for the current block and the prediction samples for the current block (S2070).

[0418] The following drawings are created to illustrate a specific example of this specification. Since the names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.

[0419] FIG. 21 is a flowchart showing the operation of a video encoding device according to an embodiment of this document.

[0420] Each step disclosed in FIG. 21 is based on a part of the content described above in FIGS. 3 to 19. Therefore, specific content that overlaps with the content described above in FIGS. 1 and 3 to 19 will be omitted from the description or simplified.

[0421] An encoding apparatus 100 according to an embodiment derives a prediction sample for a current block based on an intra prediction mode applied to the current block (S2110).

[0422] When ISP is applied to the current block, the encoding apparatus performs prediction for each sub - partition conversion block.

[0423] The encoding apparatus determines whether to apply ISP coding or an ISP mode to the current block, that is, the coding block, determines in which direction the current block is divided according to the determination result, and derives the size and number of sub - blocks to be divided.

[0424] For example, as shown in FIG. 14, when the size (width × height) of the current block is 8 × 4, the current block is divided vertically into two sub - blocks. When the size (width × height) of the current block is 4 × 8, the current block is divided horizontally into two sub - blocks. Or, as shown in FIG. 15, when the size (width × height) of the current block is larger than 4 × 8 or 8 × 4, that is, when the size of the current block is 1) 4 × N or N × 4 (N ≧ 16) or 2) M × N (M ≧ 8, N ≧ 8), the current block is divided into four sub - blocks in the horizontal or vertical direction.

[0425] The same intra prediction mode is applied to the sub - partition transform blocks that are now partitioned in the current block, and the encoding device derives prediction samples for each sub - partition transform block. That is, the encoding device performs intra prediction sequentially, for example, horizontally (Horizontal) or vertically (Verticial), from left to right or from top to bottom, according to the partitioning form of the sub - partition transform block. For the left - most or top - most sub - block, the reconstructed pixels of the coded block that have already been coded in the same way as the normal intra prediction method are referred to. Also, when each side of a subsequent internal sub - partition transform block is not adjacent to the previous sub - partition transform block, the reconstructed pixels of the adjacent coded block that have already been coded in the same way as the normal intra prediction method are referred to in order to derive the reference pixels adjacent to that side.

[0426] The encoding device 100 derives a residual sample for the current block based on the prediction sample (S2120).

[0427] Also, the encoding device 100 derives a transform coefficient for the current block based on a first - order transform of the residual sample (S2130).

[0428] The first - order transform is performed by a plurality of transform kernels. In this case, the transform kernel is selected based on the intra prediction mode.

[0429] The encoding device 100 determines whether to perform a second - order transform or a non - separable transform, specifically LFNST, on the transform coefficient for the current block, and can derive a modified transform coefficient by applying LFNST to the transform coefficient.

[0430] Unlike the first - order transform that separates the coefficients to be transformed in the vertical or horizontal direction and performs the transform, LFNST is a non - separable transform that applies the transform without separating the coefficients in a specific direction. Such a non - separable transform can be a low - frequency non - separable transform that applies the transform only to the low - frequency region, rather than to the entire target block to be transformed.

[0431] When ISP is applied to the current block, the encoding device can determine whether LFNST can be applied to the height and width of the divided sub - partition block.

[0432] The encoding device can determine whether LFNST can be applied to the height and width of the divided sub - partition block. In this case, when the height and width of the sub - partition block are 4 or more, the decoding device can parse the LFNST index.

[0433] Also, the encoding device can determine whether LFNST can be applied based on the tree type and color format of the current block.

[0434] By way of example, when the tree type of the current block is dual - tree chroma, the encoding device determines that LFNST can be applied when the height and width corresponding to the chroma component block of the current block are 4 or more.

[0435] Also, by way of example, when the tree type of the current block is single - tree or dual - tree luma, the encoding device determines that LFNST can be applied when the height and width corresponding to the luma component block of the current block are 4 or more.

[0436] For example, when the tree type of the current block is dual - tree chroma and ISP is not applied, in this case, the encoding device determines that LFNST can be applied when the height and width corresponding to the chroma component block of the current block are 4 or more.

[0437] When the reverse side and the tree type of the current block are dual tree chroma or single tree, the encoding device determines whether LFNST can be applied according to whether ISP is applied to the current block. If the height and width of the sub-partition block for the luma component block of the current block or the height and width of the current block are 4 or more, it is determined that LFNST can be applied.

[0438] Also, in one example, when the current block is a coding unit and the width and height of the coding unit are less than or equal to the size of the maximum luma conversion that can be converted, the encoding device determines that LFNST can be applied.

[0439] When it is determined to perform LFNST, the encoding device 100 derives a modified conversion coefficient for the current block or the sub-partition conversion block based on the LFNST set mapped to the intra prediction mode and the LFNST matrix included in the LFNST set (S2140).

[0440] The encoding device 100 determines the LFNST set based on the mapping relationship by the intra prediction mode applied to the current block, and can perform LFNST, that is, non-separable conversion, based on any one of the two LFNST matrices included in the LFNST set.

[0441] At this time, the same LFNST set and the same LFNST index are applied to the sub-partition conversion blocks divided in the current block. That is, since the same intra prediction mode is applied to the sub-partition conversion blocks, the LFNST sets determined based on the intra prediction mode are also applied identically to all sub-partition conversion blocks. Also, since the LFNST index is encoded in units of coding units, the same LFNST matrix is applied to the sub-partition conversion blocks divided in the current block.

[0442] As described above, the conversion set is determined by the intra prediction mode of the conversion block to be converted. The matrix applied to the LFNST has a transpose relationship with the matrix used for the reverse LFNST.

[0443] In one example, the LFNST matrix can be a non-square matrix with the number of rows less than the number of columns.

[0444] The region where the conversion coefficients used as the input data of the LFNST are located is derived based on the size of the sub-partition conversion block. For example, when the height and width of the sub-partition conversion block are 8 or more, the region is the upper left, upper right, and lower left 4×4 regions of the upper left 8×8 region of the sub-partition conversion block as shown on the left side of FIG. 6, and in the remaining cases where the height and width of the sub-partition conversion block are less than 8, the region can be the upper left 4×4 region of the current block as shown on the right side of FIG. 6.

[0445] The conversion coefficients of the region are read vertically or horizontally according to the intra prediction mode of the sub-partition conversion block for the multiplication operation with the LFNST matrix to form a one-dimensional vector.

[0446] Forty-eight corrected conversion coefficients or sixteen corrected conversion coefficients are read vertically or horizontally according to the intra prediction mode of the sub-partition conversion block and arranged in one dimension. For example, when the intra prediction mode is horizontal (modes 2 to 34 in FIG. 3) based on the diagonal direction (mode 34 in FIG. 3), the conversion coefficients are arranged horizontally, that is, in row-major order as shown in FIG. 6(a), and when the intra prediction mode is vertical (modes 35 to 66 in FIG. 3) based on the diagonal direction, the conversion coefficients are arranged horizontally, that is, in column-major order as shown in FIG. 6(b).

[0447] In one embodiment, the encoding device determines whether it meets the conditions for applying LFNST, generates and encodes an LFNST index based on the determination, selects a conversion kernel matrix, and when it meets the conditions for applying LFNST, applies LFNST to the residual samples based on the selected conversion kernel matrix and / or a simplification factor. At this time, the size of the simplified conversion kernel matrix is determined based on the simplification factor.

[0448] On the other hand, by way of an example, the encoding device can zero out the second region of the current block where there are no modified conversion coefficients (S2150).

[0449] As shown in FIGS. 11 and 12, the remaining regions of the current block where there are no modified conversion coefficients can all be processed as 0. Such zeroing out can reduce the computational amount required for the execution of the overall conversion process, reduce the amount of operations required for the entire conversion process, and thus reduce the power consumption required for the execution of the conversion. Also, the latency associated with the conversion process can be reduced, and the efficiency of image coding can be increased.

[0450] Also, the encoding device can configure the image information such that the LFNST index is signaled based on that zeroing out is performed on all of the plurality of sub-partition blocks into which the current block is divided (S2160).

[0451] Also, by way of an example, the encoding device can configure the image information such that the LFNST index indicating the LFNST matrix is signaled based on the existence of conversion coefficients in the region excluding the DC position of the current block, and based on the current block being divided into a plurality of sub-partition blocks, can configure the image information such that the LFNST index is signaled regardless of whether there are conversion coefficients in the region excluding the DC position.

[0452] The encoding device configures the video information so that the video information shown in Table 8 can be parsed by the decoding device.

[0453] That is, when the current block is not divided into a plurality of sub - partition blocks and it is indicated that transform coefficients exist in the region excluding the DC position of the current block, the encoding device configures the video information so that the LFNST index is parsed. If the current block is divided into a plurality of sub - partition blocks, the encoding device configures the video information so that the LFNST index is parsed without checking whether transform coefficients exist in the region excluding the DC position.

[0454] That is, when ISP is applied to the current block, the encoding device configures the video information so that the LFNST index is signaled even if the position of the last non - zero coefficient for all sub - partition blocks is at the DC position.

[0455] By way of an example, when the index of the sub - block including the last valid coefficient in the current block (or sub - partition block) is 0 and the position of the last valid coefficient in the sub - block is greater than 0, the encoding device determines that the valid coefficients exist in the region excluding the DC position, and configures the video information so that the LFNST index is signaled. In this document, the first position in the scan order can be 0.

[0456] Also, by way of an example, when the index of the sub - block including the last valid coefficient in the current block (or sub - partition block) is greater than 0 and the width and height of the current block are 4 or more, the encoding device determines that it is certain that LFNST is not applied, and configures the video information so that the LFNST index is not signaled.

[0457] Also, in one example, when the size of the current block (or sub - partition block) is 4×4 or 8×8 and the start of the position in the scan order is from 0, if the position of the last valid coefficient is greater than 7, the encoding device determines that it is certain that LFNST is not applied, and configures the video information so that the LFNST index is not signaled.

[0458] That is, after the variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag are derived in the decoding device, the encoding device configures the video information so that the LFNST index is parsed according to the values of the derived variables.

[0459] The encoding device performs quantization based on the modified transform coefficients for the current block to derive quantized transform coefficients, and can encode and output image information including information on the quantized transform coefficients and LFNST index information indicating the LFNST matrix when LFNST can be applied (S2170).

[0460] The encoding device can generate residual information including information on the quantized transform coefficients. The residual information can include the above - mentioned transform - related information / syntax elements. The encoding device can encode the image / video information including the residual information and output it in the form of a bitstream.

[0461] More specifically, the encoding device 200 can generate information on the quantized transform coefficients and encode the generated information on the quantized transform coefficients.

[0462] The syntax element of the LFNST index according to this embodiment can indicate whether (inverse) LFNST is applied and any one of the LFNST matrices included in the LFNST set. When the LFNST set includes two transform kernel matrices, the value of the syntax element of the LFNST index can be three.

[0463] According to one example, when the split tree structure for the current block is of the dual tree type, an LFNST index is encoded for each of the luma block and the chroma block.

[0464] According to one embodiment, the syntax element value for the transform index is derived as 0 indicating that (inverse) LFNST is not applied to the current block, 1 indicating the first LFNST matrix among the LFNST matrices, and 2 indicating the second LFNST matrix among the LFNST matrices.

[0465] In this document, at least one of quantization / inverse quantization and / or transform / inverse transform may be omitted. When the quantization / inverse quantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may sometimes be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistency of representation.

[0466] Also, in this document, the quantized transform coefficients and the transform coefficients may be respectively referred to as transform coefficients and scaled transform coefficients. In this case, the residual information can include information about the transform coefficients, and the information about the transform coefficients can be signaled via the residual coding syntax. The transform coefficients can be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients can be derived via the inverse transform (scaling) for the transform coefficients. The residual samples can be derived based on the inverse transform (transform) for the scaled transform coefficients. This can be applied / expressed similarly in another part of this document.

[0467] In the foregoing embodiments, the method has been described based on a flowchart as a series of steps or blocks, but this document is not limited to the order of the steps, and a certain step may occur in a different order from the steps described above, or simultaneously with different steps. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of this document.

[0468] The method according to the foregoing document can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0469] In this document, when an embodiment is embodied in software, the foregoing method can be embodied by modules (processes, functions, etc.) that perform the foregoing functions. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), RAM (random access memory), flash memory, memory card, storage medium, and / or other storage devices. That is, the embodiments described in this document can be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, processor, microprocessor, controller, or chip.

[0470] In addition, the decoding device and encoding device to which this document is applicable may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videophone video device, and a medical video device, etc., and may be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.

[0471] In addition, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by the computer. Multimedia data having a data structure related to this document can also be stored in a recording medium readable by the computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by the computer is stored. The recording medium readable by the computer can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Further, the recording medium readable by the computer includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by the computer or can be transmitted via a wired or wireless communication network. Further, an embodiment of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to an embodiment of this document. The program code can be stored on a carrier readable by a computer.

[0472] FIG. 22 schematically shows an example of a video / image coding system to which this document is applicable.

[0473] Referring to FIG. 22, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or a stream.

[0474] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can be composed of a separate device or an external component.

[0475] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by the process of generating related data.

[0476] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0477] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in the form of a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0478] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0479] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0480] FIG. 23 exemplarily shows a structural diagram of a content streaming system to which this document is applied.

[0481] Also, the content streaming system to which this document is applied can largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0482] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by the encoding method or the bitstream generation method to which this document is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0483] The streaming server transmits multimedia data to the user device based on a user request via a web server, and the web server plays the role of a medium for informing the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At that time, the content streaming system can include another control server, and in this case, the control server plays the role of controlling commands / responses between each device in the content streaming system.

[0484] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0485] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, and the like. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

[0486] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and embodied as a device, and the technical features of the device claims in this specification can be combined and embodied as a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a method.

Claims

1. An image decoding method performed by a decoding device, obtaining residual information from the bitstream; deriving transform coefficients for a current block based on the residual information; deriving modified transform coefficients by applying a low-frequency non-separable transform (LFNST) to the transform coefficients; deriving residual samples for the current block based on an inverse linear transform of the modified transform coefficients; generating a reconstructed picture based on the residual samples; The step of deriving the modified transform coefficients comprises: determining whether there are any non-zero transform coefficients in a second region other than the first region, the second region being a region covering samples arranged in scan order from the top left position; parsing the LFNST index based on the result of said determining; deriving the modified transform coefficients based on the LFNST matrix indicated by the LFNST index; When an Intra Sub-Partition (ISP) is applied, the LFNST index is parsed based on the fact that the non-zero transform coefficient does not exist in the second region of each sub-partition block for all sub-partition blocks divided from the current block; the second region of each sub-partition block is a region other than the first region, which is the region covering samples arranged in the scan order from the upper left position of each sub-partition block; The LFNST index is parsed based on the absence of the non-zero transform coefficients in the second region of the current block even if the ISP is not applied; The second region of the current block is a region other than the first region, which is the region covering samples arranged in the scan order from the top-left position of the current block.

2. An image encoding method performed by an image encoding device, deriving a predicted sample for the current block; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients for the current block based on a linear transform for the residual samples; deriving modified transform coefficients for transform coefficients of a first region, the first region covering samples arranged in scan order from an upper left position based on a low-frequency non-separable transform (LFNST) matrix; zeroing out a second region where the modified transform coefficients are not present; encoding image information including an LFNST index associated with the LFNST matrix based on zeroing out the second region of each sub-partition block divided from the current block to which an Intra Sub-Partition (ISP) is applied; outputting the image information including the LFNST index and residual information related to quantized transform coefficients derived by quantizing the modified transform coefficients; the second region of each sub-partition block is a region other than the first region, which is the region covering samples arranged in the scan order from the upper left position of each sub-partition block; The image information including the LFNST index is encoded based on the zeroing out for the second region of the current block even if the ISP is not applied; The second region of the current block is a region other than the first region, which is the region covering samples arranged in the scan order from the top-left position of the current block.

3. A method for transmitting data for an image, comprising the steps of: Obtaining a bitstream for the image, comprising: The bitstream comprises: deriving a predicted sample for the current block; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients for the current block based on a linear transform for the residual samples; deriving modified transform coefficients for transform coefficients of a first region covering samples arranged in scan order from an upper-left position based on a low-frequency non-separable transform (LFNST) matrix; zeroing out a second region where the modified transform coefficients are not present; and encoding image information including an LFNST index associated with the LFNST matrix based on zeroing out the second region of each sub-partition block divided from the current block to which an Intra Sub-Partition (ISP) is applied; outputting the image information including the LFNST index and residual information related to quantized transform coefficients derived by quantizing the modified transform coefficients; transmitting the data including the bitstream; the second region of each sub-partition block is a region other than the first region, which is the region covering samples arranged in the scan order from the upper left position of each sub-partition block; The image information including the LFNST index is encoded based on the zeroing out for the second region of the current block even if the ISP is not applied; The second region of the current block is a region other than the first region, which is the region covering samples arranged in the scan order from the top-left position of the current block.