Transform-based image coding method, and device therefor

The video coding method and apparatus address the challenges of high-resolution video data by employing Multiple Transform Selection (MTS) and its index, resulting in improved coding efficiency and reduced costs.

JP2025096468AActive Publication Date: 2025-06-26LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025064169
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2025-04-09
Publication Date
2025-06-26
Estimated Expiration
2040-10-05

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video data, such as 4K or 8K UHD video, poses challenges in efficient compression, transmission, storage, and reproduction, leading to higher costs and complexities.

Method used

A video coding method and apparatus that utilizes Multiple Transform Selection (MTS) and its index to enhance coding efficiency, specifically by determining the applicability of MTS indices based on block types and zero-out conditions, and signaling these indices at the coding unit level.

Benefits of technology

The proposed solution increases overall video compression efficiency, improves transform index coding efficiency, and enables efficient use of MTS and its index in video coding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096468000001_ABST
    Figure 2025096468000001_ABST
Patent Text Reader

Abstract

To provide an image decoding method.SOLUTION: An image decoding method, according to the present document, comprises the steps of: determining whether to parse an MTS index for applying an MTS to the current block; deriving residual samples for the current block by applying the MTS to the current block on the basis of the MTS index; and generating a reconstructed picture on the basis of the residual samples, where the step of determining whether to parse the MTS index includes determining the tree type of the current block, the division type of the current block, and whether zero-out for the MTS has been performed in the current block.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video coding technology, and more particularly, to a video coding method and apparatus based on transform in a video coding system.

Background Art

[0002] Recently, the demand for high-resolution and high-quality video / video such as 4K or 8K or higher UHD (Ultra High Definition) video has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from those of real-world video, such as game video, has been increasing.

[0004] Accordingly, in order to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality video / video having various characteristics as described above, a highly efficient video / video compression technology is required.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and apparatus for increasing video coding efficiency.

[0006] Another technical problem of this document is to provide a method and apparatus for increasing the efficiency of transform index coding.

[0007] Another technical problem of this document is to provide a video coding method and apparatus utilizing MTS.

[0008] Another technical problem of this document is to provide a video coding method and apparatus utilizing an MTS index.

Means for Solving the Problem

[0009] According to an embodiment of this document, a video decoding method executed by a decoding apparatus is provided. The method includes: determining whether an MTS index for applying MTS to a current block can be parsed; deriving a residual sample for the current block by applying the MTS to the current block based on the MTS index; and generating a restored picture based on the residual sample. The step of determining whether the MTS index can be parsed is characterized by determining a tree type of the current block, a split type of the current block, and whether zero-out for the MTS has been executed on the current block.

[0010] When the tree type of the current block is not a dual-tree chroma and an LFNST index indicating an LFNST kernel applied to the current block is 0, the MTS index is parsed.

[0011] When a larger value of the width and height of the current block is less than or equal to 32, the MTS index is parsed.

[0012] When the current block is not divided into a plurality of sub-partition blocks and sub-block conversion for dividing a coding unit in the current block and performing conversion is not applied, the MTS index is parsed.

[0013] Whether zeroing out for the MTS is executed is determined by whether there is a valid coefficient in a second region excluding a first region at the upper left end where a valid transform coefficient can exist within the current block. If there is no such valid coefficient in the second region, the MTS index is parsed.

[0014] The LFNST index and the MTS index are signaled at the coding unit level, and the MTS index is signaled immediately after the signaling of the LFNST index.

[0015] According to an embodiment of this document, a video encoding method executed by an encoding device is provided. The method includes deriving a transform coefficient for the current block based on the MTS for the residual samples, and encoding a residual information derived through quantization of the transform coefficient and an MTS index indicating an MTS kernel. The MTS index is encoded based on the tree type of the current block, the split type of the current block, and whether zeroing out for the MTS is executed for the current block.

[0016] According to another embodiment of this document, a digital storage medium storing video data including encoded video information and a bitstream generated by a video encoding method executed by an encoding device is provided.

[0017] According to another embodiment of this document, a digital storage medium storing video data including encoded video information and a bitstream for causing a decoding device to execute the video decoding method is provided.

Advantages of the Invention

[0018] According to this document, the overall video / video compression efficiency can be increased.

[0019] According to this document, the efficiency of transform index coding can be improved.

[0020] According to this document, a video coding method and apparatus using MTS can be provided.

[0021] According to this document, a video coding method and apparatus using an MTS index can be provided.

[0022] The effects obtainable through a specific example of this specification are not limited to the effects listed above. For example, there can be various technical effects that can be understood or induced from this specification by a person having ordinary skill in the related art. Accordingly, the specific effects of this specification are not limited to what is explicitly described in this specification, and can include various effects that can be understood or induced from the technical features of this specification.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Modes for Carrying Out the Invention

[0024] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms used in this specification are merely used to explain specific embodiments and are not intended to limit the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In this specification, terms such as "including" or "having" are used to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc. are excluded in advance.

[0025] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among each configuration, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0026] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and repeated descriptions of the same components will be omitted.

[0027] This document relates to video / video coding. For example, the methods / examples disclosed in this document are associated with the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (such as the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0028] This document presents various examples related to video / video coding, and unless otherwise mentioned, the examples can also be executed in combination with each other.

[0029] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.

[0030] A pixel or pel can mean the smallest unit that makes up a picture (or video). Also, the term "sample" can be used as the term corresponding to a pixel. A sample can generally indicate a pixel or the value of a pixel, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Or, a sample can also mean the pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean the conversion coefficient in the frequency domain.

[0031] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0032] In this document, " / " and "、" are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" also means "at least one of A, B, and / or C."

[0033] Further, in this document, "or" is interpreted as "and / or". For example, "A or B" can mean 1) only "A", or 2) only "B", or 3) "A and B". As another expression, "or" in this document can mean "additionally or alternatively".

[0034] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0035] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0036] Also, the parentheses used in this specification can mean "for example". Specifically, when it is shown as "prediction (intra-prediction)", "intra-prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "intra-prediction", but "intra-prediction" is proposed as an example of "prediction". Also, when it is shown as "prediction (i.e., intra-prediction)", "intra-prediction" is proposed as an example of "prediction".

[0037] In this specification, the technical features separately described within one drawing can be embodied individually or simultaneously.

[0038] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0039] Referring to FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in a file or streaming form via a digital storage medium or a network.

[0040] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.

[0041] The video source can obtain video / image through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / image, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / image. For example, virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by the process of generating related data.

[0042] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0043] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0044] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0045] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0046] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the video encoding device can include the image encoding device.

[0047] Referring to FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware components can further include the memory 270 as an internal / external component.

[0048] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and the binary-tree structure and / or the ternary structure can be applied thereafter. Or, the binary-tree structure can also be applied first. The coding procedure according to this document can be executed based on the final coding unit that cannot be divided further. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit is a unit of sample prediction, and the conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0049] The unit can, in some cases, be used interchangeably with terms such as a block or an area. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).

[0050] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0051] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0052] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be executed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal is not transmitted.In the case of the motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled, so that the motion vector of the current block can be indicated.

[0053] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for predicting a block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can execute an intra block copy (IBC) for predicting a block. The intra block copy can be used for content video / moving video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed to be similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0054] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square, or can be applied to a block of variable size that is not square.

[0055] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.

[0056] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0057] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.

[0058] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0059] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0060] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information within the current picture has been derived (or encoded) and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 222.

[0061] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.

[0062] Referring to FIG. 3, the decoding apparatus 300 may be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 332 and an intra predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Also, the memory 360 may include a DPB (decoded picture buffer) and may also be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0063] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit by a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.

[0064] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded adjacent to and the decoding information of the block to be decoded or the symbol / bin information decoded in the previous step, predicts the occurrence probability of the bin by the determined context model, and generates a symbol corresponding to the value of each syntax element by performing arithmetic decoding of the bin.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can also be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0065] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0066] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).

[0067] The prediction unit can perform prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0068] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed to be similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0069] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0070] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0071] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0072] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, it can also be output after filtering, or can be used for inter prediction of the next picture.

[0073] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.

[0074] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can generate a modified restored picture by applying various filtering methods to the restored picture, and can send the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0075] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block from which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0076] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied to be the same or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200, respectively.

[0077] As described above, in performing video coding, prediction is performed to increase the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is also derived in the encoding apparatus and the decoding apparatus, and the encoding apparatus can increase the video coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0078] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0079] FIG. 4 exemplarily shows a content streaming system structure diagram to which this document is applied.

[0080] Also, the content streaming system to which this document is applied can greatly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0081] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by an encoding method or a bitstream generation method to which this document is applicable, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0082] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server plays the role of a medium that informs the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device within the content streaming system.

[0083] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0084] Examples of the user device include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed distributively.

[0085] Figure 5 schematically shows the multiple conversion technique according to this document.

[0086] Referring to Figure 5, the conversion unit can correspond to the conversion unit in the encoding device of Figure 2 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of Figure 2 or the inverse conversion unit in the decoding device of Figure 3 described above.

[0087] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S510). Such a primary conversion can be called a core transform. Here, the primary conversion can be based on Multiple Transform Selection (MTS), and when multiple conversion is applied as the primary conversion, it can be called a multiple core transform.

[0088] The multi-core transform can be shown as a method of performing a transform by additionally using DCT (Discrete Cosine Transform) type 2 and DST (Discrete Sine Transform) type 7, DCT type 8, and / or DST type 1. That is, the multi-core transform can be shown as a transform method that transforms a residual signal (or a residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on a plurality of selected transform kernels from among DCT type 2, DST type 7, DCT type 8, and DST type 1. Here, the primary transform coefficients can be referred to as temporary transform coefficients from the perspective of the transform unit.

[0089] That is, when an existing transform method is applied, based on DCT type 2, a transform from the spatial domain to the frequency domain for the residual signal (or the residual block) can be applied to generate transform coefficients. In contrast, when the multi-core transform is applied, based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., a transform from the spatial domain to the frequency domain for the residual signal (or the residual block) can be applied to generate transform coefficients (or primary transform coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc., can be referred to as transform types, transform kernels, or transform cores. Such DCT / DST transform types can be defined based on basis functions.

[0090] When the multi-core transformation is executed, a vertical transformation kernel and a horizontal transformation kernel for a target block can be selected from among the transformation kernels, a vertical transformation for the target block can be executed based on the vertical transformation kernel, and a horizontal transformation for the target block can be executed based on the horizontal transformation kernel. Here, the horizontal transformation can indicate a transformation for a horizontal component of the target block, and the vertical transformation can indicate a transformation for a vertical component of the target block. The vertical transformation kernel / horizontal transformation kernel can be adaptively determined based on a prediction mode and / or a transformation index of a target block (CU or sub-block) including a residual block.

[0091] Also, according to an example, when applying MTS to execute a primary transformation, a specific basis function is set to a predetermined value, and when it is a vertical transformation or a horizontal transformation, a mapping relationship for the transformation kernel can be set by combining which basis function is applied. For example, when the horizontal direction transformation kernel is represented by trTypeHor and the vertical direction transformation kernel is represented by trTypeVer, the trTypeHor or trTypeVer value 0 can be set to DCT2, the trTypeHor or trTypeVer value 1 can be set to DST7, and the trTypeHor or trTypeVer value 2 can be set to DCT8.

[0092] In this case, in order to indicate any one of a number of conversion kernel sets, MTS index information can be encoded and signaled to the decoding device. For example, when the MTS index is 0, it indicates that both the trTypeHor and trTypeVer values are 0; when the MTS index is 1, it indicates that both the trTypeHor and trTypeVer values are 1; when the MTS index is 2, it indicates that the trTypeHor value is 2 and the trTypeVer value is 1; when the MTS index is 3, it indicates that the trTypeHor value is 1 and the trTypeVer value is 2; when the MTS index is 4, it can indicate that both the trTypeHor and trTypeVer values are 2.

[0093] By way of an example, the conversion kernel sets according to the MTS index information are shown in a table as follows.

[0094] [Table 1]

[0095] The conversion unit can derive a corrected (secondary) conversion coefficient by performing a secondary conversion based on the (primary) conversion coefficient (S520). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion means converting in a more compressive representation by utilizing the correlation existing between the (primary) conversion coefficients. The secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can represent a transform that performs a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, based on the non-separable transform matrix, the vertical conversion and the horizontal conversion can be applied to the (primary) conversion coefficient at once without separating them (or independently applying the horizontal and vertical conversions). That is, the non-separable secondary transform is not separately applied to the (primary) conversion coefficient in the vertical and horizontal directions. For example, after rearranging a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specified direction (e.g., row-first direction or column-first direction), it can represent a conversion method that generates a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix. For example, the row-first order is to arrange the rows of an M×N block in a column in the order of the first row, the second row,..., the Nth row, and the column-first order is to arrange the columns of an M×N block in a column in the order of the first column, the second column,..., the Mth column. The non-separable secondary transform can be applied to the top-left region of a block composed of (primary) conversion coefficients (hereinafter, which can be called a conversion coefficient block).For example, when both the width (W) and height (H) of the conversion coefficient block are 8 or more, an 8×8 non-separable second-order conversion can be applied to the upper left 8×8 region of the conversion coefficient block. Also, when both the width (W) and height (H) of the conversion coefficient block are 4 or more and the width (W) or height (H) of the conversion coefficient block is less than 8, a 4×4 non-separable second-order conversion can be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that both the width (W) or height (H) of the conversion coefficient block is 4 or more is satisfied, a 4×4 non-separable second-order conversion can also be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block.

[0096] Specifically, for example, when a 4×4 input block is used, the non-separable second-order conversion can be performed as follows.

[0097] The 4×4 input block X is shown as follows.

[0098]

Equation

[0099] When representing the above X in vector form, the vector JPEG2025096468000004.jpg84 is shown as follows.

[0100]

Equation

[0101] As shown in Equation 2, the vector JPEG2025096468000006.jpg84 rearranges the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.

[0102] In this case, the secondary non-separable transform can be calculated as follows.

[0103]

Equation

[0104] Here, JPEG2025096468000008.jpg75 represents the transform coefficient vector, and T represents the 16×16 (non-separable) transform matrix.

[0105] The 16×1 transform coefficient vector JPEG2025096468000009.jpg75 can be derived through the above Equation 3, and the JPEG2025096468000010.jpg75 can be re-organized in 4×4 blocks through the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is only an example, and HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of the non-separable secondary transform to reduce the computational complexity of the non-separable secondary transform.

[0106] On the other hand, for the non-separable secondary transform, a mode-based transform kernel (or transform core, transform type) can be selected. Here, the mode can include the intra prediction mode and / or the inter prediction mode.

[0107] As described above, the non-separable second-order transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 area included within the corresponding transform coefficient block when both W and H are greater than or equal to 8, and the corresponding 8×8 area is the upper-left 8×8 area within the corresponding transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 area included within the corresponding transform coefficient block when both W and H are greater than or equal to 4, and the corresponding 4×4 area is the upper-left 4×4 area within the corresponding transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.

[0108] At this time, for mode-based transform kernel selection, two non-separable second-order transform kernels can be configured for each set of non-separable second-order transforms for both the 8×8 transform and the 4×4 transform, and the number of transform sets is four. That is, four transform sets can be configured for the 8×8 transform, and four transform sets can be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform can include two 8×8 transform kernels, and in this case, each of the four transform sets for the 4×4 transform can include two 4×4 transform kernels.

[0109] However, the size of the transform, that is, the size of the area to which the transform is applied, is merely exemplary, and sizes other than 8×8 or 4×4 can be used, the number of the sets is n, and the number of transform kernels within each set is k.

[0110] The transformation set can be referred to as an NSST set or an LFNST set. The selection of a specific set from the transformation sets can be performed, for example, based on the intra prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) is an example of a reduced non-separable transform described later and represents a non-separable transform for low-frequency components.

[0111] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is only an example, and this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can represent the LM (linear model) mode.

[0112] FIG. 6 exemplarily shows the intra-directional modes of 65 prediction directions.

[0113] Referring to FIG. 6, around the 34th intra prediction mode having a lower right diagonal prediction direction, an intra prediction mode having horizontal directionality and an intra prediction mode having vertical directionality can be distinguished. H and V in FIG. 6 respectively represent horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in units of 1 / 32 on the sample grid position. This can indicate an offset with respect to the mode index value. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. On the other hand, the 34th intra prediction mode can be regarded as not strictly having either horizontal or vertical directionality, but can be classified as belonging to the horizontal directionality from the perspective of determining the conversion set of the secondary conversion. This is because for the symmetric vertical direction modes centered around the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the input data alignment method for the horizontal direction mode is used. Transposing the input data means that for the two-dimensional block data M×N, the rows become columns and the columns become rows to form N×M data. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode has a left reference pixel and predicts in the upper right direction, so it can be called an upper right diagonal intra prediction mode. In the same context, the 34th intra prediction mode is called a lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called a lower left diagonal intra prediction mode.

[0114] By way of example, the mapping of four conversion sets by the intra prediction mode is shown, for example, as in the following table.

[0115]

Table 2

[0116] As shown in Table 2, according to the intra prediction mode, any one of the four transform sets, that is, lfnstTrSetIdx can be mapped to any one of 0 to 3, that is, any one of the four.

[0117] On the other hand, when it is determined that a specific set is used for the non-separable transform, one of the k transform kernels in the specific set can be selected via the non-separable second-order transform index. The encoding device can derive a non-separable second-order transform index indicating a specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable second-order transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable second-order transform index. For example, the lfnst index value 0 can indicate the first non-separable second-order transform kernel, the lfnst index value 1 can indicate the second non-separable second-order transform kernel, and the lfnst index value 2 can indicate the third non-separable second-order transform kernel. Or, the lfnst index value 0 can indicate that the first non-separable second-order transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate the three transform kernels.

[0118] The transform unit can execute the non-separable second-order transform based on the selected transform kernel to obtain the modified (second-order) transform coefficients. The modified transform coefficients can be derived as the transform coefficients quantized via the quantization unit as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0119] On the one hand, as described above, when the secondary conversion is omitted, the (primary) conversion coefficient, which is the output of the primary (separation) conversion, can be derived as the quantization coefficient quantized via the quantization unit as described above, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse conversion unit in the encoding device.

[0120] The inverse conversion unit can execute a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficient, executes the secondary (inverse) conversion to derive the (primary) conversion coefficient (S550), and can execute the primary (inverse) conversion on the (primary) conversion coefficient to obtain the residual block (residual samples, etc.) (S560). Here, the primary conversion coefficient can be called the modified conversion coefficient from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0121] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). The secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion is NSST, RST, or LFNST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed from the bitstream. As another example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficient of the residual block.

[0122] The second inverse transform determination unit can determine the second inverse transform. At this time, the second inverse transform determination unit can determine the second inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra prediction mode. Also, as an example, the second transform determination method can be determined depending on the first transform determination method. Various combinations of the first transform and the second transform can be determined by the intra prediction mode. Also, as an example, the second inverse transform determination unit can also determine the area to which the second inverse transform is applied based on the size of the current block.

[0123] On the other hand, as described above, when the second (inverse) transform is omitted, a residual block (residual sample) can be obtained by receiving the (inverse quantized) transform coefficient and performing the first (separation) inverse transform. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0124] On the other hand, in this document, in order to reduce the amount of calculation and memory requirements due to the non-separable second transform, an RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced with the concept of NSST can be applied.

[0125] On the one hand, the conversion kernel, conversion matrix, and coefficients constituting the conversion kernel matrix described in this document, that is, the kernel coefficients or matrix coefficients, can be represented in 8 bits. This is one of the conditions for implementation in a decoding device and an encoding device, and it can reduce the memory requirement for storing the conversion kernel while reasonably accepting a performance degradation compared to existing 9-bit or 10-bit ones. Also, by representing the kernel matrix in 8 bits, a small multiplier can be used, and it can be more suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.

[0126] In this specification, RST can mean a conversion performed on the residual samples for the target block based on a transform matrix whose size has been reduced by a simplification factor. When performing a simplified conversion, the amount of computation required during conversion can be reduced due to the reduction in the size of the transform matrix. That is, RST can be used to solve the problem of computational complexity that occurs during the conversion of large-sized blocks or non-separable conversions.

[0127] RST can be called by various terms such as reduced conversion, reduced transform, reduced secondary transform, reduction transform, simplified transform, simple transform, etc., and the name called RST is not limited to the listed examples. Or, since RST is mainly performed in the low-frequency region including non-zero coefficients in the conversion block, it can also be called LFNST (Low-Frequency Non-Separable Transform). The conversion index can be named the LFNST index.

[0128] On the other hand, when the second inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 may include an inverse RST unit that derives a modified transformation coefficient based on the inverse RST for the transformation coefficient, and an inverse primary transformation unit that derives a residual sample for the target block based on the inverse primary transformation for the modified transformation coefficient. The inverse primary transformation means the inverse transformation of the primary transformation applied to the residue. In this document, deriving a transformation coefficient based on a transformation can mean applying the corresponding transformation to derive the transformation coefficient.

[0129] FIG. 7 is a diagram for explaining the RST according to an embodiment of this document.

[0130] In this specification, the "target block" can mean the current block or the residual block or the transformation block on which coding is performed.

[0131] In the RST according to an embodiment, a reduced transformation matrix can be determined in which an N-dimensional vector is mapped to an R-dimensional vector located in a different space, where R is smaller than N. N can mean the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value. The simplification factor can be called by various terms such as a reduced factor, a reduction factor, a reduced coefficient, a reduction coefficient, a simplified factor, a simple factor, etc. On the other hand, R can be called a reduced coefficient, but in some cases, the simplification factor can also mean R. Also, in some cases, the simplification factor can mean the N / R value.

[0132] In one embodiment, the simplification factor or coefficient can be signaled via a bitstream, but the embodiments are not limited thereto. For example, there may be cases where predefined values for the simplification factor or coefficient are stored in each encoding device 200 and decoding device 300. In this case, the simplification factor or coefficient is not signaled separately.

[0133] The size of the simplification transform matrix according to one embodiment is R×N, which is smaller than the size N×N of the normal transform matrix, and can be defined as in Equation 4 below.

[0134]

Equation

[0135] The matrix T in the Reduced Transform block shown in FIG. 7(a) can represent the matrix TR×N of Equation 4. As shown in FIG. 7(a), when the simplification transform matrix TR×N is multiplied by the residual samples for the target block, the transform coefficients for the target block can be derived.

[0136] In one embodiment, when the size of the block to which the transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7(a) can be expressed as a matrix operation as in Equation 5 below. In this case, the multiplication operation with the memory can be reduced by approximately 1 / 4 due to the simplification factor.

[0137] In this document, matrix multiplication can be understood as an operation of obtaining a column vector by placing a matrix on the left side of a column vector and multiplying the matrix and the column vector.

[0138]

Equation

[0139] In Equation 5, r1 to r 64 can represent the residual samples for the target block, and more specifically, they are the transform coefficients generated by applying a first-order transform. The operation result of Equation 5, the transform coefficient c i for the target block can be derived, and the derivation process of c i is as shown in Equation 6.

[0140] [Number]

[0141] The operation result of Equation 6, the transform coefficients c1 to c R for the target block can be derived. That is, when R = 16, the transform coefficients c1 to c 16 for the target block can be derived. If a normal (regular) transform instead of RST is applied and a transform matrix with a size of 64×64 (N×N) is multiplied by a residual sample with a size of 64×1 (N×1), 64 (N) transform coefficients for the target block are derived. However, because RST is applied, only 16 (R) transform coefficients for the target block are derived. Since the total number of transform coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 and the decoding device 300 can be increased.

[0142] From the perspective of the size of the transform matrix, the size of the normal transform matrix is 64×64 (N×N), and the size of the simplified transform matrix decreases to 16×64 (R×N). Therefore, when compared with the case of performing a normal transform, the memory usage can be decreased by the ratio of R / N when performing RST. Also, when compared with the number of multiplication operations N×N when using a normal transform matrix, when using a simplified transform matrix, the number of multiplication operations can be decreased by the ratio of R / N (R×N).

[0143] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual samples for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300, and the inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive residual samples for the target block based on an inverse primary conversion for the modified conversion coefficients.

[0144] Inverse RST matrix T according to one embodiment N×R has a size of N×R, which is smaller than the size N×N of the normal inverse conversion matrix, and is in a transpose relationship with the simplified conversion matrix T shown in Equation 4. R×N

[0145] The matrix T in the Reduced Inv.Transform block shown in FIG. 7(b) t can mean the inverse RST matrix T R×N T (the superscript T means transpose). As shown in FIG. 7(b), when the inverse RST matrix T R×N T is multiplied by the conversion coefficients for the target block, modified conversion coefficients for the target block or residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R

[0146] More specifically, when inverse RST is applied as the secondary inverse conversion, the inverse RST matrix T R×N TWhen applied, a modified conversion coefficient for the target block can be derived. On the other hand, inverse RST can be applied as an inverse linear transformation. In this case, the inverse RST matrix T R×N T is multiplied to derive the residual samples for the target block.

[0147] In one embodiment, when the size of the block to which the inverse transformation is applied is 8×8 and R = 16 (i.e., when R / N = 16 / 64 = 1 / 4), the RST according to FIG. 7(b) can be expressed by matrix operations as shown in Equation 7 below.

[0148]

Equation

[0149] In Equation 7, c1 to c 16 can indicate the conversion coefficients for the target block. The calculation result of Equation 7, r i indicating the modified conversion coefficient for the target block or the residual samples for the target block can be derived, and the derivation process of r i is as shown in Equation 8.

[0150]

Equation

[0151] r1 to r NIt can be derived. Considering from the perspective of the size of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64×64 (N×N), and the size of the simplified inverse transformation matrix is reduced to 64×16 (N×R). Therefore, when comparing with the time of performing the normal inverse transformation, the memory usage can be reduced by the ratio of R / N when performing the inverse RST. Also, when comparing with the number of multiplication operations N×N when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced by the ratio of R / N (N×R).

[0152] On the other hand, for the 8×8 RST as well, a conversion set configuration as shown in Table 2 can be applied. That is, the corresponding 8×8 RST can be applied by the conversion set in Table 2. Since one conversion set is composed of two or three conversions (kernels) depending on the prediction mode within the screen, it can be configured to select one from a maximum of four conversions including not applying the secondary conversion. The conversion when the secondary conversion is not applied can be regarded as the identity matrix being applied. When indexes 0, 1, 2, and 3 are assigned to each of the four conversions (for example, the 0th index can be assigned to the identity matrix, that is, when the secondary conversion is not applied), a syntax element called the conversion index or lfnst index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, through the conversion index, for the 8×8 upper left block, in the RST configuration, 8×8 RST can be specified, or when LFNST is applied, 8×8 lfnst can be specified. 8×8 lfnst and 8×8 RST refer to conversions that can be applied to the 8×8 area included within the corresponding conversion coefficient block when both the W and H of the target block to be converted are greater than or equal to 8, and the corresponding 8×8 area is the upper left 8×8 area within the corresponding conversion coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to conversions that can be applied to the 4×4 area included within the corresponding conversion coefficient block when both the W and H of the target block are greater than or equal to 4, and the corresponding 4×4 area is the upper left 4×4 area within the corresponding conversion coefficient block.

[0153] In one aspect, according to an embodiment of this document, in the conversion of the encoding process, for 64 data constituting an 8×8 region, only 48 data that are not a 16×64 conversion kernel matrix can be selected and a maximum 16×48 conversion kernel matrix can be applied. Here, "maximum" means that the maximum value of m for an m×48 conversion kernel matrix that can generate m coefficients is 16. That is, when performing RST by applying an m×48 conversion kernel matrix (m≦16) to an 8×8 region, an input of 48 data can be received and m coefficients can be generated. When m is 16, an input of 48 data is received and 16 coefficients are generated. That is, when 48 data form a 48×1 vector, a 16×1 vector can be generated by multiplying a 16×48 matrix and a 48×1 vector in sequence. At this time, 48 data forming an 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be formed based on 48 data constituting a region excluding the lower right 4×4 region of an 8×8 region. At this time, when performing matrix operation by applying a maximum 16×48 conversion kernel matrix, 16 modified conversion coefficients are generated, and the 16 modified conversion coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with 0s.

[0154] For the inverse transformation in the decoding process, the transposed matrix of the above-described conversion kernel matrix can be used. That is, when inverse RST or LFNST is performed as an inverse transformation process executed in a decoding device, the input coefficient data to which inverse RST is applied is composed of a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the inverse RST matrix corresponding to the one-dimensional vector on the left side can be arranged in a two-dimensional block according to a predetermined arrangement order.

[0155] When sorted out, in the conversion process, when RST or LFNST is applied to an 8×8 area, among the conversion coefficients of the 8×8 area, matrix multiplication of 48 conversion coefficients in the upper left, upper right, and lower left areas excluding the lower right area of the 8×8 area and a 16×48 conversion kernel matrix is performed. For the matrix multiplication, the 48 conversion coefficients are input in a one-dimensional array. When such matrix multiplication is performed, 16 modified conversion coefficients are derived, and the modified conversion coefficients can be arranged in the upper left area of the 8×8 area.

[0156] Conversely, in the inverse conversion process, when inverse RST or LFNST is applied to an 8×8 area, 16 conversion coefficients corresponding to the upper left of the 8×8 area among the conversion coefficients of the 8×8 area can be input in a one-dimensional array form according to the scanning order and matrix-multiplied with a 48×16 conversion kernel matrix. That is, the matrix multiplication in such a case can be represented as (48×16 matrix)*(16×1 conversion coefficient vector)=(48×1 modified conversion coefficient vector). Here, since the n×1 vector can be interpreted in the same sense as an n×1 matrix, it can also be denoted as an n×1 column vector. Also, * means matrix multiplication operation. When such matrix multiplication is performed, 48 modified conversion coefficients can be derived, and the 48 modified conversion coefficients can be arranged in the upper left, upper right, and lower left areas excluding the lower right area of the 8×8 area.

[0157] On the other hand, when the second inverse conversion is performed based on RST, the inverse conversion unit 235 of the encoding device 200 and the inverse conversion unit 322 of the decoding device 300 can include an inverse RST unit that derives modified conversion coefficients based on inverse RST for the conversion coefficients, and an inverse primary conversion unit that derives residual samples for the target block based on inverse primary conversion for the modified conversion coefficients. The inverse primary conversion means the inverse conversion of the primary conversion applied to the residue. In this document, deriving conversion coefficients based on a conversion can mean deriving conversion coefficients by applying the corresponding conversion.

[0158] Looking specifically at the described non-separable transform, LFNST, it is as follows. LFNST can include a forward transform by an encoding device and an inverse transform by a decoding device.

[0159] The encoding device takes as input the result (or a part of the result) derived after applying a forward primary (core) transform, and applies a forward secondary transform.

[0160] [Equation 9] y = G T x

[0161] In Equation 9 above, x and y are the input and output of the secondary transform respectively, G is a matrix representing the secondary transform, and the transform basis vector is composed of column vectors. In the case of the inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of the forward LFNST, G is the transpose of the matrix G. T becomes the dimension.

[0162] In the case of the inverse LFNST, the dimension of the matrix G is [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are submatrices obtained by sampling 8 transform basis vectors from the left side of the [48×16] matrix and the [16×16] matrix respectively.

[0163] On the other hand, in the case of the forward LFNST, the matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are submatrices obtained by sampling 8 transform basis vectors from the upper side of the [16×48] matrix and the [16×16] matrix respectively.

[0164] Therefore, in the case of the forward LFNST, the input x can be a [48×1] vector or a [16×1] vector, and the output y can be a [16×1] vector or an [8×1] vector. Since the output of the forward first-order transform in video coding and decoding is two-dimensional (2D) data, in order to form a [48×1] vector or a [16×1] vector as the input x, the 2D data that is the output of the forward transform must be appropriately arranged to form a one-dimensional vector.

[0165] FIG. 8 is a diagram showing the order of arranging the output data of the forward first-order transform into a one-dimensional vector by way of an example. The left diagrams of FIGS. 8(a) and 8(b) show the order for creating a [48×1] vector, and the right diagrams of FIGS. 8(a) and 8(b) show the order for creating a [16×1] vector. In the case of LFNST, the 2D data can be sequentially arranged in the order as shown in FIGS. 8(a) and 8(b) to obtain a one-dimensional vector x.

[0166] The arrangement direction of the output data of such a forward first-order transform can be determined by the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of FIG. 8(a), and when the intra prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of FIG. 8(b).

[0167] By way of an example, an arrangement order different from the arrangement orders of FIGS. 8(a) and 8(b) can be applied. When attempting to derive the same result (y vector) as when the arrangement orders of FIGS. 8(a) and 8(b) are applied, the column vectors of the matrix G can be rearranged according to the corresponding arrangement order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.

[0168] Since the output y derived through Equation 9 is a one-dimensional vector, if a configuration that processes the result of the forward second-order transformation as input, for example, a configuration that performs quantization or residual coding, requires two-dimensional data as input data, the output y vector of Equation 9 must be appropriately arranged as two-dimensional data again.

[0169] Figure 9 is a diagram showing the order of arranging the output data of the forward second-order transformation into a two-dimensional block by way of an example.

[0170] In the case of LFNST, it can be arranged in a 2D block according to a determined scan order. (a) of Figure 9 shows that when the output y is a [16×1] vector, the output values are arranged in the 16 positions of the two-dimensional block in a diagonal scan order. (b) of Figure 9 shows that when the output y is an [8×1] vector, the output values are arranged in the 8 positions of the two-dimensional block in a diagonal scan order, and the remaining 8 positions are filled with 0s. X in (b) of Figure 9 indicates being filled with 0s.

[0171] In other examples, since the order in which the output vector y is processed by a configuration that performs quantization or residual coding can be executed according to a preset order, the output vector y may not be arranged in a 2D block as shown in Figure 9. However, in the case of residual coding, data coding can be executed in units of 2D blocks such as CG (Coefficient Group) (for example, 4×4), and in this case, the data can be arranged in a specific order such as the diagonal scan order of Figure 9.

[0172] On the other hand, the decoding device can arrange the two-dimensional data output through an inverse quantization process or the like for inverse transformation in a preset scan order to form a one-dimensional input vector y. The input vector y can be output as an input vector x according to the following equation.

[0173] 〔Equation 10〕 x = Gy

[0174] In the case of the reverse-direction LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. In the case of the reverse-direction LFNST, the output vector x is a [48×1] vector or a [16×1] vector.

[0175] The output vector x is arranged in 2D blocks and arrayed as 2D data in the order shown in FIG. 8, and such 2D data becomes the input data (or part of the input data) for the inverse 1D transform.

[0176] Therefore, the inverse 2D transform is overall the opposite of the forward 2D transform process. In the case of the inverse transform, different from the forward direction, the inverse 1D transform is applied after applying the inverse 2D transform first.

[0177] In the reverse-direction LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transformation matrix G. Which matrix of the [48×16] matrix and the [16×16] matrix to apply is determined by the block size as well.

[0178] Also, the eight matrices can be derived from four transformation sets as in Table 2 described above, and each transformation set can be composed of two matrices. Which of the four transformation sets to use is determined by the intra prediction mode, and more specifically, the transformation set is determined based on the intra prediction mode value extended considering up to the Wide Angle Intra Prediction (WAIP). Which matrix to select from the two matrices constituting the selected transformation set is derived via index signaling. More specifically, the index values that can be transmitted are 0, 1, and 2. 0 indicates not to apply LFNST, and 1 and 2 can indicate either one of the two transformation matrices constituting the transformation set selected based on the intra prediction mode value.

[0179] FIG. 10 is a diagram showing the wide angle intra prediction mode according to an embodiment of the present document.

[0180] General intra prediction mode values can have values from 0 to 66 and from 81 to 83. As shown in the figure, the intra prediction mode values extended by WAIP can have values from -14 to 83. The values from 81 to 83 refer to the CCLM (Cross Compoonent Linear Model) mode, and the values from -14 to -1 and from 67 to 80 refer to the intra prediction mode values extended by applying WAIP.

[0181] When the width of the predicted current block is larger than the height, generally the upper reference pixels are closer to the position inside the block to be predicted. Therefore, predicting in the bottom-left direction is more accurate than predicting in the top-right direction. On the contrary, when the height of the block is larger than the width, the left reference pixels are generally closer to the position inside the block to be predicted. Therefore, predicting in the top-right direction is more accurate than predicting in the bottom-left direction. Therefore, it is advantageous to apply remapping with the index of the wide-angle intra prediction mode, i.e., mode index conversion.

[0182] When wide-angle intra prediction is applied, information for existing intra prediction can be signaled, and after the information is parsed, the information can be remapped with the index of the wide-angle intra prediction mode. Therefore, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) is not changed, i.e., the total number of intra prediction modes is 67, and the intra prediction mode coding for the specific block is not changed.

[0183] Table 3 below shows the process of deriving the intra mode modified by remapping the intra prediction mode to the wide-angle intra prediction mode.

[0184]

Table 3

[0185] In Table 3, the intra prediction mode value finally extended to the predModeIntra variable is stored. ISP_NO_SPLIT indicates that the CU block is not split into sub - partitions by the Intra Sub Partitions (ISP) technology currently adopted in the VVC standard. The fact that the cIdx variable value is 0, 1, or 2 refers to the cases of the luma, Cb, and Cr components respectively. The Log2 function shown in Table 3 returns a logarithm value with a base of 2, and the Abs function returns an absolute value.

[0186] As input values for the wide - angle intra prediction mode mapping process, variables such as predModeIntra indicating the intra prediction mode, the height and width of the transform block, etc. are used, and the output value is the modified intra prediction mode (predModeIntra). The height and width of the transform block or coding block can be the height and width of the current block for the remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to height can be set to Abs(Log2(nW / nH)).

[0187] For non - square blocks, the intra prediction mode can be classified and modified in two cases.

[0188] First, when all of the following conditions are satisfied: (1) the width of the current block is greater than the height, (2) the intra prediction mode before correction is greater than or equal to 2, and (3) the intra prediction mode is less than (whRatio>1)?(8 + 2*whRatio):8, the intra prediction mode is set to a value 65 greater than the intra prediction mode [predModeIntra is set equal to (predModeIntra + 65)].

[0189] Otherwise, when all of the following conditions are satisfied: (1) the height of the current block is greater than the width, (2) the intra prediction mode before correction is less than or equal to 66, and (3) the intra prediction mode is greater than (whRatio>1)?(60 - 2*whRatio):60, the intra prediction mode is set to a value 67 less than the intra prediction mode [predModeIntra is set equal to (predModeIntra - 67)].

[0190] Table 2 described above shows how the conversion set is selected based on the intra prediction mode values extended by WAIP in LFNST. As shown in FIG. 10, the modes from 14 to 33 and the modes from 35 to 80 are symmetric with respect to the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric around the direction corresponding to mode 34. Therefore, the same conversion set is applied to the modes located in the symmetric directions, and such symmetry is also reflected in Table 2.

[0191] However, it is assumed that the forward LFNST input data for mode 54 is symmetric with the forward LFNST input data for mode 14. For example, for mode 14 and mode 54, the two-dimensional data is rearranged into one-dimensional data according to the array orders shown in FIGS. 8(a) and 8(b), respectively, and it can be known that the pattern of the orders shown in FIGS. 8(a) and 8(b) is symmetric about the direction (diagonal term) indicated by mode 34.

[0192] On the other hand, as described above, which transformation matrix among the [48×16] matrix and the [16×16] matrix is applied to the LFNST is determined according to the size of the block to be transformed.

[0193] FIG. 11 is a diagram showing the blocks to which the LFNST is applied. FIG. 11(a) shows a 4×4 block, FIG. 11(b) shows 4×8 and 8×4 blocks, FIG. 11(c) shows 4×N or N×4 blocks where N is 16 or more, FIG. 11(d) shows an 8×8 block, and FIG. 11(e) shows an M×N block where M≧8, N≧8, and N〉8 or M〉8.

[0194] In FIG. 11, the blocks with thick frames indicate the regions to which the LFNST is applied. For the blocks in FIGS. 11(a) and 11(b), the LFNST is applied to the top-left 4×4 region, and for the block in FIG. 11(c), the LFNST is applied to each of the two continuously arranged top-left 4×4 regions. In FIGS. 11(a), 11(b), and 11(c), since the LFNST is applied in units of 4×4 regions, such an LFNST is hereinafter named "4×4 LFNST", and in the corresponding transformation matrix, a matrix dimension for G in Formulas 9 and 10 can be applied based on a [16×16] or [16×8] matrix.

[0195] More specifically, for the 4×4 block ((a) in FIG. 11, 4×4 TU or 4×4 CU), a [16×8] matrix is applied, and for the blocks in FIGS. 11(b) and (c), a [16×16] matrix is applied. This is to match the computational complexity in the worst case to 8 multiplications per sample.

[0196] For FIGS. 11(d) and (e), LFNST is applied to the upper left 8×8 region, and such LFNST is hereinafter named "8×8 LFNST". In the corresponding transformation matrix, a [48×16] or [48×8] matrix can be applied. In the case of the forward LFNST, since a [48×1] vector (the x vector in Equation 9) is input as the input data, all sample values in the upper left 8×8 region are not used as the input values of the forward LFNST. That is, as can be seen in the left order in FIG. 8(a) or the left order in FIG. 8(b), the bottom - right 4×4 block is left as it is, and a [48×1] vector can be constructed based on the samples belonging to the remaining three 4×4 blocks.

[0197] For the 8×8 block (8×8 TU or 8×8 CU) in FIG. 11(d), a [48×8] matrix can be applied, and for the 8×8 block in FIG. 11(e), a [48×16] matrix can be applied. This is also to match the computational complexity in the worst case to 8 multiplications per sample.

[0198] Similarly for the block, when the corresponding forward LFNST (4×4 LFNST or 8×8 LFNST) is applied, 8 or 16 output data (the y vector in Equation 9, [8×1] or [16×1] vector) are generated. Due to the characteristics of the matrix GT in the forward LFNST, the number of output data is the same as or less than the number of input data.

[0199] FIG. 12 shows an array of output data of the forward LFNST by way of an example, and shows a block in which the output data of the forward LFNST is arranged by way of a block.

[0200] The shaded area processed at the upper left end of the block shown in FIG. 12 corresponds to the area where the output data of the forward LFNST is located. The positions denoted by 0 indicate samples filled with 0 values, and the remaining areas indicate areas not changed by the forward LFNST. In the areas not changed by the LFNST, the output data of the forward first-order transform exists as it is without being changed.

[0201] As described above, since the dimension of the transformation matrix applied by way of a block changes, the number of output data also changes. As in FIG. 12, there may be a case where the output data of the forward LFNST cannot fill all of the upper left 4×4 blocks. In the cases of FIGS. 12(a) and (d), a [16×8] matrix and a [48×8] matrix are applied to the blocks indicated by thick lines or partial areas inside the blocks, respectively, to generate an [8×1] vector as the output of the forward LFNST. That is, only 8 output data are filled as shown in FIGS. 12(a) and (d) according to the scan order shown in FIG. 9(b), and the remaining 8 positions can be filled with 0. In the case of the LFNST application block of FIG. 11(d), the two 4×4 blocks at the upper right end and the lower left end adjacent to the upper left 4×4 block as in FIG. 12(d) are also filled with 0 values.

[0202] As described above, basically, the LFNST index is signaled to specify whether the LFNST is applicable and the transformation matrix to be applied. As shown in FIG. 12, when the LFNST is applied, since the number of output data of the forward LFNST may be the same as or less than the number of input data, areas filled with 0 values are generated as follows.

[0203] 1) Positions after the 8th position in the scan order within the upper left 4×4 block as in FIG. 12(a), that is, samples from the 9th to the 16th

[0204] 2) As shown in (d) and (e) of FIG. 12, a [16×48] matrix or an [8×48] matrix is applied to two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scan order.

[0205] Therefore, when checking the areas of 1) and 2) and non-zero data exists, since it is certain that LFNST is not applied, the signaling of the corresponding LFNST index can be omitted.

[0206] By way of example, for instance, in the case of LFNST adopted in the VVC standard, since the signaling of the LFNST index is executed after residual coding, the encoding device can know the presence or absence of non-zero data (valid coefficients) for all positions inside the TU or CU block via residual coding. Therefore, the encoding device can determine whether to execute the signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can determine whether to parse the LFNST index. If there is no non-zero data in the areas specified in 1) and 2), the signaling of the LFNST index will be executed.

[0207] To apply the truncated unary code to the LFNST index in the binary evolution method, the LFNST index is composed of a maximum of 2 bins, and for the binary codes corresponding to the possible LFNST index values 0, 1, and 2, 0, 10, and 11 are assigned respectively. In the case of the LFNST currently adopted in VVC, context-based CABAC coding (regular coding) is applied to the first bin, and bypass coding is applied to the second bin. The total number of contexts for the first bin is 2, and the primary transform pair (DCT-2, DCT-2) is applied for the horizontal and vertical directions. When the luma component and the chroma component are coded in a dual-tree type, one context is assigned, and for other cases, another context is applied. Representing the coding of such an LFNST index in a table is as follows.

[0208]

Table 4

[0209] On the other hand, for the adopted LFNST, the following simplification method can be applied.

[0210] (i) By way of example, the number of output data for the forward LFNST can be limited to a maximum of 16.

[0211] In the case of (c) in FIG. 11, 4×4 LFNSTs can be applied to two 4×4 regions adjacent to the upper left end, and at this time, up to 32 LFNST output data can be generated. If the number of output data for the forward LFNST is limited to a maximum of 16, the 4×4 LFNST is applied only to one 4×4 region existing at the upper left end for a 4×N / N×4 (N≥16) block (TU or CU), and the LFNST can be applied only once to all the blocks in FIG. 11. Through this, the implementation for video coding becomes simple.

[0212] FIG. 13 is a diagram showing an example in which the number of output data for the forward LFNST is limited to a maximum of 16. As shown in FIG. 13, when the LFNST is applied to the leftmost upper 4×4 region in a 4×N or N×4 block where N is 16 or more, the output data of the forward LFNST becomes 16.

[0213] (ii) By way of an example, additional zero-out can be applied to a region where the LFNST is not applied. In this document, zero-out can mean filling the values of all positions belonging to a specific region with 0 values. That is, zero-out can be applied to a region that is not changed by the LFNST and maintains the result of the forward first-order transformation. As described above, since the LFNST is classified into 4×4 LFNST and 8×8 LFNST, zero-out can be classified into two types ((ii)-(A) and (ii)-(B)) as follows.

[0214] (ii)-(A) When the 4×4 LFNST is applied, the region where the 4×4 LFNST is not applied can be zeroed out. FIG. 14 is a diagram showing zero-out in a block where the 4×4 LFNST is applied by way of an example.

[0215] As shown in FIG. 14, for the block to which the 4×4 LFNST is applied, that is, all the regions where the LFNST is not applied to the blocks (a), (b), and (c) in FIG. 12 can be filled with 0.

[0216] On the other hand, (d) of FIG. 14 shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 as in FIG. 13, zeroing out is performed on the remaining blocks to which the 4×4 LFNST is not applied.

[0217] (ii)-(B) When the 8×8 LFNST is applied, the area where the 8×8 LFNST is not applied can be zeroed out. FIG. 15 is a diagram showing zeroing out in a block to which the 8×8 LFNST is applied as an example.

[0218] As shown in FIG. 15, for the block to which the 8×8 LFNST is applied, that is, the areas in (d) and (e) of FIG. 12 where the LFNST is not applied can all be filled with 0s.

[0219] (iii) When the LFNST is applied by the zeroing out presented in (ii) above, the area filled with 0s can be changed. Therefore, it is possible to check whether there is non-zero data in a wider area than in the case of the LFNST of FIG. 12 with respect to the zeroing out proposed in (ii) above.

[0220] For example, when applying (ii)-(B), after checking whether there is non-zero data up to the area additionally filled with 0s in FIG. 15 in addition to the area filled with 0s in (d) and (e) of FIG. 12, signaling for the LFNST index can be performed only when there is no non-zero data.

[0221] Of course, even if the zero-out proposed in (ii) above is applied, it is possible to check whether there is non-zero data as in the existing LFNST index signaling. That is, it is possible to check whether there is non-zero data for the blocks filled with 0 in FIG. 12 and apply the LFNST index signaling. In such a case, the zero-out is executed only in the encoding device, and in the decoding device, without assuming the corresponding zero-out, that is, only check whether there is non-zero data for the regions explicitly marked with 0 in FIG. 12 and execute the LFNST index parsing.

[0222] Alternatively, according to another example, zero-out can also be executed as shown in FIG. 16. FIG. 16 is a diagram showing zero-out in a block to which 8×8 LFNST is applied according to another example.

[0223] As shown in FIGS. 14 and 15, zero-out can be applied to all regions other than the regions to which LFNST is applied, and it is also possible to apply zero-out only to partial regions as shown in FIG. 16. Zero-out is applied only to the regions other than the upper left 8×8 region in FIG. 16, and zero-out is not applied to the lower right 4×4 block inside the upper left 8×8 region.

[0224] A variety of embodiments can be derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) for the LFNST. Of course, the combinations for the simplification methods are not limited to the following embodiments, and any combination can be applied to the LFNST.

[0225] Embodiment

[0226] - Limit the number of output data for the forward LFNST to a maximum of 16 → (i)

[0227] - When 4×4 LFNST is applied, zero out all regions where 4×4 LFNST is not applied → (ii)-(A)

[0228] - When 8×8 LFNST is applied, all areas where 8×8 LFNST is not applied are zeroed out → (ii)-(B)

[0229] - For areas filled with existing 0 values and areas filled with 0 due to additional zeroing out ((ii)-(A), (ii)-(B)), after checking whether there is non-zero data, LFNST indexing signaling is performed only when there is no non-zero data → (iii)

[0230] In the case of the above embodiments, when LFNST is applied, the area where non-zero output data can exist is limited to the inner part of the upper left 4×4 area. More specifically, in the cases of (a) in FIG. 14 and (a) in FIG. 15, in the scanning order, the 8th position is the last position where non-zero data can exist. In the cases of (b) and (d) in FIG. 14 and (b) in FIG. 15, in the scanning order, the 16th position (i.e., the right lower end most position of the upper left 4×4 block) is the last position where non-zero data can exist.

[0231] Therefore, when LFNST is applied, after checking whether there is non-zero data at a position where the residual coding process is not allowed (a position beyond the last position), it is possible to determine whether LFNST index signaling is possible.

[0232] (In the case of the zeroing out method proposed in (ii), in order to reduce the number of data that will finally occur when all primary conversions and LFNST are applied, the computational amount required when performing the overall conversion process can be reduced. That is, when LFNST is applied, zeroing out is also applied to the forward primary conversion output data existing in the area where LFNST is not applied. Therefore, it is not necessary to generate data for the area that will be zeroed out when performing the forward primary conversion. Therefore, the amount of computation required for generating the corresponding data can be saved. The additional effects of the zeroing out method proposed in (ii) are summarized as follows.)

[0233] First, as described above, the amount of computation required for the execution of the overall conversion process is reduced.

[0234] In particular, when applying (ii)-(B), the amount of computation for the worst case can be reduced, and the conversion process can be lightweighted. To elaborate, generally, a large amount of operations are required for the execution of a primary conversion of a large size. When (ii)-(B) is applied, the number of data derived as the result of the forward LFNST execution can be reduced to 16 or less. As the size of the overall block (TU or CU) increases, the conversion operation reduction effect is further increased.

[0235] Second, the amount of operations required for the entire conversion process is reduced, and the power consumption required for the conversion execution can be reduced.

[0236] Third, the latency associated with the conversion process is reduced.

[0237] A secondary conversion such as LFNST adds computation to the existing primary conversion, thus increasing the overall latency associated with the conversion execution. Especially in the case of intra prediction, since the restored data of adjacent blocks is used in the prediction process, the increase in latency due to the secondary conversion during encoding can lead to an increase in latency until reconstruction, which can lead to an overall increase in latency of the intra prediction encoding.

[0238] However, when the zeroing out presented in (ii) is applied, the latency of the primary conversion execution can be significantly reduced when LFNST is applied. Therefore, the latency for the entire conversion execution is maintained or reduced, and the encoding device can be implemented more easily.

[0239] On the one hand, conventional intra prediction performs encoding without division by treating the block to be currently encoded as one encoding unit. However, ISP (Intra Sub-Paritions) coding means performing intra prediction encoding by dividing the block to be currently encoded horizontally or vertically. At this time, encoding / decoding is performed in units of the divided blocks to generate restored blocks, and the restored blocks can then be used as reference blocks for the next divided blocks. By way of example, one coding block can be divided into 2 or 4 sub-blocks and coded during ISP coding, and for one sub-block in ISP, intra prediction is performed by referring to the restored pixel values of the sub-blocks located adjacent to the left or adjacent to the upper side. Hereinafter, the "coding" used can be used in the concept including all the coding executed by the encoding device and the decoding executed by the decoding device.

[0240] Table 5 shows the number of sub-blocks divided according to the block size when ISP is applied, and the sub-partitions divided by ISP can be called transform blocks (TUs).

[0241]

Table 5

[0242] ISP is to divide the block predicted by luma intra in the vertical or horizontal direction into 2 or 4 sub-partitionings according to the block size. For example, the minimum block size to which ISP can be applied is 4×8 or 8×4. If the block size is larger than 4×8 or 8×4, the block is divided into 4 sub-partitionings.

[0243] FIG. 17 and FIG. 18 show an example of sub-blocks into which one coding block is divided. More specifically, FIG. 17 illustrates an example of division when the coding block (width (W) × height (H)) is a 4×8 block or an 8×4 block, and FIG. 18 illustrates an example of division when the coding block is not a 4×8 block, an 8×4 block, or a 4×4 block.

[0244] When applying ISP, the sub-blocks are coded sequentially, for example, horizontally or vertically, from left to right or from top to bottom, depending on the division form. After performing inverse transformation and intra prediction for one sub-block and proceeding to the restoration process, coding for the next sub-block can be carried out. For the leftmost or topmost sub-block, the restored pixels of the already coded coding block are referred to, similar to the normal intra prediction method. Also, when a side of an internal sub-block that follows is not adjacent to the previous sub-block, the restored pixels of the adjacent coded coding block are referred to, similar to the normal intra prediction method, to derive the reference pixels adjacent to the corresponding side.

[0245] In the ISP coding mode, all sub-blocks can be coded in the same intra prediction mode, and flags indicating whether to use ISP coding and in which direction (horizontal or vertical) to divide can be signaled. As shown in FIGS. 17 and 18, the number of sub-blocks can be adjusted to 2 or 4 depending on the block. When the size (width × height) of one sub-block is less than 16, division into the corresponding sub-block can be prohibited, or the application of ISP coding itself can be restricted.

[0246] On the other hand, in the case of the ISP prediction mode, one coding unit is divided into 2 or 4 partition blocks, i.e., sub-blocks, and predicted. The same in-picture prediction mode is applied to the corresponding 2 or 4 partition blocks after division.

[0247] As described above, in the splitting direction, there are both a horizontal direction (when an M×N coding unit with horizontal length M and vertical length N is split horizontally, if it is split into two, it is split into M×(N / 2) blocks, and if it is split into four, it is split into M×(N / 4) blocks) and a vertical direction (when an M×N coding unit is split vertically, if it is split into two, it is split into (M / 2)×N blocks, and if it is split into four, it is split into (M / 4)×N blocks). When split horizontally, the partition blocks are coded in the order from the upper side to the lower side, and when split vertically, the partition blocks are coded in the order from the left side to the right side. The currently coded partition block can be predicted by referring to the restored pixel values of the upper (left) partition block in the case of horizontal (vertical) direction splitting.

[0248] The conversion can be applied to the residual signal generated by the ISP prediction method in units of partition blocks. Based on the forward direction, not only the existing DCT-2 but also the DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technology can be applied to the primary transform (core transform or primary transform), and the forward LFNST (Low Frequency Non-Separable Transform) can be applied to the transform coefficients generated by the primary transform to generate the final corrected transform coefficients.

[0249] That is, the LFNST can also be applied to the partition blocks that are divided when the ISP prediction mode is applied. As described above, the same intra prediction mode is applied to the divided partition blocks. Therefore, when selecting the LFNST set derived based on the intra prediction mode, the LFNST set derived for all partition blocks can be applied. That is, since the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks accordingly.

[0250] On the other hand, by way of example, the LFNST can be applied only to a transform block whose horizontal and vertical lengths are both 4 or more. Therefore, when the horizontal or vertical length of a partition block divided by the ISP prediction method is less than 4, the LFNST is not applied and the LFNST index is not signaled either. Also, when applying the LFNST to each partition block, the corresponding partition block can be regarded as one transform block. Of course, when the ISP prediction method is not applied, the LFNST can be applied to the coding block.

[0251] Looking specifically at applying the LFNST to each partition block, it is as follows.

[0252] By way of example, after applying the forward LFNST to an individual partition block, only up to 16 (8 or 16) coefficients are left in the left-upper 4×4 area according to the transform coefficient scanning order, and then zero-out that fills all the remaining positions and areas with 0 values can be applied.

[0253] Or, by way of example, when the length of one side of the partition block is 4, the LFNST is applied only to the left-upper 4×4 area, and when the lengths of all sides of the partition block, that is, the width and height, are 8 or more, the LFNST can be applied to the remaining 48 coefficients excluding the right-lower 4×4 area inside the left-upper 8×8 area.

[0254] Alternatively, by way of an example, in order to combine the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4×4 or 8×8, only 8 transform coefficients can be output after the forward LFNST application. That is, when the partition block is 4×4, an 8×16 matrix can be applied with the transform matrix, and when the partition block is 8×8, an 8×48 matrix can be applied with the transform matrix.

[0255] On the other hand, in the current VVC standard, the LFNST index signaling is performed on a coding unit basis. Therefore, in the ISP prediction mode, when the LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition block. That is, when the LFNST index value is transmitted once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks within the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2. 0 indicates the case where the LFNST is not applied, and 1 and 2 refer to two transform matrices existing within one LFNST set when the LFNST is applied.

[0256] As described above, the LFNST set is determined by the intra prediction mode. In the case of the ISP prediction mode, since all partition blocks within the coding unit are predicted in the same intra prediction mode, the partition block can refer to the same LFNST set.

[0257] As another example, although the LFNST indexing is still performed on a coding unit basis, in the case of the ISP prediction mode, instead of uniformly determining whether to apply LFNST to all partition blocks, it is possible to determine whether to apply the LFNST index value signaled at the coding unit level to each partition block or not to apply LFNST to each partition block through separate conditions. Here, the separate conditions can be signaled in a flag form for each partition block via the bitstream. When the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST is not applied.

[0258] On the other hand, in the coding unit to which the ISP mode is applied, when the length of one side of the partition block is less than 4, for an example of applying LFNST, it is as follows.

[0259] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper left M×2 (2×M) region (where M≦N). For example, when M = 8, the corresponding upper left region becomes 8×2 (2×8), so the region where 16 residual signals exist can be the input of the forward LFNST, and an R×16 (R≦16) forward transform matrix can be applied.

[0260] Here, the forward LFNST matrix is a separate additional matrix that is not the matrix currently included in the VVC standard. Also, for complexity adjustment in the worst case, an 8×16 matrix sampled only from the upper 8 row vectors of the 16×16 matrix can be used for the transform. The complexity adjustment method will be described in detail later.

[0261] Second, when the size of the partition block is N×1 (1×N), the LFNST can be applied to the upper left M×1 (1×M) region (where M≦N). For example, when M = 16, the corresponding upper left region is 16×1 (1×16), so the region where 16 residual signals exist can be the input of the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.

[0262] Here, the corresponding forward LFNST matrix is a separate additional matrix that is not the matrix currently included in the VVC standard. Also, for worst-case complexity adjustment, an 8×16 matrix sampled from only the upper 8 row vectors of the 16×16 matrix can be used for the transformation. The complexity adjustment method will be described in detail later.

[0263] The first embodiment and the second embodiment can be applied simultaneously, or only one of the two embodiments can be applied. In particular, in the case of the second embodiment, by considering a one-dimensional transformation in the LFNST, it has been observed through experiments that the improvement in compression performance that can be obtained with the existing LFNST is not relatively large compared to the LFNST index signaling cost. However, in the case of the first embodiment, an improvement in compression performance similar to that obtained with the existing LFNST has been observed. That is, in the case of ISP, it can be confirmed through experiments that the application of LFNST for 2×N and N×2 contributes to the actual compression performance.

[0264] Currently, symmetry between intra prediction modes is applied in the LFNST of VVC. The same LFNST set is applied to two directional modes arranged around mode 34 (prediction in the lower right 45-degree diagonal direction). For example, the same LFNST set is applied to mode 18 (horizontal prediction mode) and mode 50 (vertical prediction mode). However, for modes 35 to 66, when applying the forward LFNST, the input data is transposed and then the LFNST is applied.

[0265] On the other hand, in VVC, the Wide Angle Intra Prediction (WAIP) mode is supported, and the LFNST set is derived based on the intra prediction mode modified in consideration of the WAIP mode. For the modes extended by WAIP, the symmetry is utilized in the same way as the general intra prediction direction mode to determine the LFNST set. For example, since mode - 1 is symmetric with mode 67, the same LFNST set is applied, and since mode - 14 is symmetric with mode 80, the same LFNST set is applied. For modes from 67 to 80, after transposing the input data before applying the forward LFNST, the LFNST transform is applied.

[0266] In the case of the LFNST applied to the upper - left - hand M×2 (M×1) block, the symmetry for the aforementioned LFNST cannot be applied. The reason is that the block to which the LFNST is applied is non - square. Therefore, instead of applying the symmetry based on the intra prediction mode like the LFNST in Table 2, the symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.

[0267] FIG. 19 is a diagram showing the symmetry between an M×2 (M×1) block and a 2×M (1×M) block according to an example.

[0268] As shown in FIG. 19, since mode 2 in the M×2 (M×1) block can be regarded as symmetric with mode 66 in the 2×M (1×M) block, the same LFNST set can be applied to the 2×M (1×M) block and the M×2 (M×1) block.

[0269] At this time, in order to apply the LFNST set applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on Mode 2 instead of Mode 66. That is, after transposing the input data of the 2×M (1×M) block before applying the forward LFNST, the LFNST can be applied.

[0270] FIG. 20 is a diagram showing an example of transposing a 2×M block.

[0271] FIG. 20(a) is a drawing for explaining applying the LFNST by reading the input data in column-first order for the 2×M block, and FIG. 20(b) is a drawing for explaining applying the LFNST by reading the input data in row-first order for the M×2 (M×1) block. Organizing the method of applying the LFNST to the upper left M×2 (M×1) or 2×M (M×1) block, it is as follows.

[0272] 1. First, as shown in FIGS. 20(a) and 20(b), the input data is arranged to form the input vector of the forward LFNST. For example, referring to FIG. 19, for the M×2 block predicted in Mode 2, it follows the order in FIG. 20(b), and for the 2×M block predicted in Mode 66, after arranging the input data according to the order in FIG. 20(a), the LFNST set for Mode 2 can be applied.

[0273] 2. For the M×2 (M×1) block, the LFNST set is determined based on the modified intra prediction mode considering WAIP. As described above, a preset mapping relationship is established between the intra prediction mode and the LFNST set, and this can be represented by a mapping table as shown in Table 2.

[0274] For the 2×M (1×M) block, after obtaining the modes that are symmetric about the prediction mode in the 45-degree diagonal direction towards the lower right (mode 34 in the case of the VVC standard) from the intra prediction modes modified considering WAIP, the LFNST set is determined based on the corresponding symmetric modes and the mapping table. The mode (y) that is symmetric about mode 34 can be derived through the following formula. The matter regarding the mapping table will be further specifically described below.

[0275] [Equation 11] if 2≦x≦66, y = 68 - x, otherwise (x≦ -1 or x≧67), y = 66 - x

[0276] When applying the forward LFNST, the input data prepared through process 1 can be multiplied by the LFNST kernel to derive the conversion coefficients. The LFNST kernel can be selected from the LFNST set determined in process 2 and the pre-specified LFNST index.

[0277] For example, when M = 8 and a 16×16 matrix is applied in the LFNST kernel, 16 conversion coefficients can be generated by multiplying the corresponding matrix by 16 input data. The generated conversion coefficients can be arranged in the left upper 8×2 or 2×8 region according to the scanning order used in the VVC standard.

[0278] Figure 21 shows the scanning order for an 8×2 or 2×8 region according to an example.

[0279] For regions other than the left upper 8×2 or 2×8 region, they can be filled with all 0 values (zero-out), or the existing conversion coefficients to which a first-order transformation is applied can be maintained as they are. The pre-specified LFNST index is one of the LFNST index values (0, 1, 2) that are tried when calculating the RD cost while changing the LFNST index value during the encoding process.

[0280] In the case of a configuration that sets the computational complexity for the worst case to below a certain level (for example, 8 multiplications / sample), for example, after multiplying an 8×16 matrix that only takes the upper 8 rows of the 16×16 matrix to generate only 8 conversion coefficients, the 8 conversion coefficients can be arranged in the scanning order as shown in FIG. 21, and zero-out can also be applied to the remaining coefficient area. The complexity adjustment for the worst case will be described later.

[0281] When applying the inverse-direction LFNST, place the pre-set number (for example, 16) of conversion coefficients in the input vector, select the LFNST kernel (for example, a 16×16 matrix) derived from the LFNST set obtained in the second process and the parsed LFNST index, and then multiply the LFNST kernel and the corresponding input vector to derive the output vector.

[0282] In the case of an M×2 (M×1) block, the output vector can be arranged in the row-major order as shown in FIG. 20(b), and in the case of a 2×M (1×M) block, the output vector can be arranged in the column-major order as shown in FIG. 20(a).

[0283] Excluding the area where the corresponding output vector is arranged inside the upper left M×2 (M×1) or 2×M (M×2) area, for the remaining area and the area outside the upper left M×2 (M×1) or 2×M (M×2) area in the partition block, it can be configured to be filled with all 0 values (zero-out), or to maintain the conversion coefficients restored through the residual coding and inverse quantization process as they are.

[0284] When configuring the input vector in the same way as in the third step, the input data can be arranged in the scanning order as shown in FIG. 21, and in order to set the computational complexity for the worst case to below a certain level, the number of input data can be reduced (for example, 8 instead of 16) to configure the input vector.

[0285] For example, when M = 8 and eight input data are used, only the left 16×8 matrix is taken from the corresponding 16×16 matrix and multiplied, and then 16 output data can be obtained. The complexity adjustment for the worst case will be described later.

[0286] In the above embodiment, the case where symmetry is applied between the M×2 (M×1) block and the 2×M (1×M) block when LFNST is applied is presented. However, according to other examples, different LFNST sets can also be applied to the two blocks respectively.

[0287] Hereinafter, various examples of the LFNST set configuration for the ISP mode and the mapping method using the intra prediction mode will be described.

[0288] In the case of the ISP mode, the LFNST set configuration may be different from the existing LFNST set. That is, a kernel different from the existing LFNST kernel can also be applied, and another mapping table different from the mapping table between the intra prediction mode index currently applied to the VVC standard and the LFNST set can be applied. The mapping table currently applied to the VVC standard is as shown in Table 2.

[0289] In Table 2, the preModeIntra value means the intra prediction mode value changed in consideration of WAIP, and the lfnstTrSetIdx value is the index value indicating a specific LFNST set. Each LFNST set is composed of two LFNST kernels.

[0290] When the ISP prediction mode is applied, when the horizontal length and the vertical length of each partition block are all greater than or equal to 4, the same kernel as the LFNST kernel currently applied in the VVC standard can be applied, and the mapping table can also be applied as it is. Of course, the current VVC standard, other LFNST kernels, and other mapping tables can also be applied.

[0291] When the ISP prediction mode is applied, if the horizontal or vertical length of each partition block is less than 4, the current VVC standard, other LFNST kernels, and other mapping tables can be applied. Hereinafter, Tables 6 to 8 show the mapping tables between the intra prediction mode values (intra prediction mode values changed considering WAIP) applicable to M×2 (M×1) blocks or 2×M (1×M) blocks and the LFNST sets.

[0292]

Table 6

[0293]

Table 7

[0294]

Table 8

[0295] The mapping table in Table 6 is composed of 7 LFNST sets, the mapping table in Table 7 is composed of 4 LFNST sets, and the mapping table in Table 8 is composed of 2 LFNST sets. As another example, when composed of 1 LFNST set, the lfnstTrSetIdx value can be fixed to 0 for the preModeIntra value.

[0296] Hereinafter, a method for maintaining the computational complexity in the worst case when applying LFNST in the ISP mode will be described.

[0297] In the case of the ISP mode, when applying LFNST, the application of LFNST can be restricted to maintain the number of multiplications per sample (or per coefficient, per position) below a certain value. Depending on the size of the partition block, LFNST can be applied as follows to maintain the number of multiplications per sample (or per coefficient, per position) below 8.

[0298] 1. When the horizontal and vertical lengths of the partition block are both 4 or more, the same calculation complexity adjustment method as that for the worst case of LFNST in the current VVC standard can be applied.

[0299] That is, when the partition block is a 4×4 block, instead of a 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix can be applied in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix can be applied in the reverse direction. Also, when the partition block is an 8×8 block, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the upper 8 rows from a 16×48 matrix can be applied, and in the reverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns from a 48×16 matrix can be applied.

[0300] In the case of a 4×N or N×4 (N>4) block, when performing the forward transform, only for the upper left 4×4 block, after applying a 16×16 matrix, the 16 coefficients generated can be arranged in the upper left 4×4 region, and the other regions can be filled with 0 values. Also, when performing the inverse transform, after arranging the 16 coefficients located in the upper left 4×4 block in the scanning order to form an input vector, multiplying by a 16×16 matrix can generate 16 output data. The generated output data can be arranged in the upper left 4×4 region, and the remaining regions excluding the upper left 4×4 region can be filled with 0.

[0301] In the case of an 8×N or N×8 (N > 8) block, when performing a forward transform, for the ROI region (the remaining region after excluding the lower-right 4×4 block from the upper-left 8×8 block) within the upper-left 8×8 block, only after applying a 16×48 matrix, the 16 coefficients generated can be arranged in the upper-left 4×4 region, and the remaining regions can all be filled with 0 values. Also, when performing an inverse transform, after arranging the 16 coefficients located in the upper-left 4×4 block in scanning order to form an input vector, multiplying by a 48×16 matrix can generate 48 output data. The generated output data can fill the ROI region, and the remaining regions can all be filled with 0 values.

[0302] 2. When the size of the partition block is N×2 or 2×N, and applying LFNST to the upper-left M×2 or 2×M region (M ≤ N), a matrix sampled according to the N value can be applied.

[0303] When M = 8, for a partition block with N = 8, that is, an 8×2 or 2×8 block, in the case of a forward transform, instead of applying a 16×16 matrix, an 8×16 matrix sampled from the upper 8 rows of a 16×16 matrix can be applied; in the case of an inverse transform, instead of applying a 16×16 matrix, a 16×8 matrix sampled from the left 8 columns of a 16×16 matrix can be applied.

[0304] When N is greater than 8, in the case of a forward transform, for the upper-left 8×2 or 2×8 block, after applying a 16×16 matrix, the 16 output data generated can be arranged in the upper-left 8×2 or 2×8 block, and the remaining regions can be filled with 0 values. In the case of an inverse transform, after arranging the 16 coefficients located in the upper-left 8×2 or 2×8 block in scanning order to form an input vector, multiplying by the corresponding 16×16 matrix can generate 16 output data. The generated output data can be arranged in the upper-left 8×2 or 2×8 block, and the remaining regions can all be filled with 0 values.

[0305] 3. When the size of the partition block is N×1 or 1×N and the LFNST is applied to the upper left M×1 or 1×M region (M≦N), a matrix sampled according to the N value can be applied.

[0306] When M = 16, for a partition block with N = 16, that is, a 16×1 or 1×16 block, in the case of the forward transform, instead of the 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from the 16×16 matrix can be applied; in the case of the inverse transform, instead of the 16×16 matrix, a 16×8 matrix obtained by sampling the left 8 columns from the 16×16 matrix can be applied.

[0307] When N is greater than 16, in the case of the forward transform, after applying the 16×16 matrix to the upper left 16×1 or 1×16 block, the 16 output data generated can be arranged in the upper left 16×1 or 1×16 block, and the remaining region can be filled with 0 values. In the case of the inverse transform, after arranging the 16 coefficients located in the upper left 16×1 or 1×16 block in the scanning order to form an input vector, the corresponding 16×16 matrix can be multiplied to generate 16 output data. The generated output data can be arranged in the upper left 16×1 or 1×16 block, and the remaining region can all be filled with 0 values.

[0308] As another example, in order to keep the multiplication count per sample (or per coefficient, per position) below a certain value, the multiplication count per sample (or per coefficient, per position) can be kept below 8 based on the size of the ISP coding unit rather than the size of the ISP partition block. If there is only one block among the ISP partition blocks that satisfies the condition for applying LFNST, the complexity calculation for the worst case of LFNST can be applied based on the size of the corresponding coding unit that is not the size of the partition block. For example, if the luma coding block for a certain coding unit is divided into 4 partition blocks of size 4×4 and coded by ISP, and there are no non-zero transform coefficients for 2 of the partition blocks, the other 2 partition blocks can be set to generate 16 transform coefficients each (based on the encoder standard) that are not 8 each.

[0309] In the following, when in the ISP mode, a method of signaling the LFNST index will be considered.

[0310] As described above, the LFNST index can have values of 0, 1, and 2. 0 indicates that LFNST is not applied, and 1 and 2 indicate one of the two LFNST kernel matrices included in the selected LFNST set. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method of transmitting the LFNST index in the current VVC standard will be described as follows.

[0311] 1. The LFNST index can be transmitted once per coding unit (CU). When it is a dual-tree, individual LFNST indexes can be signaled for the luma block and the chroma block respectively.

[0312] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.

[0313] A. When it is a mode where no conversion is applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0314] B. When the first conversion is not DCT-2 (DST7 or DCT8), i.e., when the horizontal or vertical conversion is not DCT-2

[0315] C. When the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma conversion that can be performed. For example, when the size of the maximum luma conversion that can be performed is 64, and the size of the luma block of the coding block is 128×16, etc., LFNST cannot be applied.

[0316] In the case of dual tree, it is determined whether each of the coding unit for the luma component and the coding unit for the chroma component exceeds the size of the maximum luma conversion. That is, it is checked whether the horizontal or vertical length of the luma block exceeds the size of the maximum luma conversion that can be performed, and it is checked whether the horizontal / vertical length of the corresponding luma block for the color format and the size of the maximum luma conversion that can be performed exceed the size of the maximum luma conversion for the chroma block. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice that of the corresponding chroma block, and the conversion size of the corresponding luma block is twice that of the corresponding chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length and the conversion size of the corresponding luma block are the same as those of the corresponding chroma block.

[0317] 64-length conversion or 32-length conversion means a conversion applied horizontally or vertically with lengths of 64 or 32 respectively, and "conversion size" can mean 64 or 32 which is the corresponding length.

[0318] In the case of a single tree, after checking whether the horizontal length or the vertical length of the luma block exceeds the maximum luma transform block size that can be transformed, if it exceeds, the LFNST index signaling can be omitted.

[0319] The LFNST index can be transmitted only when the horizontal length and the vertical length of the coding unit D are all 4 or more.

[0320] In the case of a dual tree, the LFNST index can be signaled only when the horizontal length and the vertical length of the corresponding component (i.e., luma or chroma component) are all 4 or more.

[0321] In the case of a single tree, the LFNST index can be signaled when the horizontal length and the vertical length of the luma component are all 4 or more.

[0322] E. When the position of the last non-zero coefficient is not the DC position (the upper left corner position of the block), in the case of a dual-tree type luma block, the LFNST index is transmitted when the position of the last non-zero coefficient is not the DC position. In the case of a dual-tree type chroma block, the corresponding LNFST index is transmitted when the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position.

[0323] In the case of a single-tree type, the LFNST index is transmitted when the position of the last non-zero coefficient of one of the luma component, Cb component, and Cr component is not the DC position.

[0324] Here, when the CBF (coded block flag) value indicating the existence of a conversion coefficient for a single conversion block is 0, the position of the last non-zero coefficient for the corresponding conversion block is not checked to determine whether LFNST index signaling is possible. That is, when the corresponding CBF value is 0, since no conversion is applied to the corresponding block, the position of the last non-zero coefficient is not considered when checking the conditions for LFNST index signaling.

[0325] For example, 1) in the case of the dual-tree type and for the luma component, when the corresponding CBF value is 0, the LFNST index is not signaled; 2) in the case of the dual-tree type and for the chroma component, when the CBF value for Cb is 0 and the CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is transmitted; 3) in the case of the single-tree type, the position of the last non-zero coefficient is only checked for components where the respective CBF values are 1 for all of luma, Cb, and Cr.

[0326] If it is confirmed that a conversion coefficient exists at a position where the F.LFNST conversion coefficient cannot exist, LFNST index signaling can be omitted. In the case of 4×4 and 8×8 conversion blocks, according to the conversion coefficient scanning order in the VVC standard, the LFNST conversion coefficient can exist at 8 positions starting from the DC position, and the remaining positions are all filled with 0. Also, in the case of blocks other than 4×4 and 8×8 conversion blocks, according to the conversion coefficient scanning order in the VVC standard, the LFNST conversion coefficient can exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0.

[0327] Therefore, after performing residual coding, if a non-zero conversion coefficient exists in the area where the 0 value should be filled, LFNST index signaling can be omitted.

[0328] On the one hand, the ISP mode is applicable only when it is a luma block, or it can also be applied to both luma and chroma blocks. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks for prediction, and the transformation can also be applied to each corresponding partition block. Therefore, when determining the conditions for signaling the LFNST index in units of coding units, the fact that LFNST can be applied to each corresponding partition block must be considered. Also, when the ISP prediction mode is applied only to a specific component (e.g., a luma block), the LFNST index must be signaled considering the fact that it is divided into partition blocks only for the corresponding component. Organizing the possible LFNST index signaling methods in the ISP mode, it is as follows.

[0329] 1. The LFNST index can be transmitted once for each coding unit (CU). In the case of a dual-tree, individual LFNST indexes can be signaled for the luma and chroma blocks respectively.

[0330] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.

[0331] A. In the case of a mode where transformation is not applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0332] B. When the horizontal or vertical length of the luma block of the coding unit exceeds the maximum luma transform size that can be transformed, for example, when the maximum luma transform size that can be transformed is 64, and the size of the luma block of the coding block is 128×16, etc., LFNST cannot be applied.

[0333] It is also possible to determine whether the LFNST index can be signaled based on the size of the partition block instead of the coding unit. That is, when the horizontal or vertical length of the partition block for the corresponding luma block exceeds the maximum luma transform size that can be transformed, the LFNST index signaling can be omitted and the LFNST index value can be analogized to 0.

[0334] In the case of a dual tree, it is determined whether each of the coding unit or partition block for the luma component and the coding unit or partition block for the chroma component exceeds the maximum transform block size. That is, the horizontal and vertical lengths of the coding unit or partition block for luma are each compared with the maximum luma transform size. If even one of them is larger than the maximum luma transform size, LFNST is not applied. For the coding unit or partition block for chroma, the horizontal / vertical length of the corresponding luma block for the color format is compared with the maximum luma transform size that can be transformed. For example, when the color format is 4:2:0, the horizontal / vertical lengths of the corresponding luma blocks are each twice that of the corresponding chroma block, and the transform size of the corresponding luma block is twice that of the corresponding chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical lengths and the transform size of the corresponding luma blocks are the same as those of the corresponding chroma blocks.

[0335] In the case of a single tree, after checking whether the horizontal or vertical length of a luma block (coding unit or partition block) exceeds the maximum luma transform block size that can be transformed, if it exceeds, the LFNST index signaling can be omitted.

[0336] C. If applying the LFNST included in the current VVC standard, the LFNST index can be transmitted only when the horizontal and vertical lengths of the partition block are both 4 or more.

[0337] If applying the LFNST for 2×M (1×M) or M×2 (M×1) blocks in addition to the LFNST included in the current VVC standard, the LFNST index can be transmitted only when the size of the partition block is larger than or equal to the 2×M (1×M) or M×2 (M×1) block. Here, the meaning that a P×Q block is larger than or equal to an R×S block means P≧R and Q≧S.

[0338] To summarize, the LFNST index can be transmitted only when the partition block is larger than or equal to the minimum size for which LFNST is applicable. In the case of a dual tree, the LFNST index can be signaled only when the partition block for the luma or chroma component is larger than or equal to the minimum size for which LFNST is applicable. In the case of a single tree, the LFNST index can be signaled only when the partition block for the luma component is larger than or equal to the minimum size for which LFNST is applicable.

[0339] In this document, when an M×N block is larger than or equal to a K×L block, it means that M is greater than or equal to K and N is greater than or equal to L. When an M×N block is larger than a K×L block, it means that M is greater than or equal to K, N is greater than or equal to L, and either M is greater than K or N is greater than L. When an M×N block is smaller than or equal to a K×L block, it means that M is less than or equal to K and N is less than or equal to L. When an M×N block is smaller than a K×L block, it means that M is less than or equal to K, N is less than or equal to L, and either M is less than K or N is less than L.

[0340] D. When the position of the last non-zero coefficient (last non-zero coefficient position) is not the DC position (the upper left corner position of the block), in the case of a dual-tree type luma block, the LFNST index can be transmitted when the position of the last non-zero coefficient of any one of all partition blocks is not the DC position. In the case of a dual-tree type and a chroma block, when the position of the last non-zero coefficient of all partition blocks for Cb (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) and the position of the last non-zero coefficient of all partition blocks for Cr (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) are not the DC position, the corresponding LNFST index can be transmitted.

[0341] In the case of the single-tree type, the corresponding LFNST index can be transmitted when the position of the last non-zero coefficient of any one of all partition blocks for the luma component, Cb component, and Cr component is not the DC position.

[0342] Here, when the CBF (coded block flag) value indicating the existence of a conversion coefficient for each partition block is 0, the position of the last non-zero coefficient for the corresponding partition block is not checked to determine whether LFNST index signaling is possible. That is, when the corresponding CBF value is 0, since no conversion is applied to the corresponding block, the position of the last non-zero coefficient for the corresponding partition block is not considered when checking the conditions for LFNST index signaling.

[0343] For example, 1) in the case of a dual-tree type and a luma component, when the corresponding CBF value for each partition block is 0, the corresponding partition block is excluded when determining whether LFNST index signaling is possible. 2) In the case of a dual-tree type and a chroma component, when the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, the position of the last non-zero coefficient for Cr is checked to determine whether the corresponding LFNST index signaling is possible. 3) In the case of a single-tree type, the position of the last non-zero coefficient is checked only for the blocks with a CBF value of 1 for all partition blocks of the luma component, Cb component, and Cr component to determine whether LFNST index signaling is possible.

[0344] In the case of the ISP mode, the video information can also be configured so as not to check the position of the last non-zero coefficient, and an example thereof is as follows.

[0345] i. In the case of the ISP mode, it is possible to allow LFNST index signaling by omitting the check of the position of the last non-zero coefficient for both the luma block and the chroma block. That is, even if the position of the last non-zero coefficient for all partition blocks is the DC position or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.

[0346] ii. When in ISP mode, only for luma blocks, the check for the position of the last non-zero coefficient can be omitted, and for chroma blocks, the check for the position of the last non-zero coefficient in the aforementioned method can be executed. For example, in the case of a dual-tree type and a luma block, LFNST index signaling is allowed without checking the position of the last non-zero coefficient, and in the case of a dual-tree type and a chroma block, the presence or absence of the DC position corresponding to the position of the last non-zero coefficient in the aforementioned method can be checked to determine whether the corresponding LFNST index can be signaled.

[0347] iii. When in ISP mode and of single-tree type, the method i or ii can be applied. That is, when in ISP mode and applying method i to a single-tree type, for both luma and chroma blocks, the check for the position of the last non-zero coefficient can be omitted to allow LFNST index signaling. Or, applying method ii, for the partition blocks of the luma component, the check for the position of the last non-zero coefficient can be omitted, and for the partition blocks of the chroma component (when ISP is not applied to the chroma component, the number of partition blocks can be regarded as 1), the check for the position of the last non-zero coefficient in the aforementioned method can be executed to determine whether the corresponding LFNST index can be signaled.

[0348] E. If it is confirmed that a transform coefficient exists at a position where the LFNST transform coefficient cannot exist for even one of all the partition blocks, the LFNST index signaling can be omitted.

[0349] For example, in the case of 4×4 partition blocks and 8×8 partition blocks, depending on the conversion coefficient scanning order in the VVC standard, LFNST conversion coefficients can exist at 8 positions starting from the DC position, and the remaining positions are all filled with 0. Also, when it is larger than or equal to 4×4 and is not a 4×4 partition block or an 8×8 partition block, depending on the conversion coefficient scanning order in the VVC standard, LFNST conversion coefficients can exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0.

[0350] Therefore, after performing residual coding, if there are non-zero conversion coefficients in the area to be filled with 0 values, LFNST index signaling can be omitted.

[0351] If LFNST can also be applied to the case where the partition block is 2×M (1×M) or M×2 (M×1), the area where the LFNST conversion coefficients can be located can be specified as follows. The area outside the area where the conversion coefficients can be located can be filled with 0. If there are non-zero conversion coefficients in the area to be filled with 0 when assuming that LFNST is applied, LFNST index signaling can be omitted.

[0352] i. LFNST can be applied to 2×M or M×2 blocks. When M = 8, only 8 LFNST conversion coefficients can be generated for 2×8 or 8×2 partition blocks. When the conversion coefficients are arranged in the scanning order as shown in Figure 20, 8 conversion coefficients are arranged in the scanning order starting from the DC position, and the remaining 8 positions can be filled with 0.

[0353] For a 2×N or N×2 (N>8) partition block, 16 LFNST transform coefficients can be generated. When the transform coefficients are arranged in the scanning order as shown in FIG. 20, 16 transform coefficients are arranged in the scanning order starting from the DC position, and the remaining area can be filled with 0s. That is, for a 2×N or N×2 (N>8) partition block, the area other than the upper left 2×8 or 8×2 block can be filled with 0s. For a 2×8 or 8×2 partition block, 16 transform coefficients can also be generated instead of 8 LFNST transform coefficients, and in this case, there is no area to be filled with 0s. As described above, when LFNST is applied, if it is detected that there are non-zero transform coefficients in the area determined to be filled with 0s even in one partition block, the LFNST index signaling is omitted, and the LFNST index can be inferred to be 0.

[0354] ii. LFNST can be applied to a 1×M or M×1 block. When M = 16, only 8 LFNST transform coefficients can be generated for a 1×16 or 16×1 partition block. When the transform coefficients are arranged in the scanning order from left to right or from top to bottom, 8 transform coefficients are arranged in the corresponding scanning order starting from the DC position, and the remaining 8 positions can be filled with 0s.

[0355] For a 1×N or N×1 (N>16) partition block, 16 LFNST transform coefficients can be generated. When the transform coefficients are arranged in the scanning order from left to right or from top to bottom, 16 transform coefficients are arranged in the corresponding scanning order starting from the DC position, and the remaining area can be filled with 0s. That is, for a 1×N or N×1 (N>16) partition block, the area other than the upper left 1×16 or 16×1 block can be filled with 0s.

[0356] For a 1×16 or 16×1 partition block, 16 transformation coefficients can be generated instead of 8 LFNST transformation coefficients, and in this case, no area to be filled with 0 is generated. As described above, when LFNST is applied, if it is detected that there is a non-zero transformation coefficient in an area determined to be filled with 0 even in one partition block, LFNST index signaling can be omitted and the LFNST index can be inferred to be 0.

[0357] On the other hand, in the case of the ISP mode, in the current VVC standard, for the horizontal and vertical directions, DST-7 is applied instead of DCT-2 without signaling for the MTS index by looking at the length conditions independently for each direction. It is determined whether the horizontal or vertical length is greater than or equal to 4 and less than or equal to 16, and the primary transformation kernel is determined according to the determination result. Therefore, for the case where it is in the ISP mode and LFNST can be applied, the following transformation combination configurations are possible.

[0358] 1. When the LFNST index is 0 (including the case where the LFNST index is inferred to be 0), it is possible to follow the primary transformation determination conditions in the ISP included in the current VVC standard. That is, it is checked whether the length conditions (greater than or equal to 4 and less than or equal to 16) are satisfied independently for the horizontal and vertical directions. If satisfied, DST-7 can be applied instead of DCT-2 for the primary transformation, and if not satisfied, DCT-2 can be applied.

[0359] 2. When the LFNST index is greater than 0, the following two configurations are possible for the primary transformation.

[0360] A. DCT-2 can be applied to both the horizontal and vertical directions.

[0361] B. It can comply with the primary transformation determination conditions when it is the ISP currently included in the VVC standard. That is, it checks whether the length conditions (greater than or equal to 4 and less than or equal to 16) are satisfied independently for the horizontal and vertical directions respectively. If satisfied, DST-7 can be applied instead of DCT-2; if not satisfied, DCT-2 can be applied.

[0362] In the ISP mode, the LFNST index can be configured such that the video information is transmitted not for each coding unit but for each partition block. In such a case, it can be determined whether the LFNST index can be signaled by regarding that there is only one partition block in the unit where the LFNST index is signaled by the above-described LFNST index signaling method.

[0363] On the other hand, the signaling order of the LFNST index and the MTS index will be considered below.

[0364] By way of example, the LFNST index signaled in residual coding can be coded next to the coding position for the last non-zero coefficient position, and the MTS index can be coded immediately after the LFNST index. In the case of such a configuration, the LFNST index can be signaled for each transform unit. Or, even if not signaled in residual coding, the LFNST index can be coded next to the coding for the last valid coefficient position, and the MTS index can be coded next to the LFNST index.

[0365] The syntax of the residual coding according to an example is as follows.

[0366]

Table 9-1

[0367]

Table 9-2

[0368] The meanings of the main variables shown in Table 9 are as follows.

[0369] 1. cbWidth, cbHeight: The width and height of the current coding block

[0370] 2. log2TbWidth, log2TbHeight: The base-2 logarithm values for the width and height of the current transform block. The zero-out is reflected, and it can be reduced to the upper left area where non-zero coefficients can exist.

[0371] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is applicable. When the flag value is 0, it indicates that LFNST is not applicable, and when the flag value is 1, it indicates that LFNST is applicable. It is defined in the Sequence Parameter Set (SPS).

[0372] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the (x0, y0) position. chType can have values of 0 and 1. 0 indicates the luma component, and 1 indicates the chroma component. The (x0, y0) position indicates the position on the picture, and the CuPredMode[chType][x0][y0] value can be MODE_INTRA (intra prediction) or MODE_INTER (inter prediction).

[0373] 5.IntraSubPartitionsSplit[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. It indicates what ISP split is applied at the position (x0, y0), and ISP_NO_SPLIT indicates that the coding unit corresponding to the position (x0, y0) is not split into partition blocks.

[0374] 6.intra_mip_flag[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. intra_mip_flag is a flag indicating whether the MIP (Matrix-based Intra Prediction) prediction mode is applied. When the flag value is 0, it indicates that MIP is not applicable, and when the flag value is 1, it indicates that MIP is applied.

[0375] 7.cIdx: A value of 0 indicates luma, and values of 1 and 2 indicate the chroma components Cb and Cr, respectively.

[0376] 8.treeType: It refers to single-tree, dual-tree, etc. (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luma component, DUAL_TREE_CHROMA: dual-tree for chroma component)

[0377] 9.tu_cbf_cb[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. It indicates the CBF (Coded Block Flag) for the Cb component. When its value is 0, it means that no non-zero coefficients exist in the corresponding transform unit for the Cb component, and when it is 1, it indicates that non-zero coefficients exist in the corresponding transform unit for the Cb component.

[0378] 10. lastSubBlock: Indicates the position in the scan order of the sub-block (Coefficient Group (CG)) where the last non-zero coefficient is located. 0 refers to the sub-block containing the DC component, and if it is greater than 0, it means it is not the sub-block containing the DC component.

[0379] 11. lastScanPos: Indicates the position in the scan order of the last non-zero coefficient within a sub-block. If a sub-block is composed of 16 positions, values from 0 to 15 are possible.

[0380] 12. lfnst_idx[x0][y0]: It is the LFNST index syntax element to be parsed. If not parsed, it is analogous to a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.

[0381] 13. LastSignificantCoeffX, LastSignificantCoeffY: Indicate the x-coordinate and y-coordinate where the last non-zero coefficient is located within the transform block. The x-coordinate starts from 0 and increases from left to right, and the y-coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means the last non-zero coefficient is located at the DC.

[0382] 14. cu_sbt_flag: It is a flag indicating whether the SubBlock Transform (SBT) currently included in the VVC standard is applicable. If the flag value is 0, it indicates that SBT is not applicable, and if the flag value is 1, it indicates that SBT is applied.

[0383] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These are flags indicating whether explicit MTS is applied to the inter-CU and intra-CU, respectively. When the corresponding flag value is 0, it indicates that MTS is not applicable to the inter-CU or intra-CU, and when it is 1, it indicates that it is applicable.

[0384] 16. tu_mts_idx[x0][y0]: It is the MTS index syntax element to be parsed. If not parsed, it is analogous to a value of 0. That is, the default value is set to 0, indicating that DCT-2 is applied to all in the horizontal and vertical directions.

[0385] As shown in Table 9, in the case of a single tree, it is possible to determine whether to signal the LFNST index based only on the condition of the last valid coefficient position for luma. That is, when the last valid coefficient position is not DC and the last valid coefficient exists inside the upper-left sub-block (CG), for example, a 4×4 block, the LFNST index is signaled. At this time, in the case of 4×4 and 8×8 transform blocks, the LFNST index is signaled only when the last valid coefficient exists at positions from 0 to 7 inside the upper-left sub-block.

[0386] In the case of a dual tree, for luma and chroma, the LFNST index is signaled independently, and in the case of chroma, the LFNST index can be signaled by applying only the condition of the last valid coefficient position to the Cb component. The corresponding condition is not checked for the Cr component. If the CBF value for Cb is 0, the LFNST index can be signaled by applying the condition of the last valid coefficient position to the Cr component.

[0387] "Min(log2TbWidth, log2TbHeight) >= 2" in Table 9 can be expressed as "Min(tbWidth, tbHeight) >= 4", and "Min(log2TbWidth, log2TbHeight) >= 4" can be expressed as "Min(tbWidth, tbHeight) >= 16".

[0388] In Table 9, log2ZoTbWidth and log2ZoTbHeight respectively mean the base-2 log values of the width and height with respect to the upper left region where the last valid coefficient can exist due to zeroing out.

[0389] As in Table 9, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before the MTS index or LFNST index value is parsed, and the second is after the parsing of the MTS index.

[0390] Since the first update is before the MTS index (tu_mts_idx[x0][y0]) value is parsed, log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.

[0391] After the MTS index is parsed, log2ZoTbWidth and log2ZoTbHeight will be set when the MTS index value is greater than 0 (in the case of the DST-7 / DCT-8 combination). When DST-7 / DCT-8 is applied independently in the horizontal and vertical directions in the first transformation, up to 16 valid coefficients can exist for each row or column in each direction. That is, after applying DST-7 / DCT-8 with a length of 32 or more, up to 16 transformation coefficients can be derived for each row or column from the left or upper side. Therefore, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions of a two-dimensional block, valid coefficients can exist only in the maximum upper left 16×16 region.

[0392] Also, when DCT-2 is currently applied independently for the horizontal and vertical directions in the first transformation, up to 32 valid coefficients can exist for each row or column in each direction. That is, when applying DCT-2 with a length of 64 or more, up to 32 transformation coefficients can be derived for each row or column starting from the left side or the upper side. Therefore, when DCT-2 is applied to both the horizontal and vertical directions of a two-dimensional block, valid coefficients can exist only up to the maximum upper-left 32×32 region.

[0393] Also, when DST-7 / DCT-8 is applied in one direction and DCT-2 is applied in the other direction for the horizontal and vertical directions, 16 valid coefficients can exist in the former direction and 32 valid coefficients can exist in the latter direction. For example, in the case of a 64×8 transformation block, where DCT-2 is applied horizontally and DST-7 is applied vertically (which can occur in a situation where implicit MTS is applied), valid coefficients can exist in the maximum upper-left 32×8 region.

[0394] If log2ZoTbWidth and log2ZoTbHeight are updated at two locations as shown in Table 9, that is, if they are updated before MTS index parsing, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the following table.

[0395]

Table 10

[0396] Also, in such a case, in the binary process for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values.

[0397]

Table 11

[0398] On the other hand, in one example, when it is in the ISP mode and LFNST is applied, when the signaling in Table 9 is applied, the spec text can be configured as shown in Table 12. When compared with Table 9, the condition where only the LFNST index is signaled for the case where it is not in the ISP mode (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 9) is deleted.

[0399] In the case of a single tree, when reusing the LFNST index transmitted when it is luma (when cIdx = 0) when it is chroma, the LFNST index transmitted for the first ISP partition block where valid coefficients exist can be applied to the chroma conversion block. Or, even if it is in the case of a single tree, the LFNST index can be signaled separately for the chroma component from the luma component. The explanations for the variables described in Table 12 are as in Table 9.

[0400]

Table 12

[0401] If, in another example, in Table 12, when it is in the ISP and it is allowed for the last valid coefficient to be located at the DC position, the parsing condition of the LFNST index can be changed as follows.

[0402]

Table 13

[0403] On the one hand, according to an example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values of 0, 1, and 2. 0 indicates not applying LFNST, and 1 and 2 indicate the first candidate and the second candidate among the two LFNST kernel candidates included in the selected LFNST set, respectively. The LFNST index is coded via truncated unary binarization, and the 0, 1, and 2 values can be coded with bin strings 0, 10, and 11, respectively.

[0404] According to an example, LFNST can be applied only when DCT-2 is applied to both the horizontal and vertical directions in the first-order transform. Therefore, if the MTS index is signaled after the LFNST index signaling, the MTS index can be signaled only when the LFNST index value is 0. When the LFNST index is not 0, the first-order transform can be performed by applying DCT-2 to both the horizontal and vertical directions without signaling the MTS index.

[0405] The MTS index value can have values of 0, 1, 2, 3, and 4. 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied to the horizontal and vertical directions, respectively. Also, the MTS index can be coded via truncated unary binarization, and the 0, 1, 2, 3, and 4 values can be coded with bin strings 0, 10, 110, 1110, and 1111, respectively.

[0406] The LFNST index and the MTS index can be signaled at the coding unit level, and the MTS index can be coded following the LFNST index at the coding unit level. The coding unit syntax table for this is as follows.

[0407]

Table 14

[0408] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 14 can be set as shown in Table 15 below.

[0409] The variable LfnstDcOnly becomes 1 when the last valid coefficient is located at all DC positions (the upper left corner position) for a transform block where the corresponding CBF (Coded Block Flag, 1 if there is at least one valid coefficient in the corresponding block, 0 otherwise) value is 1, and becomes 0 otherwise. More specifically, in the case of dual-tree luma, the position of the last valid coefficient is checked for one luma transform block, and in the case of dual-tree chroma, the position of the last valid coefficient is checked for both the transform block for Cb and the transform block for Cr. In the case of single-tree, the position of the last valid coefficient can be checked for the transform blocks for luma, Cb, and Cr.

[0410] The variable LfnstZeroOutSigCoeffFlag is 0 when there is a valid coefficient at the position where zeroing out occurs when LFNST is applied, and 1 otherwise.

[0411] The lfnst_idx[x0][y0] included in Table 14 and the following tables indicates the LFNST index for the corresponding coding unit, and tu_mts_idx[x0][y0] indicates the MTS index for the corresponding coding unit.

[0412] As shown in Table 14, the condition for checking whether the transform_skip_flag[x0][y0] value is 0 can be included in the condition for signaling lfnst_idx[x0][y0] (!transform_skip_flag[x0][y0]). In this case, the condition for checking whether the existing tu_mts_idx[x0][y0] value is 0 (i.e., checking whether it is DCT-2 in both the horizontal and vertical directions) can be omitted.

[0413] transform_skip_flag[x0][y0] indicates whether the coding unit is coded in the transform skip mode where the transform is omitted, and the flag is signaled prior to the MTS index and the LFNST index. That is, since lfnst_idx[x0][y0] is signaled before signaling the tu_mtx_idx[x0][y0] value, only the condition for the transform_skip_flag[x0][y0] value can be checked.

[0414] As shown in Table 14, when coding tu_mts_idx[x0][y0], various conditions are checked. As described above, tu_mts_idx[x0][y0] is signaled only when the lfnst_idx[x0][y0] value is 0.

[0415] Also, tu_cbf_luma[x0][y0] is a flag indicating whether there are valid coefficients for the luma component, and cbWidth and cbHeight indicate the width and height of the coding unit for the luma component, respectively.

[0416] According to Table 14, when both the width and height of the coding unit for the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, that is, whether MTS can be applied is determined by the width and height of the coding unit for the luma component.

[0417] In other examples, when transform block tiling (TU tiling) occurs (for example, when the maximum transform size is set to 32, a 64×64 coding unit is divided into 4 32×32 transform blocks for coding), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or less, the same MTS index value can be applied to all transform blocks within the coding unit and the same first transform can be applied. Also, when transform block tiling occurs, the tu_cbf_luma[x0][y0] value in Table 14 is the CBF value for the top-left transform block, or can be set to 1 if the corresponding CBF value is 1 for any one transform block among all transform blocks.

[0418] Also, according to Table 14, in the case of the ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP partition blocks.

[0419] On the other hand, tu_mts_idx[x0][y0] can be signaled only when not in the ISP mode (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT).

[0420] As shown in Table 14, when signaling the MTS index immediately after the LFNST index, information about the primary transform cannot be known when performing residual coding. That is, the MTS index is signaled after residual coding. Therefore, the part that zeros out leaving only 16 coefficients for the 32-length DST-7 or DCT-8 in the residual coding part can be changed as shown in Table 15 below.

[0421]

Table 15-1

[0422]

Table 15-2

[0423] The part that checks the tu_mts_idx[x0][y0] value can be omitted in the process of determining log2ZoTbWidth and log2ZoTbHeight as shown in Table 15 (where log2ZoTbWidth and log2ZoTbHeight respectively represent the base-2 log values of the width and height for the upper left region remaining after zeroing out).

[0424] The binary conversion for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 15 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 11.

[0425] Also, as shown in Table 15, when determining log2ZoTbWidth and log2ZoTbHeight in residual coding, a condition to check the sps_mts_enable_flag can be added.

[0426] TR in Table 11 indicates the Truncated Rice binarization method, and based on cMax and cRiceParam defined in Table 11, the last significant coefficient information can be binarized by the method described in the following table.

[0427]

Table 16

[0428] For example, if information about the position of the last significant coefficient for the luma transform block is recorded in the residual coding process, the MTS index can also be signaled as shown in Table 17.

[0429]

Table 17

[0430] In Table 17, LumaLastSignificantCoeffX and LumaLastSignificantCoeffY indicate the X coordinate and Y coordinate of the position of the last significant coefficient for the luma transform block, respectively. The condition that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY should be smaller than 16 is added to Table 17. If either one of them is 16 or more, since DCT-2 is applied in both the horizontal and vertical directions, signaling for tu_mts_idx[x0][y0] is omitted, and it can be inferred that DCT-2 is applied to both the horizontal and vertical directions.

[0431] The fact that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY are less than 16 means that the last significant coefficient exists within the top-left 16×16 region. When a 32-length DST-7 or DCT-8 is applied in the current VVC standard, it indicates that there may be a zero-out applied that leaves only 16 transform coefficients from the leftmost or topmost side. Therefore, tu_mts_idx[x0][y0] can be signaled to indicate the transform kernel used for the first-level transform.

[0432] On the other hand, in another example, the coding unit syntax table and the residual coding syntax table are as shown in the following table.

[0433]

Table 18

[0434]

Table 19

[0435] In Table 18, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 19. The variable MtsZeroOutSigCoeffFlag has its value changed from 1 to 0 if there is a significant coefficient in the region to be filled with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15). In this case, the MTS index is not signaled as in Table 19.

[0436] On the other hand, in one example, when determining log2ZoTbWidth and log2ZoTbHeight in residual coding, as shown in the following table, a condition to check sps_mts_enable_flag can be added.

[0437]

Table 20

[0438] As shown in Table 20, when tu_cbf_luma[x0][y0] is 1, MtsZeroOutSigCoeffFlag can be set to 1, and when tu_cbf_luma[x0][y0] is 0, the existing MtsZeroOutSigCoeffFlag value can be maintained. Therefore, when tu_cbf_luma[x0][y0] is 0 and the MtsZeroOutSigCoeffFlag value is maintained at 0, the coding of mts_idx[x0][y0] can be omitted. That is, when the CBF value of the luma component is 0, since no transformation is applied, the MTS index is meaningless and the coding of the MTS index can be omitted.

[0439] The following drawings are created to illustrate a specific example of this specification. Since the names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.

[0440] FIG. 22 is a flowchart showing the operation of a video decoding apparatus according to an embodiment of this document.

[0441] Each step disclosed in FIG. 22 is based on a part of the content detailed in FIGS. 5 to 21. Therefore, the specific content overlapping with the content detailed in FIGS. 3, 5 to 21 is omitted from the description or simplified.

[0442] A decoding apparatus 300 according to an embodiment can receive residual information from a bitstream (S2210).

[0443] More specifically, the decoding device 300 can decode information on the quantized transform coefficients for the current block from the bitstream, and can derive the quantized transform coefficients for the target block based on the information on the quantized transform coefficients for the current block. The information on the quantized transform coefficients for the target block can be included in the SPS (Sequence Parameter Set) or the slice header, and includes at least one of information on whether simplified transform (RST) is applied, information on the simplification factor, information on the minimum transform size to which the simplified transform is applied, information on the maximum transform size to which the simplified transform is applied, the simplified inverse transform size, and information on the transform index indicating any one of the transform kernel matrices included in the transform set.

[0444] In addition, the decoding device can further receive information on the intra prediction mode for the current block and information on whether ISP is applied to the current block. The decoding device can derive whether the current block is divided into a predetermined number of sub-partition transform blocks by receiving and parsing the flag information indicating whether to apply ISP coding or the ISP mode. Here, the current block is a coding block. Also, the decoding device can derive the size and number of the sub-partition blocks to be divided through the flag information indicating the direction in which the current block is divided.

[0445] The decoding device 300 can perform inverse quantization on the residual information for the current block, that is, the quantized transform coefficients, to derive the transform coefficients (S2220).

[0446] The derived conversion coefficients can be arranged in a reverse diagonal scan order in units of 4×4 blocks, and the conversion coefficients within a 4×4 block can also be arranged in a reverse diagonal scan order. That is, the conversion coefficients after inverse quantization can be arranged according to the reverse scan order applied in video codecs such as VVC and HEVC.

[0447] The conversion coefficients derived based on such residual information may be the conversion coefficients inverse quantized as described above, or may be the quantized conversion coefficients. That is, the conversion coefficients may be any data that can check whether the data in the current block is non-zero regardless of whether quantization is possible.

[0448] The decoding device can apply an inverse transform to the quantized conversion coefficients to derive residual samples.

[0449] As described above, the decoding device can apply LFNST, which is a non-separable transform, or MTS, which is a separable transform, to derive residual samples, and such transforms can be executed based on an LFNST kernel, that is, an LFNST index indicating an LFNST matrix, and an MTS index indicating an MTS kernel, respectively.

[0450] The decoding device can determine whether it is possible to parse an MTS index for applying MTS to the current block, and by way of example, can determine the tree type of the current block, the split type of the current block, and whether zeroing out for MTS has been executed on the current block (S2230).

[0451] When the tree type of the current block is not dual-tree chroma and the LFNST index indicating the LFNST kernel applied to the current block is 0, the decoding device determines that the MTS index is to be parsed and can parse the MTS index.

[0452] That is, when the tree type of the current block is single trigger or dual tree Luma and the LFNST index is 0, that is, when LFNST is not applied to the current block, the MTS index can be parsed.

[0453] However, when certain conditions are met even if the MTS index is not parsed, MTS can be implicitly applied. For example, when the current block is divided into sub-partition blocks, or when sub-block transform (SBT) is applied, or when the MIP (matrix-based intra prediction) mode is not applied to the intra prediction of the current block, implicit MTS can be applied.

[0454] Also, by way of an example, the decoding device can parse the MTS index when the larger value of the width and height of the current block is less than or equal to 32. That is, when the width or height of the current block is greater than 32, MTS cannot be applied.

[0455] Also, by way of an example, when the current block is not divided into a plurality of sub-partition blocks and sub-block transform that divides the coding unit of the current block and executes conversion is not applied, the MTS index can be parsed. As described above, when ISP or SBT is applied to the current block, MTS can be implicitly executed and the MTS index is not signaled.

[0456] In addition, the decoding device can parse the MTS index according to whether zeroing out for MTS has been executed. When there are valid coefficients in a second region excluding the first region at the upper left corner where valid transform coefficients can exist within the current block, it can be determined that zeroing out has not been executed. That is, when the valid coefficients do not exist in the second region, it is determined that zeroing out has been executed and the MTS index can be parsed.

[0457] The first region is a 16×16 region at the upper left corner of the current block.

[0458] The decoding device can derive a variable MtsZeroOutSigCoeffFlag that can indicate that zeroing out has been executed when applying MTS. The variable MtsZeroOutSigCoeffFlag indicates whether there are transform coefficients in the upper left region where the last valid coefficient can exist after zeroing out during MTS execution, that is, in a region other than the 16×16 region at the upper left corner. It is initially set to 1, and when there are transform coefficients in a region other than the 16×16 region, its value can be changed from 1 to 0. When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.

[0459] By way of an example, the decoding device checks the value of the transform skip flag, and when the value is 0, it can parse the MTS index.

[0460] The conditions under which the MTS index is parsed can be combined by at least two or more AND conditions. By way of example, the decoding device, when the tree type of the current block is not dual-tree chroma, the LFNST index indicating the LFNST kernel is 0, and the larger value of the width and height of the current block is less than or equal to 32, the current block is not divided into sub-partition blocks, sub-block conversion is not applied to the current block, and zeroing out by MTS execution is performed, can parse the MTS index.

[0461] The decoding device can receive and parse at least one of the LFNST index or the MTS index at the coding unit level, and can parse the LFNST index indicating the LFNST kernel prior to, that is, immediately before, the MTS index indicating the MTS kernel.

[0462] After parsing the MTS index, the decoding device can apply MTS to the current block based on the MTS index to derive residual samples for the current block (S2240).

[0463] Next, the decoding device 300 can generate restored samples based on the residual samples for the current block and the predicted samples for the current block (S2250).

[0464] The following drawings are created to illustrate a specific example of this specification. Since the names of the specific devices and the names of the specific signals / messages / fields described in the drawings are presented by way of example, the technical features of this specification are not limited to the specific names used in the following drawings.

[0465] FIG. 23 is a flowchart showing the operation of a video encoding device according to an embodiment of this document.

[0466] Each step disclosed in FIG. 23 is based on part of the content detailed in FIGS. 5 to 21. Therefore, specific content overlapping with that detailed in FIGS. 2, 5 to 21 will be omitted from the description or simplified.

[0467] An encoding device 200 according to an embodiment can derive a prediction sample for a current block based on an intra prediction mode applied to the current block (S2310).

[0468] When ISP is applied to the current block, the encoding device can perform prediction for each sub - partition conversion block.

[0469] The encoding device can determine whether to apply ISP coding or an ISP mode to the current block, i.e., the coding block. Based on the determination result, it can determine in which direction the current block is divided, and derive the size and number of sub - blocks to be divided.

[0470] The encoding device 200 can derive a residual sample for the current block based on the prediction sample (S2320).

[0471] The encoding device 200 can apply at least one of LFNST or MTS to the residual sample to derive a transform coefficient for the current block, and can arrange the transform coefficients in a predetermined scanning order. For example, based on MTS for the residual sample, the transform coefficient for the current block can be derived (S2330).

[0472] The first - order transformation can be performed via a plurality of transformation kernels such as MTS. In this case, the transformation kernel can be selected based on the intra prediction mode.

[0473] After deriving the conversion coefficient by applying MTS, the encoding device can zero out the remaining area of the current block excluding the specific area at the upper left corner of the current block, for example, a 16×16 area.

[0474] The encoding device can encode the MTS index and encode the residual information derived through quantization of the conversion coefficient based on the tree type of the current block, the split type of the current block, and whether zeroing out for MTS has been performed on the current block (S2340).

[0475] When the tree type of the current block is not dual tree chroma and the LFNST index indicating the LFNST kernel applied to the current block is 0, the encoding device can configure the video information so that the MTS index is signaled and signal the MTS index.

[0476] That is, when the tree type of the current block is single tree or dual tree luma, when the LFNST index is 0, that is, when LFNST is not applied to the current block, the encoding device can signal the MTS index.

[0477] However, when specific conditions are met even if the MTS index is not signaled, MTS can be implicitly applied. For example, when the current block is split into sub-partition blocks, or when sub-block transform (SBT) is applied, or when the MIP (matrix-based intra prediction) mode is not applied to the intra prediction of the current block, implicit MTS can be applied.

[0478] Also, according to an example, when the larger value of the width and height of the current block is less than or equal to 32, the encoding device can configure the video information so that the MTS index is signaled and can signal the MTS index. That is, when the width or height of the current block is greater than 32, MTS cannot be applied.

[0479] Also, according to an example, when sub-block conversion that divides the current block into coding units and performs conversion without dividing the current block into a plurality of sub-partition blocks is not applied, the MTS index can be signaled. As described above, when ISP or SBT is applied to the current block, MTS can be implicitly executed and the MTS index is not signaled.

[0480] Also, the encoding device can signal the MTS index according to whether zero-out for MTS has been executed. When valid coefficients exist in a second region excluding the first region at the upper left corner where valid transform coefficients can exist within the current block, it can be determined that zero-out has not been executed. That is, when the valid coefficients do not exist in the second region, it is determined that zero-out has been executed and the MTS index can be signaled.

[0481] The first region is a 16×16 region at the upper left corner of the current block.

[0482] The encoding device derives a variable MtsZeroOutSigCoeffFlag that can indicate that zeroing out has been performed when MTS is applied, and this can be composed of video information for MTS index signaling. The variable MtsZeroOutSigCoeffFlag indicates whether conversion coefficients exist in the upper left corner region where the last valid coefficient can exist due to zeroing out after MTS execution, that is, in a region other than the upper left 16×16 region. It is initially set to 1, and if conversion coefficients exist in a region other than the 16×16 region, its value can be changed from 1 to 0. When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.

[0483] By way of an example, the encoding device can check the conversion skip flag value, and if the value is 0, it can signal the MTS index.

[0484] The conditions under which the MTS index is encoded can be combined with at least two or more in an AND condition. By way of an example, the encoding device can signal the MTS index when the tree type of the current block is not dual tree chroma, the LFNST index indicating the LFNST kernel is 0, and the larger value of the width and height of the current block is less than or equal to 32, the current block is not divided into sub - partition blocks, sub - block conversion is not applied to the current block, and zeroing out by MTS execution has been performed.

[0485] The encoding device can signal at least one of the LFNST index or the MTS index at the coding unit level, and the video information can be configured such that the LFNST index indicating the LFNST kernel is signaled immediately before, that is, immediately before, the MTS index indicating the MTS kernel.

[0486] The encoding device can generate residual information including information on the quantized transform coefficients. The residual information can include the above-described transform-related information / syntax elements. The encoding device can encode video / video information including the residual information and output it in the form of a bitstream.

[0487] More specifically, the encoding device 200 can generate information on the quantized transform coefficients and encode the generated information on the quantized transform coefficients.

[0488] The video information can include an LFNST index indicating an LFNST matrix when the LFNST can be applied.

[0489] The syntax elements of the LFNST index according to this embodiment can indicate whether (inverse) LFNST is applied and any one of the LFNST matrices included in the LFNST set. When the LFNST set includes two transform kernel matrices, the value of the syntax elements of the LFNST index is three.

[0490] By way of an example, when the split tree structure for the current block is of the dual tree type, the LFNST index can be encoded for each of the luma block and the chroma block.

[0491] According to one embodiment, the syntax element value for the transform index can be derived as 0 indicating that (inverse) LFNST is not applied to the current block, 1 indicating the first LFNST matrix among the LFNST matrices, and 2 indicating the second LFNST matrix among the LFNST matrices.

[0492] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation can be omitted. When the quantization / inverse quantization is omitted, the quantized transform coefficient can be called a transform coefficient. When the transformation / inverse transformation is omitted, the transform coefficient can also be called a coefficient or a residual coefficient, or, for the sake of uniformity of expression, can still be called a transform coefficient.

[0493] Also, in this document, the quantized transform coefficient and the transform coefficient can each be called a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information can include information regarding the transform coefficient(s), and the information regarding the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information regarding the transform coefficient(s)), and the scaled transform coefficient can be derived via an inverse transformation (scaling) with respect to the transform coefficient. The residual sample can be derived based on an inverse transformation (transformation) with respect to the scaled transform coefficient. This can be similarly applied / expressed in other parts of this document.

[0494] In the above-described embodiments, the method is described based on a flowchart in a series of steps or blocks, but this document is not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from those described above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of this document.

[0495] The method according to the above-described document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0496] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.

[0497] In addition, the decoding device and the encoding device to which this document is applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a picture phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.

[0498] In addition, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (e.g., transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network. Also, the embodiments of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0499] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and embodied as a device, and the technical features of the device claims in this specification can be combined and embodied as a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and embodied as a method.

Claims

1. 1. A video decoding method performed by a decoding device, comprising: receiving residual information from a bitstream; deriving transform coefficients for a current block based on the residual information; determining whether an MTS index for applying MTS to the current block can be parsed; applying the MTS to the current block based on the MTS index to derive a residual sample for the current block; generating a reconstructed picture based on the residual samples; The step of determining whether the MTS index can be parsed includes: determining a tree type of the current block, whether the current block is divided into a plurality of sub-partition blocks, whether LFNST is performed on the current block, and whether zero-out is performed on the MTS of the current block; Whether the LFNST is performed on the current block is determined based on an LFNST index associated with an LFNST kernel to be applied to the current block; The LFNST index and the MTS index are signaled in a coding unit level syntax; The MTS index is signaled immediately after the LFNST index is signaled in the coding unit level syntax.

2. The video decoding method of claim 1 , wherein the MTS index is parsed based on the fact that a tree type of the current block is not dual tree chroma and the LFNST index is 0.

3. The video decoding method of claim 2 , wherein the MTS index is parsed based on whether a larger value of a width and a height of the current block is less than or equal to 32.

4. The video decoding method of claim 3, wherein the MTS index is parsed based on the fact that the current block is not divided into the plurality of sub-partition blocks and that a sub-block transformation that divides a coding unit and performs a transformation is not applied to the current block.

5. Determining whether zeroing out has been performed on the MTS includes: determining whether a valid coefficient exists in a second region excluding a first region at an upper left corner of the current block in which a valid coefficient can exist; The video decoding method of claim 4 , wherein the MTS index is parsed based on the absence of the significant coefficients in the second region.

6. A video encoding method performed by a video encoding apparatus, comprising: deriving a predicted sample for the current block; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients for the current block based on an MTS for the residual samples; and encoding residual information derived via quantization of the transform coefficients and an MTS index associated with an MTS kernel; The MTS index is encoded based on a tree type of the current block, whether the current block is divided into a plurality of sub-partition blocks, whether LFNST is performed on the current block, and whether zero-out for the MTS is performed on the current block; Whether the LFNST is performed on the current block is determined based on an LFNST index associated with an LFNST kernel to be applied to the current block; The LFNST index and the MTS index are signaled in a coding unit level syntax; The video encoding method, wherein the MTS index is signaled immediately after the signaling of the LFNST index in the coding unit level syntax.

7. The video encoding method of claim 6 , wherein the MTS index is encoded based on the fact that a tree type of the current block is not dual tree chroma and the LFNST index is 0.

8. The video encoding method of claim 7 , wherein the MTS index is encoded based on whether a larger value of a width and a height of the current block is smaller than or equal to 32.

9. the current block is not divided into the plurality of sub-partition blocks, Since a sub-block transformation is not applied to the current block, the sub-block transformation is performed by dividing the current block into coding units. The video encoding method of claim 8 , wherein the MTS index is encoded.

10. Determining whether zeroing out for the MTS is performed includes: determining whether a valid coefficient exists in a second region excluding a first region at an upper left corner of the current block in which a valid coefficient can exist; The video encoding method of claim 9 , wherein the MTS index is encoded based on the absence of the significant coefficients in the second region.

11. A method for generating and transmitting data for video information, comprising: generating a bitstream of the video information including residual information, the residual information being generated by deriving a prediction sample for a current block, deriving a residual sample for the current block based on the prediction sample, deriving a transform coefficient for the current block based on an MTS for the residual sample, and encoding the derived residual information via a quantization of the transform coefficient and an MTS index associated with an MTS kernel to generate the bitstream; transmitting said data including a bitstream of said video information; The MTS index is encoded based on a tree type of the current block, whether the current block is divided into a plurality of sub-partition blocks, whether LFNST is performed on the current block, and whether zero-out for the MTS is performed on the current block; Whether the LFNST is performed on the current block is determined based on an LFNST index associated with an LFNST kernel to be applied to the current block; The LFNST index and the MTS index are signaled in a coding unit level syntax; The method of claim 1, wherein the MTS index is signaled immediately after signaling the LFNST index in the coding unit level syntax.