Video coding method and apparatus based on transformation

The video coding method employs LFNST and MTS to enhance coding efficiency for high-resolution video data, addressing the challenges of efficient compression and transmission of high-quality video content.

JP7697070B2Active Publication Date: 2025-06-23LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024001253
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2024-01-09
Publication Date
2025-06-23
Estimated Expiration
2040-10-05

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video data, particularly in fields like VR, AR, and game video, has led to a need for more efficient video coding technologies to effectively compress, transmit, store, and reproduce such data.

Method used

A video coding method and apparatus that utilizes LFNST (Low-Frequency Non-Separable Transform) and MTS (Multiple Transform Selection) to improve coding efficiency. This involves deriving conversion coefficients for a current block, determining the presence of valid coefficients in specific regions, parsing MTS and LFNST indices, and applying these indices to derive residual samples.

Benefits of technology

The proposed solution enhances overall video compression efficiency, improves the efficiency of transform index coding, and provides a method for effective LFNST and MTS index signaling, thereby addressing the challenges of high-resolution video data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697070000036
    Figure 0007697070000036
  • Figure 0007697070000037
    Figure 0007697070000037
  • Figure 0007697070000038
    Figure 0007697070000038
Patent Text Reader

Abstract

To provide a method and a device that increase image coding efficiency.SOLUTION: An image decoding method according to the present document includes steps of: deriving a transform coefficient for a current block on the basis of residual information; determining whether an effective coefficient is in a second area that excludes an upper left first area of the current block; parsing an MTS index from a bitstream if the effective coefficient is not in the second area; and deriving a residual sample by applying, to transform coefficients of the first area, an MTS kernel derived on the basis of the MTS index.SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to video coding technology, and more particularly, to a video coding method and apparatus based on transform in a video coding system.

Background Art

[0002] Recently, the demand for high-resolution and high-quality video / video such as 4K or 8K and above UHD (Ultra High Definition) video / video has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] Also, recently, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from those of real-world video, such as game video, has been increasing.

[0004] Accordingly, in order to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality video / video having various characteristics as described above, a highly efficient video / video compression technology is required.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The technical problem of this document is to provide a method and apparatus for increasing video coding efficiency.

[0006] Another technical problem of this document is to provide a method and apparatus for increasing the efficiency of transform index coding.

[0007] Another technical problem of this document is to provide a video coding method and apparatus utilizing LFNST and MTS.

[0008] Another technical problem of this document is to provide a video coding method and apparatus for LFNST index and MTS index signaling.

Means for Solving the Problems

[0009] According to an embodiment of this document, a video decoding method executed by a decoding apparatus is provided. The method includes steps of deriving a conversion coefficient for a current block based on residual information, determining whether there are valid coefficients in a second region excluding a first region at the upper left corner of the current block, parsing an MTS index from the bitstream when there are no valid coefficients in the second region, and deriving residual samples by applying an MTS kernel derived based on the MTS index to the conversion coefficients of the first region.

[0010] The first region is a 16×16 region at the upper left corner of the current block.

[0011] The step of determining whether there are valid coefficients in the second region includes deriving a variable value indicating whether there are valid coefficients in the second region in the decoding process of the residual coding level. The variable value is initially set to 1. When there are valid coefficients in the second region, the variable is changed to 0 and the MTS index is not parsed.

[0012] The MTS index is parsed at the coding unit level.

[0013] Further comprising the step of parsing an LFNST index indicating the LFNST kernel applied to the current block, wherein the LFNST index and the MTS index are signaled at the coding unit level, and the MTS index is signaled immediately after the signaling of the LFNST index.

[0014] When the tree type of the current block is dual tree luma or single tree, and the value of the LFNST index is 0, the MTS index is signaled.

[0015] According to one embodiment of the present document, there is provided a video encoding method executed by an encoding device. The method includes the steps of deriving a residual sample for the current block based on a prediction sample, deriving a transform coefficient for the current block based on an MTS for the residual sample, zeroing out a second region of the current block excluding a first region at the upper left end of the current block, and encoding residual information derived through quantization of the transform coefficient and an MTS index indicating an MTS kernel.

[0016] According to another embodiment of the present document, there is provided a digital storage medium storing video data including encoded video information and a bitstream generated by a video encoding method executed by an encoding device.

[0017] According to another embodiment of the present document, there is provided a digital storage medium storing video data including encoded video information and a bitstream for causing a decoding device to execute the video decoding method.

Advantages of the Invention

[0018] According to the present document, the overall video / video compression efficiency can be improved.

[0019] According to this document, the efficiency of conversion index coding can be improved.

[0020] According to this document, a video coding method and apparatus using LFNST and MTS can be provided.

[0021] According to this document, a video coding method and apparatus for LFNST index and MTS index signaling can be provided.

[0022] The effects obtainable through a specific example of this specification are not limited to the effects listed above. For example, there can be various technical effects that a person having ordinary skill in the related art can understand or derive from this specification. Accordingly, the specific effects of this specification are not limited to what is explicitly described in this specification, but can include various effects that can be understood or derived from the technical features of this specification.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Embodiments for Carrying Out the Invention

[0024] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms used in this specification are merely used to explain specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "including" or "having" are used to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should not be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc., are precluded in advance.

[0025] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among each configuration, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0026] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0027] This document relates to video / video coding. For example, the methods / examples disclosed in this document are associated with the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0028] This document presents various examples related to video / video coding, and unless otherwise mentioned, the examples can be combined and executed with each other.

[0029] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can include one or more tiles.

[0030] A pixel or pel can mean the smallest unit that constitutes a picture (or video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, can also indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component. Or, a sample can also mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can also mean a conversion coefficient in the frequency domain.

[0031] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In a general case, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0032] In this document, " / " and "、" are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" also means "at least one of A, B, and / or C".

[0033] Further, in this document, "or" is interpreted as "and / or". For example, "A or B" can mean 1) only "A", or 2) only "B", or 3) "A and B". As another expression, "or" in this document can mean "additionally or alternatively".

[0034] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0035] Also, in this specification, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "at least one of A, B, or C" and "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0036] Also, the parentheses used in this specification can mean "for example". Specifically, when it is shown as "prediction (intra-prediction)", "intra-prediction" is proposed as an example of "prediction". As another expression, "prediction" in this specification is not limited to "intra-prediction", but "intra-prediction" is proposed as an example of "prediction". Also, when it is shown as "prediction (i.e., intra-prediction)", "intra-prediction" is proposed as an example of "prediction".

[0037] In this specification, the technical features separately described within one drawing can be embodied separately or simultaneously.

[0038] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0039] Referring to FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in a file or streaming form.

[0040] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can be composed of a separate device or an external component.

[0041] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced by the process of generating related data.

[0042] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0043] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or through a digital storage medium in streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file through a pre-determined file format and can include elements for transmission through a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0044] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0045] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0046] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the video encoding device can include the image encoding device.

[0047] Referring to FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0048] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to this document can be executed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the final coding unit described above.The prediction unit is a unit of sample prediction, and the conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0049] A unit can, in some cases, be used interchangeably with terms such as a block or an area. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).

[0050] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform a prediction on the block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0051] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the level of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0052] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can configure a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, no residual signal is transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and by signaling the motion vector difference, the motion vector of the current block can be indicated.

[0053] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in a way similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0054] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square or can be applied to a block of variable size that is not square.

[0055] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The signaling / transmitted information and / or syntax elements described later in this document can be encoded through the above-described encoding procedure and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmitting unit (not shown) for transmission and / or a storing unit (not shown) for storage, or the transmitting unit can also be included in the entropy encoding unit 240.

[0056] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and can also be used for inter prediction of the next picture after being filtered as described later.

[0057] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0058] The filtering unit 260 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. As will be described later in the description of each filtering method, the filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.

[0059] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.

[0060] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.

[0061] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.

[0062] Referring to FIG. 3, the decoding apparatus 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 may include an inter-prediction unit 332 and an intra-prediction unit 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filtering unit 350 described above may be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 may include a DPB (decoded picture buffer) and may also be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0063] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block splitting related information obtained from the bitstream. The decoding device 300 can execute decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be split from a coding tree unit or a maximum coding unit by a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.

[0064] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / video information can further include general constraint information. The decoding device can decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the information of the syntax element to be decoded adjacent to and the decoding information of the block to be decoded or the symbol / bin information decoded in the previous step, predicts the occurrence probability of the bin by the determined context model, and generates a symbol corresponding to the value of each syntax element by performing arithmetic decoding of the bin.At this time, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the information regarding the residual for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0065] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output the transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient by using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.

[0066] In the inverse conversion unit 322, the conversion coefficient is inversely converted to obtain a residual signal (residual block, residual sample array).

[0067] The prediction unit can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0068] The prediction unit can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can perform intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed to be similar to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document.

[0069] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to or away from the current block depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0070] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0071] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0072] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and as will be described later, can also be output after filtering, or can be used for inter prediction of the next picture.

[0073] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0074] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0075] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block from which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0076] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding apparatus 300 can be applied so as to be identical or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding apparatus 200, respectively.

[0077] As described above, in performing video coding, prediction is performed to increase the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is similarly derived in the encoding apparatus and the decoding apparatus, and the encoding apparatus can increase the video coding efficiency by signaling information (residual information) regarding the residual between the original block, which is not the original sample value of the original block itself, and the predicted block to the decoding apparatus. The decoding apparatus can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0078] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse transform the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.

[0079] Figure 4 schematically shows the multiple conversion techniques according to this document.

[0080] Referring to Figure 4, the conversion unit can correspond to the conversion unit in the encoding device of Figure 2 described above, and the inverse conversion unit can correspond to the inverse conversion unit in the encoding device of Figure 2 or the inverse conversion unit in the decoding device of Figure 3 described above.

[0081] The conversion unit can perform a primary conversion based on the residual samples (residual sample array) in the residual block to derive (primary) conversion coefficients (S410). Such a primary conversion can be called a core transform. Here, the primary conversion can be based on Multiple Transform Selection (MTS), and when a multiple conversion is applied as the primary conversion, it can be called a multiple core transform.

[0082] The multiple core transform can be shown as a method of performing conversion by additionally using DCT (Discrete Cosine Transform) type 2, DST (Discrete Sine Transform) type 7, DCT type 8, and / or DST type 1. That is, the multiple core transform can be shown as a conversion method for converting a residual signal (or residual block) in the spatial domain into conversion coefficients (or primary conversion coefficients) in the frequency domain based on a plurality of conversion kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the primary conversion coefficients can be called temporary conversion coefficients from the perspective of the conversion unit.

[0083] That is, when the existing conversion method is applied, based on DCT type 2, a conversion from the spatial domain to the frequency domain for the residual signal (or residual block) can be applied to generate conversion coefficients. In contrast, when the multi-core conversion is applied, based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc., a conversion from the spatial domain to the frequency domain for the residual signal (or residual block) can be applied to generate conversion coefficients (or primary conversion coefficients). Here, DCT type 2, DST type 7, DCT type 8, and DST type 1, etc. can be called conversion types, conversion kernels or conversion cores. Such DCT / DST conversion types can be defined based on basis functions.

[0084] When the multi-core conversion is executed, a vertical conversion kernel and a horizontal conversion kernel for the target block can be selected from among the conversion kernels, and a vertical conversion for the target block can be executed based on the vertical conversion kernel, and a horizontal conversion for the target block can be executed based on the horizontal conversion kernel. Here, the horizontal conversion can indicate a conversion for the horizontal component of the target block, and the vertical conversion can indicate a conversion for the vertical component of the target block. The vertical conversion kernel / horizontal conversion kernel can be adaptively determined based on the prediction mode and / or conversion index of the target block (CU or sub-block) including the residual block.

[0085] Also, according to an example, when applying MTS to perform a primary transformation, specific basis functions are set to predetermined values, and when it is a vertical transformation or a horizontal transformation, a mapping relationship for the transformation kernel can be set by combining which basis functions are applied. For example, when the horizontal transformation kernel is represented by trTypeHor and the vertical transformation kernel is represented by trTypeVer, the trTypeHor or trTypeVer value 0 can be set to DCT2, the trTypeHor or trTypeVer value 1 can be set to DST7, and the trTypeHor or trTypeVer value 2 can be set to DCT8.

[0086] In this case, in order to indicate any one of a number of transformation kernel sets, MTS index information can be encoded and signaled to the decoding device. For example, when the MTS index is 0, it indicates that both the trTypeHor and trTypeVer values are 0; when the MTS index is 1, it indicates that both the trTypeHor and trTypeVer values are 1; when the MTS index is 2, it indicates that the trTypeHor value is 2 and the trTypeVer value is 1; when the MTS index is 3, it indicates that the trTypeHor value is 1 and the trTypeVer value is 2; when the MTS index is 4, it can indicate that both the trTypeHor and trTypeVer values are 2.

[0087] According to an example, when showing the transformation kernel set by the MTS index information in a table, it is as follows.

[0088] [Table 1]

[0089] The conversion unit can derive a corrected (secondary) conversion coefficient by performing a secondary conversion based on the (primary) conversion coefficient (S420). The primary conversion is a conversion from the spatial domain to the frequency domain, and the secondary conversion means converting in a more compressed representation by utilizing the correlation existing between the (primary) conversion coefficients. The secondary conversion can include a non-separable transform. In this case, the secondary conversion can be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can indicate a conversion that performs a secondary conversion on the (primary) conversion coefficient derived through the primary conversion based on a non-separable transform matrix to generate a corrected conversion coefficient (or secondary conversion coefficient) for the residual signal. Here, based on the non-separable transform matrix, the conversion can be applied at once without separating the vertical conversion and the horizontal conversion (or independently applying the horizontal and vertical conversions) to the (primary) conversion coefficient. That is, the non-separable secondary transform does not apply separately in the vertical and horizontal directions to the (primary) conversion coefficient. For example, after rearranging a two-dimensional signal (conversion coefficient) into a one-dimensional signal through a specified direction (e.g., row-first direction or column-first direction), it can indicate a conversion method of generating a corrected conversion coefficient (or secondary conversion coefficient) based on the non-separable transform matrix. For example, the row-first order is to arrange in a column in the order of the first row, the second row,..., the Nth row for an M×N block, and the column-first order is to arrange in a column in the order of the first column, the second column,..., the Mth column for an M×N block. The non-separable secondary transform can be applied to the top-left region of a block composed of (primary) conversion coefficients (hereinafter, which can be called a conversion coefficient block).For example, when both the width (W) and height (H) of the conversion coefficient block are 8 or more, an 8×8 non-separable second-order conversion can be applied to the upper left 8×8 region of the conversion coefficient block. Also, when both the width (W) and height (H) of the conversion coefficient block are 4 or more and the width (W) or height (H) of the conversion coefficient block is less than 8, a 4×4 non-separable second-order conversion can be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block. However, the embodiments are not limited thereto. For example, even when only the condition that both the width (W) and height (H) of the conversion coefficient block are 4 or more is satisfied, a 4×4 non-separable second-order conversion can also be applied to the upper left min(8, W)×min(8, H) region of the conversion coefficient block.

[0090] Specifically, for example, when a 4×4 input block is used, the non-separable second-order conversion can be performed as follows.

[0091] The 4×4 input block X is shown as follows.

[0092]

Number

[0093] When representing X in vector form, the vector TIFF0007697070000003.tif84 is shown as follows.

[0094]

Number

[0095] As in Equation 2, the vector TIFF0007697070000005.tif84 rearranges the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.

[0096] In this case, the secondary non-separable transformation can be calculated as follows.

[0097] [Number]

[0098] Here, TIFF0007697070000007.tif75 represents the transformation coefficient vector, and T represents the 16×16 (non-separable) transformation matrix.

[0099] The 16×1 transformation coefficient vector TIFF0007697070000008.tif75 can be derived through the above Equation 3, and the TIFF0007697070000009.tif75 can be re-organized in 4×4 blocks through the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is only an example, and for reducing the computational complexity of the non-separable secondary transformation, HyGT (Hypercube-Givens Transform) etc. can also be used for the calculation of the non-separable secondary transformation.

[0100] On the other hand, for the non-separable secondary transformation, a mode-dependent transformation kernel (or transformation core, transformation type) can be selected. Here, the mode can include an intra prediction mode and / or an inter prediction mode.

[0101] As described above, the non-separable second-order transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined based on the width (W) and height (H) of the transformation coefficient block. The 8×8 transformation refers to a transformation that can be applied to an 8×8 region included within the corresponding transformation coefficient block when both W and H are greater than or equal to 8, and the corresponding 8×8 region is the upper left 8×8 region within the corresponding transformation coefficient block. Similarly, the 4×4 transformation refers to a transformation that can be applied to a 4×4 region included within the corresponding transformation coefficient block when both W and H are greater than or equal to 4, and the corresponding 4×4 region is the upper left 4×4 region within the corresponding transformation coefficient block. For example, the 8×8 transformation kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transformation kernel matrix can be a 16×16 / 8×16 matrix.

[0102] At this time, for mode-based transformation kernel selection, two non-separable second-order transformation kernels can be configured for each transformation set for non-separable second-order transformation for both the 8×8 transformation and the 4×4 transformation, and the number of transformation sets is four. That is, four transformation sets can be configured for the 8×8 transformation, and four transformation sets can be configured for the 4×4 transformation. In this case, each of the four transformation sets for the 8×8 transformation can include two 8×8 transformation kernels, and in this case, each of the four transformation sets for the 4×4 transformation can include two 4×4 transformation kernels.

[0103] However, the size of the transformation, that is, the size of the region to which the transformation is applied, is merely exemplary, and sizes other than 8×8 or 4×4 can be used, the number of the sets is n, and the number of transformation kernels within each set is k.

[0104] The transformation set can be called an NSST set or an LFNST set. The selection of a specific set from the transformation sets can be performed, for example, based on the intra prediction mode of the current block (CU or sub-block). LFNST (Low-Frequency Non-Separable Transform) is an example of a reduced non-separable transform described later and represents a non-separable transform for low-frequency components.

[0105] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is only an example, and this document can also be applied when the number of intra prediction modes is different. On the other hand, in some cases, the 67th intra prediction mode can be further used, and the 67th intra prediction mode can represent the LM (linear model) mode.

[0106] FIG. 5 exemplarily shows the intra-directional modes of 65 prediction directions.

[0107] Referring to FIG. 5, it is possible to distinguish an intra prediction mode having horizontal directionality and an intra prediction mode having vertical directionality, centered around the 34th intra prediction mode having a downward right diagonal prediction direction. H and V in FIG. 5 respectively represent horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. This can indicate an offset with respect to the mode index value. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. On the other hand, the 34th intra prediction mode can be regarded as not strictly horizontal or vertical directionality, but can be classified as belonging to the horizontal directionality from the perspective of determining the conversion set of the secondary conversion. This is because for the vertical direction mode that is symmetric around the 34th intra prediction mode, the input data is transposed and used, and for the 34th intra prediction mode, the input data alignment method for the horizontal direction mode is used. Transposing the input data means that for the two-dimensional block data M×N, the rows become columns and the columns become rows to form N×M data. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode has a left reference pixel and predicts in the upward right direction, so it can be called an upward right diagonal intra prediction mode. In the same context, the 34th intra prediction mode is called a downward right diagonal intra prediction mode, and the 66th intra prediction mode can be called a downward left diagonal intra prediction mode.

[0108] By way of example, the mapping of four conversion sets by the intra prediction mode is shown, for example, as in the following table.

[0109]

Table 2

[0110] As shown in Table 2, depending on the intra prediction mode, it can be mapped to any one of the four transform sets, that is, lfnstTrSetIdx can be mapped to any one of 0 to 3, that is, any one of the four.

[0111] On the other hand, when it is determined that a specific set is used for the non-separable transform, one of the k transform kernels in the specific set can be selected via the non-separable second-order transform index. The encoding device can derive a non-separable second-order transform index indicating a specific transform kernel based on rate-distortion (RD) checking, and can signal the non-separable second-order transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable second-order transform index. For example, the lfnst index value 0 can indicate the first non-separable second-order transform kernel, the lfnst index value 1 can indicate the second non-separable second-order transform kernel, and the lfnst index value 2 can indicate the third non-separable second-order transform kernel. Or, the lfnst index value 0 can indicate that the first non-separable second-order transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate the three transform kernels.

[0112] The transform unit can execute the non-separable second-order transform based on the selected transform kernel to obtain the corrected (second-order) transform coefficients. The corrected transform coefficients can be derived as the transform coefficients quantized via the quantization unit as described above, and can be encoded and signaled to the decoding device and transmitted to the inverse quantization / inverse transform unit in the encoding device.

[0113] On the one hand, as described above, when the secondary conversion is omitted, the (primary) conversion coefficient, which is the output of the primary (separation) conversion, can be derived as the quantization coefficient quantized through the quantization unit as described above, encoded, and signaled to the decoding device and transmitted to the inverse quantization / inverse conversion unit in the encoding device.

[0114] The inverse conversion unit can execute a series of procedures in the reverse order of the procedures executed by the conversion unit described above. The inverse conversion unit receives the (inverse quantized) conversion coefficient, executes the secondary (inverse) conversion to derive the (primary) conversion coefficient (S450), and can execute the primary (inverse) conversion on the (primary) conversion coefficient to obtain the residual block (residual sample) (S460). Here, the primary conversion coefficient can be called the modified conversion coefficient from the perspective of the inverse conversion unit. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0115] On the other hand, the decoding device can further include a secondary inverse conversion applicability determination unit (or an element that determines the applicability of the secondary inverse conversion) and a secondary inverse conversion determination unit (or an element that determines the secondary inverse conversion). The secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion. For example, the secondary inverse conversion is NSST, RST, or LFNST, and the secondary inverse conversion applicability determination unit can determine the applicability of the secondary inverse conversion based on the secondary conversion flag parsed from the bitstream. As another example, the secondary inverse conversion applicability determination unit can also determine the applicability of the secondary inverse conversion based on the conversion coefficient of the residual block.

[0116] The second inverse transform determination unit can determine the second inverse transform. At this time, the second inverse transform determination unit can determine the second inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified by the intra prediction mode. Also, as an example, the second transform determination method can be determined depending on the first transform determination method. Various combinations of the first transform and the second transform can be determined by the intra prediction mode. Also, as an example, the second inverse transform determination unit can also determine the area to which the second inverse transform is applied based on the size of the current block.

[0117] On the other hand, as described above, when the second (inverse) transform is omitted, a residual block (residual sample) can be obtained by receiving the (inverse quantized) transform coefficient and performing the first (separation) inverse transform. As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based on this.

[0118] On the other hand, in this document, in order to reduce the amount of calculation and memory requirements by the non-separable second transform, RST (reduced secondary transform) in which the size of the transform matrix (kernel) is reduced in the concept of NSST can be applied.

[0119] On the one hand, the conversion kernel, conversion matrix, and coefficients constituting the conversion kernel matrix described in this document, i.e., kernel coefficients or matrix coefficients, can be represented in 8 bits. This is one of the conditions for implementation in a decoding device and an encoding device. Along with a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit ones, the memory requirement for storing the conversion kernel can be reduced. Also, by representing the kernel matrix in 8 bits, a small multiplier can be used, and it can be more compatible with SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.

[0120] In this specification, RST can be meant to be the conversion performed on the residual samples for the target block based on a transform matrix whose size has been reduced by a simplification factor. When performing the simplified conversion, the amount of computation required during the conversion can be reduced due to the reduction in the size of the transform matrix. That is, RST can be utilized to solve the problem of computational complexity that occurs during the conversion of large-sized blocks or non-separable conversions.

[0121] RST can be called by various terms such as reduced conversion, reduction transform, reduced secondary transform, simplified transform, simple transform, etc. The name called RST is not limited to the listed examples. Or, since RST is mainly performed in the low-frequency region containing non-zero coefficients in the conversion block, it can also be called LFNST (Low-Frequency Non-Separable Transform). The conversion index can be named as the LFNST index.

[0122] On the other hand, when the second inverse transformation is performed based on RST, the inverse transformation unit 235 of the encoding device 200 and the inverse transformation unit 322 of the decoding device 300 may include an inverse RST unit that derives a modified transformation coefficient based on the inverse RST for the transformation coefficient, and an inverse primary transformation unit that derives a residual sample for the target block based on the inverse primary transformation for the modified transformation coefficient. The inverse primary transformation means the inverse transformation of the primary transformation applied to the residual. In this document, deriving a transformation coefficient based on a transformation can mean deriving the transformation coefficient by applying the corresponding transformation.

[0123] FIG. 6 is a diagram for explaining RST according to an embodiment of this document.

[0124] In this specification, the "target block" can mean the current block or the residual block or the transformation block on which coding is executed.

[0125] In the RST according to an embodiment, a reduced transformation matrix can be determined in which an N - dimensional vector is mapped to an R - dimensional vector located in a different space, where R is smaller than N. N can mean the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value. The simplification factor can be called by various terms such as reduced factor, reduction factor, simplified factor, simple factor, etc. On the other hand, R can be called a reduced coefficient, but in some cases, the simplification factor can also mean R. Also, in some cases, the simplification factor can mean the N / R value.

[0126] In one embodiment, the simplification factor or coefficient can be signaled via a bitstream, but the embodiment is not limited thereto. For example, there may be cases where predefined values for the simplification factor or coefficient are stored in each encoding device 200 and decoding device 300. In this case, the simplification factor or coefficient is not signaled separately.

[0127] The size of the simplification transform matrix according to one embodiment is R×N, which is smaller than the size N×N of the normal transform matrix, and can be defined as in Equation 4 below.

[0128]

Equation

[0129] The matrix T in the Reduced Transform block shown in FIG. 6(a) can represent the matrix TR×N of Equation 4. As shown in FIG. 6(a), when the simplification transform matrix TR×N is multiplied by the residual samples for the target block, the transform coefficients for the target block can be derived.

[0130] In one embodiment, when the size of the block to which the transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6(a) can be expressed by the matrix operation as in Equation 5 below. In this case, the multiplication operation with the memory can be reduced to approximately 1 / 4 by the simplification factor.

[0131] In this document, matrix multiplication can be understood as an operation of multiplying a matrix by a column vector with the matrix placed on the left side of the column vector to obtain a column vector.

[0132]

Equation

[0133] In Equation 5, r1 to r 64 can represent the residual samples for the target block, and more specifically, are the conversion coefficients generated by applying a primary conversion. The calculation result of Equation 5, the conversion coefficient c i for the target block can be derived, and the derivation process of c i is as shown in Equation 6.

[0134] [Number]

[0135] The calculation result of Equation 6, the conversion coefficients c1 to c R for the target block can be derived. That is, when R = 16, the conversion coefficients c1 to c 16 for the target block can be derived. If a normal (regular) conversion instead of RST is applied and a conversion matrix with a size of 64×64 (N×N) is multiplied by a residual sample with a size of 64×1 (N×1), 64 (N) conversion coefficients for the target block are derived. However, because RST is applied, only 16 (R) conversion coefficients for the target block are derived. Since the total number of conversion coefficients for the target block decreases from N to R and the amount of data transmitted from the encoding device 200 to the decoding device 300 decreases, the transmission efficiency between the encoding device 200 - decoding device 300 can be increased.

[0136] Considering from the perspective of the size of the conversion matrix, the size of the normal conversion matrix is 64×64 (N×N), and the size of the simplified conversion matrix decreases to 16×64 (R×N). Therefore, when compared with the time of performing a normal conversion, the memory usage can be decreased by the ratio of R / N when performing RST. Also, when compared with the number of multiplication operations N×N when using a normal conversion matrix, when using a simplified conversion matrix, the number of multiplication operations can be decreased by the ratio of R / N (R×N).

[0137] In one embodiment, the conversion unit 232 of the encoding device 200 can derive conversion coefficients for a target block by performing a primary conversion and an RST-based secondary conversion on the residual samples for the target block. Such conversion coefficients can be transmitted to the inverse conversion unit of the decoding device 300. The inverse conversion unit 322 of the decoding device 300 can derive modified conversion coefficients based on an inverse RST (reduced secondary transform) for the conversion coefficients, and can derive residual samples for the target block based on an inverse primary conversion for the modified conversion coefficients.

[0138] Inverse RST matrix T according to one embodiment N×R has a size of N×R, which is smaller than the size N×N of a normal inverse conversion matrix, and is in a transpose relationship with the simplified conversion matrix T shown in Equation 4. R×N

[0139] The matrix T in the Reduced Inv.Transform block shown in FIG. 6(b) t can mean the inverse RST matrix T R×N T (the superscript T means transpose). As shown in FIG. 6(b), when the inverse RST matrix T R×N T is multiplied by the conversion coefficients for the target block, modified conversion coefficients for the target block or residual samples for the target block can be derived. The inverse RST matrix T R×N T can also be expressed as (T R×N ) T N×R .

[0140] More specifically, when inverse RST is applied as the secondary inverse conversion, the inverse RST matrix T R×N TWhen applied, a modified conversion coefficient for the target block can be derived. On the other hand, inverse RST can be applied as an inverse linear transformation. In this case, when the inverse RST matrix T R×N T is applied, residual samples for the target block can be derived.

[0141] In one embodiment, when the size of the block to which the inverse transformation is applied is 8×8 and R = 16 (i.e., when R / N = 16 / 64 = 1 / 4), the RST according to FIG. 6(b) can be expressed by matrix operations such as the following Equation 7.

[0142]

Equation

[0143] In Equation 7, c1 to c 16 can represent the conversion coefficients for the target block. The calculation result of Equation 7, r i indicating the modified conversion coefficient for the target block or the residual samples for the target block, can be derived, and the derivation process of r i is as shown in Equation 8.

[0144]

Equation

[0145] The calculation result of Equation 8, r1 to r NIt can be derived. Considering from the perspective of the size of the inverse transformation matrix, the size of the normal inverse transformation matrix is 64×64 (N×N), and the size of the simplified inverse transformation matrix is reduced to 64×16 (N×R). Therefore, when compared with the time of performing the normal inverse transformation, the memory usage can be reduced by the ratio of R / N when performing the inverse RST. Also, when compared with the number of multiplication operations N×N when using the normal inverse transformation matrix, when using the simplified inverse transformation matrix, the number of multiplication operations can be reduced by the ratio of R / N (N×R).

[0146] On the other hand, for 8×8 RST as well, a conversion set configuration as shown in Table 2 can be applied. That is, the corresponding 8×8 RST can be applied by the conversion set in Table 2. Since one conversion set is composed of two or three conversions (kernels) depending on the prediction mode within the screen, it can be configured to select one from a maximum of four conversions including not applying the secondary conversion. The conversion when the secondary conversion is not applied can be regarded as the identity matrix being applied. When indexes 0, 1, 2, and 3 are assigned to each of the four conversions (for example, the 0th index can be assigned when the identity matrix, i.e., when the secondary conversion is not applied), a syntax element called the conversion index or lfnst index can be signaled for each conversion coefficient block to specify the conversion to be applied. That is, through the conversion index, for the 8×8 upper left block, in the RST configuration, 8×8 RST can be specified, or when LFNST is applied, 8×8 lfnst can be specified. 8×8 lfnst and 8×8 RST refer to conversions that can be applied to the 8×8 region contained within the corresponding conversion coefficient block when both the W and H of the target block to be converted are greater than or equal to 8, and the corresponding 8×8 region is the upper left 8×8 region within the corresponding conversion coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to conversions that can be applied to the 4×4 region contained within the corresponding conversion coefficient block when both the W and H of the target block are greater than or equal to 4, and the corresponding 4×4 region is the upper left 4×4 region within the corresponding conversion coefficient block.

[0147] On the one hand, according to one embodiment of this document, in the conversion of the encoding process, for 64 data constituting an 8×8 region, only 48 data that are not a 16×64 conversion kernel matrix can be selected, and a maximum 16×48 conversion kernel matrix can be applied. Here, "maximum" means that the maximum value of m for an m×48 conversion kernel matrix that can generate m coefficients is 16. That is, when performing RST by applying an m×48 conversion kernel matrix (m≤16) to an 8×8 region, 48 data inputs can be received to generate m coefficients. When m is 16, 48 data inputs are received to generate 16 coefficients. That is, when 48 data form a 48×1 vector, a 16×1 vector can be generated by multiplying a 16×48 matrix and a 48×1 vector in sequence. At this time, 48 data forming an 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be formed based on 48 data constituting a region excluding the lower right 4×4 region of the 8×8 region. At this time, when performing matrix operations by applying a maximum 16×48 conversion kernel matrix, 16 modified conversion coefficients are generated, and the 16 modified conversion coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with 0s.

[0148] For the inverse transformation in the decoding process, the transposed matrix of the conversion kernel matrix described above can be used. That is, when inverse RST or LFNST is performed as the inverse transformation process executed in the decoding device, the input coefficient data to which the inverse RST is applied is composed of a one-dimensional vector in a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the inverse RST matrix corresponding to the one-dimensional vector on the left side can be arranged in a two-dimensional block in a predetermined arrangement order.

[0149] Upon sorting, in the conversion process, when RST or LFNST is applied to an 8×8 region, among the conversion coefficients of the 8×8 region, matrix multiplication is performed between the 48 conversion coefficients in the upper left, upper right, and lower left regions excluding the lower right region of the 8×8 region and a 16×48 conversion kernel matrix. For the matrix multiplication, the 48 conversion coefficients are input in a one-dimensional array. When such matrix multiplication is performed, 16 modified conversion coefficients are derived, and the modified conversion coefficients can be arranged in the upper left region of the 8×8 region.

[0150] Conversely, in the inverse conversion process, when inverse RST or LFNST is applied to an 8×8 region, among the conversion coefficients of the 8×8 region, the 16 conversion coefficients corresponding to the upper left of the 8×8 region can be input in a one-dimensional array form in scanning order and subjected to matrix multiplication with a 48×16 conversion kernel matrix. That is, the matrix multiplication in such a case can be represented as (48×16 matrix) * (16×1 conversion coefficient vector) = (48×1 modified conversion coefficient vector). Here, since an n×1 vector can be interpreted in the same sense as an n×1 matrix, it can also be denoted as an n×1 column vector. Also, * means matrix multiplication operation. When such matrix multiplication is performed, 48 modified conversion coefficients can be derived, and the 48 modified conversion coefficients can be arranged in the upper left, upper right, and lower left regions excluding the lower right region of the 8×8 region.

[0151] On the other hand, when the secondary inverse conversion is performed based on RST, the inverse conversion unit 235 of the encoding device 200 and the inverse conversion unit 322 of the decoding device 300 can include an inverse RST unit that derives modified conversion coefficients based on inverse RST for the conversion coefficients, and an inverse primary conversion unit that derives residual samples for the target block based on inverse primary conversion for the modified conversion coefficients. The inverse primary conversion means the inverse conversion of the primary conversion applied to the residue. In this document, deriving conversion coefficients based on a conversion can mean deriving conversion coefficients by applying the corresponding conversion.

[0152] Looking specifically at the detailed non-separable transform, LFNST, it is as follows. LFNST can include a forward transform by an encoding device and an inverse transform by a decoding device.

[0153] The encoding device takes as input the result (or a part of the result) derived after applying a forward primary transform and applies a forward secondary transform.

[0154] [Equation 9] y = G T x

[0155] In Equation 9 above, x and y are the input and output of the secondary transform respectively, and G is a matrix representing the secondary transform, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows × number of columns], in the case of the forward LFNST, G is the transpose of matrix G. T becomes the dimension.

[0156] In the case of the inverse LFNST, the dimension of matrix G is [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices obtained by sampling 8 transform basis vectors from the left side of the [48×16] matrix and the [16×16] matrix respectively.

[0157] On the other hand, in the case of the forward LFNST, matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transform basis vectors from the upper side of the [16×48] matrix and the [16×16] matrix respectively.

[0158] Therefore, in the case of the forward LFNST, the input x can be a [48×1] vector or a [16×1] vector, and the output y can be a [16×1] vector or an [8×1] vector. Since the output of the forward first-order transform in video coding and decoding is two-dimensional (2D) data, in order to form a [48×1] vector or a [16×1] vector as the input x, the 2D data that is the output of the forward transform must be appropriately arranged to form a one-dimensional vector.

[0159] Figure 7 is a diagram showing the order of arranging the output data of the forward first-order transform into a one-dimensional vector by way of an example. The left diagrams of (a) and (b) in Figure 7 show the order for creating a [48×1] vector, and the right diagrams of (a) and (b) in Figure 7 show the order for creating a [16×1] vector. In the case of LFNST, the 2D data can be sequentially arranged in the order as shown in (a) and (b) of Figure 7 to obtain a one-dimensional vector x.

[0160] The arrangement direction of the output data of such a forward first-order transform can be determined by the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of (a) in Figure 7, and when the intra prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward first-order transform can be arranged in the order of (b) in Figure 7.

[0161] By way of an example, an arrangement order different from the arrangement orders of (a) and (b) in Figure 7 can be applied. When trying to derive the same result (y vector) as when applying the arrangement orders of (a) and (b) in Figure 7, the column vectors of the matrix G can be rearranged according to the corresponding arrangement order. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.

[0162] Since the output y derived through Equation 9 is a one-dimensional vector, if a configuration that processes the result of the forward second-order transformation as input, for example, a configuration that performs quantization or residual coding, requires two-dimensional data as input data, the output y vector of Equation 9 must be properly arranged again in 2D data.

[0163] Figure 8 is a diagram showing the order of arranging the output data of the forward second-order transformation in a two-dimensional block by way of an example.

[0164] In the case of LFNST, it can be arranged in a 2D block according to a determined scan order. (a) of Figure 8 shows that when the output y is a [16×1] vector, the output values are arranged in the 16 positions of the two-dimensional block in a diagonal scan order. (b) of Figure 8 shows that when the output y is an [8×1] vector, the output values are arranged in the 8 positions of the two-dimensional block in a diagonal scan order, and the remaining 8 positions are filled with 0. X in (b) of Figure 8 indicates that it is filled with 0.

[0165] In other examples, since the order in which the output vector y is processed by a configuration that performs quantization or residual coding can be executed according to a preset order, the output vector y may not be arranged in a 2D block as shown in Figure 8. However, in the case of residual coding, data coding can be executed in units of 2D blocks such as CG (Coefficient Group) (for example, 4×4), and in this case, the data can be arranged in a specific order such as the diagonal scan order in Figure 8.

[0166] On the other hand, the decoding device can arrange the two-dimensional data output through an inverse quantization process or the like for inverse transformation in a preset scan order to form a one-dimensional input vector y. The input vector y can be output as the input vector x according to the following equation.

[0167] [Equation 10] x = Gy

[0168] In the case of the reverse-direction LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. In the case of the reverse-direction LFNST, the output vector x is a [48×1] vector or a [16×1] vector.

[0169] The output vector x is arranged in 2D blocks and arrayed as 2D data in the order shown in FIG. 7, and such 2D data becomes the input data (or a part of the input data) for the inverse 1D transform.

[0170] Therefore, the inverse 2D transform is overall the opposite of the forward 2D transform process. In the case of the inverse transform, unlike the forward direction, the inverse 1D transform is applied after the inverse 2D transform is applied first.

[0171] In the reverse-direction LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transformation matrix G. Which matrix of the [48×16] matrix and the [16×16] matrix to apply is determined by the block size as well.

[0172] Also, the eight matrices can be derived from four transformation sets as shown in Table 2 described above, and each transformation set can be composed of two matrices. Which transformation set among the four transformation sets to use is determined by the intra prediction mode. More specifically, the transformation set is determined based on the intra prediction mode value extended considering up to the Wide Angle Intra Prediction (WAIP). Which matrix to select from the two matrices constituting the selected transformation set is derived through index signaling. More specifically, the index values that can be transmitted are 0, 1, and 2. 0 indicates not to apply LFNST, and 1 and 2 can indicate either one of the two transformation matrices constituting the transformation set selected based on the intra prediction mode value.

[0173] FIG. 9 is a diagram showing the wide angle intra prediction mode according to an embodiment of the present document.

[0174] General intra prediction mode values can have values from 0 to 66 and from 81 to 83. As shown in the figure, the intra prediction mode values extended by WAIP can have values from -14 to 83. The values from 81 to 83 indicate the CCLM (Cross Compoonent Linear Model) mode, and the values from -14 to -1 and from 67 to 80 indicate the intra prediction mode values extended by the application of WAIP.

[0175] When the width of the predicted current block is larger than the height, generally the upper reference pixels are closer to the position inside the block to be predicted. Therefore, predicting in the bottom - left direction is more accurate than predicting in the top - right direction. On the contrary, when the height of the block is larger than the width, generally the left reference pixels are closer to the position inside the block to be predicted. Therefore, predicting in the top - right direction is more accurate than predicting in the bottom - left direction. Therefore, it is advantageous to apply remapping with the index of the wide - angle intra - prediction mode, that is, mode index conversion.

[0176] When wide - angle intra - prediction is applied, information for existing intra - prediction can be signaled. After the information is parsed, the information can be remapped with the index of the wide - angle intra - prediction mode. Therefore, the total number of intra - prediction modes for a specific block (for example, a non - square block of a specific size) does not change, that is, the total number of intra - prediction modes is 67, and the intra - prediction mode coding for the specific block does not change.

[0177] Table 3 below shows the process of deriving the intra - mode modified by remapping the intra - prediction mode to the wide - angle intra - prediction mode.

[0178]

Table 3

[0179] In Table 3, the intra prediction mode value finally extended to the predModeIntra variable is stored. ISP_NO_SPLIT indicates that the CU block is not split into sub - partitions by the Intra Sub Partitions (ISP) technology currently adopted in the VVC standard. The fact that the cIdx variable value is 0, 1, or 2 refers to the cases of the luma, Cb, and Cr components respectively. The Log2 function shown in Table 3 returns a log value with base 2, and the Abs function returns the absolute value.

[0180] As input values for the wide angle intra prediction mode mapping process, variables such as predModeIntra indicating the intra prediction mode, the height and width of the transform block, etc. are used, and the output value is the modified intra prediction mode (predModeIntra). The height and width of the transform block or coding block can be the height and width of the current block for the remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to height can be set to Abs(Log2(nW / nH)).

[0181] For non - square blocks, the intra prediction mode can be classified and modified in two cases.

[0182] First, when all of the following conditions are met: (1) the width of the current block is greater than the height, (2) the intra prediction mode before correction is greater than or equal to 2, and (3) the intra prediction mode is less than (whRatio>1)?(8+2*whRatio):8, the intra prediction mode is set equal to (predModeIntra+65).

[0183] Otherwise, when all of the following conditions are met: (1) the height of the current block is greater than the width, (2) the intra prediction mode before correction is less than or equal to 66, and (3) the intra prediction mode is greater than (whRatio>1)?(60-2*whRatio):60, the intra prediction mode is set equal to (predModeIntra-67).

[0184] Table 2 described above shows how the conversion set is selected based on the intra prediction mode values extended by WAIP in LFNST. As shown in FIG. 9, the modes from 14 to 33 and the modes from 35 to 80 are symmetric with respect to the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric around the direction corresponding to mode 34. Therefore, the same conversion set is applied to the modes located in symmetric directions, and such symmetry is also reflected in Table 2.

[0185] However, it is assumed that the forward LFNST input data for mode 54 is symmetric with the forward LFNST input data for mode 14. For example, for mode 14 and mode 54, the two-dimensional data is rearranged into one-dimensional data according to the array orders shown in FIGS. 7(a) and 7(b) respectively, and it can be known that the pattern of the orders shown in FIGS. 7(a) and 7(b) is symmetric about the direction (diagonal term) indicated by mode 34.

[0186] On the other hand, as described above, which conversion matrix of the [48×16] matrix and the [16×16] matrix is applied to the LFNST is also determined according to the size of the block to be converted.

[0187] FIG. 10 is a diagram showing the appearance of the block to which the LFNST is applied. FIG. 10(a) shows a 4×4 block, FIG. 10(b) shows 4×8 and 8×4 blocks, FIG. 10(c) shows 4×N or N×4 blocks where N is 16 or more, FIG. 10(d) shows an 8×8 block, and FIG. 10(e) shows an M×N block where M≧8, N≧8, and N〉8 or M〉8.

[0188] In FIG. 10, the blocks with thick frames indicate the areas to which the LFNST is applied. For the blocks in FIGS. 10(a) and 10(b), the LFNST is applied to the top-left 4×4 area, and for the block in FIG. 10(c), the LFNST is applied to each of the two continuously arranged top-left 4×4 areas. In FIGS. 10(a), 10(b), and 10(c), the LFNST is applied in units of 4×4 areas, so such an LFNST is hereinafter named "4×4 LFNST", and in the corresponding conversion matrix, a matrix dimension for G in Formulas 9 and 10 can be applied based on a [16×16] or [16×8] matrix.

[0189] More specifically, for the 4×4 block ((a) in FIG. 10, 4×4 TU or 4×4 CU), a [16×8] matrix is applied, and for the blocks in FIGS. 10(b) and 10(c), a [16×16] matrix is applied. This is to match the computational complexity for the worst case to 8 multiplications per sample.

[0190] For FIGS. 10(d) and 10(e), LFNST is applied to the upper left 8×8 region, and such LFNST is hereinafter named "8×8 LFNST". In the corresponding transformation matrix, a [48×16] or [48×8] matrix can be applied. In the case of the forward LFNST, since a [48×1] vector (the x vector in Equation 9) is input as the input data, all sample values in the upper left 8×8 region are not used as the input values of the forward LFNST. That is, as can be seen in the left order of FIG. 7(a) or the left order of FIG. 7(b), the bottom-right 4×4 block is left as it is, and a [48×1] vector can be constructed based on the samples belonging to the remaining three 4×4 blocks.

[0191] For the 8×8 block (8×8 TU or 8×8 CU) in FIG. 10(d), a [48×8] matrix can be applied, and for the 8×8 block in FIG. 10(e), a [48×16] matrix can be applied. This is also to match the computational complexity for the worst case to 8 multiplications per sample.

[0192] Similarly for the block, when the corresponding forward LFNST (4×4 LFNST or 8×8 LFNST) is applied, 8 or 16 output data (the y vector in Equation 9, [8×1] or [16×1] vector) are generated. Due to the characteristics of the matrix GT in the forward LFNST, the number of output data is the same as or less than the number of input data.

[0193] FIG. 11 shows an array of output data of the forward LFNST by way of an example, and is a diagram showing a block in which the output data of the forward LFNST is arranged by way of a block pattern.

[0194] The shaded area processed at the upper left end of the block shown in FIG. 11 corresponds to the area where the output data of the forward LFNST is located. The positions indicated by 0 indicate samples filled with 0 values, and the remaining areas indicate areas that are not changed by the forward LFNST. In the areas not changed by the LFNST, the output data of the forward first-order transform exists as it is without being changed.

[0195] As described above, since the dimension of the transformation matrix applied by way of a block pattern changes, the number of output data also changes. As in FIG. 11, there may be cases where the output data of the forward LFNST cannot fill all of the upper left 4×4 blocks. In the cases of FIGS. 11(a) and 11(d), a [16×8] matrix and a [48×8] matrix are applied to the blocks indicated by thick lines or partial areas inside the blocks, respectively, to generate an [8×1] vector as the output of the forward LFNST. That is, only 8 output data are filled as shown in FIGS. 11(a) and 11(d) according to the scan order shown in FIG. 8(b), and the remaining 8 positions can be filled with 0. In the case of the LFNST application block of FIG. 10(d), the two 4×4 blocks at the upper right end and the lower left end adjacent to the upper left 4×4 block as in FIG. 11(d) are also filled with 0 values.

[0196] As described above, basically, the LFNST index is signaled to specify whether the LFNST is applicable and the transformation matrix to be applied. As shown in FIG. 11, when the LFNST is applied, since the number of output data of the forward LFNST may be the same as or less than the number of input data, areas filled with 0 values are generated as follows.

[0197] 1) Positions after the 8th position in the scan order within the upper left 4×4 block as in FIG. 11(a), that is, samples from the 9th to the 16th

[0198] 2) As shown in (d) and (e) of FIG. 11, a [16×48] matrix or an [8×48] matrix is applied to two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scan order.

[0199] Therefore, when checking the regions of 1) and 2) and non-zero data exists, it is certain that LFNST is not applied, so the signaling of the corresponding LFNST index can be omitted.

[0200] By way of example, for instance, in the case of LFNST adopted in the VVC standard, since the signaling of the LFNST index is executed after residual coding, the encoding device can know the presence or absence of non-zero data (valid coefficients) for all positions inside the TU or CU block via residual coding. Therefore, the encoding device can determine whether to execute the signaling for the LFNST index based on the presence or absence of non-zero data, and the decoding device can determine whether to parse the LFNST index. If no non-zero data exists in the regions specified in 1) and 2), the signaling of the LFNST index will be executed.

[0201] To apply the truncated unary code to the LFNST index in the binary evolution method, the LFNST index is composed of a maximum of two bins, and the binary codes for the possible LFNST index values of 0, 1, and 2 are assigned 0, 10, and 11 respectively. In the case of the LFNST currently adopted in VVC, context-based CABAC coding (regular coding) is applied to the first bin, and bypass coding is applied to the second bin. The total number of contexts for the first bin is two, and the primary transform pair (DCT-2, DCT-2) is applied for the horizontal and vertical directions. When the luma and chroma components are coded in a dual-tree type, one context is assigned, and for other cases, another context is applied. Representing the coding of such an LFNST index in a table is as follows.

[0202]

Table 4

[0203] On the other hand, for the adopted LFNST, the following simplification method can be applied.

[0204] (i) By way of example, the number of output data for the forward LFNST can be limited to a maximum of 16.

[0205] In the case of (c) in FIG. 10, 4×4 LFNST can be applied to two 4×4 regions adjacent to the upper left end, and at this time, up to 32 LFNST output data can be generated. If the number of output data for the forward LFNST is limited to a maximum of 16, the 4×4 LFNST is applied only to one 4×4 region existing at the upper left end for a 4×N / N×4 (N≧16) block (TU or CU), and the LFNST can be applied only once to all the blocks in FIG. 10. Through this, the implementation for video coding becomes simple.

[0206] FIG. 12 is a diagram showing an example in which the number of output data for the forward LFNST is limited to a maximum of 16. As shown in FIG. 12, when the LFNST is applied to the leftmost upper 4×4 region in a 4×N or N×4 block where N is 16 or more, the output data of the forward LFNST becomes 16.

[0207] (ii) Additionally, zero-out can be applied to the regions where LFNST is not applied, by way of example. In this document, zero-out can be meant to fill the values of all positions belonging to a specific region with 0 values. That is, zero-out can be applied to the regions that are not changed by LFNST and maintain the result of the forward first-order transformation. As described above, since LFNST is classified into 4×4 LFNST and 8×8 LFNST, zero-out can be classified into two types ((ii)-(A) and (ii)-(B)) as follows.

[0208] (ii)-(A) When 4×4 LFNST is applied, the regions where 4×4 LFNST is not applied can be zeroed out. FIG. 13 is a diagram showing zero-out in a block where 4×4 LFNST is applied, by way of example.

[0209] As shown in FIG. 13, for the block to which 4×4 LFNST is applied, that is, all the regions where LFNST is not applied to the blocks (a), (b), and (c) in FIG. 11 can be filled with 0.

[0210] On the other hand, (d) of FIG. 13 shows that zeroing out is performed on the remaining blocks to which the 4×4 LFNST is not applied when the maximum value of the number of output data of the forward LFNST is limited to 16 as shown in FIG. 12.

[0211] (ii)-(B) When the 8×8 LFNST is applied, the area where the 8×8 LFNST is not applied can be zeroed out. FIG. 14 is a diagram showing zeroing out in a block to which the 8×8 LFNST is applied by way of example.

[0212] As shown in FIG. 14, for the block to which the 8×8 LFNST is applied, that is, the areas in (d) and (e) of FIG. 11 where the LFNST is not applied can all be filled with 0s.

[0213] (iii) When the LFNST is applied by the zeroing out presented in (ii) above, the area filled with 0s can be changed. Therefore, it is possible to check whether there is non-zero data in a wider area than in the case of the LFNST of FIG. 11 with respect to the zeroing out proposed in (ii) above.

[0214] For example, when applying (ii)-(B), after checking whether there is non-zero data up to the area additionally filled with 0s in FIG. 14 in addition to the areas filled with 0 values in (d) and (e) of FIG. 11, signaling for the LFNST index can be performed only when there is no non-zero data.

[0215] Of course, even if the zero-out proposed in (ii) above is applied, it is possible to check whether there is non-zero data as in the existing LFNST index signaling. That is, it is possible to check whether there is non-zero data for the blocks filled with 0 in FIG. 11 and apply the LFNST index signaling. In such a case, the zero-out is executed only in the encoding device, and in the decoding device, without assuming the corresponding zero-out, that is, it is possible to execute the LFNST index parsing by checking only whether there is non-zero data for the regions explicitly marked with 0 in FIG. 11.

[0216] Or, according to another example, zero-out can also be executed as shown in FIG. 15. FIG. 15 is a diagram showing zero-out in a block to which 8×8 LFNST is applied according to another example.

[0217] As shown in FIGS. 13 and 14, zero-out can be applied to all regions other than the regions to which LFNST is applied, and it is also possible to apply zero-out only to partial regions as shown in FIG. 15. Zero-out is applied only to the regions other than the upper left 8×8 region in FIG. 15, and zero-out is not applied to the lower right 4×4 block inside the upper left 8×8 region.

[0218] Various embodiments can be derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) for the LFNST. Of course, the combinations for the simplification methods are not limited to the following embodiments, and any combination can be applied to the LFNST.

[0219] Embodiment

[0220] - Limit the number of output data for the forward LFNST to a maximum of 16 → (i)

[0221] - When 4×4 LFNST is applied, zero out all regions where 4×4 LFNST is not applied → (ii)-(A)

[0222] - When 8×8 LFNST is applied, all areas where 8×8 LFNST is not applied are zeroed out → (ii)-(B)

[0223] - For areas filled with existing 0 values and areas zeroed out additionally ((ii)-(A), (ii)-(B)), after checking whether there is non-zero data in the areas filled with 0, LFNST indexing signaling is performed only when there is no non-zero data → (iii)

[0224] In the case of the above embodiment, when LFNST is applied, the area where non-zero output data can exist is limited to the inner part of the upper left 4×4 area. More specifically, in the cases of (a) in FIG. 13 and (a) in FIG. 14, in the scan order, the 8th position is the last position where non-zero data can exist, and in the cases of (b) and (d) in FIG. 13 and (b) in FIG. 14, in the scan order, the 16th position (i.e., the outermost position at the lower right end of the upper left 4×4 block) is the last position where non-zero data can exist.

[0225] Therefore, when LFNST is applied, after checking whether there is non-zero data at a position where the residual coding process is not allowed (a position beyond the last position), it is possible to determine whether LFNST index signaling is possible.

[0226] (In the case of the zeroing-out method proposed in (ii), in order to reduce the number of data that will finally occur when all primary conversions and LFNST are applied, the computational amount required when executing the overall conversion process can be reduced. That is, when LFNST is applied, zeroing out is also applied to the forward primary conversion output data existing in the area where LFNST is not applied, so there is no need to generate data for the area that will be zeroed out from the time of executing the forward primary conversion. Therefore, the amount of calculation required for generating the corresponding data can be saved. Summarizing the additional effects of the zeroing-out method proposed in (ii), it is as follows.)

[0227] First, as described above, the amount of computation required for the execution of the overall conversion process is reduced.

[0228] In particular, when applying (ii)-(B), the amount of computation for the worst case can be reduced, and the conversion process can be lightweighted. To elaborate, generally, a large amount of operations are required for the execution of a primary conversion of a large size. When (ii)-(B) is applied, the number of data derived as the result of the forward LFNST execution can be reduced to 16 or less. As the size of the overall block (TU or CU) increases, the effect of reducing the conversion operation amount is further increased.

[0229] Second, the amount of operations required for the entire conversion process is reduced, and the power consumption required for the conversion execution can be reduced.

[0230] Third, the latency associated with the conversion process is reduced.

[0231] A secondary conversion such as LFNST adds computation to the existing primary conversion, thus increasing the overall latency associated with the conversion execution. Especially in the case of intra prediction, since the restored data of adjacent blocks is used in the prediction process, the increase in latency due to the secondary conversion during encoding can lead to an increase in latency until reconstruction, which can lead to an overall increase in latency of the intra prediction encoding.

[0232] However, when applying the zeroing out presented in (ii), the latency of the primary conversion execution can be significantly reduced when applying LFNST. Therefore, the latency for the entire conversion execution is maintained or reduced, and the encoding device can be implemented more easily.

[0233] On the one hand, conventional intra prediction performs encoding without division by treating the block to be currently encoded as one encoding unit. However, ISP (Intra Sub-Paritions) coding means performing intra prediction encoding by dividing the block to be currently encoded in the horizontal or vertical direction. At this time, encoding / decoding is performed in units of the divided blocks to generate restored blocks, and the restored blocks can then be used as reference blocks for the next divided blocks. By way of example, during ISP coding, one coding block can be divided into 2 or 4 sub-blocks for coding, and for one sub-block in ISP, intra prediction is performed by referring to the restored pixel values of the sub-blocks located adjacent to the left or adjacent to the upper side. Hereinafter, the "coding" used can be used in a concept that includes all the coding performed by the encoding device and the decoding performed by the decoding device.

[0234] Table 5 shows the number of sub-blocks divided according to the block size when ISP is applied, and the sub-partitions divided by ISP can be called transform blocks (TUs).

[0235]

Table 5

[0236] ISP is to divide the block predicted by luma intra in the vertical or horizontal direction into 2 or 4 sub-partitionings according to the block size. For example, the minimum block size to which ISP can be applied is 4×8 or 8×4. If the block size is larger than 4×8 or 8×4, the block is divided into 4 sub-partitionings.

[0237] FIG. 16 and FIG. 17 show an example of sub - blocks into which one coding block is divided. More specifically, FIG. 16 is an illustration of the division when the coding block (width (W)×height (H)) is a 4×8 block or an 8×4 block, and FIG. 17 is an illustration of the division when the coding block is not a 4×8 block, an 8×4 block, or a 4×4 block.

[0238] When applying ISP, the sub - blocks are coded sequentially, for example, horizontally or vertically, from left to right or from top to bottom according to the division form. After performing the inverse transformation and intra - prediction for one sub - block and reaching the restoration process, the coding for the next sub - block can proceed. For the left - most or top - most sub - block, it refers to the restored pixels of the already - coded coding block in the same way as the normal intra - prediction method. Also, when each side of the subsequent internal sub - block is not adjacent to the previous sub - block, in order to derive the reference pixels adjacent to the corresponding side, it refers to the restored pixels of the adjacent already - coded coding block in the same way as the normal intra - prediction method.

[0239] In the ISP coding mode, all sub - blocks can be coded in the same intra - prediction mode, and flags indicating whether to use ISP coding and in which direction (horizontal or vertical) to divide can be signaled. As shown in FIGS. 16 and 17, the number of sub - blocks can be adjusted to 2 or 4 according to the block. When the size (width×height) of one sub - block is less than 16, the division into the corresponding sub - block can be prohibited, or the application of ISP coding itself can be restricted.

[0240] On the other hand, in the case of the ISP prediction mode, one coding unit is divided into 2 or 4 partition blocks, that is, sub - blocks, and predicted. The same in - picture prediction mode is applied to the corresponding 2 or 4 divided partition blocks.

[0241] As described above, in the splitting direction, there are both a horizontal direction (when an M×N coding unit with horizontal length M and vertical length N is split horizontally, if it is split into two, it is split into M×(N / 2) blocks, and if it is split into four, it is split into M×(N / 4) blocks) and a vertical direction (when an M×N coding unit is split vertically, if it is split into two, it is split into (M / 2)×N blocks, and if it is split into four, it is split into (M / 4)×N blocks). When split horizontally, the partition blocks are coded in the order from the upper side to the lower side, and when split vertically, the partition blocks are coded in the order from the left side to the right side. The currently coded partition block can be predicted by referring to the restored pixel values of the upper (left) partition block in the case of horizontal (vertical) direction splitting.

[0242] The conversion can be applied to the residual signal generated by the ISP prediction method in units of partition blocks. Based on the forward direction, not only the existing DCT-2 but also the DST-7 / DCT-8 combination-based MTS (Multiple Transform Selection) technology can be applied to the primary transform (core transform or primary transform), and the forward LFNST (Low Frequency Non-Separable Transform) can be applied to the transform coefficients generated by the primary transform to generate the final corrected transform coefficients.

[0243] That is, LFNST can also be applied to the partition blocks that are divided when the ISP prediction mode is applied. As described above, the same intra prediction mode is applied to the divided partition blocks. Therefore, when selecting the LFNST set derived based on the intra prediction mode, the LFNST sets derived for all partition blocks can be applied. That is, since the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks accordingly.

[0244] On the other hand, by way of an example, LFNST can be applied only to transform blocks where both the horizontal and vertical lengths are 4 or more. Therefore, when the horizontal or vertical length of the partition block divided by the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled either. Also, when applying LFNST to each partition block, the corresponding partition block can be regarded as one transform block. Of course, when the ISP prediction method is not applied, LFNST can be applied to the coding block.

[0245] Looking specifically at applying LFNST to each partition block, it is as follows.

[0246] By way of an example, after applying the forward LFNST to an individual partition block, only up to 16 (8 or 16) coefficients are left in the left-upper 4×4 region according to the transform coefficient scanning order, and then zeroing out where the remaining positions and regions are all filled with 0 values can be applied.

[0247] Or, by way of an example, when the length of one side of the partition block is 4, LFNST is applied only to the left-upper 4×4 region, and when the lengths of all sides of the partition block, that is, the width and height, are 8 or more, LFNST can be applied to the remaining 48 coefficients excluding the right-lower 4×4 region inside the left-upper 8×8 region.

[0248] Alternatively, by way of an example, in order to combine the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4×4 or 8×8, only 8 transform coefficients can be output after the forward LFNST application. That is, when the partition block is 4×4, an 8×16 matrix can be applied with the transform matrix, and when the partition block is 8×8, an 8×48 matrix can be applied with the transform matrix.

[0249] On the other hand, in the current VVC standard, LFNST index signaling is performed on a coding unit basis. Therefore, in the ISP prediction mode, when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition block. That is, when the LFNST index value is transmitted once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks within the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates the case where LFNST is not applied, and 1 and 2 refer to two transform matrices existing within one LFNST set when LFNST is applied.

[0250] As described above, the LFNST set is determined by the intra prediction mode. In the case of the ISP prediction mode, since all partition blocks within the coding unit are predicted in the same intra prediction mode, the partition block can refer to the same LFNST set.

[0251] As another example, although the LFNST indexing is still performed on a coding unit basis, in the case of the ISP prediction mode, instead of uniformly determining whether to apply LFNST to all partition blocks, it is possible to determine whether to apply the LFNST index value signaled at the coding unit level to each partition block or not to apply LFNST to each partition block via separate conditions. Here, the separate conditions can be signaled in flag form for each partition block via the bitstream. When the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST is not applied.

[0252] On the other hand, in the coding unit to which the ISP mode is applied, when the length of one side of the partition block is less than 4, for an example of applying LFNST, it is as follows.

[0253] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper left M×2 (2×M) region (where M≦N). For example, when M = 8, since the corresponding upper left region is 8×2 (2×8), the region where 16 residual signals exist can be the input of the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.

[0254] Here, the forward LFNST matrix is a separate additional matrix that is not the matrix currently included in the VVC standard. Also, for complexity adjustment in the worst case, an 8×16 matrix sampled only from the upper 8 row vectors of the 16×16 matrix can be used for the transformation. The complexity adjustment method will be described in detail later.

[0255] Second, when the size of the partition block is N×1 (1×N), the LFNST can be applied to the upper left M×1 (1×M) region (where M≦N). For example, when M = 16, the corresponding upper left region is 16×1 (1×16), so the region where 16 residual signals exist can be the input of the forward LFNST, and an R×16 (R≦16) forward transformation matrix can be applied.

[0256] Here, the corresponding forward LFNST matrix is a separate additional matrix that is not the matrix currently included in the VVC standard. Also, for worst-case complexity adjustment, an 8×16 matrix sampled from only the upper 8 row vectors of the 16×16 matrix can be used for the transformation. The complexity adjustment method will be described in detail later.

[0257] The first embodiment and the second embodiment can be applied simultaneously, or only one of the two embodiments can be applied. In particular, in the case of the second embodiment, by considering a one-dimensional transformation for the LFNST, it has been observed through experiments that the improvement in compression performance that can be obtained with the existing LFNST is not relatively large compared to the LFNST index signaling cost. However, in the case of the first embodiment, an improvement in compression performance similar to that obtained with the existing LFNST has been observed. That is, in the case of ISP, it can be confirmed through experiments that the application of LFNST for 2×N and N×2 contributes to the actual compression performance.

[0258] Currently, symmetry between intra prediction modes is applied in the LFNST of VVC. The same LFNST set is applied to two directional modes arranged around mode 34 (prediction in the lower right 45-degree diagonal direction). For example, the same LFNST set is applied to mode 18 (horizontal prediction mode) and mode 50 (vertical prediction mode). However, for modes 35 to 66, when applying the forward LFNST, the input data is transposed and then the LFNST is applied.

[0259] On the other hand, in VVC, the Wide Angle Intra Prediction (WAIP) mode is supported, and the LFNST set is derived based on the intra prediction mode modified in consideration of the WAIP mode. For the modes extended by WAIP, the symmetry is utilized in the same way as the general intra prediction direction mode to determine the LFNST set. For example, since mode -1 is symmetric with mode 67, the same LFNST set is applied, and since mode -14 is symmetric with mode 80, the same LFNST set is applied. For modes from 67 to 80, after transposing the input data before applying the forward LFNST, the LFNST transform is applied.

[0260] In the case of the LFNST applied to the upper left M×2 (M×1) block, the symmetry for the aforementioned LFNST cannot be applied. The reason is that the block to which the LFNST is applied is non-square. Therefore, instead of applying the symmetry based on the intra prediction mode like the LFNST in Table 2, the symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.

[0261] FIG. 18 is a diagram showing the symmetry between an M×2 (M×1) block and a 2×M (1×M) block according to an example.

[0262] As shown in FIG. 18, since mode 2 in the M×2 (M×1) block can be regarded as symmetric with mode 66 in the 2×M (1×M) block, the same LFNST set can be applied to the 2×M (1×M) block and the M×2 (M×1) block.

[0263] At this time, in order to apply the LFNST set applied to the M×2 (M×1) block to the 2×M (1×M) block, the LFNST set is selected based on Mode 2 instead of Mode 66. That is, after transposing the input data of the 2×M (1×M) block before applying the forward LFNST, the LFNST can be applied.

[0264] FIG. 19 is a diagram showing an example of transposing a 2×M block.

[0265] FIG. 19(a) is a drawing for explaining that the input data is read in column-first order for the 2×M block and the LFNST is applied, and FIG. 19(b) is a diagram for explaining that the input data is read in row-first order for the M×2 (M×1) block and the LFNST is applied. Organizing the method of applying the LFNST to the upper left M×2 (M×1) or 2×M (M×1) block, it is as follows.

[0266] 1. First, as shown in FIGS. 19(a) and 19(b), the input data is arranged to form the input vector of the forward LFNST. For example, referring to FIG. 18, for the M×2 block predicted in Mode 2, it follows the order in FIG. 19(b), and for the 2×M block predicted in Mode 66, after arranging the input data according to the order in FIG. 19(a), the LFNST set for Mode 2 can be applied.

[0267] 2. For the M×2 (M×1) block, the LFNST set is determined based on the modified intra prediction mode considering WAIP. As described above, a preset mapping relationship is established between the intra prediction mode and the LFNST set, which can be represented by a mapping table as shown in Table 2.

[0268] For the 2×M (1×M) block, after obtaining the modes symmetric about the prediction mode in the 45-degree diagonal direction to the lower right (mode 34 in the case of the VVC standard) from the intra prediction modes modified considering WAIP, the LFNST set is determined based on the corresponding symmetric modes and the mapping table. The mode (y) symmetric about mode 34 can be derived through the following formula. The mapping table will be further specifically described below.

[0269] [Equation 11] if 2≤x≤66, y = 68 - x, otherwise (x≤ - 1 or x≥67), y = 66 - x

[0270] When applying the forward LFNST, the input data prepared through process 1 can be multiplied by the LFNST kernel to derive the conversion coefficients. The LFNST kernel can be selected from the LFNST set determined in process 2 and the pre-specified LFNST index.

[0271] For example, when M = 8 and a 16×16 matrix is applied in the LFNST kernel, 16 conversion coefficients can be generated by multiplying the corresponding matrix with 16 input data. The generated conversion coefficients can be arranged in the left upper 8×2 or 2×8 region according to the scanning order used in the VVC standard.

[0272] Figure 20 shows the scanning order for an 8×2 or 2×8 region according to an example.

[0273] For regions other than the left upper 8×2 or 2×8 region, they can be filled with all 0 values (zero - out), or the existing conversion coefficients to which a linear transformation is applied can be maintained as they are. The pre - specified LFNST index is one of the LFNST index values (0, 1, 2) that are tried when calculating the RD cost while changing the LFNST index value in the encoding process.

[0274] In the case of a configuration that limits the computational complexity for the worst case to a certain level (e.g., 8 multiplications / sample), for example, after multiplying an 8×16 matrix that only takes the upper 8 rows of the 16×16 matrix to generate only 8 transform coefficients, the 8 transform coefficients can be arranged according to the scanning order as shown in FIG. 20, and zero-out can also be applied to the remaining coefficient area. The complexity adjustment for the worst case will be described later.

[0275] When applying the inverse-direction LFNST, place the preset number (e.g., 16) of transform coefficients in the input vector, select the LFNST kernel (e.g., 16×16 matrix) derived from the LFNST set obtained in the second process and the parsed LFNST index, and then multiply the LFNST kernel and the corresponding input vector to derive the output vector.

[0276] For the M×2 (M×1) block, the output vector can be arranged in row-major order as shown in FIG. 19(b), and for the 2×M (1×M) block, the output vector can be arranged in column-major order as shown in FIG. 19(a).

[0277] Excluding the area where the corresponding output vector is arranged within the upper left M×2 (M×1) or 2×M (M×2) area, for the remaining area and the area other than the upper left M×2 (M×1) or 2×M (M×2) area within the partition block, it can be configured to fill them all with 0 values (zero-out), or maintain the transform coefficients restored through the residual coding and inverse quantization process as they are.

[0278] When constructing the input vector in the same way as in step 3, the input data can be arranged according to the scanning order in FIG. 20. To limit the computational complexity for the worst case to a certain level, the number of input data can be reduced (e.g., 8 instead of 16) to construct the input vector.

[0279] For example, when M = 8 and eight input data are used, only the left 16×8 matrix from the corresponding 16×16 matrix is taken and multiplied, and then 16 output data can be obtained. The complexity adjustment for the worst case will be described later.

[0280] In the above embodiment, the case of applying symmetry between the M×2 (M×1) block and the 2×M (1×M) block when applying LFNST is presented. However, according to other examples, different LFNST sets can also be applied to the two blocks respectively.

[0281] Hereinafter, various examples of the LFNST set configuration for the ISP mode and the mapping method using the intra prediction mode will be described.

[0282] In the case of the ISP mode, the LFNST set configuration may be different from the existing LFNST set. That is, a kernel different from the existing LFNST kernel can also be applied, and another mapping table different from the mapping table between the intra prediction mode index currently applied to the VVC standard and the LFNST set can be applied. The mapping table currently applied to the VVC standard is as shown in Table 2.

[0283] In Table 2, the preModeIntra value means the intra prediction mode value changed in consideration of WAIP, and the lfnstTrSetIdx value is the index value indicating a specific LFNST set. Each LFNST set is composed of two LFNST kernels.

[0284] When the ISP prediction mode is applied, when the horizontal length and the vertical length of each partition block are all greater than or equal to 4, the same kernel as the LFNST kernel currently applied in the VVC standard can be applied, and the mapping table can also be applied as it is. Of course, the current VVC standard, other LFNST kernels, and other mapping tables can also be applied.

[0285] When the ISP prediction mode is applied, if the horizontal or vertical length of each partition block is less than 4, the current VVC standard, other LFNST kernels, and other mapping tables can be applied. Hereinafter, Tables 6 to 8 show the mapping tables between the intra prediction mode values (intra prediction mode values changed considering WAIP) applicable to M×2 (M×1) blocks or 2×M (1×M) blocks and the LFNST sets.

[0286]

Table 6

[0287]

Table 7

[0288]

Table 8

[0289] The mapping table in Table 6 is composed of 7 LFNST sets, the mapping table in Table 7 is composed of 4 LFNST sets, and the mapping table in Table 8 is composed of 2 LFNST sets. As another example, when composed of 1 LFNST set, the lfnstTrSetIdx value can be fixed to 0 for the preModeIntra value.

[0290] Hereinafter, a method for maintaining the computational complexity in the worst case when applying LFNST in the ISP mode will be described.

[0291] In the case of the ISP mode, when applying LFNST, the application of LFNST can be restricted to maintain the multiplication count per sample (or per coefficient, per position) below a certain value. Depending on the size of the partition block, LFNST can be applied as follows to maintain the multiplication count per sample (or per coefficient, per position) below 8.

[0292] 1. When the horizontal length and vertical length of the partition block are both 4 or more, the same calculation complexity adjustment method as the worst-case calculation complexity adjustment method for LFNST in the current VVC standard can be applied.

[0293] That is, when the partition block is a 4×4 block, instead of a 16×16 matrix, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix can be applied in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix can be applied in the reverse direction. Also, when the partition block is an 8×8 block, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the upper 8 rows from a 16×48 matrix can be applied, and in the reverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns from a 48×16 matrix can be applied.

[0294] In the case of a 4×N or N×4 (N>4) block, when performing the forward transform, for the left-upper 4×4 block only, the 16 coefficients generated after applying a 16×16 matrix can be arranged in the left-upper 4×4 area, and the other areas can be filled with 0 values. Also, when performing the inverse transform, after arranging the 16 coefficients located in the left-upper 4×4 block in the scanning order to form an input vector, a 16×16 matrix can be multiplied to generate 16 output data. The generated output data can be arranged in the left-upper 4×4 area, and the remaining area excluding the left-upper 4×4 area can be filled with 0.

[0295] In the case of an 8×N or N×8 (N>8) block, when performing the forward transform, only for the ROI region (the remaining region in the upper left 8×8 block excluding the lower right 4×4 block) within the upper left 8×8 block, after applying a 16×48 matrix, the 16 generated coefficients are arranged in the upper left 4×4 region, and the remaining regions can all be filled with 0 values. Also, when performing the inverse transform, after arranging the 16 coefficients located in the upper left 4×4 block in scanning order to form an input vector, multiplying by a 48×16 matrix can generate 48 output data. The generated output data fills the ROI region, and the remaining regions can all be filled with 0 values.

[0296] 2. When the size of the partition block is N×2 or 2×N and the LFNST is applied to the upper left M×2 or 2×M region (M≦N), a matrix sampled according to the N value can be applied.

[0297] When M = 8, for a partition block with N = 8, that is, an 8×2 or 2×8 block, in the case of the forward transform, instead of a 16×16 matrix, an 8×16 matrix sampled from the upper 8 rows of a 16×16 matrix can be applied, and in the case of the inverse transform, instead of a 16×16 matrix, a 16×8 matrix sampled from the left 8 columns of a 16×16 matrix can be applied.

[0298] When N is greater than 8, in the case of the forward transform, after applying a 16×16 matrix to the upper left 8×2 or 2×8 block, the 16 generated output data are arranged in the upper left 8×2 or 2×8 block, and the remaining regions can be filled with 0 values. In the case of the inverse transform, after arranging the 16 coefficients located in the upper left 8×2 or 2×8 block in scanning order to form an input vector, multiplying by the corresponding 16×16 matrix can generate 16 output data. The generated output data are arranged in the upper left 8×2 or 2×8 block, and the remaining regions can all be filled with 0 values.

[0299] 3. When the size of the partition block is N×1 or 1×N and the LFNST is applied to the upper left M×1 or 1×M area (M≦N), a matrix sampled by the N value can be applied.

[0300] When M = 16, for a partition block with N = 16, that is, a 16×1 or 1×16 block, in the case of forward transformation, instead of a 16×16 matrix, an 8×16 matrix sampled from the upper 8 rows of the 16×16 matrix can be applied; in the case of inverse transformation, instead of a 16×16 matrix, a 16×8 matrix sampled from the left 8 columns of the 16×16 matrix can be applied.

[0301] When N is greater than 16, in the case of forward transformation, after applying a 16×16 matrix to the upper left 16×1 or 1×16 block, the 16 output data generated can be arranged in the upper left 16×1 or 1×16 block, and the remaining area can be filled with 0 values. In the case of inverse transformation, after arranging the 16 coefficients located in the upper left 16×1 or 1×16 block in the scanning order to form an input vector, the corresponding 16×16 matrix can be multiplied to generate 16 output data. The generated output data can be arranged in the upper left 16×1 or 1×16 block, and the remaining area can all be filled with 0 values.

[0302] As another example, in order to maintain the multiplication count per sample (or per coefficient, per position) below a certain value, the multiplication count per sample (or per coefficient, per position) can be maintained at 8 or less based on the size of the ISP coding unit rather than the size of the ISP partition block. If there is only one block among the ISP partition blocks that satisfies the condition for applying LFNST, the complexity calculation for the worst case of LFNST can be applied based on the size of the corresponding coding unit that is not the size of the partition block. For example, if the luma coding block for a certain coding unit is divided into 4 partition blocks of size 4×4 and coded by ISP, and there are no non-zero transform coefficients for 2 of the partition blocks, the other 2 partition blocks can be set to generate 16 transform coefficients each (based on the encoder standard) that are not 8.

[0303] In the following, when in the ISP mode, the method of signaling the LFNST index will be considered.

[0304] As described above, the LFNST index can have values of 0, 1, and 2. 0 indicates that LFNST is not applied, and 1 and 2 indicate one of the two LFNST kernel matrices included in the selected LFNST set respectively. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method of transmitting the LFNST index in the current VVC standard will be described as follows.

[0305] 1. The LFNST index can be transmitted once per coding unit (CU). When it is a dual-tree, individual LFNST indexes can be signaled for the luma block and the chroma block respectively.

[0306] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.

[0307] A. In the case of a mode where no transformation is applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0308] B. When the first transformation is not DCT-2 (DST7 or DCT8), i.e., when the horizontal or vertical transformation is not DCT-2

[0309] C. When the horizontal or vertical length of the luma block of the coding unit exceeds the size of the maximum luma transformation that can be performed. For example, when the size of the maximum luma transformation that can be performed is 64, and the size of the luma block of the coding block is 128×16, etc., LFNST cannot be applied.

[0310] In the case of a dual tree, it is determined whether the size of the maximum luma transformation is exceeded for each of the coding unit for the luma component and the coding unit for the chroma component. That is, it is checked whether the size of the maximum luma transformation that can be performed for the luma block is exceeded, and it is checked whether the horizontal / vertical length of the corresponding luma block for the color format and the size of the maximum transformation that can be performed exceed the size of the maximum luma transformation that can be performed for the chroma block. For example, when the color format is 4:2:0, the horizontal / vertical length of the corresponding luma block is twice that of the corresponding chroma block, respectively, and the transformation size of the corresponding luma block is twice that of the corresponding chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length and the transformation size of the corresponding luma block are the same as those of the corresponding chroma block.

[0311] 64-length conversion or 32-length conversion each means a conversion applied horizontally or vertically having 64 or 32 lengths respectively, and "conversion size" can mean 64 or 32 which is the corresponding length.

[0312] In the case of a single tree, after checking whether the horizontal length or vertical length of the luma block exceeds the maximum luma transform block size for which conversion is possible, if it exceeds, the LFNST index signaling can be omitted.

[0313] The LFNST index can be sent only when the horizontal length and vertical length of all coding units D are 4 or more.

[0314] In the case of a dual tree, the LFNST index can be signaled only when the horizontal length and vertical length of the corresponding component (i.e., luma or chroma component) are both 4 or more.

[0315] In the case of a single tree, the LFNST index can be signaled when the horizontal length and vertical length of the luma component are both 4 or more.

[0316] E. When the position of the last non-zero coefficient (last non-zero coefficient position) is not the DC position (the upper left corner position of the block), in the case of a dual tree type luma block, the LFNST index is sent when the position of the last non-zero coefficient is not the DC position. In the case of a dual tree type chroma block, the corresponding LNFST index is sent when the position of the last non-zero coefficient for Cb or the position of the last non-zero coefficient for Cr is not the DC position.

[0317] In the case of a single tree type, the LFNST index is sent when the position of the last non-zero coefficient of any one of the luma component, Cb component, and Cr component is not the DC position.

[0318] Here, when the CBF (coded block flag) value indicating the existence of a transform coefficient for a single transform block is 0, the position of the last non-zero coefficient for the corresponding transform block is not checked to determine whether LFNST index signaling is possible. That is, when the corresponding CBF value is 0, since no transform is applied to the corresponding block, the position of the last non-zero coefficient is not considered when checking the conditions for LFNST index signaling.

[0319] For example, 1) in the case of the dual-tree type and for the luma component, when the corresponding CBF value is 0, no LFNST index is signaled; 2) in the case of the dual-tree type and for the chroma component, when the CBF value for Cb is 0 and the CBF value for Cr is 1, only the position of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is transmitted; 3) in the case of the single-tree type, the position of the last non-zero coefficient is only checked for components where the CBF value for all of luma, Cb, and Cr is 1.

[0320] If it is confirmed that a transform coefficient exists at a position where an F.LFNST transform coefficient cannot exist, LFNST index signaling can be omitted. In the case of 4×4 and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, an LFNST transform coefficient can exist at 8 positions starting from the DC position, and the remaining positions are all filled with 0s. Also, in the case of blocks other than 4×4 and 8×8 transform blocks, according to the transform coefficient scanning order in the VVC standard, an LFNST transform coefficient can exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0s.

[0321] Therefore, after performing residual coding, if a non-zero transform coefficient exists in the region where the 0 value should be filled, LFNST index signaling can be omitted.

[0322] On the one hand, the ISP mode is applicable only when it is a luma block, or it can also be applied to both luma and chroma blocks. As described above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks for prediction, and the transformation can also be applied to each corresponding partition block. Therefore, when determining the conditions for signaling the LFNST index in units of coding units, the fact that LFNST can be applied to each corresponding partition block must be considered. Also, when the ISP prediction mode is applied only to a specific component (e.g., a luma block), the LFNST index must be signaled considering the fact that it is divided into partition blocks only for the corresponding component. When sorting out the possible LFNST index signaling methods in the ISP mode, it is as follows.

[0323] 1. The LFNST index can be transmitted once for each coding unit (CU). When it is a dual-tree, individual LFNST indexes can be signaled for the luma and chroma blocks respectively.

[0324] 2. When the LFNST index is not signaled, the LFNST index value is determined (inferred) to be the default value of 0. When the LFNST index value is inferred to be 0, it is as follows.

[0325] A. In the case of a mode where transformation is not applied (e.g., transform skip, BDPCM, lossless coding, etc.)

[0326] When the horizontal or vertical length of the luma block of the coding unit exceeds the maximum luma transform size that can be transformed, for example, when the maximum luma transform size that can be transformed is 64, and the size of the luma block of the coding block is 128×16, etc., LFNST cannot be applied.

[0327] It is also possible to determine whether the LFNST index can be signaled based on the size of the partition block instead of the coding unit. That is, when the horizontal or vertical length of the partition block for the corresponding luma block exceeds the maximum luma transform size that can be transformed, the LFNST index signaling can be omitted and the LFNST index value can be analogized to 0.

[0328] In the case of a dual tree, it is determined whether each of the coding unit or partition block for the luma component and the coding unit or partition block for the chroma component exceeds the maximum transform block size. That is, the horizontal and vertical lengths of the coding unit or partition block for luma are each compared with the maximum luma transform size. If even one of them is larger than the maximum luma transform size, LFNST is not applied. In the case of the coding unit or partition block for chroma, the horizontal / vertical length of the corresponding luma block for the color format is compared with the maximum luma transform size that can be transformed. For example, when the color format is 4:2:0, the horizontal / vertical lengths of the corresponding luma blocks are each twice that of the corresponding chroma block, and the transform size of the corresponding luma block is twice that of the corresponding chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical lengths and transform sizes of the corresponding luma blocks are the same as those of the corresponding chroma blocks.

[0329] In the case of a single tree, after checking whether the horizontal or vertical length of a luma block (coding unit or partition block) exceeds the maximum luma transform block size that can be transformed, if it exceeds, the LFNST index signaling can be omitted.

[0330] C. If applying the LFNST included in the current VVC standard, the LFNST index can be sent only when the horizontal and vertical lengths of the partition block are both 4 or more.

[0331] If applying the LFNST for 2×M (1×M) or M×2 (M×1) blocks in addition to the LFNST included in the current VVC standard, the LFNST index can be sent only when the size of the partition block is larger than or equal to the 2×M (1×M) or M×2 (M×1) block. Here, the meaning that a P×Q block is larger than or equal to an R×S block means P≧R and Q≧S.

[0332] To summarize, the LFNST index can be sent only when the partition block is larger than or equal to the minimum size for which the LFNST is applicable. In the case of a dual tree, the LFNST index can be signaled only when the partition block for the luma or chroma component is larger than or equal to the minimum size for which the LFNST is applicable. In the case of a single tree, the LFNST index can be signaled only when the partition block for the luma component is larger than or equal to the minimum size for which the LFNST is applicable.

[0333] In this document, the fact that an M×N block is larger than or equal to a K×L block means that M is greater than or equal to K and N is greater than or equal to L. The fact that an M×N block is larger than a K×L block means that M is greater than or equal to K, N is greater than or equal to L, and either M is greater than K or N is greater than L. The fact that an M×N block is smaller than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L, and the fact that an M×N block is smaller than a K×L block means that M is less than or equal to K, N is less than or equal to L, and either M is less than K or N is less than L.

[0334] D. When the position of the last non-zero coefficient (last non-zero coefficient position) is not the DC position (the upper left corner position of the block), in the case of a dual-tree type luma block, if the position of the last non-zero coefficient of any one of all the partition blocks is not the DC position, the corresponding LFNST index can be transmitted. In the case of a dual-tree type and a chroma block, when the position of the last non-zero coefficient of all the partition blocks for Cb (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) and the position of the last non-zero coefficient of all the partition blocks for Cr (when the ISP mode is not applied to the chroma component, the number of partition blocks is regarded as 1) are not the DC position, the corresponding LNFST index can be transmitted.

[0335] In the case of the single-tree type, when the position of the last non-zero coefficient of any one of all the partition blocks for the luma component, Cb component, and Cr component is not the DC position, the corresponding LFNST index can be transmitted.

[0336] Here, when the CBF (coded block flag) value indicating the existence of a conversion coefficient for each partition block is 0, the position of the last non-zero coefficient for the corresponding partition block is not checked to determine whether LFNST index signaling is possible. That is, when the corresponding CBF value is 0, since no conversion is applied to the corresponding block, the position of the last non-zero coefficient for the corresponding partition block is not considered when checking the conditions for LFNST index signaling.

[0337] For example, 1) in the case of the dual-tree type and for the luma component, when the corresponding CBF value for each partition block is 0, the corresponding partition block is excluded when determining whether LFNST index signaling is possible. 2) In the case of the dual-tree type and for the chroma component, when the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, the position of the last non-zero coefficient for Cr is checked to determine whether the corresponding LFNST index signaling is possible. 3) In the case of the single-tree type, the position of the last non-zero coefficient is checked only for the blocks with a CBF value of 1 for all partition blocks of the luma component, Cb component, and Cr component to determine whether LFNST index signaling is possible.

[0338] In the case of the ISP mode, the video information can also be configured so as not to check the position of the last non-zero coefficient, and the embodiments thereof are as follows.

[0339] i. In the case of the ISP mode, it is possible to allow LFNST index signaling by omitting the check of the position of the last non-zero coefficient for both the luma block and the chroma block. That is, even if the position of the last non-zero coefficient for all partition blocks is the DC position or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.

[0340] ii. In the case of the ISP mode, only for the luma block, the check for the position of the last non-zero coefficient can be omitted. In the case of the chroma block, the check for the position of the last non-zero coefficient in the above-described method can be executed. For example, in the case of the dual-tree type and the luma block, the LFNST index signaling can be allowed without checking the position of the last non-zero coefficient. In the case of the dual-tree type and the chroma block, the DC position existence of the last non-zero coefficient in the above-described method can be checked to determine whether the corresponding LFNST index can be signaled.

[0341] iii. In the case of the ISP mode and the single-tree type, the method i or ii can be applied. That is, when applying the method i to the single-tree type in the ISP mode, the check for the position of the last non-zero coefficient can be omitted for both the luma block and the chroma block, and the LFNST index signaling can be allowed. Or, when applying the method ii, the check for the position of the last non-zero coefficient can be omitted for the partition block for the luma component, and for the partition block for the chroma component (when ISP is not applied to the chroma component, the number of partition blocks can be regarded as 1), the check for the position of the last non-zero coefficient in the above-described method can be executed to determine whether the corresponding LFNST index can be signaled.

[0342] E. If it is confirmed that there is a transform coefficient at a position where the LFNST transform coefficient does not exist in any one of all the partition blocks, the LFNST index signaling can be omitted.

[0343] For example, in the case of 4×4 partition blocks and 8×8 partition blocks, according to the conversion coefficient scanning order in the VVC standard, LFNST conversion coefficients can exist at 8 positions starting from the DC position, and the remaining positions are all filled with 0. Also, when it is larger than or equal to 4×4 and is not a 4×4 partition block or an 8×8 partition block, according to the conversion coefficient scanning order in the VVC standard, LFNST conversion coefficients can exist at 16 positions starting from the DC position, and the remaining positions are all filled with 0.

[0344] Therefore, after performing residual coding, if there are non-zero conversion coefficients in the area to be filled with 0 values, LFNST index signaling can be omitted.

[0345] If LFNST can also be applied to the case where the partition block is 2×M (1×M) or M×2 (M×1), the area where the LFNST conversion coefficients can be located can be specified as follows. The area outside the area where the conversion coefficients can be located can be filled with 0. If there are non-zero conversion coefficients in the area to be filled with 0 when assuming that LFNST is applied, LFNST index signaling can be omitted.

[0346] i. LFNST can be applied to 2×M or M×2 blocks. When M = 8, only 8 LFNST conversion coefficients can be generated for 2×8 or 8×2 partition blocks. When the conversion coefficients are arranged in the scanning order as shown in Figure 20, 8 conversion coefficients are arranged in the scanning order starting from the DC position, and the remaining 8 positions can be filled with 0.

[0347] For 2×N or N×2 (N>8) partition blocks, 16 LFNST transform coefficients can be generated. When the transform coefficients are arranged in the scanning order as shown in FIG. 20, 16 transform coefficients are arranged in the scanning order from the DC position, and the remaining area can be filled with 0s. That is, for the area other than the upper left 2×8 or 8×2 block in the 2×N or N×2 (N>8) partition block, it can be filled with 0s. For 2×8 or 8×2 partition blocks, 16 transform coefficients can also be generated instead of 8 LFNST transform coefficients. In this case, there is no area that should be filled with 0s. As described above, when LFNST is applied, if it is detected that there is a non-zero transform coefficient in the area determined to be filled with 0s even in one partition block, the LFNST index signaling is omitted and the LFNST index can be inferred to be 0.

[0348] ii. LFNST can be applied to 1×M or M×1 blocks. When M = 16, only 8 LFNST transform coefficients can be generated for 1×16 or 16×1 partition blocks. When the transform coefficients are arranged in the scanning order from left to right or from top to bottom, 8 transform coefficients are arranged in the corresponding scanning order from the DC position, and the remaining 8 positions can be filled with 0s.

[0349] For 1×N or N×1 (N>16) partition blocks, 16 LFNST transform coefficients can be generated. When the transform coefficients are arranged in the scanning order from left to right or from top to bottom, 16 transform coefficients are arranged in the corresponding scanning order from the DC position, and the remaining area can be filled with 0s. That is, for the area other than the upper left 1×16 or 16×1 block in the 1×N or N×1 (N>16) partition block, it can be filled with 0s.

[0350] For a 1×16 or 16×1 partition block, 16 transformation coefficients can be generated instead of 8 LFNST transformation coefficients, and in this case, no area to be filled with 0 is generated. As described above, when LFNST is applied, if it is detected that there is a non-zero transformation coefficient in the area determined to be filled with 0 even in one partition block, the LFNST index signaling can be omitted, and the LFNST index can be inferred to be 0.

[0351] On the other hand, in the case of the ISP mode, in the current VVC standard, for the horizontal and vertical directions, the DST-7 is applied instead of DCT-2 without signaling for the MTS index by looking at the length conditions independently. It is determined whether the horizontal or vertical length is greater than or equal to 4 and less than or equal to 16, and the primary transformation kernel is determined according to the determination result. Therefore, for the case where it is in the ISP mode and LFNST can be applied, the following conversion combination configurations are possible.

[0352] 1. When the LFNST index is 0 (including the case where the LFNST index is inferred to be 0), the primary transformation determination conditions in the ISP included in the current VVC standard can be followed. That is, it is checked whether the length conditions (greater than or equal to 4 and less than or equal to 16) are satisfied independently for the horizontal and vertical directions. If satisfied, DST-7 can be applied instead of DCT-2 for the primary transformation, and if not satisfied, DCT-2 can be applied.

[0353] 2. When the LFNST index is greater than 0, the following two configurations are possible for the primary transformation.

[0354] A. DCT-2 can be applied to both the horizontal and vertical directions.

[0355] B. It can comply with the primary transformation determination conditions when it is the ISP currently included in the VVC standard. That is, it checks whether the length conditions (greater than or equal to 4 and less than or equal to 16) are satisfied independently for the horizontal and vertical directions respectively. If satisfied, DST-7 can be applied instead of DCT-2; if not satisfied, DCT-2 can be applied.

[0356] When in the ISP mode, the LFNST index can be configured such that the video information is sent not for each coding unit but for each partition block. In such a case, it is possible to determine whether the LFNST index can be signaled by regarding that there is only one partition block within the unit where the LFNST index is signaled by the aforementioned LFNST index signaling method.

[0357] On the other hand, the signaling order of the LFNST index and the MTS index is considered below.

[0358] By way of example, the LFNST index signaled in residual coding can be coded next to the coding position for the last non-zero coefficient position, and the MTS index can be coded immediately after the LFNST index. In the case of such a configuration, the LFNST index can be signaled for each transform unit. Or, even if not signaled in residual coding, the LFNST index can be coded next to the coding for the last valid coefficient position, and the MTS index can be coded next to the LFNST index.

[0359] The syntax of the residual coding according to an example is as follows.

[0360]

Table 9-1

[0361]

Table 9-2

[0362] The meanings of the main variables shown in Table 9 are as follows.

[0363] 1. cbWidth, cbHeight: The width and height of the current coding block

[0364] 2. log2TbWidth, log2TbHeight: The base-2 logarithm values for the width and height of the current transform block. It can be reduced to the upper left area where zero-out is reflected and non-zero coefficients can exist.

[0365] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is applicable. When the flag value is 0, it indicates that LFNST is not applicable; when the flag value is 1, it indicates that LFNST is applicable. It is defined in the Sequence Parameter Set (SPS).

[0366] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the (x0, y0) position. chType can have values of 0 and 1. 0 indicates the luma component, and 1 indicates the chroma component. The (x0, y0) position indicates the position on the picture, and the CuPredMode[chType][x0][y0] value can be MODE_INTRA (intra prediction) or MODE_INTER (inter prediction).

[0367] 5.IntraSubPartitionsSplit[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. It indicates what kind of ISP split is applied at the position (x0, y0), and ISP_NO_SPLIT indicates that the coding unit corresponding to the position (x0, y0) is not split into partition blocks.

[0368] 6.intra_mip_flag[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. intra_mip_flag is a flag indicating whether the MIP (Matrix-based Intra Prediction) prediction mode is applied. When the flag value is 0, it indicates that MIP is not applicable, and when the flag value is 1, it indicates that MIP is applied.

[0369] 7.cIdx: A value of 0 indicates luma, and values of 1 and 2 indicate the chroma components Cb and Cr, respectively.

[0370] 8.treeType: It refers to single-tree, dual-tree, etc. (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luma component, DUAL_TREE_CHROMA: dual-tree for chroma component)

[0371] 9.tu_cbf_cb[x0][y0]: The content for the position (x0, y0) is the same as that of the above 4. It indicates the CBF (Coded Block Flag) for the Cb component. When its value is 0, it means that there are no non-zero coefficients in the corresponding transform unit for the Cb component, and when it is 1, it indicates that there are non-zero coefficients in the corresponding transform unit for the Cb component.

[0372] 10. lastSubBlock: Indicates the position in the scan order of the sub-block (Coefficient Group (CG)) where the last non-zero coefficient is located. 0 refers to the sub-block containing the DC component, and if it is greater than 0, it is not the sub-block containing the DC component.

[0373] 11. lastScanPos: Indicates the position in the scan order of the last non-zero coefficient within a sub-block. If a sub-block is composed of 16 positions, values from 0 to 15 are possible.

[0374] 12. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is analogized to a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.

[0375] 13. LastSignificantCoeffX, LastSignificantCoeffY: Indicate the x and y coordinates where the last non-zero coefficient is located within the transform block. The x coordinate starts from 0 and increases from left to right, and the y coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means that the last non-zero coefficient is located at the DC.

[0376] 14. cu_sbt_flag: A flag indicating whether the SubBlock Transform (SBT) currently included in the VVC standard is applicable. If the flag value is 0, it indicates that SBT is not applicable, and if the flag value is 1, it indicates that SBT is applied.

[0377] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These are flags indicating whether explicit MTS is applied to the inter-CU and intra-CU respectively. When the corresponding flag value is 0, it indicates that MTS is not applicable to the inter-CU or intra-CU, and when it is 1, it indicates that it is applicable.

[0378] 16. tu_mts_idx[x0][y0]: It is the MTS index syntax element to be parsed. When not parsed, it is analogous to the value 0. That is, the default value is set to 0, indicating that DCT-2 is applied to all in the horizontal and vertical directions.

[0379] As shown in Table 9, in the case of a single tree, it is possible to determine whether to signal the LFNST index only based on the condition of the last valid coefficient position for luma. That is, when the last valid coefficient position is not DC and the last valid coefficient exists inside the upper left sub-block (CG), for example, a 4×4 block, the LFNST index is signaled. At this time, in the case of 4×4 and 8×8 transform blocks, the LFNST index is signaled only when the last valid coefficient exists at positions from 0 to 7 inside the upper left sub-block.

[0380] In the case of a dual tree, for luma and chroma, the LFNST index is signaled independently. For chroma, the LFNST index can be signaled by applying the condition of the last valid coefficient position only to the Cb component. For the Cr component, the corresponding condition is not checked. If the CBF value for Cb is 0, the LFNST index can be signaled by applying the condition of the last valid coefficient position to the Cr component.

[0381] "Min(log2TbWidth, log2TbHeight) >= 2" in Table 9 can be expressed as "Min(tbWidth, tbHeight) >= 4", and "Min(log2TbWidth, log2TbHeight) >= 4" can be expressed as "Min(tbWidth, tbHeight) >= 16".

[0382] In Table 9, log2ZoTbWidth and log2ZoTbHeight respectively mean the base-2 log values of the width and height for the upper left region where the last valid coefficient can exist due to zeroing out.

[0383] As in Table 9, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before the MTS index or LFNST index value is parsed, and the second is after the parsing of the MTS index.

[0384] Since the first update is before the MTS index (tu_mts_idx[x0][y0]) value is parsed, log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.

[0385] After the MTS index is parsed, log2ZoTbWidth and log2ZoTbHeight will be set when the MTS index value is greater than 0 (in the case of the DST-7 / DCT-8 combination). When applying DST-7 / DCT-8 independently for the horizontal and vertical directions in the first transformation, up to 16 valid coefficients can exist for each row or column in each direction. That is, after applying DST-7 / DCT-8 with a length of 32 or more, up to 16 transformation coefficients can be derived for each row or column starting from the left or top. Therefore, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions for a two-dimensional block, valid coefficients can exist only in the upper left 16×16 region at most.

[0386] Also, when DCT-2 is currently applied independently to the horizontal and vertical directions in the first conversion, up to 32 valid coefficients can exist for each row or column in each direction. That is, when applying DCT-2 of a length of 64 or more, up to 32 conversion coefficients can be derived for each row or column starting from the left or upper side. Therefore, when DCT-2 is applied to both the horizontal and vertical directions for a two-dimensional block, valid coefficients can exist only up to the maximum upper-left 32×32 region.

[0387] Also, when DST-7 / DCT-8 is applied to one of the horizontal and vertical directions and DCT-2 is applied to the other, 16 valid coefficients can exist in the former direction and 32 valid coefficients can exist in the latter direction. For example, in the case of a 64×8 conversion block where DCT-2 is applied horizontally and DST-7 is applied vertically (which can occur in a situation where implicit MTS is applied), valid coefficients can exist in the maximum upper-left 32×8 region.

[0388] If log2ZoTbWidth and log2ZoTbHeight are updated at two locations as shown in Table 9, that is, if they are updated before MTS index parsing, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the following table.

[0389]

Table 10

[0390] Also, in such a case, in the binary process for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values.

[0391]

Table 11

[0392] On the other hand, in one example, when it is in the ISP mode and LFNST is applied, when the signaling in Table 9 is applied, the spec text can be configured as shown in Table 12. When compared with Table 9, the condition where only the LFNST index is signaled for the case where it is not in the ISP mode (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 9) is deleted.

[0393] In the case of a single tree, when reusing the LFNST index transmitted when it is luma (when cIdx = 0) when it is chroma, the LFNST index transmitted for the first ISP partition block where valid coefficients exist can be applied to the chroma conversion block. Or, even if it is the case of a single tree, for the case of chroma components, the LFNST index can be signaled separately from the luma components. The explanations for the variables described in Table 12 are as in Table 9.

[0394]

Table 12

[0395] If, according to another example, in Table 12, when it is in the ISP and it is allowed for the last valid coefficient to be located at the DC position, the parsing condition of the LFNST index can be changed as follows.

[0396]

Table 13

[0397] On the one hand, according to one example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values of 0, 1, and 2. 0 indicates not applying LFNST, and 1 and 2 indicate the first candidate and the second candidate among the two LFNST kernel candidates included in the selected LFNST set, respectively. The LFNST index is coded via truncated unary binarization, and the 0, 1, and 2 values can be coded with bin strings 0, 10, and 11, respectively.

[0398] According to one example, LFNST can be applied only when DCT-2 is applied to both the horizontal and vertical directions in the first transformation. Therefore, if the MTS index is signaled after the LFNST index signaling, the MTS index can be signaled only when the LFNST index value is 0. When the LFNST index is not 0, the first transformation can be performed by applying DCT-2 to both the horizontal and vertical directions without signaling the MTS index.

[0399] The MTS index value can have values of 0, 1, 2, 3, and 4. 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied to the horizontal and vertical directions, respectively. Also, the MTS index can be coded via truncated unary binarization, and the 0, 1, 2, 3, and 4 values can be coded with bin strings 0, 10, 110, 1110, and 1111, respectively.

[0400] The LFNST index and the MTS index can be signaled at the coding unit level, and the MTS index can be coded following the LFNST index at the coding unit level. The coding unit syntax table for this is as follows.

[0401]

Table 14

[0402] When comparing Table 14 with Table 13, the condition for checking whether the tu_mts_idx[x0][y0] value is 0 under the condition of signaling lfnst_idx[x0][y0] (i.e., checking whether it is DCT-2 in both the horizontal and vertical directions) is changed to the condition of checking whether the transform_skip_flag[x0][y0] value is 0 (!transform_skip_flag[x0][y0]). The transform_skip_flag[x0][y0] indicates whether the coding unit is coded in the transform skip mode where the transformation is omitted, and the flag is signaled prior to the MTS index and the LFNST index. That is, since lfnst_idx[x0][y0] is signaled before signaling the tu_mtx_idx[x0][y0] value, only the condition for the transform_skip_flag[x0][y0] value can be checked.

[0403] As shown in Table 14, when coding tu_mts_idx[x0][y0], various conditions are checked, and as described above, tu_mts_idx[x0][y0] is signaled only when the lfnst_idx[x0][y0] value is 0.

[0404] Also, tu_cbf_luma[x0][y0] is a flag indicating whether there is a valid coefficient for the luma component, and cbWidth and cbHeight indicate the width and height of the coding unit for the luma component, respectively.

[0405] Also, in Table 14, (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT) indicates a case where the ISP mode is not applied, and (!cu_sbt_flag) indicates a case where SBT is not applied.

[0406] According to Table 14, when both the width and height of the coding unit for the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, that is, whether MTS is applicable is determined by the width and height of the coding unit for the luma component.

[0407] In other examples, when transform block tiling (TU tiling) occurs (for example, when the maximum transform size is set to 32, a 64×64 coding unit is divided into 4 32×32 transform blocks for coding), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or less, the same MTS index value can be applied to all transform blocks within the coding unit and the same primary transform can be applied. Also, when transform block tiling occurs, the tu_cbf_luma[x0][y0] value in Table 14 is the CBF value for the upper-left transform block, or can be set to 1 if the corresponding CBF value is 1 for even one transform block among all transform blocks.

[0408] In one example, when the ISP mode is applied to the current block, LFNST can be applied, and in this case, Table 14 can be changed as shown in Table 15.

[0409]

Table 15

[0410] As shown in Table 15, even in the case of the ISP mode (IntraSubPartitionsSplitType!=ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP partition blocks.

[0411] Also, as shown in Table 15, since tu_mts_idx[x0][y0] can be signaled only when not in the ISP mode, the MTS index coding part is as shown in Table 14.

[0412] As shown in Tables 14 and 15, when the MTS index is signaled immediately after the LFNST index, information about the primary transform cannot be known when performing residual coding. That is, the MTS index is signaled after the residual coding. Therefore, the part that zeros out leaving only 16 coefficients for the 32-length DST-7 or DCT-8 in the residual coding part can be changed as shown in Table 16 below.

[0413]

Table 16-1

[0414]

Table 16-2

[0415] In the process of determining log2ZoTbWidth and log2ZoTbHeight as shown in Table 16 (where log2ZoTbWidth and log2ZoTbHeight respectively represent the base-2 log values of the width and height of the upper left region remaining after zeroing out), the part of checking the tu_mts_idx[x0][y0] value can be omitted.

[0416] The binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 16 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 11.

[0417] Also, when determining log2ZoTbWidth and log2ZoTbHeight by residual coding as shown in Table 16, a condition for checking the sps_mts_enable_flag can be added.

[0418] TR in Table 11 indicates the Truncated Rice binarization method, and based on cMax and cRiceParam defined in Table 11, the last significant coefficient information can be binarized by the method described in the following table.

[0419] As an example, if the information about the position of the last significant coefficient for the luma transform block is recorded during the residual coding process, the MTS index can also be signaled as shown in Table 17.

[0420]

Table 17

[0421] In Table 17, LumaLastSignificantCoeffX and LumaLastSignificantCoeffY indicate the X coordinate and Y coordinate of the position of the last significant coefficient for the luma transform block, respectively. A condition that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY should be smaller than 16 was added to Table 17. If either one of them is 16 or more, since DCT-2 is applied in both the horizontal and vertical directions, the signaling for tu_mts_idx[x0][y0] can be omitted and it can be analogized that DCT-2 is applied to both the horizontal and vertical directions.

[0422] The fact that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY are smaller than 16 means that the last significant coefficient exists within the upper-left 16×16 region. When the current VVC standard applies DST-7 or DCT-8 of length 32, it indicates that there may be a possibility that zeroing out is applied to leave only 16 transform coefficients from the leftmost or uppermost side. Therefore, tu_mts_idx[x0][y0] can be signaled to indicate the transform kernel used for the first-stage transform.

[0423] On the other hand, according to another example, the coding unit syntax table and the residual coding syntax table are as shown in the following table.

[0424]

Table 18

[0425]

Table 19

[0426] In Table 18, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 19. The variable MtsZeroOutSigCoeffFlag has its value changed from 1 to 0 if there are significant coefficients in the area to be filled with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15). In this case, the MTS index is not signaled as in Table 19.

[0427] On the other hand, an example of the syntax table at the conversion unit level can be represented as follows.

[0428]

Table 20

[0429] As shown in Table 20, MtsZeroOutSigCoeffFlag can be set to 1 when tu_cbf_luma[x0][y0] is 1, and the existing MtsZeroOutSigCoeffFlag value can be maintained when tu_cbf_luma[x0][y0] is 0. Therefore, when tu_cbf_luma[x0][y0] is 0 and the MtsZeroOutSigCoeffFlag value is maintained at 0, the coding of mts_idx[x0][y0] can be omitted. That is, when the CBF value of the luma component is 0, since no conversion is applied, the MTS index is meaningless and the coding of the MTS index can be omitted.

[0430] The following drawings are created to illustrate a specific example of this specification. The names of the specific devices and the names of specific signals / messages / fields described in the drawings are presented by way of example, so the technical features of this specification are not limited to the specific names used in the following drawings.

[0431] FIG. 21 is a flowchart showing the operation of a video decoding apparatus according to an embodiment of this document.

[0432] Each step disclosed in FIG. 21 is based on a part of the content detailed in FIGS. 4 to 20. Therefore, specific content overlapping with the content detailed in FIGS. 3 to 20 will be omitted from the description or simplified.

[0433] A decoding apparatus 300 according to an embodiment can receive residual information from a bitstream (S2110).

[0434] More specifically, the decoding apparatus 300 can decode information regarding quantized transform coefficients for a current block from the bitstream, and can derive quantized transform coefficients for a target block based on the information regarding quantized transform coefficients for the current block. The information regarding quantized transform coefficients for the target block can be included in an SPS (Sequence Parameter Set) or a slice header, and can include at least one of information regarding whether simplified transform (RST) is applied, information regarding a simplification factor, information regarding a minimum transform size to which simplified transform is applied, information regarding a maximum transform size to which simplified transform is applied, a simplified inverse transform size, and information regarding a transform index indicating any one of the transform kernel matrices included in a transform set.

[0435] In addition, the decoding device can further receive information regarding the intra prediction mode for the current block and information regarding whether ISP is applied to the current block. The decoding device can derive whether the current block is divided into a predetermined number of sub-partition conversion blocks by receiving and parsing flag information indicating whether to apply ISP coding or the ISP mode. Here, the current block is a coding block. Also, the decoding device can derive the size and number of sub-partition blocks to be divided through flag information indicating the direction in which the current block is divided.

[0436] The decoding device 300 can perform inverse quantization on the residual information for the current block, that is, the quantized transform coefficients, to derive the transform coefficients (S2120).

[0437] The derived transform coefficients can be arranged in the reverse diagonal scan order in 4×4 block units, and the transform coefficients within the 4×4 block can also be arranged in the reverse diagonal scan order. That is, the transform coefficients after inverse quantization can be arranged according to the reverse scan order applied in video codecs such as VVC and HEVC.

[0438] The transform coefficients derived based on such residual information may be the transform coefficients after inverse quantization as described above, or may be the quantized transform coefficients. That is, the transform coefficients may be any data that can check whether the data in the current block is non-zero regardless of whether quantization is possible.

[0439] The decoding device can apply an inverse transform to the quantized transform coefficients to derive residual samples.

[0440] As described above, the decoding device can derive residual samples by applying non-separable conversion LFNST or separable conversion MTS, and such conversions can each be performed based on an LFNST index indicating an LFNST kernel, i.e., an LFNST matrix, and an MTS index indicating an MTS kernel.

[0441] By way of example, the decoding device can receive and parse at least one of the LFNST index or the MTS index at the coding unit level, and can parse the LFNST index indicating the LFNST kernel immediately before, i.e., immediately after, the MTS index indicating the MTS kernel.

[0442] Depending on a specific condition for the LFNST index, for example, when the LFNST index value is 0, the MTS index can be parsed. Also, by way of example, when the tree type of the current block is dual-tree luma or single-tree and the LFNST index value is 0, the MTS index can be parsed.

[0443] On the other hand, when the value of the LFNST index is 0, the residual coding level can include the syntax for the last significant coefficient position information, and the LFNST index can be parsed after the last significant coefficient position information has been parsed.

[0444] By way of example, when the current block is a luma block and the LFNST index is 0, the MTS index can be parsed. That is, when the current block is a luma block, if the LFNST index is greater than 0, the MTS index is not parsed.

[0445] In one example, when the tree type of the current block is a dual tree, the LFNST index for each of the luma block and the chroma block can be parsed.

[0446] On the other hand, in the step of deriving the conversion coefficient, the width and height for the upper left region where the last valid coefficient can exist due to zeroing out within the current block can be derived, and the width and height for the upper left region can be derived before parsing the MTS index.

[0447] On the other hand, the position of the last valid coefficient can be derived from the width and height for the upper left region, and the last valid coefficient position information can be binary-coded based on the width and height for the upper left region.

[0448] Alternatively, in one example, when the tree type of the current block is a single tree type, the decoding device can parse the LFNST index after performing residual coding on the luma block and the chroma block of the current block.

[0449] When the LFNST index is parsed at the coding unit level after residual coding is performed at a level other than the transform block level or the residual coding level, an LFNST index that reflects the position of the complete transform coefficients for either the luma block or the chroma block that is not transform coefficient information and the zeroing-out information involved in the transform process can be received.

[0450] Also, when the tree type of the current block is a dual tree type and the chroma component is coded, the decoding device can parse the LFNST index after performing residual coding on the Cb component and the Cr component of the chroma block.

[0451] If the LFNST index is parsed at the coding unit level after residual coding is performed at a level other than the transform block level or the residual coding level, it is possible to receive the LFNST index reflecting the positions of the complete transform coefficients for the Cb and Cr components that are not transform coefficient information for either the Cb component or the Cr component of the chroma block, and the zero-out information involved in the transform process.

[0452] Also, when the current block is divided into a plurality of sub-partition blocks, the decoding apparatus can parse the LFNST index after performing residual coding on the plurality of sub-partition blocks.

[0453] Similarly to the above, if the LFNST index is parsed at the coding unit level after residual coding is performed at a level other than the transform block level or the residual coding level, it is possible to receive the LFNST index reflecting the positions of the complete transform coefficients for all sub-partition blocks that are not transform coefficient information for some or individual sub-partition blocks, and the zero-out information involved in the transform process.

[0454] On the other hand, by way of an example, when the current block is divided into the plurality of sub-partition blocks, the LFNST index can be parsed regardless of whether there are transform coefficients in the region excluding the DC positions of each of the plurality of sub-partition blocks. That is, when ISP is applied to the current block, if it is allowed that the last valid coefficient is located only at the DC position for all sub-partition blocks, signaling of the LFNST index can be allowed.

[0455] On the one hand, the decoding device can derive a first variable indicating whether there is a transform coefficient in the region excluding the DC position of the current block in the residual coding step, and can also derive a second variable indicating whether there is a transform coefficient in the second region excluding the upper left first region of the current block or a sub-partition block divided from the current block.

[0456] When there is a transform coefficient in the region excluding the DC position and there is no transform coefficient in the second region, the decoding device can parse the LFNST index.

[0457] Specifically, in order to determine whether the LFNST index can be parsed, the decoding device can derive a first variable indicating whether there is the transform coefficient, i.e., a valid coefficient, in the region excluding the DC position of the current block.

[0458] The first variable is the variable LfnstDcOnly that can be derived in the residual coding process. When the index of the sub-block including the last valid coefficient in the current block is 0 and the position of the last valid coefficient in the sub-block is greater than 0, the first variable can be derived to be 0. When the first variable is 0, the LFNST index can be parsed. A sub-block means a 4×4 block used as a coding unit in residual coding and can also be named as CG (Coefficient Group). That the index of the sub-block is 0 refers to the upper left 4×4 sub-block.

[0459] The first variable can initially be set to 1, and can be maintained as 1 or changed to 0 depending on whether there is a valid coefficient in the region excluding the DC position.

[0460] The variable LfnstDcOnly indicates whether there is a non-zero coefficient at a position other than the DC component for at least one transform block within one coding unit. If there is a non-zero coefficient at a position other than the DC component for at least one transform block within one coding unit, it becomes 0, and if there is no non-zero coefficient at a position other than the DC component for all transform blocks within one coding unit, it can become 1.

[0461] Also, the decoding device can check whether zeroing out for the second region has been executed by deriving a second variable indicating whether there is a valid coefficient in the second region excluding the first region at the upper left end of the current block.

[0462] The second variable is the variable LfnstZeroOutSigCoeffFlag that can indicate that zeroing out has been executed when LFNST is applied. The second variable is initially set to 1, and if there is a valid coefficient in the second region, the second variable can be changed to 0.

[0463] The variable LfnstZeroOutSigCoeffFlag can be derived to be 0 when the index of the sub-block where the last non-zero coefficient exists is greater than 0, the width and height of the transform block are both the same as or greater than 4, or the last position of the non-zero coefficient within the sub-block where the last non-zero coefficient exists is greater than 7, and the size of the transform block is 4×4 or 8×8. A sub-block means a 4×4 block used as a coding unit in residual coding and can also be named CG (Coefficient Group). The index of the sub-block being 0 means the upper left 4×4 sub-block.

[0464] That is, when non-zero coefficients are derived in regions other than the upper left region where the LFNST conversion coefficient can exist in the conversion block, or when non-zero coefficients exist outside the 8th position in the scan order for 4×4 blocks and 8×8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0465] For example, when ISP is applied to a coding unit, if it is confirmed that conversion coefficients exist at positions where the LFNST conversion coefficient cannot exist even for one sub-partition block among all sub-partition blocks, the LFNST index signaling can be omitted. That is, when valid coefficients exist in the second region without zeroing out being performed in one sub-partition block, the LFNST index is not signaled.

[0466] On the other hand, the first region can be derived based on the size of the current block.

[0467] For example, when the size of the current block is 4×4 or 8×8, the first region is from the upper left end of the current block to the 8th sample position in the scan direction. When the current block is divided, when the size of the sub-partition block is 4×4 or 8×8, the first region is from the upper left end of the sub-partition block to the 8th sample position in the scan direction.

[0468] When the size of the current block is 4×4 or 8×8, since 8 pieces of data are output via the forward LFNST, the 8 conversion coefficients received by the decoding device can be arranged from the upper left end of the current block to the 8th sample position in the scan direction as shown in FIGS. 13(a) and 14(a).

[0469] Also, in the remaining cases where the current block size is not 4×4 or 8×8, the first region is a 4×4 region at the upper left corner of the current block. When the current block size is not 4×4 or 8×8, 16 data are output through the forward LFNST. Therefore, the 16 conversion coefficients received by the decoding device can be arranged in the 4×4 region at the upper left corner of the current block as shown in FIGS. 13(b) to (d) and FIG. 14(b).

[0470] On the other hand, the conversion coefficients that can be arranged in the first region can be arranged in the diagonal scan direction as shown in FIG. 8.

[0471] As described above, when the current block is divided into sub-partition blocks, the decoding device can parse the LFNST index when there are no conversion coefficients in all of the individual second regions for the plurality of sub-partition blocks. When there are conversion coefficients in the second region for any one of the sub-partition blocks, the LFNST index is not parsed.

[0472] As described above, LFNST can be applied to sub-partition blocks with a width and height of 4 or more, and the LFNST index for the current block, which is a coding block, can be applied to a plurality of sub-partition blocks.

[0473] On the other hand, the zero-out reflected by LFNST (including all zero-outs that can be associated with the application of LFNST) is also applied as it is to the sub-partition blocks. Therefore, it is also applied to the first region and the sub-partition blocks. That is, when the divided sub-partition block is a 4×4 block or an 8×8 block, LFNST is applied to the conversion coefficients from the upper left corner to the 8th in the scan direction of the sub-partition block. When the sub-partition block is not a 4×4 block or an 8×8 block, LFNST can be applied to the conversion coefficients in the 4×4 region at the upper left corner of the sub-partition block.

[0474] Further, in the residual coding step, the decoding device determines whether there are transform coefficients in the area excluding the upper left 16×16 area of the current block (S2130). If there are no transform coefficients in the area excluding the 16×16 area, the MTS index can be parsed (S2140).

[0475] For this purpose, the decoding device can derive a third variable indicating whether there are transform coefficients in the area excluding the upper left 16×16 area.

[0476] The third variable is the variable MtsZeroOutSigCoeffFlag which can indicate that zeroing has been performed when MTS is applied. The variable MtsZeroOutSigCoeffFlag indicates whether there are transform coefficients in the upper left area where the last valid coefficient can exist after zeroing, that is, the area other than the upper left 16×16 area, is initially set to 1, and if there are transform coefficients in the area other than the 16×16 area, its value can be changed from 1 to 0. When the value of the third variable is 0, the MTS index is not signaled.

[0477] The decoding device can derive residual samples by applying at least one of LFNST executed based on the LFNST index or MTS executed based on the MTS index. By way of example, residual samples can be derived by applying the MTS kernel derived based on the MTS index to the transform coefficients in the upper left 16×16 area (S2150).

[0478] Next, the decoding device 300 can generate restored samples based on the residual samples for the current block and the predicted samples for the current block (S2160).

[0479] The following drawings are created to illustrate a specific example of this specification. Since the names of the specific devices and the names of the specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.

[0480] FIG. 22 is a flowchart showing the operation of a video encoding apparatus according to an embodiment of this document.

[0481] Each step disclosed in FIG. 22 is based on a part of the content detailed in FIGS. 4 to 20. Therefore, specific content that overlaps with the content detailed in FIGS. 2 and 4 to 20 is omitted from the description or simplified.

[0482] An encoding apparatus 200 according to an embodiment can derive a prediction sample for a current block based on the intra prediction mode applied to the current block (S2210).

[0483] When ISP is applied to the current block, the encoding apparatus can perform prediction for each sub - partition conversion block.

[0484] The encoding apparatus can determine whether to apply ISP coding or the ISP mode to the current block, i.e., the coding block, determine in which direction the current block is divided according to the determination result, and derive the size and number of the sub - blocks to be divided.

[0485] The encoding apparatus 200 can derive a residual sample for the current block based on the prediction sample (S2220).

[0486] The encoding device 200 can derive the conversion coefficients for the current block by applying at least one of LFNST or MTS to the residual samples, and can arrange the conversion coefficients in a predetermined scanning order. By way of example, the conversion coefficients for the current block can be derived based on MTS for the residual samples (S2230).

[0487] The first conversion can be performed via a plurality of conversion kernels such as MTS, and in this case, the conversion kernel can be selected based on the intra prediction mode.

[0488] After deriving the conversion coefficients by applying MTS, the encoding device can zero out the remaining region of the current block excluding the left upper end specific region of the current block, for example, the 16×16 region (S2240).

[0489] By way of example, when the encoding device applies MTS to the first conversion of the current block, zeroing out can be performed. The encoding device can perform zeroing out to fill the region excluding the left upper end 16×16 region of the current block or the sub - partition block with 0, and can encode the MTS index with a third variable indicating whether there are conversion coefficients in the zeroing out region.

[0490] The third variable is the variable MtsZeroOutSigCoeffFlag that can indicate that zeroing out has been performed when MTS is applied. The variable MtsZeroOutSigCoeffFlag indicates whether there are conversion coefficients in the left upper end region where the last valid coefficient can exist after zeroing out after MTS execution, that is, the region other than the left upper end 16×16 region. It is initially set to 1, and when there are conversion coefficients in the region other than the 16×16 region, its value can be changed from 1 to 0. When the value of the third variable is 0, the MTS index is not encoded and signaled.

[0491] According to one example, the encoding device can derive the last valid coefficient position based on the width and height with respect to such a top left region, and binaryize the last valid coefficient position information.

[0492] According to one example, the width and height with respect to the top left region can be derived before the signaling of the MTS index.

[0493] In addition, the encoding device 200 can determine whether to perform a secondary transformation or a non-separable transformation, specifically LFNST, on the transformation coefficients for the current block, and can derive the modified transformation coefficients by applying LFNST to the transformation coefficients.

[0494] Unlike the primary transformation that separates and transforms the coefficients to be transformed in the vertical or horizontal direction, LFNST is a non-separable transformation that applies the transformation without separating the coefficients in a specific direction. Such a non-separable transformation is a low-frequency non-separable transformation that applies the transformation only to the low-frequency region that is not the entire target block to be transformed.

[0495] The encoding device can determine whether it is possible to apply LFNST to the height and width of the divided sub-partition blocks when ISP is applied to the current block.

[0496] The encoding device can determine whether it is possible to apply LFNST to the height and width of the divided sub-partition blocks. In this case, when the height and width of the sub-partition block are 4 or more, the decoding device can parse the LFNST index.

[0497] The encoding device can configure and output video information such that at least one of the LFNST index and the MTS index is signaled at the coding unit level and the MTS index is signaled immediately after the signaling of the LFNST index, and can encode the residual information derived through quantization of the transform coefficients (S2250).

[0498] Depending on specific conditions for the LFNST index, for example, when the LFNST index value is 0, the MTS index can be encoded. Also, by way of example, when the tree type of the current block is dual tree luma or single tree and the LFNST index value is 0, the MTS index can be encoded.

[0499] By way of example, when the current block is a luma block and the LFNST index indicates 0, the encoding device can encode the MTS index.

[0500] By way of example, when the tree type of the current block is dual tree, the encoding device can encode the LFNST index for each of the luma block and the chroma block.

[0501] Or, by way of example, when the tree type of the current block is single tree type, the encoding device can encode the LFNST index at the coding unit level after deriving all the transform coefficients for the luma block and the chroma block of the current block.

[0502] When the LFNST index is encoded at the coding unit level after all the transform coefficients that are not at the transform block level or the residual coding level have been derived, an LFNST index reflecting the positions of the complete transform coefficients for the luma blocks and chroma blocks that are not transform coefficient information for either one of the luma block or chroma block, and the zero-out information involved in the transform process, can be encoded.

[0503] Also, when the tree type of the current block is the dual tree type and the chroma component is coded, after all the transform coefficients for the Cb component and Cr component of the chroma block have been derived, the LFNST index can be encoded at the coding unit level.

[0504] When the LFNST index is encoded at the coding unit level after all the transform coefficients that are not at the transform block level or the residual coding level have been derived, an LFNST index reflecting the positions of the complete transform coefficients for the Cb component and Cr component that are not transform coefficient information for either one of the Cb component and Cr component, and the zero-out information involved in the transform process, can be encoded.

[0505] Also, when the current block is divided into a plurality of sub-partition blocks, after all the transform coefficients for the plurality of sub-partition blocks have been derived, the encoding device can encode the LFNST index at the coding unit level.

[0506] Similar to the above, when the LFNST index is parsed at the coding unit level after all transform coefficients that are not at the transform block level or the residual coding level are derived, a complete LFNST index reflecting the positions of the transform coefficients for all sub - partition blocks that are not transform coefficient information for some or individual sub - partition blocks and the zero - out information involved in the transform process can be encoded.

[0507] On the other hand, when the current block is divided into the plurality of sub - partition blocks, the encoding device can encode the LFNST index regardless of whether there are transform coefficients in the region excluding the DC positions of each of the plurality of sub - partition blocks. That is, when ISP is applied to the current block, if it is allowed that the last valid coefficient is located only at the DC position for all sub - partition blocks, signaling of the LFNST index can be allowed.

[0508] In the process of deriving the transform coefficients, the encoding device can derive a first variable indicating whether there are transform coefficients in the region excluding the DC position of the current block and a second variable indicating whether there are transform coefficients in the second region excluding the upper - left first region of the current block or the sub - partition block divided from the current block.

[0509] When there are transform coefficients in the region excluding the DC position and there are no transform coefficients in the second region, the encoding device can encode the LFNST index.

[0510] Specifically, the first variable is the variable LfnstDcOnly, and when the index of the sub - block including the last valid coefficient in the current block is 0 and the position of the last valid coefficient in the sub - block is greater than 0, it can be derived to be 0. When the first variable is 0, the LFNST index can be encoded.

[0511] The first variable can be initially set to 1, and can be maintained as 1 or changed to 0 depending on whether there is a valid coefficient in the area excluding the DC position.

[0512] The variable LfnstDcOnly indicates whether there is a non-zero coefficient at a position other than the DC component for at least one transform block within one coding unit. If there is a non-zero coefficient at a position other than the DC component for at least one transform block within one coding unit, it becomes 0, and if there is no non-zero coefficient at a position other than the DC component for all transform blocks within one coding unit, it can become 1.

[0513] Also, after LFNST is executed, the encoding device can zero out a second area of the current block where there are no modified transform coefficients, and can derive a second variable indicating whether there are the transform coefficients in the second area.

[0514] As shown in FIGS. 13 and 14, the remaining areas of the current block where there are no modified transform coefficients can all be processed to be 0. Such zeroing out reduces the computational amount required for the execution of the overall transformation process, and by reducing the amount of arithmetic operations required for the entire transformation process, the power consumption required for the transformation execution can be reduced. Also, the latency associated with the transformation process can be reduced and the video coding efficiency can be increased.

[0515] The second variable is the variable LfnstZeroOutSigCoeffFlag that can indicate that zeroing out has been executed when LFNST is applied. The second variable is initially set to 1, and if there is a valid coefficient in the second area, the second variable can also be changed to 0.

[0516] The variable LfnstZeroOutSigCoeffFlag can be derived to 0 when the index of the sub-block where the last non-zero coefficient exists is greater than 0, both the width and height of the transform block are 4 or more, or the position of the last non-zero coefficient within the sub-block where the last non-zero coefficient exists is greater than 7, and the size of the transform block is 4×4 or 8×8.

[0517] That is, when non-zero coefficients are derived in areas other than the upper left area where LFNST transform coefficients can exist in the transform block, or when non-zero coefficients exist outside the 8th position in the scan order for 4×4 blocks and 8×8 blocks, the variable LfnstZeroOutSigCoeffFlag is set to 0.

[0518] Since the description of the first area and the zero-out when ISP is applied are the same as the description of the decoding method, duplicate descriptions are omitted.

[0519] In addition, the encoding device can perform quantization based on the transform coefficients or the corrected transform coefficients for the current block to derive quantized transform coefficients, and encode and output video information including information on the quantized transform coefficients.

[0520] The encoding device can generate residual information including information on the quantized transform coefficients. The residual information can include the above-described transform-related information / syntax elements. The encoding device can encode video / video information including the residual information and output it in the form of a bitstream.

[0521] More specifically, the encoding device can generate information on the quantized transform coefficients and encode the generated information on the quantized transform coefficients.

[0522] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation can be omitted. When the quantization / inverse quantization is omitted, the quantized transform coefficient can be called a transform coefficient. When the transformation / inverse transformation is omitted, the transform coefficient can also be called a coefficient or a residual coefficient, or, for the sake of uniformity of expression, can still be called a transform coefficient.

[0523] Also, in this document, the quantized transform coefficient and the transform coefficient can be called a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information can include information about the transform coefficient(s), and the information about the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information about the transform coefficient(s)), and the scaled transform coefficient can be derived via an inverse transform (scaling) for the transform coefficient. The residual sample can be derived based on an inverse transform (transformation) for the scaled transform coefficient. This can be applied / expressed similarly in other parts of this document.

[0524] In the above-described embodiments, the method is described based on a flowchart in a series of steps or blocks, but this document is not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of this document.

[0525] The method according to the above-described document can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0526] In this document, when an embodiment is implemented by software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.

[0527] In addition, the decoding device and the encoding device to which this document is applied can be included in a multimedia broadcast transceiver device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a picture phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.

[0528] Also, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to this document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Also, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or can be transmitted via a wired or wireless communication network. Also, the embodiments of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0529] FIG. 23 exemplarily shows a content streaming system structure diagram to which this document is applicable.

[0530] Also, the content streaming system to which this document is applicable can greatly include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.

[0531] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by an encoding method or bitstream generation method to which this document applies, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0532] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server serves to control commands / responses between each device in the content streaming system.

[0533] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0534] Examples of the user device include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.

[0535] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.

Claims

1. 1. A video decoding method performed by a decoding device, comprising: setting an initial value of a variable to 1, said variable relating to whether there are valid coefficients in a second region excluding a first region at the top left corner of the current block; obtaining residual information from the bitstream; deriving transform coefficients for the current block based on the residual information; determining whether a significant coefficient exists in the second region during a residual coding level decoding process; parsing a multiple transform selection (MTS) index from the bitstream based on the absence of the significant coefficients in the second region; deriving residual samples by applying an MTS kernel derived based on the MTS index to the transform coefficients of the first region; generating a reconstructed picture based on the residual samples; determining whether the significant coefficient exists in the second region includes deriving a modified value, which is zero, for the variable in a residual coding level decoding process based on the significant coefficient existing in the second region; Based on the changed value being 0 for the variable, the MTS index is not parsed.

2. The image decoding method of claim 1 , wherein the first region is a 16×16 region at a top left corner of the current block.

3. The video decoding method of claim 2 , wherein the MTS index is parsed at a coding unit level.

4. and parsing an LFNST index associated with an LFNST kernel to be applied to the current block; The video decoding method of claim 3 , wherein the LFNST index and the MTS index are signaled at a coding unit level, and the MTS index is signaled immediately after the signaling of the LFNST index.

5. The tree type of the current block is a dual tree luma or a single tree, The video decoding method of claim 4 , wherein the MTS index is signaled based on the LFNST index being a value of 0.

6. A video encoding method performed by a video encoding apparatus, comprising: deriving a predicted sample for the current block; deriving a residual sample for the current block based on the predicted sample; deriving transform coefficients for the current block based on a multiple transform selection (MTS) for the residual samples; zeroing out a second region of the current block excluding a first region at a top left corner of the current block; encoding residual information derived via quantization of the transform coefficients and an MTS index associated with an MTS kernel; setting a first value of a variable to 1, said variable relating to whether there is a valid coefficient in said second region; a modified value of 0 for a variable is derived during a residual coding level encoding process based on the presence of the significant coefficient in the second region, and the MTS index is not encoded based on the modified value of 0 for the variable.

7. The method of claim 6 , wherein the first region is a 16×16 region at a top left corner of the current block.

8. The video encoding method of claim 7 , wherein the MTS index is signaled at a coding unit level.

9. signaling an LFNST index associated with an LFNST kernel to be applied to the current block; The video encoding method of claim 8 , wherein the MTS index is signaled immediately after the LFNST index is signaled.

10. The tree type of the current block is a dual tree luma or a single tree, The video encoding method of claim 9 , wherein the MTS index is signaled based on the LFNST index having a value of 0.

11. 1. A method for transmitting data for video information, comprising: obtaining a bitstream of the video information including residual information, the residual information being generated by deriving predicted samples for a current block, deriving residual samples for the current block based on the predicted samples, deriving transform coefficients for the current block based on a multiple transform selection (MTS) for the residual samples, zeroing out a second region of the current block excluding a first region at a top left corner of the current block, and encoding the residual information derived through quantization of the transform coefficients and an MTS index associated with an MTS kernel to generate the bitstream; transmitting the data including the bitstream of the video information; setting a first value of a variable to 1, said variable relating to whether there is a valid coefficient in said second region; a modified value of 0 for a variable is derived in a residual coding level encoding process based on the presence of the significant coefficient in the second region, and based on the modified value of 0 for the variable, the MTS index is not encoded.

Citation Information

Patent Citations

  • JPP7418561B